What size net gives valid generalization?
Baum, E. and Haussler, D · 1988
Earlier work this paper cites.
The pascal recognising textual entailment challenge
Dagan, I., Glickman, O., and Magnini, B · 2005
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Dolan, W. B. and Brockett, C · 2005
Earlier work this paper cites.
The second pascal recognising textual entailment challenge
Haim, R. B., Dagan, I., Dolan, B., Ferro, L., Giampiccolo, D., Magnini, B., and Szpektor, I · 2006
Earlier work this paper cites.
The third pascal recognizing textual entailment challenge
Giampiccolo, D., Magnini, B., Dagan, I., and Dolan, B · 2007
Earlier work this paper cites.
The fifth pascal recognizing textual entailment challenge
Bentivogli, L., Clark, P., Dagan, I., and Giampiccolo, D · 2009
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A. Y., and Potts, C · 2013
Earlier work this paper cites.
Learning to learn by gradient descent by gradient descent
Andrychowicz, M., Denil, M., Colmenarejo, S. G., Hoffman, M. W., Pfau, D., Schaul, T., and de Freitas, N · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z · 2016
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Earlier work this paper cites.
Toward controlled generation of text
Hu, Z., Yang, Z., Liang, X., Salakhutdinov, R., and Xing, E. P · 2017
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
Kirkpatrick, J., Pascanu, R., Rabinowitz, N., Veness, J., Desjardins, G., Rusu, A. A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., et al · 2017
Earlier work this paper cites.
Temporal ensembling for semi-supervised learning
Laine, S. and Aila, T · 2017
Earlier work this paper cites.
Adversarial training methods for semi-supervised text classification
Miyato, T., Dai, A. M., and Goodfellow, I. J · 2017
Earlier work this paper cites.
First Quora dataset release: Question pairs, 2017
Shankar, I., Nikhil, D., and Kornél, C · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Learning to model the tail
Wang, Y.-X., Ramanan, D., and Hebert, M · 2017
Earlier work this paper cites.
Bilevel programming for hyperparameter optimization and meta-learning
Franceschi, L., Frasconi, P., Salzo, S., Grazzi, R., and Pontil, M · 2018
Earlier work this paper cites.
Using trusted data to train deep networks on labels corrupted by severe noise
Hendrycks, D., Mazeika, M., Wilson, D., and Gimpel, K · 2018
Earlier work this paper cites.
Marian: Fast neural machine translation in C++
Junczys-Dowmunt, M., Grundkiewicz, R., Dwojak, T., Hoang, H. T., Heafield, K., Neckermann, T., Seide, F., Germann, U., Aji, A. F., Bogoychev, N., Martins, A. F. T., and Birch, A · 2018
Earlier work this paper cites.
Weakly-supervised neural text classification
Meng, Y., Shen, J., Zhang, C., and Han, J · 2018
Earlier work this paper cites.
Style transfer through back-translation
Prabhumoye, S., Tsvetkov, Y., Salakhutdinov, R., and Black, A. W · 2018
Earlier work this paper cites.
Learning to reweight examples for robust deep learning
Ren, M., Zeng, W., Yang, B., and Urtasun, R · 2018
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Williams, A., Nangia, N., and Bowman, S · 2018
Earlier work this paper cites.
Learning to teach with dynamic loss functions
Wu, L., Tian, F., Xia, Y., Fan, Y., Qin, T., Lai, J., and Liu, T.-Y · 2018
Earlier work this paper cites.