Automatically constructing a corpus of sentential paraphrases
Dolan, W. B. and Brockett, C · 2005
Earlier work this paper cites.
The pascal recognizing textual entailment challenge
Dagan, I., Glickman, O., and Magnini, B · 2006
Earlier work this paper cites.
The pascal recognising textual entailment challenge
Haim, R. B., Dagan, I., Dolan, B., Ferro, L., Giampiccolo, D., Magnini, B., and Szpektor, I · 2006
Earlier work this paper cites.
The pascal recognizing textual entailment challenge
Giampiccolo, D., Magnini, B., Dagan, I., and Dolan, B · 2007
Earlier work this paper cites.
The fifth pascal recognizing textual entailment challenge
Bentivogli, L., Clark, P., Dagan, I., and Giampiccolo, D · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Predictive coding under the free-energy principle
Friston, K. and Kiebel, S · 2009
Earlier work this paper cites.
The free-energy principle: a unified brain theory?
Friston, K · 2010
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A., and Potts, C · 2013
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Original
Zhu, Y., Kiros, R., Zemel, R. S., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S · 2015
Earlier work this paper cites.
Layer normalization
Original
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Deep networks with stochastic depth
Original
Huang, G., Sun, Y., Liu, Z., Sedra, D., and Weinberger, K · 2016
Earlier work this paper cites.
SGDR: stochastic gradient descent with restarts
Original
Loshchilov, I. and Hutter, F · 2016
Earlier work this paper cites.
Squad: 100, 000+ questions for machine comprehension of text
Original
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A · 2016
Earlier work this paper cites.
Instance normalization: The missing ingredient for fast stylization
Original
Ulyanov, D., Vedaldi, A., and Lempitsky, V. S · 2016
Earlier work this paper cites.
See, hear, and read: Deep aligned representations, 2017
Aytar, Y., Vondrick, C., and Torralba, A · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Gemmeke, J. F., Ellis, D. P. W., Freedman, D., Jansen, A., Lawrence, W., Moore, R. C., Plakal, M., and Ritter, M · 2017
Earlier work this paper cites.
Learned in translation: Contextualized word vectors
Original
McCann, B., Bradbury, J., Xiong, C., and Socher, R · 2017
Earlier work this paper cites.
Learning from between-class examples for deep sound recognition
Original
Tokozume, Y., Ushiku, Y., and Harada, T · 2017
Earlier work this paper cites.
Neural discrete representation learning
van den Oord, A., Vinyals, O., et al · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Semeval-2017 task 1: Semantic textual similarity - multilingual and cross-lingual focused evaluation
Cer, D. M., Diab, M. T., Agirre, E., Lopez-Gazpio, I., and Specia, L · 2018
Earlier work this paper cites.
Unsupervised machine translation using monolingual corpora only
Lample, G., Denoyer, L., and Ranzato, M · 2018
Earlier work this paper cites.
Deep contextualized word representations
Peters, M. E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., and Zettlemoyer, L · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I · 2018
Earlier work this paper cites.
Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results, 2018
Tarvainen, A. and Valpola, H · 2018
Earlier work this paper cites.