Connecting modalities: Semi-supervised segmentation and annotation of images using unaligned text corpora
R. Socher and L. Fei-Fei · 2010
Earlier work this paper cites.
Large scale image annotation: learning to rank with joint word-image embeddings
J. Weston, S. Bengio, and N. Usunier · 2010
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, M. Ranzato, and T. Mikolov · 2013
Earlier work this paper cites.
Microsoft COCO captions: Data collection and evaluation server
Original
X. Chen, H. Fang, T. Lin, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Earlier work this paper cites.
Layer normalization
Original
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Earlier work this paper cites.
Learning visual features from large weakly supervised data
A. Joulin, L. van der Maaten, A. Jabri, and N. Vasilache · 2016
Earlier work this paper cites.
A simple but tough-to-beat baseline for sentence embeddings
S. Arora, Y. Liang, and T. Ma · 2017
Earlier work this paper cites.
A closer look at memorization in deep networks
D. Arpit, S. Jastrzebski, N. Ballas, D. Krueger, E. Bengio, M. S. Kanwal, T. Maharaj, A. Fischer, A. C. Courville, Y. Bengio, and S. Lacoste-Julien · 2017
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2017
Earlier work this paper cites.
Learning visual n-grams from web data
A. Li, A. Jabri, A. Joulin, and L. van der Maaten · 2017
Earlier work this paper cites.
Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results
A. Tarvainen and H. Valpola · 2017
Earlier work this paper cites.
Word translation without parallel data
G. Lample, A. Conneau, M. Ranzato, L. Denoyer, and H. Jégou · 2018
Earlier work this paper cites.
All-but-the-top: Simple and effective postprocessing for word representations
J. Mu and P. Viswanath · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Original
A. van den Oord, Y. Li, and O. Vinyals · 2018
Earlier work this paper cites.
A convergence theory for deep learning via over-parameterization
Z. Allen-Zhu, Y. Li, and Z. Song · 2019
Earlier work this paper cites.
How contextual are contextualized word representations? comparing the geometry of bert, elmo, and GPT-2 embeddings
K. Ethayarajh · 2019
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
J. Frankle and M. Carbin · 2019
Earlier work this paper cites.