Fetching the paper…
Reading the bibliography…
Humans learn language by interaction with their environment and listening to other humans.
M. D. S. Braine and M. Bowerman, “Children’s first word combinations,”
1976
Earlier work this paper cites.
J. M. Pine and E. Lieven, “Reanalysing rote-learned phrases: individual differences in the transition to multi-word speech,”
1993
Earlier work this paper cites.
M. Tomasello, “First steps toward a usage-based theory of language acquisition,”
2000
Earlier work this paper cites.
E. Lieven, H. Behrens, J. Speares, and M. Tomasello, “Early syntactic creativity: a usage-based approach,”
2003
Earlier work this paper cites.
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean, “Distributed representations of words and phrases and their compositionality,” in
2013
Earlier work this paper cites.
M. Hodosh, P. Young, and J. Hockenmaier, “Framing image description as a ranking task: Data, models and evaluation metrics,”
2013
Earlier work this paper cites.
J. Pennington, R. Socher, and C. Manning, “Glove: Global vectors for word representation,” in
2014
Earlier work this paper cites.
Q. B. Nguyen, J. Gehring, M. Müller, S. Stücker, and A. Waibel, “Multilingual shifting deep bottleneck features for low-resource asr,” in
2014
Earlier work this paper cites.
J. Chung, C. Gulcehre, K. Cho, and Y. Bengio, “Empirical evaluation of gated recurrent neural networks on sequence modeling,” in
2014
Earlier work this paper cites.
R. Kiros, Y. Zhu, R. R. Salakhutdinov, R. Zemel, R. Urtasun, A. Torralba, and S. Fidler, “Skip-thought vectors,” in
2015
Earlier work this paper cites.
A. Karpathy and L. Fei-Fei, “Deep visual-semantic alignments for generating image descriptions,” in
2015
Cited alongside, same era.
D. Harwath and J. Glass, “Deep multimodal semantic embeddings for speech and images,” in
2015
Cited alongside, same era.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in
2015
Cited alongside, same era.
Z.-Q. Wang and D. Wang, “A joint training framework for robust automatic speech recognition,”
2016
Cited alongside, same era.
E. Agirre, C. Banea, D. Cer, M. Diab, A. Gonzalez-Agirre, R. Mihalcea, G. Rigau, and J. Wiebe, “Semeval-2016 task 1: Semantic textual similarity, monolingual and cross-lingual evaluation,” in
2016
Cited alongside, same era.
G. Chrupała, L. Gelderloos, and A. Alishahi, “Representations of language in a model of visually grounded speech signal,” in
2017
Later among the works it cites.
R. Fer, P. Matejka, F. Grezl, O. Plchot, K. Vesely, and J. H. Cernocky, “Multilingually trained bottleneck features in spoken language recognition,”
2017
Later among the works it cites.
L. N. Smith, “Cyclical learning rates for training neural networks,” in
2017
Later among the works it cites.
G. Huang, Y. Li, G. Pleiss, Z. Liu, J. E. Hopcroft, and K. Q. Weinberger, “Snapshot Ensembles: Train 1, get M for free,” in
2017
Later among the works it cites.
H. Kamper, S. Settle, G. Shakhnarovich, and K. Livescu, “Visually grounded learning of keyword prediction from untranscribed speech,”
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in
2016
Cited alongside, same era.
A. Conneau, D. Kiela, H. Schwenk, L. Barrault, and A. Bordes, “Supervised Learning of Universal Sentence Representations from Natural Language Inference Data,” in
2017
Cited alongside, same era.
F. Faghri, D. J. Fleet, R. Kiros, and S. Fidler, “VSE++: improved visual-semantic embeddings,”
2017
Cited alongside, same era.
D. Merkx and S. L. Frank, “Learning semantic sentence representations from visually grounded language without lexical knowledge,”
Cited in the paper.
J. Drexler and J. Glass, “Analysis of audio-visual features for unsupervised speech recognition,” in
2017
Later among the works it cites.
X. Wang, H. Pham, P. Yin, and G. Neubig, “A tree-based decoder for neural machine translation,” in
2018
Later among the works it cites.
D. Kiela, A. Conneau, A. Jabri, and M. Nickel, “Learning visually grounded sentence representations,” in
2018
Later among the works it cites.
D. Harwath, A. Recasens, D. Surís, G. Chuang, A. Torralba, and J. Glass, “Jointly discovering visual objects and spoken words from raw sensory input,”
2018
Later among the works it cites.