Fetching the paper…
Reading the bibliography…
This is a lecture note for the course DS-GA 3001 <Natural Language Understanding with Distributed Representation> at the Center for Data Science , New York University in Fall, 2015.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Translation
W. Weaver · 1955
Earlier work this paper cites.
The bandwagon (edtl.)
C. Shannon · 1956
Earlier work this paper cites.
A synopsis of linguistic theory 1930-1955
J. R. Firth · 1957
Earlier work this paper cites.
A review of B. F. skinner’s verbal behavior
N. Chomsky · 1959
Earlier work this paper cites.
Principles of neurodynamics: perceptrons and the theory of brain mechanisms
F. Rosenblatt · 1962
Earlier work this paper cites.
Geometrical and statistical properties of systems of linear inequalities with applications in pattern recognition
T. M. Cover · 1965
Earlier work this paper cites.
Linguistic contributions to the study of mind (future)
N. Chomsky · 1968
Earlier work this paper cites.
Understanding natural language
T. Winograd · 1972
Earlier work this paper cites.
Statistical analysis of non-lattice data
J. Besag · 1975
Earlier work this paper cites.
Practical Methods of Optimization
R. Fletcher · 1987
Earlier work this paper cites.
Training stochastic model recognition algorithms as networks can lead to maximum mutual information estimation of parameters
J. S. Bridle · 1990
Earlier work this paper cites.
A statistical approach to machine translation
P. F. Brown, J. Cocke, S. A. D. Pietra, V. J. D. Pietra, F. Jelinek, J. D. Lafferty, R. L. Mercer, and P. S. Roossin · 1990
Earlier work this paper cites.
Transforming neural-net output levels to probability distributions
J. Denker and Y. Lecun · 1991
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Y. Bengio, P. Simard, and P. Frasconi · 1994
Earlier work this paper cites.
Mixture density networks
C. M. Bishop · 1994
Earlier work this paper cites.
Improved backing-off for m-gram language modeling
R. Kneser and H. Ney · 1995
Earlier work this paper cites.
Artificial intelligence: a modern approach
S. Russell and P. Norvig · 1995
Earlier work this paper cites.
The Nature of Statistical Learning Theory
V. Vapnik · 1995
Earlier work this paper cites.
An empirical study of smoothing techniques for language modeling
S. F. Chen and J. Goodman · 1996
Earlier work this paper cites.
Recursive hetero-associative memories for translation
M. L. Forcada and R. P. Ñeco · 1997
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Online algorithms and stochastic approximations
L. Bottou · 1998
Earlier work this paper cites.
Efficient BackProp
Y. LeCun, L. Bottou, G. Orr, and K. R. Müller · 1998
Earlier work this paper cites.
Object recognition from local scale-invariant features
D. G. Lowe · 1999
Earlier work this paper cites.
Foundations of statistical natural language processing
C. D. Manning and H. Schütze · 1999
Earlier work this paper cites.
Learning to forget: Continual prediction with lstm
F. A. Gers, J. Schmidhuber, and F. Cummins · 2000
Earlier work this paper cites.
Statistical modeling: The two cultures (with comments and a rejoinder by the author)
L. Breiman et al · 2001
Earlier work this paper cites.
Gradient flow in recurrent nets: the difficulty of learning long-term dependencies
S. Hochreiter, Y. Bengio, P. Frasconi, and J. Schmidhuber · 2001
Earlier work this paper cites.
Syntactic structures
N. Chomsky · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Statistical phrase-based translation
P. Koehn, F. J. Och, and D. Marcu · 2003
Earlier work this paper cites.
The web as a parallel corpus
P. Resnik and N. A. Smith · 2003
Earlier work this paper cites.
Pharaoh: a beam search decoder for phrase-based statistical machine translation models
P. Koehn · 2004
Earlier work this paper cites.
Limited discrepancy beam search
D. Furcy and S. Koenig · 2005
Earlier work this paper cites.
Europarl: A parallel corpus for statistical machine translation
P. Koehn · 2005
Earlier work this paper cites.
Beam-stack search: Integrating backtracking with beam search
R. Zhou and E. A. Hansen · 2005
Earlier work this paper cites.
Neural probabilistic language models
Y. Bengio, H. Schwenk, J.-S. Senécal, F. Morin, and J.-L. Gauvain · 2006
Earlier work this paper cites.
Pattern recognition and machine learning
C. M. Bishop · 2006
Cited alongside, same era.
Re-evaluation the role of bleu in machine translation research
C. Callison-Burch, M. Osborne, and P. Koehn · 2006
Cited alongside, same era.
Semi-Supervised Learning
O. Chapelle, B. Schölkopf, and A. Zien, editors · 2006
Cited alongside, same era.
Extreme learning machine: Theory and applications
G.-B. Huang, Q.-Y. Zhu, and C.-K. Siew · 2006
Cited alongside, same era.
Poverty of the stimulus? a rational approach
A. Perfors, J. Tenenbaum, and T. Regier · 2006
Cited alongside, same era.
A study of translation edit rate with targeted human annotation
M. Snover, B. Dorr, R. Schwartz, L. Micciulla, and J. Makhoul · 2006
Cited alongside, same era.
Meteor universal: Language specific translation evaluation for any target language
M. Denkowski and A. Lavie · 2014
Later among the works it cites.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2014
Later among the works it cites.
Edinburgh’s phrase-based machine translation systems for WMT-14
N. Durrani, B. Haddow, P. Koehn, and K. Heafield · 2014
Later among the works it cites.
From captions to visual concepts and back
H. Fang, S. Gupta, F. Iandola, R. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. C. Platt, C. L. Zitnick, and G. Zweig · 2014
Later among the works it cites.
On using very large target vocabulary for neural machine translation
S. Jean, K. Cho, R. Memisevic, and Y. Bengio · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Automatic acquisition of chinese–english parallel corpus from the web
Y. Zhang, K. Wu, J. Gao, and P. Vines · 2006
Cited alongside, same era.
Continuous space language models
H. Schwenk · 2007
Cited alongside, same era.
The matrix cookbook
K. B. Petersen, M. S. Pedersen, et al · 2008
Cited alongside, same era.
Visualizing data using t-SNE
L. van der Maaten and G. E. Hinton · 2008
Cited alongside, same era.
Theano: a cpu and gpu math expression compiler
J. Bergstra, O. Breuleux, F. Bastien, P. Lamblin, R. Pascanu, G. Desjardins, J. Turian, D. Warde-Farley, and Y. Bengio · 2010
Cited alongside, same era.
A tutorial on Bayesian optimization of expensive cost functions, with application to active user modeling and hierarchical reinforcement learning
E. Brochu, V. M. Cora, and N. de Freitas · 2010
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and F.-F. Li · 2014
Later among the works it cites.
Learning image embeddings using convolutional neural networks for improved multi-modal semantics
D. Kiela and L. Bottou · 2014
Later among the works it cites.
Multimodal neural language models
R. Kiros, R. Salakhutdinov, and R. Zemel · 2014
Later among the works it cites.
Neural word embedding as implicit matrix factorization
O. Levy and Y. Goldberg · 2014
Later among the works it cites.
Explain images with multimodal recurrent neural networks
J. Mao, W. Xu, Y. Yang, J. Wang, and A. L. Yuille · 2014
Later among the works it cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Later among the works it cites.
Verbal behavior
B. F. Skinner · 2014
Later among the works it cites.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Later among the works it cites.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2014
Later among the works it cites.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2014
Later among the works it cites.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Closest in time.
Automatic differentiation in machine learning: a survey
A. G. Baydin, B. A. Pearlmutter, and A. A. Radul · 2015
Closest in time.
Large-scale simple question answering with memory networks
A. Bordes, N. Usunier, S. Chopra, and J. Weston · 2015
Closest in time.
Describing multimedia content using attention-based encoder–decoder networks
K. Cho, A. Courville, and Y. Bengio · 2015
Closest in time.
Multi-task learning for multiple language translation
D. Dong, H. Wu, W. He, D. Yu, and H. Wang · 2015
Closest in time.
A primer on neural network models for natural language processing
Y. Goldberg · 2015
Closest in time.
K. Greff, R. K. Srivastava, J. Koutník, B. R. Steunebrink, and J. Schmidhuber · 2015
Closest in time.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Closest in time.
A fast variational approach for learning markov random field language models
Y. Jernite, A. M. Rush, and D. Sontag · 2015
Closest in time.
Document context language models
Y. Ji, T. Cohn, L. Kong, C. Dyer, and J. Eisenstein · 2015
Closest in time.
An empirical exploration of recurrent network architectures
R. Jozefowicz, W. Zaremba, and I. Sutskever · 2015
Closest in time.
Character-aware neural language models
Y. Kim, Y. Jernite, D. Sontag, and A. M. Rush · 2015
Closest in time.
A simple way to initialize recurrent networks of rectified linear units
Q. V. Le, N. Jaitly, and G. E. Hinton · 2015
Closest in time.
Deep learning
Y. LeCun, Y. Bengio, and G. Hinton · 2015
Closest in time.
Character-based neural machine translation
W. Ling, I. Trancoso, C. Dyer, and A. W. Black · 2015
Closest in time.
Multi-task sequence to sequence learning
M.-T. Luong, Q. V. Le, I. Sutskever, O. Vinyals, and L. Kaiser · 2015
Closest in time.
Effective approaches to attention-based neural machine translation
M.-T. Luong, H. Pham, and C. D. Manning · 2015
Closest in time.
Deep learning in neural networks: An overview
J. Schmidhuber · 2015
Closest in time.
Neural machine translation of rare words with subword units
R. Sennrich, B. Haddow, and A. Birch · 2015
Closest in time.
From feedforward to recurrent lstm neural networks for language modeling
M. Sundermeyer, H. Ney, and R. Schluter · 2015
Closest in time.
Larger-context language modelling
T. Wang and K. Cho · 2015
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhutdinov, R. Zemel, and Y. Bengio · 2015
Closest in time.