Fetching the paper…
Reading the bibliography…
Over the past few years, neural networks have re-emerged as powerful machine-learning models, yielding state-of-the-art results in fields such as image recognition and speech processing.
Distributional Structure
Harris, Z. (1954) · 1954
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Polyak, B. T. (1964) · 1964
Earlier work this paper cites.
Learning representations by back-propagating errors
Rumelhart, D. E., Hinton, G. E., & Williams, R. J. (1986) · 1986
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
Cybenko, G. (1989) · 1989
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Hornik, K., Stinchcombe, M., & White, H. (1989) · 1989
Earlier work this paper cites.
Finding Structure in Time
Elman, J. L. (1990) · 1990
Earlier work this paper cites.
Recursive Distributed Representations
Pollack, J. B. (1990) · 1990
Earlier work this paper cites.
Backpropagation through time: What it does and how to do it.
Werbos, P. J. (1990) · 1990
Earlier work this paper cites.
Convolutional Networks for Images, Speech, and Time-Series
LeCun, Y., & Bengio, Y. (1995) · 1995
Earlier work this paper cites.
Learning Task-Dependent Distributed Representations by Backpropagation Through Structure
Goller, C., & Küchler, A. (1996) · 1996
Earlier work this paper cites.
Hierarchical Recurrent Neural Networks for Long-Term Dependencies
Hihi, S. E., & Bengio, Y. (1996) · 1996
Earlier work this paper cites.
Long short-term memory
Hochreiter, S., & Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
Bidirectional recurrent neural networks
Schuster, M., & Paliwal, K. K. (1997) · 1997
Earlier work this paper cites.
Purely Functional Data Structures
Okasaki, C. (1999) · 1999
Earlier work this paper cites.
Maximum Entropy Markov Models for Information Extraction and Segmentation.
McCallum, A., Freitag, D., & Pereira, F. C. (2000) · 2000
Earlier work this paper cites.
Discriminative Training Methods for Hidden Markov Models: Theory and Experiments with Perceptron Algorithms
Collins, M. (2002) · 2002
Earlier work this paper cites.
On the algorithmic implementation of multiclass kernel-based vector machines
Crammer, K., & Singer, Y. (2002) · 2002
Earlier work this paper cites.
A Neural Probabilistic Language Model
Bengio, Y., Ducharme, R., Vincent, P., & Janvin, C. (2003) · 2003
Earlier work this paper cites.
Fast Methods for Kernel-based Text Analysis
Kudo, T., & Matsumoto, Y. (2003) · 2003
Earlier work this paper cites.
SVMTool: A general POS tagger generator based on Support Vector Machines
Giménez, J., & Màrquez, L. (2004) · 2004
Earlier work this paper cites.
Kernel Methods for Pattern Analysis
Shawe-Taylor, J., & Cristianini, N. (2004) · 2004
Earlier work this paper cites.
Coarse-to-Fine n-Best Parsing and MaxEnt Discriminative Reranking
Charniak, E., & Johnson, M. (2005) · 2005
Earlier work this paper cites.
Discriminative Reranking for Natural Language Parsing
Collins, M., & Koo, T. (2005) · 2005
Earlier work this paper cites.
Loss functions for discriminative training of energybased models.
LeCun, Y., & Huang, F. (2005) · 2005
Earlier work this paper cites.
A tutorial on energy-based learning
LeCun, Y., Chopra, S., Hadsell, R., Ranzato, M., & Huang, F. (2006) · 2006
Earlier work this paper cites.
Unsupervised Models for Morpheme Segmentation and Morphology Learning
Creutz, M., & Lagus, K. (2007) · 2007
Earlier work this paper cites.
A unified architecture for natural language processing: Deep neural networks with multitask learning
Collobert, R., & Weston, J. (2008) · 2008
Earlier work this paper cites.
Supervised sequence labelling with recurrent neural networks
Graves, A. (2008) · 2008
Earlier work this paper cites.
Algorithms for Deterministic Incremental Dependency Parsing
Nivre, J. (2008) · 2008
Earlier work this paper cites.
Search-based Structured Prediction
Hal Daumé III, Langford, J., & Marcu, D. (2009) · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X., & Bengio, Y. (2010) · 2010
Earlier work this paper cites.
An Efficient Algorithm for Easy-First Non-Directional Dependency Parsing
Goldberg, Y., & Elhadad, M. (2010) · 2010
Earlier work this paper cites.
Recurrent neural network based language model.
Mikolov, T., Karafiát, M., Burget, L., Cernocky, J., & Khudanpur, S. (2010) · 2010
Earlier work this paper cites.
Introduction to Automatic Differentiation and MATLAB Object-Oriented Programming
Neidinger, R. (2010) · 2010
Earlier work this paper cites.
Learning Continuous Phrase Representations and Syntactic Parsing with Recursive Neural Networks
Socher, R., Manning, C., & Ng, A. (2010) · 2010
Earlier work this paper cites.
Natural language processing (almost) from scratch
Collobert, R., Weston, J., Bottou, L., Karlen, M., Kavukcuoglu, K., & Kuksa, P. (2011) · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., & Singer, Y. (2011) · 2011
Earlier work this paper cites.
Deep sparse rectifier neural networks
Glorot, X., Bordes, A., & Bengio, Y. (2011) · 2011
Earlier work this paper cites.
Extensions of recurrent neural network language model
Mikolov, T., Kombrink, S., Lukáš Burget, Černocky, J. H., & Khudanpur, S. (2011) · 2011
Earlier work this paper cites.
Linguistic Structure Prediction
Smith, N. A. (2011) · 2011
Earlier work this paper cites.
Parsing Natural Scenes and Natural Language with Recursive Neural Networks
Socher, R., Lin, C. C.-Y., Ng, A. Y., & Manning, C. D. (2011) · 2011
Earlier work this paper cites.
Generating text with recurrent neural networks
Sutskever, I., Martens, J., & Hinton, G. E. (2011) · 2011
Earlier work this paper cites.
Practical recommendations for gradient-based training of deep architectures
Bengio, Y. (2012) · 2012
Earlier work this paper cites.
Stochastic gradient descent tricks
Bottou, L. (2012) · 2012
Earlier work this paper cites.
Improving neural networks by preventing co-adaptation of feature detectors
Hinton, G. E., Srivastava, N., Krizhevsky, A., Sutskever, I., & Salakhutdinov, R. R. (2012) · 2012
Earlier work this paper cites.
ImageNet Classification with Deep Convolutional Neural Networks
Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2012) · 2012
Earlier work this paper cites.
Statistical language models based on neural networks
Mikolov, T. (2012) · 2012
Earlier work this paper cites.
On the difficulty of training Recurrent Neural Networks
Pascanu, R., Mikolov, T., & Bengio, Y. (2012) · 2012
Cited alongside, same era.
Semantic Compositionality through Recursive Matrix-Vector Spaces
Socher, R., Huval, B., Manning, C. D., & Ng, A. Y. (2012) · 2012
Cited alongside, same era.
LSTM Neural Networks for Language Modeling.
Sundermeyer, M., Schlüter, R., & Ney, H. (2012) · 2012
Cited alongside, same era.
Lecture 6.5—RmsProp: Divide the gradient by a running average of its recent magnitude
Tieleman, T., & Hinton, G. (2012) · 2012
Cited alongside, same era.
ADADELTA: An Adaptive Learning Rate Method
Zeiler, M. D. (2012) · 2012
Cited alongside, same era.
word2vec Parameter Learning Explained
Rong, X. (2014) · 2014
Later among the works it cites.
Learning Character-level Representations for Part-of-Speech Tagging.
Santos, C. D., & Zadrozny, B. (2014) · 2014
Later among the works it cites.
Recursive Deep Learning For Natural Language Processing and Computer Vision
Socher, R. (2014) · 2014
Later among the works it cites.
Translation Modeling with Bidirectional Recurrent Neural Networks
Sundermeyer, M., Alkhouli, T., Wuebker, J., & Ney, H. (2014) · 2014
Later among the works it cites.
Sequence to Sequence Learning with Neural Networks
Sutskever, I., Vinyals, O., & Le, Q. V. V. (2014) · 2014
Later among the works it cites.
Recurrent Neural Networks for Word Alignment Model
Tamura, A., Watanabe, T., & Sumita, E. (2014) · 2014
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Combination of Recurrent Neural Networks and Factored Language Models for Code-Switching Language Modeling
Adel, H., Vu, N. T., & Schultz, T. (2013) · 2013
Cited alongside, same era.
Joint Language and Translation Modeling with Recurrent Neural Networks
Auli, M., Galley, M., Quirk, C., & Zweig, G. (2013) · 2013
Cited alongside, same era.
Improving deep neural networks for LVCSR using rectified linear units and dropout
Dahl, G., Sainath, T., & Hinton, G. (2013) · 2013
Cited alongside, same era.
Adaptation Data Selection using Neural Language Models: Experiments in Machine Translation
Duh, K., Neubig, G., Sudoh, K., & Tsukada, H. (2013) · 2013
Cited alongside, same era.
Training Deterministic Parsers with Non-Deterministic Oracles
Goldberg, Y., & Nivre, J. (2013) · 2013
Cited alongside, same era.
Efficient Implementation of Beam-Search Incremental Parsers
Goldberg, Y., Zhao, K., & Huang, L. (2013) · 2013
Cited alongside, same era.
Simple Customization of Recursive Neural Networks for Semantic Relation Classification
Hashimoto, K., Miwa, M., Tsuruoka, Y., & Chikayama, T. (2013) · 2013
Cited alongside, same era.
Later among the works it cites.
A Neural Network Approach to Selectional Preference Acquisition
Van de Cruys, T. (2014) · 2014
Later among the works it cites.
Recurrent Neural Network Regularization
Zaremba, W., Sutskever, I., & Vinyals, O. (2014) · 2014
Later among the works it cites.
Relation Classification via Convolutional Deep Neural Network
Zeng, D., Liu, K., Lai, S., Zhou, G., & Zhao, J. (2014) · 2014
Later among the works it cites.
Improved Transition-based Parsing by Modeling Characters instead of Words with LSTMs
Ballesteros, M., Dyer, C., & Smith, N. A. (2015) · 2015
Closest in time.
Automatic differentiation in machine learning: a survey
Baydin, A. G., Pearlmutter, B. A., Radul, A. A., & Siskind, J. M. (2015) · 2015
Closest in time.
Deep Learning
Bengio, Y., Goodfellow, I. J., & Courville, A. (2015) · 2015
Closest in time.
Non-Linear Text Regression with a Deep Convolutional Neural Network
Bitvai, Z., & Cohn, T. (2015) · 2015
Closest in time.
Event Extraction via Dynamic Multi-Pooling Convolutional Neural Networks
Chen, Y., Xu, L., Liu, K., Zeng, D., & Zhao, J. (2015) · 2015
Closest in time.
Fast and Accurate Preordering for SMT using Neural Networks
de Gispert, A., Iglesias, G., & Byrne, B. (2015) · 2015
Closest in time.
Question Answering over Freebase with Multi-Column Convolutional Neural Networks
Dong, L., Wei, F., Zhou, M., & Xu, K. (2015) · 2015
Closest in time.
Classifying Relations by Ranking with Convolutional Neural Networks
dos Santos, C., Xiang, B., & Zhou, B. (2015) · 2015
Closest in time.
Neural CRF Parsing
Durrett, G., & Klein, D. (2015) · 2015
Closest in time.
Transition-Based Dependency Parsing with Stack Long Short-Term Memory
Dyer, C., Ballesteros, M., Ling, W., Matthews, A., & Smith, N. A. (2015) · 2015
Closest in time.
Sentence Compression by Deletion with LSTMs
Filippova, K., Alfonseca, E., Colmenares, C. A., Kaiser, L., & Vinyals, O. (2015) · 2015
Closest in time.
Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning
Gal, Y., & Ghahramani, Z. (2015) · 2015
Closest in time.
Greff, K., Srivastava, R. K., Koutník, J., Steunebrink, B. R., & Schmidhuber, J. (2015) · 2015
Closest in time.
Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification
He, K., Zhang, X., Ren, S., & Sun, J. (2015) · 2015
Closest in time.
Deep Unordered Composition Rivals Syntactic Methods for Text Classification
Iyyer, M., Manjunatha, V., Boyd-Graber, J., & Daumé III, H. (2015) · 2015
Closest in time.
Effective Use of Word Order for Text Categorization with Convolutional Neural Networks
Johnson, R., & Zhang, T. (2015) · 2015
Closest in time.
An Empirical Exploration of Recurrent Network Architectures
Jozefowicz, R., Zaremba, W., & Sutskever, I. (2015) · 2015
Closest in time.
Visualizing and Understanding Recurrent Networks
Karpathy, A., Johnson, J., & Li, F.-F. (2015) · 2015
Closest in time.
The Forest Convolutional Network: Compositional Distributional Semantics with a Neural Chart and without Binarization
Le, P., & Zuidema, W. (2015) · 2015
Closest in time.
A Simple Way to Initialize Recurrent Networks of Rectified Linear Units
Le, Q. V., Jaitly, N., & Hinton, G. E. (2015) · 2015
Closest in time.
Improving Distributional Similarity with Lessons Learned from Word Embeddings
Levy, O., Goldberg, Y., & Dagan, I. (2015) · 2015
Closest in time.
Two/Too Simple Adaptations of Word2Vec for Syntax Problems
Ling, W., Dyer, C., Black, A. W., & Trancoso, I. (2015a) · 2015
Closest in time.
Finding Function in Form: Compositional Character Models for Open Vocabulary Word Representation
Ling, W., Dyer, C., Black, A. W., Trancoso, I., Fermandez, R., Amir, S., Marujo, L., & Luis, T. (2015b) · 2015
Closest in time.
A Dependency-Based Neural Network for Relation Classification
Liu, Y., Wei, F., Li, S., Ji, H., Zhou, M., & WANG, H. (2015) · 2015
Closest in time.
Dependency-based Convolutional Neural Networks for Sentence Embedding
Ma, M., Huang, L., Zhou, B., & Xiang, B. (2015) · 2015
Closest in time.
Multi-domain Dialog State Tracking using Recurrent Neural Networks
Mrkšić, N., Ó Séaghdha, D., Thomson, B., Gasic, M., Su, P.-H., Vandyke, D., Wen, T.-H., & Young, S. (2015) · 2015
Closest in time.
Event Detection and Domain Adaptation with Convolutional Neural Networks
Nguyen, T. H., & Grishman, R. (2015) · 2015
Closest in time.
An Effective Neural Network Model for Graph-based Dependency Parsing
Pei, W., Ge, T., & Chang, B. (2015) · 2015
Closest in time.
Learning Tag Embeddings and Tag-specific Composition Functions in Recursive Neural Network
Qian, Q., Tian, B., Huang, M., Liu, Y., Zhu, X., & Zhu, X. (2015) · 2015
Closest in time.
A Neural Network Approach to Context-Sensitive Generation of Conversational Responses
Sordoni, A., Galley, M., Auli, M., Brockett, C., Ji, Y., Mitchell, M., Nie, J.-Y., Gao, J., & Dolan, B. (2015) · 2015
Closest in time.
Improved Semantic Representations From Tree-Structured Long Short-Term Memory Networks
Tai, K. S., Socher, R., & Manning, C. D. (2015) · 2015
Closest in time.
Transition-based Neural Constituent Parsing
Watanabe, T., & Sumita, E. (2015) · 2015
Closest in time.
Structured Training for Neural Network Transition-Based Parsing
Weiss, D., Alberti, C., Collins, M., & Petrov, S. (2015) · 2015
Closest in time.
CCG Supertagging with a Recurrent Neural Network
Xu, W., Auli, M., & Clark, S. (2015) · 2015
Closest in time.
Convolutional Neural Network for Paraphrase Identification
Yin, W., & Schütze, H. (2015) · 2015
Closest in time.
A Neural Probabilistic Structured-Prediction Model for Transition-Based Dependency Parsing
Zhou, H., Zhang, Y., Huang, S., & Chen, J. (2015) · 2015
Closest in time.
Recursive Deep Models for Discourse Parsing
Li, J., Li, R., & Hovy, E. (2014) · 2069
Closest in time.