Fetching the paper…
Reading the bibliography…
Attention networks have proven to be an effective approach for embedding categorical inference within a deep neural network.
Trainable Grammars for Speech Recognition
James K. Baker · 1979
Earlier work this paper cites.
Three New Probabilistic Models for Dependency Parsing: An Exploration
Jason M. Eisner · 1996
Earlier work this paper cites.
Gradient-based Learning Applied to Document Recognition
Yann LeCun, Leon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data
John Lafferty, Andrew McCallum, and Fernando Pereira · 2001
Earlier work this paper cites.
Dependency Parsing as Belief Propagation
David A. Smith and Jason Eisner · 2008
Earlier work this paper cites.
First- and Second-Order Expectation Semirings with Applications to Minimum-Risk Training on Translation Forests
Zhifei Li and Jason Eisner · 2009
Earlier work this paper cites.
Conditional Neural Fields
Jian Peng, Liefeng Bo, and Jinbo Xu · 2009
Earlier work this paper cites.
Neural Conditional Random Fields
Trinh-Minh-Tri Do and Thierry Artiéres · 2010
Earlier work this paper cites.
Natural Language Processing (almost) from Scratch
Ronan Collobert, Jason Weston, Leon Bottou, Michael Karlen, Koray Kavukcuoglu, and Pavel Kuksa · 2011
Earlier work this paper cites.
Parameter Learning with Truncated Message-Passing
Justin Domke · 2011
Earlier work this paper cites.
Adaptive Subgradient Methods for Online Learning and Stochastic Optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Pointwise Prediction for Robust, Adaptable Japanese Morphological Analysis
Graham Neubig, Yosuke Nakata, and Shinsuke Mori · 2011
Earlier work this paper cites.
Empirical Risk Minimization of Graphical Model Parameters Given Approximate Inference, Decoding, and Model Structure
Veselin Stoyanov, Alexander Ropson, and Jason Eisner · 2011
Earlier work this paper cites.
Generic methods for optimization-based modeling
Justin Domke · 2012
Earlier work this paper cites.
Minimum-Risk Training of Approximate CRF-based NLP Systems
Veselin Stoyanov and Jason Eisner · 2012
Earlier work this paper cites.
Alex Graves, Greg Wayne, and Ivo Danihelka · 2014
Earlier work this paper cites.
Deep Structured Output Learning for Unconstrained Text Recognition
Max Jaderberg, Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman · 2014
Earlier work this paper cites.
GloVe: Global Vectors for Word Representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning · 2014
Earlier work this paper cites.
Jason Weston, Sumit Chopra, and Antoine Bordes · 2014
Earlier work this paper cites.
Neural Machine Translation by Jointly Learning to Align and Translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Cited alongside, same era.
Tree-Structured Composition in Neural Networks without Tree-Structured Architectures
Samuel R. Bowman, Christopher D. Manning, and Christopher Potts · 2015
Cited alongside, same era.
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals · 2015
Cited alongside, same era.
Learning Deep Structured Models
Liang-Chieh Chen, Alexander G. Schwing, Alan L. Yuille, and Raquel Urtasun · 2015
Cited alongside, same era.
Describing Multimedia Content using Attention-based Encoder-Decoder Networks
Kyunghyun Cho, Aaron Courville, and Yoshua Bengio · 2015
Cited alongside, same era.
Attention-Based Models for Speech Recognition
Structured Prediction Energy Networks
David Belanger and Andrew McCallum · 2016
Later among the works it cites.
A Fast Unified Model for Parsing and Sentence Understanding
Samuel R. Bowman, Jon Gauthier, Abhinav Rastogi, Raghav Gupta, Christopher D. Manning, and Christopher Potts · 2016
Later among the works it cites.
Enhancing and Combining Sequential and Tree LSTM for Natural Language Inference
Qian Chen, Xiaodan Zhu, Zhenhua Ling, Si Wei, and Hui Jiang · 2016
Later among the works it cites.
Inside-Outside and Forward-Backward Algorithms are just Backprop
Jason M. Eisner · 2016
Later among the works it cites.
Hybrid Computing Using a Neural Network with Dynamic External Memory
Alex Graves, Greg Wayne, Malcolm Reynolds, Tim Harley, Ivo Danihelka, Agnieszka Grabska-Barwinska, Sergio Gomez Colmenarejo, Edward Grefenstette, Tiago Ramalho, John Agapiou, Adria Puigdomenech Badia, Karl Moritz Hermann, Yori Zwols, Georg Ostrovski, Adam Cain, Helen King, Christopher Summerfield, Phil Blunsom, Koray Kavukcuoglu, and Demis Hassabis · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jan Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio · 2015
Cited alongside, same era.
Neural CRF Parsing
Greg Durrett and Dan Klein · 2015
Cited alongside, same era.
Approximation-Aware Dependency Parsing by Belief Propagation
Matthew R. Gormley, Mark Dredze, and Jason Eisner · 2015
Cited alongside, same era.
Learning to Transduce with Unbounded Memory
Edward Grefenstette, Karl Moritz Hermann, Mustafa Suleyman, and Phil Blunsom · 2015
Cited alongside, same era.
Teaching Machines to Read and Comprehend
Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom · 2015
Cited alongside, same era.
Adam: A Method for Stochastic Optimization
Diederik Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Effective Approaches to Attention-based Neural Machine Translation
Minh-Thang Luong, Hieu Pham, and Christopher D. Manning · 2015
Cited alongside, same era.
Simple and Accurate Dependency Parsing using Bidirectional LSTM Feature Representations
Eliyahu Kipperwasser and Yoav Goldberg · 2016
Later among the works it cites.
Segmental Recurrent Neural Networks
Lingpeng Kong, Chris Dyer, and Noah A. Smith · 2016
Later among the works it cites.
Neural Architectures for Named Entity Recognition
Guillaume Lample, Miguel Ballesteros, Sandeep Subramanian, Kazuya Kawakami, and Chris Dyer · 2016
Later among the works it cites.
Segmental Recurrent Neural Networks for End-to-End Speech Recognition
Liang Lu, Lingpeng Kong, Chris Dyer, Noah A. Smith, and Steve Renals · 2016
Later among the works it cites.
Natural language inference by tree-based convolution and heuristic matching
Lili Mou, Rui Men, Ge Li, Yan Xu, Lu Zhang, Rui Yan, and Zhi Jin · 2016
Later among the works it cites.
Neural Tree Indexers for Text Understanding
Tsendsuren Munkhdalai and Hong Yu · 2016
Later among the works it cites.
Aspec: Asian scientific paper excerpt corpus
Toshiaki Nakazawa, Manabu Yaguchi, Kiyotaka Uchimoto, Masao Utiyama, Eiichiro Sumita, Sadao Kurohashi, and Hitoshi Isahara · 2016
Later among the works it cites.
A Decomposable Attention Model for Natural Language Inference
Ankur P. Parikh, Oscar Tackstrom, Dipanjan Das, and Jakob Uszkoreit · 2016
Later among the works it cites.
Reasoning about Entailment with Neural Attention
Tim Rocktäschel, Edward Grefenstette, Karl Moritz Hermann, Tomas Kocisky, and Phil Blunsom · 2016
Later among the works it cites.
Proximal Deep Structured Models
Shenlong Wang, Sanja Fidler, and Raquel Urtasun · 2016
Later among the works it cites.
Learning Natural Language Inference with LSTM
Shuohang Wang and Jing Jiang · 2016
Later among the works it cites.
Online Segment to Segment Neural Transduction
Lei Yu, Jan Buys, and Phil Blunsom · 2016
Later among the works it cites.
Textual Entailment with Structured Attentions and Composition
Kai Zhao, Liang Huang, and Minbo Ma · 2016
Later among the works it cites.
The Neural Noisy Channel
Lei Yu, Phil Blunsom, Chris Dyer, Edward Grefenstette, and Tomas Kocisky · 2017
Closest in time.