Fetching the paper…
Reading the bibliography…
Dynamic neural network toolkits such as PyTorch, DyNet, and Chainer offer more flexibility for implementing models that cope with data of varying dimensions and structure, relative to toolkits that operate on statically declared computations (e.g., TensorFlow, CNTK, and Theano).
Automatic differentiation of algorithms
Michael Bartholomew-Briggs, Steven Brown, Bruce Christianson, and Laurence Dixon · 2000
Earlier work this paper cites.
Scheduling with batching: A review
Chris N. Potts and Mikhail Y. Kovalyov · 2000
Earlier work this paper cites.
Theano: A CPU and GPU math compiler in Python
James Bergstra, Olivier Breuleux, Frédéric Bastien, Pascal Lamblin, Razvan Pascanu, Guillaume Desjardins, Joseph Turian, David Warde-Farley, and Yoshua Bengio · 2010
Earlier work this paper cites.
Parsing natural scenes and natural language with recursive neural networks
Richard Socher, Cliff C Lin, Chris Manning, and Andrew Y Ng · 2011
Earlier work this paper cites.
Learning multilingual named entity recognition from Wikipedia
Joel Nothman, Nicky Ringland, Will Radford, Tara Murphy, and James R. Curran · 2012
Earlier work this paper cites.
Training deterministic parsers with non-deterministic oracles
Yoav Goldberg and Joakim Nivre · 2013
Earlier work this paper cites.
Vectorization technology to improve interpreter performance
Erven Rohou, Kevin Williams, and David Yuste · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher Manning, Andrew Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer · 2015
Earlier work this paper cites.
Transition-based dependency parsing with stack long short-term memory
Chris Dyer, Miguel Ballesteros, Wang Ling, Austin Matthews, and Noah A. Smith · 2015
Earlier work this paper cites.
Caffe con troll: Shallow ideas to speed up deep learning
Stefan Hadjis, Firas Abuzaid, Ce Zhang, and Christopher Ré · 2015
Earlier work this paper cites.
Bidirectional LSTM-CRF models for sequence tagging
Zhiheng Huang, Wei Xu, and Kai Yu · 2015
Cited alongside, same era.
Finding function in form: Compositional character models for open vocabulary word representation
Wang Ling, Chris Dyer, Alan W Black, Isabel Trancoso, Ramon Fermandez, Silvio Amir, Luis Marujo, and Tiago Luis · 2015
Cited alongside, same era.
Improved semantic representations from tree-structured long short-term memory networks
Kai Sheng Tai, Richard Socher, and Christopher D. Manning · 2015
Cited alongside, same era.
Chainer: a next-generation open source framework for deep learning
Seiya Tokui, Kenta Oono, Shohei Hido, and Justin Clayton · 2015
Cited alongside, same era.
Tensorflow: Large-scale machine learning on heterogeneous distributed systems
Martın Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, et al · 2016
Latticernn: Recurrent neural networks over lattices
Faisal Ladhak, Ankur Gandhe, Markus Dreyer, Lambert Matthias, Ariya Rastrow, and Björn Hoffmeister · 2016
Later among the works it cites.
Semantic object parsing with graph LSTM
Xiaodan Liang, Xiaohui Shen, Jiashi Feng, Liang Lin, and Shuicheng Yan · 2016
Later among the works it cites.
Multilingual part-of-speech tagging with bidirectional long short-term memory models and auxiliary loss
Barbara Plank, Anders Søgaard, and Yoav Goldberg · 2016
Later among the works it cites.
Neural programmer-interpreters
Scott Reed and Nando de Freitas · 2016
Later among the works it cites.
Neural program lattices
Chengtao Li, Daniel Tarlow, Alexander L. Gaunt, Marc Brockschmidt, and Nate Kushman · 2017
Closest in time.
Deep learning with dynamic computation graphs
Moshe Looks, Marcello Herreshoff, DeLesley Hutchins, and Peter Norvig · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Training with exploration improves a greedy stack LSTM parser
Miguel Ballesteros, Yoav Goldberg, Chris Dyer, and Noah A. Smith · 2016
Cited alongside, same era.
A fast unified model for parsing and sentence understanding
Samuel R. Bowman, Jon Gauthier, Abhinav Rastogi, Raghav Gupta, Christopher D. Manning, and Christopher Potts · 2016
Cited alongside, same era.
Recurrent neural network grammars
Chris Dyer, Adhiguna Kuncoro, Miguel Ballesteros, and Noah A. Smith · 2016
Cited alongside, same era.
Easy-first dependency parsing with hierarchical tree LSTMs
Eliyahu Kiperwasser and Yoav Goldberg · 2016
Cited alongside, same era.
Simple and accurate dependency parsing using bidirectional LSTM feature representations
Eliyahu Kiperwasser and Yoav Goldberg · 2016
Cited alongside, same era.
QCD-aware recursive neural networks for jet physics
Gilles Louppe, Kyunghyun Cho, Cyril Becot, and Kyle Cranmer · 2017
Closest in time.
DyNet: The dynamic neural network toolkit
Graham Neubig, Chris Dyer, Yoav Goldberg, Austin Matthews, Waleed Ammar, Antonios Anastasopoulos, Miguel Ballesteros, David Chiang, Daniel Clothiaux, Trevor Cohn, Kevin Duh, Manaal Faruqui, Cynthia Gan, Dan Garrette, Yangfeng Ji, Lingpeng Kong, Adhiguna Kuncoro, Gaurav Kumar, Chaitanya Malaviya, Paul Michel, Yusuke Oda, Matthew Richardson, Naomi Saphra, Swabha Swayamdipta, and Pengcheng Yin · 2017
Closest in time.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Closest in time.
Learning to compose words into sentences with reinforcement learning
Dani Yogatama, Phil Blunsom, Chris Dyer, Edward Grefenstette, and Wang Ling · 2017
Closest in time.