Regularizing and optimizing LSTM language models
Stephen Merity, Nitish Shirish Keskar, and Richard Socher · 2018
Later among the works it cites.
When recurrent models don’t need to be recurrent
Original
John Miller and Moritz Hardt · 2018
Later among the works it cites.
SparseMAP: Differentiable sparse structured inference
Vlad Niculae, Andre Martins, Mathieu Blondel, and Claire Cardie · 2018
Later among the works it cites.
Relational recurrent neural networks
Adam Santoro, Ryan Faulkner, David Raposo, Jack Rae, Mike Chrzanowski, Theophane Weber, Daan Wierstra, Oriol Vinyals, Razvan Pascanu, and Timothy Lillicrap · 2018
Later among the works it cites.
Learning longer-term dependencies in RNNs with auxiliary losses
Trieu H Trinh, Andrew M Dai, Thang Luong, and Quoc V Le · 2018
Later among the works it cites.
Breaking the softmax bottleneck: A high-rank RNN language model
Zhilin Yang, Zihang Dai, Ruslan Salakhutdinov, and William W. Cohen · 2018
Later among the works it cites.
Adaptive input representations for neural language modeling
Alexei Baevski and Michael Auli · 2019
Closest in time.
Trellis networks for sequence modeling
Shaojie Bai, J. Zico Kolter, and Vladlen Koltun · 2019
Closest in time.
Generating long sequences with sparse transformers
Original
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 2019
Closest in time.
Recurrent stacking of layers for compact neural machine translation models
Raj Dabre and Atsushi Fujita · 2019
Closest in time.
Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V. Le, and Ruslan Salakhutdinov · 2019
Closest in time.
Universal transformers
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Łukasz Kaiser · 2019
Closest in time.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Closest in time.
Implicit deep learning
Original
Laurent El Ghaoui, Fangda Gu, Bertrand Travacca, and Armin Askari · 2019
Closest in time.
DARTS: Differentiable architecture search
Hanxiao Liu, Karen Simonyan, and Yiming Yang · 2019
Closest in time.
PyTorch: An imperative style, high-performance deep learning library
Benoit Steiner, Zachary DeVito, Soumith Chintala, Sam Gross, Adam Paszke, Francisco Massa, Adam Lerer, Gregory Chanan, Zeming Lin, Edward Yang, et al · 2019
Closest in time.
SATNet: Bridging deep learning and logical reasoning using a differentiable satisfiability solver
Po-Wei Wang, Priya Donti, Bryan Wilder, and Zico Kolter · 2019
Closest in time.
Equilibrated recurrent neural network: Neuronal time-delayed self-feedback improves accuracy and stability
Original
Ziming Zhang, Anil Kag, Alan Sullivan, and Venkatesh Saligrama · 2019
Closest in time.