Fetching the paper…
Reading the bibliography…
Structured prediction energy networks (SPENs; Belanger & McCallum 2016) use neural network architectures to define energy functions that can capture arbitrary dependencies among parts of structured outputs.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Conditional random fields: Probabilistic models for segmenting and labeling sequence data
John D. Lafferty, Andrew McCallum, and Fernando C. N. Pereira · 2001
Earlier work this paper cites.
Discriminative training methods for hidden Markov models: Theory and experiments with perceptron algorithms
Michael Collins · 2002
Earlier work this paper cites.
Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition
Erik F. Tjong Kim Sang and Fien De Meulder · 2003
Earlier work this paper cites.
Max-margin Markov networks
Ben Taskar, Carlos Guestrin, and Daphne Koller · 2004
Earlier work this paper cites.
Large margin methods for structured and interdependent output variables
Ioannis Tsochantaridis, Thorsten Joachims, Thomas Hofmann, and Yasemin Altun · 2005
Earlier work this paper cites.
A tutorial on energy-based learning
Yann LeCun, Sumit Chopra, Raia Hadsell, Marc’Aurelio Ranzato, and Fu-Jie Huang · 2006
Earlier work this paper cites.
Probabilistic Graphical Models: Principles and Techniques
Daphne Koller and Nir Friedman · 2009
Earlier work this paper cites.
Design challenges and misconceptions in named entity recognition
Lev Ratinov and Dan Roth · 2009
Earlier work this paper cites.
Part-of-speech tagging for Twitter: annotation, features, and experiments
Kevin Gimpel, Nathan Schneider, Brendan O’Connor, Dipanjan Das, Daniel Mills, Jacob Eisenstein, Michael Heilman, Dani Yogatama, Jeffrey Flanigan, and Noah A. Smith · 2011
Earlier work this paper cites.
Efficient inference in fully connected CRFs with Gaussian edge potentials
Philipp Krähenbühl and Vladlen Koltun · 2011
Earlier work this paper cites.
Generic methods for optimization-based modeling
Justin Domke · 2012
Earlier work this paper cites.
On amortizing inference cost for structured prediction
Vivek Srikumar, Gourab Kundu, and Dan Roth · 2012
Earlier work this paper cites.
Auto-encoding variational Bayes
Diederik Kingma and Max Welling · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean · 2013
Earlier work this paper cites.
Improved part-of-speech tagging for online conversational text with word clusters
Olutobi Owoputi, Brendan O’Connor, Chris Dyer, Kevin Gimpel, Nathan Schneider, and Noah A. Smith · 2013
Earlier work this paper cites.
Do deep nets really need to be deep?
Jimmy Ba and Rich Caruana · 2014
Cited alongside, same era.
Amortized inference in probabilistic reasoning
Samuel Gershman and Noah Goodman · 2014
Cited alongside, same era.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Cited alongside, same era.
GloVe: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning · 2014
Cited alongside, same era.
Structural learning with amortized inference
Kai-Wei Chang, Shyam Upadhyay, Gourab Kundu, and Dan Roth · 2015
Cited alongside, same era.
End-to-end sequence labeling via bi-directional LSTM-CNNs-CRF
Xuezhe Ma and Eduard Hovy · 2016
Later among the works it cites.
Inference networks for sequential Monte Carlo in graphical models
Brooks Paige and Frank Wood · 2016
Later among the works it cites.
Improved techniques for training GANs
Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, Xi Chen, and Xi Chen · 2016
Later among the works it cites.
Do deep convolutional nets really need to be deep?
Gregor Urban, Krzysztof J. Geras, Samira Ebrahimi Kahou, Ozlem Aslan, Shengjie Wang, Rich Caruana, Abdel-rahman Mohamed, Matthai Philipose, and Matthew Richardson · 2016
Later among the works it cites.
Proximal deep structured models
Shenlong Wang, Sanja Fidler, and Raquel Urtasun · 2016
Later among the works it cites.
Sequence-to-sequence learning as beam-search optimization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A neural algorithm of artistic style
Leon A. Gatys, Alexander S. Ecker, and Matthias Bethge · 2015
Cited alongside, same era.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeffrey Dean · 2015
Cited alongside, same era.
DeepDream-a code example for visualizing neural networks
Alexander Mordvintsev, Christopher Olah, and Mike Tyka · 2015
Cited alongside, same era.
Structured prediction energy networks
David Belanger and Andrew McCallum · 2016
Cited alongside, same era.
Blending LSTMs into CNNs
Krzysztof J. Geras, Abdel rahman Mohamed, Rich Caruana, Gregor Urban, Shengjie Wang, Ozlem Aslan, Matthai Philipose, Matthew Richardson, and Charles Sutton · 2016
Cited alongside, same era.
Perceptual losses for real-time style transfer and super-resolution
Justin Johnson, Alexandre Alahi, and Li Fei-Fei · 2016
Cited alongside, same era.
Sam Wiseman and Alexander M. Rush · 2016
Later among the works it cites.
Energy-based generative adversarial network
Junbo Jake Zhao, Michaël Mathieu, and Yann LeCun · 2016
Later among the works it cites.
Input convex neural networks
Brandon Amos, Lei Xu, and J. Zico Kolter · 2017
Later among the works it cites.
End-to-end learning for structured prediction energy networks
David Belanger, Bishan Yang, and Andrew McCallum · 2017
Later among the works it cites.
Calibrating energy-based generative adversarial networks
Zihang Dai, Amjad Almahairi, Bachman Philip, Eduard Hovy, and Aaron Courville · 2017
Later among the works it cites.
Towards decoding as continuous optimisation in neural machine translation
Cong Duy Vu Hoang, Gholamreza Haffari, and Trevor Cohn · 2017
Later among the works it cites.
Learning what’s easy: Fully differentiable neural easy-first taggers
André F. T. Martins and Julia Kreutzer · 2017
Later among the works it cites.
Regularizing neural networks by penalizing confident output distributions
Gabriel Pereyra, George Tucker, Jan Chorowski, Lukasz Kaiser, and Geoffrey E. Hinton · 2017
Later among the works it cites.
Learning to embed words in context for syntactic tasks
Lifu Tu, Kevin Gimpel, and Karen Livescu · 2017
Later among the works it cites.
A continuous relaxation of beam search for end-to-end training of neural sequence models
Kartik Goyal, Graham Neubig, Chris Dyer, and Taylor Berg-Kirkpatrick · 2018
Closest in time.