Fetching the paper…
Reading the bibliography…
fairseq is an open-source sequence modeling toolkit that allows researchers and developers to train custom models for translation, summarization, language modeling, and other text generation tasks.
Strategies for structuring story generation
Angela Fan, Mike Lewis, and Yann Dauphin. 2019 · 1902
Earlier work this paper cites.
Mixture models for diverse machine translation: Tricks of the trade
Tianxiao Shen, Myle Ott, Michael Auli, and Marc’Aurelio Ranzato. 2019 · 1902
Earlier work this paper cites.
Rouge: a package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
multi-bleu.perl
Hieu Hoang, Philipp Koehn, Ulrich Germann, Kenneth Heafield, and Barry Haddow. 2006 · 2006
Earlier work this paper cites.
Teaching machines to read and comprehend
Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015 · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Minh-Thang Luong, Hieu Pham, and Christopher D Manning. 2015 · 2015
Earlier work this paper cites.
Improving neural language models with a continuous cache
Edouard Grave, Armand Joulin, and Nicolas Usunier. 2016 · 2016
Earlier work this paper cites.
Exploring the limits of language modeling
Rafal Józefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu. 2016 · 2016
Earlier work this paper cites.
Character-aware neural language models
Yoon Kim, Yacine Jernite, David Sontag, and Alexander M Rush. 2016 · 2016
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter. 2016 · 2016
Earlier work this paper cites.
Abstractive text summarization using sequence-to-sequence rnns and beyond
Ramesh Nallapati, Bowen Zhou, Cicero dos Santos, Caglar Gulcehre, and Bing Xiang. 2016 · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Diverse beam search: Decoding diverse solutions from neural sequence models
Ashwin K Vijayakumar, Michael Cogswell, Ramprasath R Selvaraju, Qing Sun, Stefan Lee, David Crandall, and Dhruv Batra. 2016 · 2016
Earlier work this paper cites.
Weighted transformer network for machine translation
Karim Ahmed, Nitish Shirish Keskar, and Richard Socher. 2017 · 2017
Earlier work this paper cites.
Language modeling with gated convolutional networks
Yann N. Dauphin, Angela Fan, Michael Auli, and David Grangier. 2017 · 2017
Earlier work this paper cites.
Convolutional Sequence to Sequence Learning
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N Dauphin. 2017 · 2017
Cited alongside, same era.
Efficient softmax approximation for gpus
Edouard Grave, Armand Joulin, Moustapha Cissé, David Grangier, and Hervé Jégou. 2017 · 2017
Cited alongside, same era.
OpenNMT: Open-source toolkit for neural machine translation
Guillaume Klein, Yoon Kim, Yuntian Deng, Jean Senellart, and Alexander M. Rush. 2017 · 2017
Cited alongside, same era.
Parlai: A dialog research software platform
A. H. Miller, W. Feng, A. Fisch, J. Lu, D. Batra, A. Bordes, D. Parikh, and J. Weston. 2017 · 2017
Cited alongside, same era.
An empirical study of mini-batch creation strategies for neural machine translation
Makoto Morishita, Yusuke Oda, Graham Neubig, Koichiro Yoshino, Katsuhito Sudoh, and Satoshi Nakamura. 2017 · 2017
Cited alongside, same era.
OpenSeq2Seq: Extensible Toolkit for Distributed and Mixed Precision Training of Sequence-to-Sequence Models
Oleksii Kuchaiev, Boris Ginsburg, Igor Gitman, Vitaly Lavrukhin, Carl Case, and Paulius Micikevicius. 2018 · 2018
Later among the works it cites.
Phrase-based & neural unsupervised machine translation
Guillaume Lample, Myle Ott, Alexis Conneau, Ludovic Denoyer, and Marc’Aurelio Ranzato. 2018 · 2018
Later among the works it cites.
Controlling length in abstractive summarization using a convolutional neural network
Yizhu Liu, Zhiyi Luo, and Kenny Zhu. 2018 · 2018
Later among the works it cites.
An analysis of neural language modeling at multiple scales
Stephen Merity, Nitish Shirish Keskar, and Richard Socher. 2018 · 2018
Later among the works it cites.
Mixed Precision Training
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory F. Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Romain Paulus, Caiming Xiong, and Richard Socher. 2017 · 2017
Cited alongside, same era.
Get to the point: Summarization with pointer-generator networks
Abigail See, Peter J. Liu, and Christopher D. Manning. 2017 · 2017
Cited alongside, same era.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc V. Le, Geoffrey E. Hinton, and Jeff Dean. 2017 · 2017
Cited alongside, same era.
Attention Is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Massively multilingual sentence embeddings for zero-shot cross-lingual transfer and beyond
Mikel Artetxe and Holger Schwenk. 2018 · 2018
Cited alongside, same era.
A stable and effective learning strategy for trainable greedy decoding
Yun Chen, Victor OK Li, Kyunghyun Cho, and Samuel R Bowman. 2018 · 2018
Cited alongside, same era.
A multilayer convolutional encoder-decoder neural network for grammatical error correction
Shamil Chollampatt and Hwee Tou Ng. 2018 · 2018
Cited alongside, same era.
Shashi Narayan, Shay B Cohen, and Mirella Lapata. 2018 · 2018
Later among the works it cites.
A call for clarity in reporting bleu scores
Matt Post. 2018 · 2018
Later among the works it cites.
Fast parametric learning with activation memorization
Jack W. Rae, Chris Dyer, Peter Dayan, and Timothy P. Lillicrap. 2018 · 2018
Later among the works it cites.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. 2018 · 2018
Later among the works it cites.
Adafactor: Adaptive learning rates with sublinear memory cost
Noam Shazeer and Mitchell Stern. 2018 · 2018
Later among the works it cites.
Double path networks for sequence to sequence learning
Kaitao Song, Xu Tan, Di He, Jianfeng Lu, Tao Qin, and Tie-Yan Liu. 2018 · 2018
Later among the works it cites.
Tensor2Tensor for Neural Machine Translation
A. Vaswani, S. Bengio, E. Brevdo, F. Chollet, A. N. Gomez, S. Gouws, L. Jones, Ł. Kaiser, N. Kalchbrenner, N. Parmar, R. Sepassi, N. Shazeer, and J. Uszkoreit. 2018 · 2018
Later among the works it cites.
Adaptive input representations for neural language modeling
Alexei Baevski and Michael Auli. 2019 · 2019
Closest in time.
Wizard of Wikipedia: Knowledge-powered conversational agents
Emily Dinan, Stephen Roller, Kurt Shuster, Angela Fan, Michael Auli, and Jason Weston. 2019 · 2019
Closest in time.
Pay less attention with lightweight and dynamic convlutions
Felix Wu, Angela Fan, Alexei Baevski, Yann N. Dauphin, and Michael Auli. 2019 · 2019
Closest in time.