Fetching the paper…
Reading the bibliography…
We present KERMIT, a simple insertion-based approach to generative modeling for sequences and sequence pairs.
Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
Cho, K., van Merrienboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
Sequence to Sequence Learning with Neural Networks
Sutskever, I., Vinyals, O., and Le, Q. (2014) · 2014
Earlier work this paper cites.
Neural Machine Translation by Jointly Learning to Align and Translate
Bahdanau, D., Cho, K., and Bengio, Y. (2015) · 2015
Earlier work this paper cites.
Distilling the Knowledge in a Neural Network
Hinton, G., Vinyals, O., and Dean, J. (2015) · 2015
Earlier work this paper cites.
Effective Approaches to Attention-based Neural Machine Translation
Luong, M.-T., Pham, H., and Manning, C. D. (2015) · 2015
Earlier work this paper cites.
Show and Tell: A Neural Image Caption Generator
Vinyals, O., Toshev, A., Bengio, S., and Erhan, D. (2015) · 2015
Earlier work this paper cites.
Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhutdinov, R., Zemel, R., and Bengio, Y. (2015) · 2015
Earlier work this paper cites.
End-to-End Attention-based Large Vocabulary Speech Recognition
Bahdanau, D., Chorowski, J., Serdyuk, D., Brakel, P., and Bengio, Y. (2016) · 2016
Earlier work this paper cites.
Listen, Attend and Spell: A Neural Network for Large Vocabulary Conversational Speech Recognition
Chan, W., Jaitly, N., Le, Q., and Vinyals, O. (2016) · 2016
Earlier work this paper cites.
Sequence-Level Knowledge Distillation
Kim, Y. and Rush, A. M. (2016) · 2016
Cited alongside, same era.
WaveNet: A Generative Model for Raw Audio
Oord, A., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., and Kavukcuoglu, K. (2016) · 2016
Cited alongside, same era.
SQuAD: 100,000+ Questions for Machine Comprehension of Text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P. (2016) · 2016
Cited alongside, same era.
Attention Is All You Need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. (2017) · 2017
Cited alongside, same era.
Tacotron: Towards End-to-End Speech Synthesis
Wang, Y., Skerry-Ryan, R., Stanton, D., Wu, Y., Weiss, R. J., Jaitly, N., Yang, Z., Xiao, Y., Chen, Z., Bengio, S., Le, Q., Agiomyrgiannakis, Y., Clark, R., and Saurous, R. A. (2017) · 2017
Cited alongside, same era.
Transforming Question Answering Datasets Into Natural Language Inference Datasets
Demszky, D., Guu, K., and Liang, P. (2018) · 2018
Improving Language Understanding by Generative Pre-Training
Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I. (2018) · 2018
Later among the works it cites.
Blockwise Parallel Decoding for Deep Autoregressive Models
Stern, M., Shazeer, N., and Uszkoreit, J. (2018) · 2018
Later among the works it cites.
Tensor2Tensor for Neural Machine Translation
Vaswani, A., Bengio, S., Brevdo, E., Chollet, F., Gomez, A. N., Gouws, S., Jones, L., Kaiser, L., Kalchbrenner, N., Parmar, N., Sepassi, R., Shazeer, N., and Uszkoreit, J. (2018) · 2018
Later among the works it cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2019) · 2019
Closest in time.
Language Models are Unsupervised Multitask Learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. (2019) · 2019
Closest in time.
Insertion Transformer: Flexible Sequence Generation via Insertion Operations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Non-Autoregressive Neural Machine Translation
Gu, J., Bradbury, J., Xiong, C., Li, V. O., and Socher, R. (2018) · 2018
Cited alongside, same era.
Deterministic Non-Autoregressive Neural Sequence Modeling by Iterative Refinement
Lee, J., Mansimov, E., and Cho, K. (2018) · 2018
Cited alongside, same era.
Deep contextualized word representations
Peters, M. E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., and Zettlemoyer, L. (2018) · 2018
Cited alongside, same era.
Stern, M., Chan, W., Kiros, J., and Uszkoreit, J. (2019) · 2019
Closest in time.
ERNIE: Enhanced Representation through Knowledge Integration
Sun, Y., Wang, S., Li, Y., Feng, S., Chen, X., Zhang, H., Tian, X., Zhu, D., Tian, H., and Wu, H. (2019) · 2019
Closest in time.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R. (2019) · 2019
Closest in time.