Fetching the paper…
Reading the bibliography…
We show that autoregressive language models can learn to infill text after we apply a straightforward transformation to the dataset, which simply moves a span of text from the middle of a document to its end.
KERMIT: generative insertion-based modeling for sequences
W. Chan, N. Kitaev, K. Guu, M. Stern, and J. Uszkoreit · 1906
Earlier work this paper cites.
CTRL: A conditional transformer language model for controllable generation
N. S. Keskar, B. McCann, L. R. Varshney, C. Xiong, and R. Socher · 1909
Earlier work this paper cites.
Bpe-dropout: Simple and effective subword regularization
I. Provilkov, D. Emelianenko, and E. Voita · 1910
Earlier work this paper cites.
Enabling language models to fill in the blanks
C. Donahue, M. Lee, and P. Liang · 2005
Earlier work this paper cites.
Learning to summarize from human feedback
N. Stiennon, L. Ouyang, J. Wu, D. M. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano · 2009
Earlier work this paper cites.
The Winograd Schema Challenge
H. J. Levesque, E. Davis, and L. Morgenstern · 2012
Earlier work this paper cites.
A corpus and cloze evaluation for deeper understanding of commonsense stories
N. Mostafazadeh, N. Chambers, X. He, D. Parikh, D. Batra, L. Vanderwende, P. Kohli, and J. Allen · 2016
Earlier work this paper cites.
The LAMBADA dataset: Word prediction requiring a broad discourse context
D. Paperno, G. Kruszewski, A. Lazaridou, N. Q. Pham, R. Bernardi, S. Pezzelle, M. Baroni, G. Boleda, and R. Fernández · 2016
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
QuAC: Question answering in context
E. Choi, H. He, M. Iyyer, M. Yatskar, W.-t. Yih, Y. Choi, P. Liang, and L. Zettlemoyer · 2018
Earlier work this paper cites.
Maskgan: Better text generation via filling in the______, 2018
W. Fedus, I. Goodfellow, and A. M. Dai · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever · 2018
Earlier work this paper cites.
Self-attention with relative position representations
P. Shaw, J. Uszkoreit, and A. Vaswani · 2018
Earlier work this paper cites.
Transformer-XL: Attentive language models beyond a fixed-length context
Z. Dai, Z. Yang, Y. Yang, J. Carbonell, Q. Le, and R. Salakhutdinov · 2019
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
J. Devlin, M. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs
D. Dua, Y. Wang, P. Dasigi, G. Stanovsky, S. Singh, and M. Gardner · 2019
Earlier work this paper cites.
Insertion-based decoding with automatically inferred generation order
J. Gu, Q. Liu, and K. Cho · 2019
Earlier work this paper cites.
TIGS: An inference algorithm for text infilling with gradient search
D. Liu, J. Fu, P. Liu, and J. Lv · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2019
Cited alongside, same era.
Mass: Masked sequence to sequence pre-training for language generation
K. Song, X. Tan, T. Qin, J. Lu, and T.-Y. Liu · 2019
Cited alongside, same era.
Insertion transformer: Flexible sequence generation via insertion operations
M. Stern, W. Chan, J. Kiros, and J. Uszkoreit · 2019
Cited alongside, same era.
XLNet: Generalized Autoregressive Pretraining for Language Understanding
Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. Salakhutdinov, and Q. V. Le · 2019
Cited alongside, same era.
D. Hernandez, J. Kaplan, T. Henighan, and S. McCandlish · 2021
Later among the works it cites.
DOBF: A deobfuscation pre-training objective for programming languages
M. Lachaux, B. Rozière, M. Szafraniec, and G. Lample · 2021
Later among the works it cites.
Jurassic-1: Technical details and evaluation
O. Lieber, O. Sharir, B. Lenz, and Y. Shoham · 2021
Later among the works it cites.
Scaling language models: Methods, analysis & insights from training gopher
J. W. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, F. Song, J. Aslanides, S. Henderson, R. Ring, S. Young, et al · 2021
Later among the works it cites.
Winogrande: An adversarial winograd schema challenge at scale
K. Sakaguchi, R. L. Bras, C. Bhagavatula, and Y. Choi · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
HellaSwag: Can a machine really finish your sentence?
R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi · 2019
Cited alongside, same era.
Piqa: Reasoning about physical commonsense in natural language
Y. Bisk, R. Zellers, R. L. Bras, J. Gao, and Y. Choi · 2020
Cited alongside, same era.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
The curious case of neural text degeneration
A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi · 2020
Cited alongside, same era.
SpanBERT: Improving pre-training by representing and predicting spans
M. Joshi, D. Chen, Y. Liu, D. S. Weld, L. Zettlemoyer, and O. Levy · 2020
Cited alongside, same era.
Distribution augmentation for generative modeling
H. Jun, R. Child, M. Chen, J. Schulman, A. Ramesh, A. Radford, and I. Sutskever · 2020
Cited alongside, same era.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Cited alongside, same era.
W. Zhu, Z. Hu, and E. P. Xing · 2021
Later among the works it cites.
CM3: A causal masked multimodal model of the internet
A. Aghajanyan, B. Huang, C. Ross, V. Karpukhin, H. Xu, N. Goyal, D. Okhonko, M. Joshi, G. Ghosh, M. Lewis, and L. Zettlemoyer · 2022
Closest in time.
On the role of bidirectionality in language model pre-training, 2022
M. Artetxe, J. Du, N. Goyal, L. Zettlemoyer, and V. Stoyanov · 2022
Closest in time.
Palm: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al · 2022
Closest in time.
GLM: General language model pretraining with autoregressive blank infilling
Z. Du, Y. Qian, X. Liu, M. Ding, J. Qiu, Z. Yang, and J. Tang · 2022
Closest in time.
Incoder: A generative model for code infilling and synthesis, 2022
D. Fried, A. Aghajanyan, J. Lin, S. Wang, E. Wallace, F. Shi, R. Zhong, W.-t. Yih, L. Zettlemoyer, and M. Lewis · 2022
Closest in time.
Training compute-optimal large language models
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, et al · 2022
Closest in time.
New GPT-3 Capabilities: Edit and Insert
OpenAI, M. Bavarian, A. Jiang, H. Jun, and H. Pondé · 2022
Closest in time.
Training language models to follow instructions with human feedback, 2022
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. Christiano, J. Leike, and R. Lowe · 2022
Closest in time.
Unifying language learning paradigms, 2022
Y. Tay, M. Dehghani, V. Q. Tran, X. Garcia, D. Bahri, T. Schuster, H. S. Zheng, N. Houlsby, and D. Metzler · 2022
Closest in time.
Lamda: Language models for dialog applications
R. Thoppilan, D. De Freitas, J. Hall, N. Shazeer, A. Kulshreshtha, H.-T. Cheng, A. Jin, T. Bos, L. Baker, Y. Du, et al · 2022
Closest in time.
What language model architecture and pretraining objective work best for zero-shot generalization?, 2022
T. Wang, A. Roberts, D. Hesslow, T. L. Scao, H. W. Chung, I. Beltagy, J. Launay, and C. Raffel · 2022
Closest in time.