Fetching the paper…
Reading the bibliography…
Long samples of text from neural language models can be of poor quality.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Estimation of probabilities from sparse data for the language model component of a speech recognizer
S. Katz. 1987 · 1987
Earlier work this paper cites.
A comparison of the enhanced Good-Turing and deleted estimation methods for estimating probabilities of English bigrams
Kenneth W Church and William A Gale. 1991 · 1991
Earlier work this paper cites.
Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition , 1st edition
Daniel Jurafsky and James H. Martin. 2000 · 2000
Earlier work this paper cites.
Pattern recognition and machine learning
Christopher M Bishop. 2006 · 2006
Earlier work this paper cites.
Probabilistic context-free grammar induction based on structural zeros
Mehryar Mohri and Brian Roark. 2006 · 2006
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
From softmax to sparsemax: A sparse model of attention and multi-label classification
Andre Martins and Ramon Astudillo. 2016 · 2016
Earlier work this paper cites.
An analysis of the effects of decoding algorithms on fairness in open-ended language generation
Jwala Dhamala, Varun Kumar, Rahul Gupta, Kai-Wei Chang, and Aram Galstyan. 2022 · 2018
Earlier work this paper cites.
Hierarchical neural story generation
Angela Fan, Mike Lewis, and Yann Dauphin. 2018 · 2018
Earlier work this paper cites.
ELI5: long form question answering
Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli. 2019 · 2019
Earlier work this paper cites.
Unifying human and statistical evaluation for natural language generation
Tatsunori B Hashimoto, Hugh Zhang, and Percy Liang. 2019 · 2019
Cited alongside, same era.
Stochastic beams and where to find them: The Gumbel-top-k trick for sampling sequences without replacement
Wouter Kool, Herke Van Hoof, and Max Welling. 2019 · 2019
Cited alongside, same era.
Text summarization with pretrained encoders
Yang Liu and Mirella Lapata. 2019 · 2019
Cited alongside, same era.
Sparse sequence-to-sequence models
Ben Peters, Vlad Niculae, and André FT Martins. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Cited alongside, same era.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020 · 2020
MIROSTAT: A neural text decoding algorithm that directly controls perplexity
Sourya Basu, Govardana Sachitanandam Ramachandran, Nitish Shirish Keskar, and Lav R. Varshney. 2021 · 2021
Later among the works it cites.
The Pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. 2021 · 2021
Later among the works it cites.
Conditional Poisson stochastic beams
Clara Meister, Afra Amini, Tim Vieira, and Ryan Cotterell. 2021a · 2021
Later among the works it cites.
Language model evaluation beyond perplexity
Clara Meister and Ryan Cotterell. 2021 · 2021
Later among the works it cites.
Revisiting the Uniform Information Density hypothesis
Clara Meister, Tiago Pimentel, Patrick Haller, Lena Jäger, Ryan Cotterell, and Roger Levy. 2021b · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Automatic detection of generated text is easiest when humans are fooled
Daphne Ippolito, Daniel Duckworth, Chris Callison-Burch, and Douglas Eck. 2020 · 2020
Cited alongside, same era.
Improved natural language generation via loss truncation
Daniel Kang and Tatsunori B. Hashimoto. 2020 · 2020
Cited alongside, same era.
A systematic characterization of sampling algorithms for open-ended language generation
Moin Nadeem, Tianxing He, Kyunghyun Cho, and James Glass. 2020 · 2020
Cited alongside, same era.
Beyond accuracy: Behavioral testing of NLP models with CheckList
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020 · 2020
Cited alongside, same era.
Consistency of a recurrent language model with respect to incomplete decoding
Sean Welleck, Ilia Kulikov, Jaedeok Kim, Richard Yuanzhe Pang, and Kyunghyun Cho. 2020 · 2020
Cited alongside, same era.
Typical decoding for natural language generation
Clara Meister, Tiago Pimentel, Gian Wiher, and Ryan Cotterell. 2022a
Cited in the paper.
Mauve: Measuring the gap between neural text and human text using divergence frontiers
Krishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun, Sean Welleck, Yejin Choi, and Zaid Harchaoui. 2021 · 2021
Later among the works it cites.
Maxat Tezekbayev, Vassilina Nikoulina, Matthias Gallé, and Zhenisbek Assylbekov. 2021 · 2021
Later among the works it cites.
A cognitive regularizer for language modeling
Jason Wei, Clara Meister, and Ryan Cotterell. 2021 · 2021
Later among the works it cites.
PaLM: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022 · 2022
Closest in time.
Ian Osband, Zheng Wen, Seyed Mohammad Asghari, Vikranth Dwaracherla, Morteza Ibrahimi, Xiuyuan Lu, and Benjamin Van Roy. 2022 · 2022
Closest in time.