Fetching the paper…
Reading the bibliography…
Despite their ubiquity in language generation, it remains unknown why truncation sampling heuristics like nucleus sampling are so effective.
Speech understanding systems: summary of results of the five-year research effort at Carnegie-Mellon University., 1977
Raj Reddy · 1977
Earlier work this paper cites.
The statistical analysis of compositional data
J. Aitchison · 1982
Earlier work this paper cites.
Self-organized language modeling for speech recognition
Fred Jelinek · 1990
Earlier work this paper cites.
Hierarchical neural story generation
Angela Fan, Mike Lewis, and Yann Dauphin · 2018
Earlier work this paper cites.
Breaking the softmax bottleneck: A high-rank RNN language model
Zhilin Yang, Zihang Dai, Ruslan Salakhutdinov, and William W. Cohen · 2018
Earlier work this paper cites.
Breaking the softmax bottleneck via learnable monotonic pointwise non-linearities
O. Ganea, S. Gelly, Gary Bécigneul, and Aliaksei Severyn · 2019
Earlier work this paper cites.
OpenWebText Corpus, 2019
Aaron Gokaslan, Vanya Cohen, Ellie Pavlick, and Stefanie Tellex · 2019
Earlier work this paper cites.
Unifying human and statistical evaluation for natural language generation
Tatsunori B Hashimoto, Hugh Zhang, and Percy Liang · 2019
Earlier work this paper cites.
Comparison of diverse decoding methods from conditional language models
Daphne Ippolito, Reno Kriz, João Sedoc, Maria Kustikova, and Chris Callison-Burch · 2019
Earlier work this paper cites.
On NMT search errors and model errors: Cat got your tongue?
Felix Stahlberg and Bill Byrne · 2019
Earlier work this paper cites.
Mixtape: Breaking the softmax bottleneck efficiently
Zhilin Yang, Thang Luong, Russ R Salakhutdinov, and Quoc V Le · 2019
Earlier work this paper cites.
Stolen probability: A structural weakness of neural language models
David Demeter, Gregory Kimmel, and Doug Downey · 2020
Cited alongside, same era.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi · 2020
Cited alongside, same era.
If beam search is the answer, what was the question?
Clara Meister, Ryan Cotterell, and Tim Vieira · 2020
Cited alongside, same era.
Mirostat: A perplexity-controlled neural text decoding algorithm
Sourya Basu, Govardana Sachitanandam Ramachandran, Nitish Shirish Keskar, and Lav R. Varshney · 2021
Cited alongside, same era.
Decoding methods for neural narrative generation
Alexandra DeLucia, Aaron Mueller, Xiang Lisa Li, and João Sedoc · 2021
Cited alongside, same era.
MAUVE: Measuring the gap between neural text and human text using divergence frontiers
Krishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun, Sean Welleck, Yejin Choi, and Zaid Harchaoui · 2021
RankGen: Improving text generation with large ranking models
Kalpesh Krishna, Yapei Chang, John Wieting, and Mohit Iyyer · 2022
Later among the works it cites.
Bloom: A 176b-parameter open-access multilingual language model, 2022
Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al · 2022
Later among the works it cites.
MOSEK Optimizer API for Python 9.3.22. Version 10.0. , 2023
MOSEK ApS · 2023
Closest in time.
Pythia: A suite for analyzing large language models across training and scaling
Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Softmax bottleneck makes language models unable to represent multi-mode word distributions
Haw-Shiuan Chang and Andrew McCallum · 2022
Cited alongside, same era.
PaLM: Scaling language modeling with pathways, 2022
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, et al · 2022
Cited alongside, same era.
Low-rank softmax can have unargmaxable classes in theory but rarely in practice
Andreas Grivas, Nikolay Bogoychev, and Adam Lopez · 2022
Cited alongside, same era.
Truncation sampling as language model desmoothing
John Hewitt, Christopher Manning, and Percy Liang · 2022
Cited alongside, same era.
Markus Freitag, Behrooz Ghorbani, and Patrick Fernandes · 2023
Closest in time.
Contrastive decoding: Open-ended text generation as optimization
Xiang Lisa Li, Ari Holtzman, Daniel Fried, Percy Liang, Jason Eisner, Tatsunori Hashimoto, Luke Zettlemoyer, and Mike Lewis · 2023
Closest in time.
Locally Typical Sampling
Clara Meister, Tiago Pimentel, Gian Wiher, and Ryan Cotterell · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Rico Sennrich, Jannis Vamvas, and Alireza Mohammadshahi · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models, 2023
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, et al · 2023
Closest in time.