Fetching the paper…
Reading the bibliography…
Given an input sequence (or prefix), modern language models often assign high probabilities to output sequences that are repetitive, incoherent, or irrelevant to the prefix; as such, model-generated text also contains such artifacts.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Real or fake? learning to discriminate machine from human generated text
Anton Bakhtin, Sam Gross, Myle Ott, Yuntian Deng, Marc’Aurelio Ranzato, and Arthur Szlam. 2019 · 1906
Earlier work this paper cites.
Peter: A Novel of which He is Not the Hero
Francis Hopkinson Smith. 1911 · 1911
Earlier work this paper cites.
On faithfulness and factuality in abstractive summarization
Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020 · 1919
Earlier work this paper cites.
Measuring nominal scale agreement among many raters
Joseph L Fleiss. 1971 · 1971
Earlier work this paper cites.
Coherence and coreference
Jerry R Hobbs. 1979 · 1979
Earlier work this paper cites.
Centering: a framework for modeling the local coherence of discourse
Barbara J Grosz, Scott Weinstein, and Aravind K Joshi. 1995 · 1995
Earlier work this paper cites.
Okapi at trec-3
Stephen E Robertson, Steve Walker, Susan Jones, Micheline M Hancock-Beaulieu, Mike Gatford, et al. 1995 · 1995
Earlier work this paper cites.
Discriminative reranking for machine translation
Libin Shen, Anoop Sarkar, and Franz Josef Och. 2004 · 2004
Earlier work this paper cites.
Evaluation of text generation: A survey
Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2020 · 2006
Earlier work this paper cites.
A tutorial on energy-based learning
Yann LeCun, Sumit Chopra, Raia Hadsell, Marc’Aurelio Ranzato, and Fu Jie Huang. 2006 · 2006
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. 2013 · 2013
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Chinese poetry generation with recurrent neural networks
Xingxing Zhang and Mirella Lapata. 2014 · 2014
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer. 2015 · 2015
Earlier work this paper cites.
Skip-thought vectors
Ryan Kiros, Yukun Zhu, Russ R Salakhutdinov, Richard Zemel, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Earlier work this paper cites.
Learning distributed representations of sentences from unlabelled data
Felix Hill, Kyunghyun Cho, and Anna Korhonen. 2016 · 2016
Earlier work this paper cites.
A corpus and cloze evaluation for deeper understanding of commonsense stories
Nasrin Mostafazadeh, Nathanael Chambers, Xiaodong He, Devi Parikh, Dhruv Batra, Lucy Vanderwende, Pushmeet Kohli, and James Allen. 2016 · 2016
Earlier work this paper cites.
Sequence level training with recurrent neural networks
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba. 2016 · 2016
Earlier work this paper cites.
Improved deep metric learning with multi-class n-pair loss objective
Kihyuk Sohn. 2016 · 2016
Earlier work this paper cites.
Sequence-to-sequence learning as beam-search optimization
Sam Wiseman and Alexander M. Rush. 2016 · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al. 2016 · 2016
Earlier work this paper cites.
Discourse-based objectives for fast unsupervised sentence representation learning
Yacine Jernite, Samuel R Bowman, and David Sontag. 2017 · 2017
Earlier work this paper cites.
Six challenges for neural machine translation
Philipp Koehn and Rebecca Knowles. 2017 · 2017
Earlier work this paper cites.
Dense-captioning events in videos
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles. 2017 · 2017
Earlier work this paper cites.
Learning to decode for future success
Jiwei Li, Will Monroe, and Dan Jurafsky. 2017 · 2017
Earlier work this paper cites.
ParlAI: A dialog research software platform
Alexander Miller, Will Feng, Dhruv Batra, Antoine Bordes, Adam Fisch, Jiasen Lu, Devi Parikh, and Jason Weston. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang. 2018 · 2018
Earlier work this paper cites.
HotFlip: White-box adversarial examples for text classification
Javid Ebrahimi, Anyi Rao, Daniel Lowd, and Dejing Dou. 2018 · 2018
Earlier work this paper cites.
Classical structured prediction losses for sequence to sequence learning
Sergey Edunov, Myle Ott, Michael Auli, David Grangier, and Marc’Aurelio Ranzato. 2018 · 2018
Earlier work this paper cites.
Hierarchical neural story generation
Angela Fan, Mike Lewis, and Yann Dauphin. 2018 · 2018
Earlier work this paper cites.
Learning to write with cooperative discriminators
Ari Holtzman, Jan Buys, Maxwell Forbes, Antoine Bosselut, David Golub, and Yejin Choi. 2018 · 2018
Earlier work this paper cites.
Sharp nearby, fuzzy far away: How neural language models use context
Urvashi Khandelwal, He He, Peng Qi, and Dan Jurafsky. 2018 · 2018
Earlier work this paper cites.
An efficient framework for learning sentence representations
Lajanugen Logeswaran and Honglak Lee. 2018 · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. 2018 · 2018
Cited alongside, same era.
Twin networks: Matching the future for sequence generation
Dmitriy Serdyuk, Nan Rosemary Ke, Alessandro Sordoni, Adam Trischler, Chris Pal, and Yoshua Bengio. 2018 · 2018
Cited alongside, same era.
Prediction with a short memory
Vatsal Sharan, Sham Kakade, Percy Liang, and Gregory Valiant. 2018 · 2018
Cited alongside, same era.
Tackling the story ending biases in the story cloze test
Rishi Sharma, James Allen, Omid Bakhshandeh, and Nasrin Mostafazadeh. 2018 · 2018
Cited alongside, same era.
Adafactor: Adaptive learning rates with sublinear memory cost
Noam Shazeer and Mitchell Stern. 2018 · 2018
Cited alongside, same era.
Learning neural trans-dimensional random field language models with noise-contrastive estimation
CLIPScore: A reference-free evaluation metric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. 2021 · 2021
Later among the works it cites.
Surface form competition: Why the highest probability answer isn’t always right
Ari Holtzman, Peter West, Vered Shwartz, Yejin Choi, and Luke Zettlemoyer. 2021 · 2021
Later among the works it cites.
The perils of using Mechanical Turk to evaluate open-ended text generation
Marzena Karpinska, Nader Akoury, and Mohit Iyyer. 2021 · 2021
Later among the works it cites.
Hurdles to progress in long-form question answering
Kalpesh Krishna, Aurko Roy, and Mohit Iyyer. 2021 · 2021
Later among the works it cites.
Discriminative reranking for neural machine translation
Ann Lee, Michael Auli, and Marc’Aurelio Ranzato. 2021 · 2021
Later among the works it cites.
The power of scale for parameter-efficient prompt tuning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bin Wang and Zhijian Ou. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
GLTR: Statistical detection and visualization of generated text
Sebastian Gehrmann, Hendrik Strobelt, and Alexander Rush. 2019 · 2019
Cited alongside, same era.
Bias correction of learned generative models using likelihood-free importance weighting
Aditya Grover, Jiaming Song, Ashish Kapoor, Kenneth Tran, Alekh Agarwal, Eric J Horvitz, and Stefano Ermon. 2019 · 2019
Cited alongside, same era.
Global autoregressive models for data-efficient sequence learning
Tetiana Parshakova, Jean-Marc Andreoli, and Marc Dymetman. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Compressive transformers for long-range sequence modelling
Jack W Rae, Anna Potapenko, Siddhant M Jayakumar, Chloe Hillier, and Timothy P Lillicrap. 2019 · 2019
Cited alongside, same era.
Brian Lester, Rami Al-Rfou, and Noah Constant. 2021 · 2021
Later among the works it cites.
Cut the carp: Fishing for zero-shot story evaluation
Shahbuland Matiana, JR Smith, Ryan Teehan, Louis Castricato, Stella Biderman, Leo Gao, and Spencer Frazier. 2021 · 2021
Later among the works it cites.
Webgpt: Browser-assisted question-answering with human feedback
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al. 2021 · 2021
Later among the works it cites.
Mauve: Measuring the gap between neural text and human text using divergence frontiers
Krishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun, Sean Welleck, Yejin Choi, and Zaid Harchaoui. 2021 · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021 · 2021
Later among the works it cites.
Do long-range language models actually use long-range context?
Simeng Sun, Kalpesh Krishna, Andrew Mattarella-Micke, and Mohit Iyyer. 2021 · 2021
Later among the works it cites.
Beir: A heterogeneous benchmark for zero-shot evaluation of information retrieval models
Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. 2021 · 2021
Later among the works it cites.
FUDGE: Controlled text generation with future discriminators
Kevin Yang and Dan Klein. 2021 · 2021
Later among the works it cites.
Trading off diversity and quality in natural language generation
Hugh Zhang, Daniel Duckworth, Daphne Ippolito, and Arvind Neelakantan. 2021 · 2021
Later among the works it cites.
Cont: Contrastive neural text generation
Chenxin An, Jiangtao Feng, Kai Lv, Lingpeng Kong, Xipeng Qiu, and Xuanjing Huang. 2022 · 2022
Closest in time.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, et al. 2022 · 2022
Closest in time.
The efficiency misnomer
Mostafa Dehghani, Anurag Arnab, Lucas Beyer, Ashish Vaswani, and Yi Tay. 2022 · 2022
Closest in time.
Is GPT-3 text indistinguishable from human text? scarecrow: A framework for scrutinizing machine text
Yao Dou, Maxwell Forbes, Rik Koncel-Kedziorski, Noah Smith, and Yejin Choi. 2022 · 2022
Closest in time.
Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text
Sebastian Gehrmann, Elizabeth Clark, and Thibault Sellam. 2022 · 2022
Closest in time.
Truncation sampling as language model desmoothing
John Hewitt, Christopher D Manning, and Percy Liang. 2022 · 2022
Closest in time.
Internet-augmented dialogue generation
Mojtaba Komeili, Kurt Shuster, and Jason Weston. 2022 · 2022
Closest in time.
Contrastive decoding: Open-ended text generation as optimization
Xiang Lisa Li, Ari Holtzman, Daniel Fried, Percy Liang, Jason Eisner, Tatsunori Hashimoto, Luke Zettlemoyer, and Mike Lewis. 2022 · 2022
Closest in time.
Brio: Bringing order to abstractive summarization
Yixin Liu, Pengfei Liu, Dragomir Radev, and Graham Neubig. 2022 · 2022
Closest in time.
Locally typical sampling
Clara Meister, Tiago Pimentel, Gian Wiher, and Ryan Cotterell. 2022 · 2022
Closest in time.
Mix and match: Learning-free controllable text generation using energy language models
Fatemehsadat Mireshghallah, Kartik Goyal, and Taylor Berg-Kirkpatrick. 2022 · 2022
Closest in time.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Closest in time.
Cold decoding: Energy-based constrained text generation with langevin dynamics
Lianhui Qin, Sean Welleck, Daniel Khashabi, and Yejin Choi. 2022 · 2022
Closest in time.
Scaling up models and data with t5x
Adam Roberts, Hyung Won Chung, Anselm Levskaya, Gaurav Mishra, James Bradbury, Daniel Andor, Sharan Narang, et al. 2022 · 2022
Closest in time.
Contrastive search is what you need for neural text generation
Yixuan Su and Nigel Collier. 2022 · 2022
Closest in time.
A contrastive framework for neural text generation
Yixuan Su, Tian Lan, Yan Wang, Dani Yogatama, Lingpeng Kong, and Nigel Collier. 2022 · 2022
Closest in time.
Chapterbreak: A challenge dataset for long-range language models
Simeng Sun, Katherine Thai, and Mohit Iyyer. 2022 · 2022
Closest in time.
Relic: Retrieving evidence for literary claims
Katherine Thai, Yapei Chang, Kalpesh Krishna, and Mohit Iyyer. 2022 · 2022
Closest in time.
Language modeling via stochastic processes
Rose E Wang, Esin Durmus, Noah Goodman, and Tatsunori Hashimoto. 2022 · 2022
Closest in time.
Massive-scale decoding for text generation using lattices
Jiacheng Xu, Siddhartha Jonnalagadda, and Greg Durrett. 2022 · 2022
Closest in time.