Fetching the paper…
Reading the bibliography…
Autoregressive language models are the currently dominant paradigm for text generation, but they have some fundamental limitations that cannot be remedied by scale-for example inherently sequential and unidirectional generation.
Local asymptotic minimax and admissibility in estimation
Jaroslav Hájek · 1972
Earlier work this paper cites.
Statistical analysis of non-lattice data
Julian Besag · 1975
Earlier work this paper cites.
Logarithmic sobolev inequalities for finite markov chains
Persi Diaconis and Laurent Saloff-Coste · 1996
Earlier work this paper cites.
Asymptotic statistics , volume 3
Aad W Van der Vaart · 2000
Earlier work this paper cites.
Generalized pseudo-likelihood estimates for markov random fields on lattice
Fuchun Huang and Yosihiko Ogata · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
BLEURT: learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur P. Parikh · 2004
Earlier work this paper cites.
Modified logarithmic sobolev inequalities in discrete settings
Sergey G Bobkov and Prasad Tetali · 2006
Earlier work this paper cites.
High-dimensional ising model selection using l1-regularized logistic regression
Pradeep Ravikumar, Martin J Wainwright, and John D Lafferty · 2010
Earlier work this paper cites.
Operator reverse monotonicity of the inverse
Alexis Akira Toda · 2011
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson · 2012
Earlier work this paper cites.
An inequality for relative entropy and logarithmic sobolev inequalities in euclidean spaces
Katalin Marton · 2013
Earlier work this paper cites.
Findings of the 2014 workshop on statistical machine translation
Ondrej Bojar, Christian Buck, Christian Federmann, Barry Haddow, Philipp Koehn, Johannes Leveling, Christof Monz, Pavel Pecina, Matt Post, Herve Saint-Amand, Radu Soricut, Lucia Specia, and Ale s Tamchyna · 2014
Earlier work this paper cites.
Approximate tensorization of entropy at high temperature
Pietro Caputo, Georg Menz, and Prasad Tetali · 2015
Earlier work this paper cites.
Katalin Marton · 2015
Earlier work this paper cites.
Findings of the 2016 conference on machine translation
Ond rej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Aurelie Neveol, Mariana Neves, Martin Popel, Matt Post, Raphael Rubino, Carolina Scarton, Lucia Specia, Marco Turchi, Karin Verspoor, and Marcos Zampieri · 2016
Earlier work this paper cites.
Generating sentences from a continuous space
Samuel R. Bowman, Luke Vilnis, Oriol Vinyals, Andrew Dai, Rafal Jozefowicz, and Samy Bengio · 2016
Earlier work this paper cites.
Sequence-level knowledge distillation
Yoon Kim and Alexander M. Rush · 2016
Earlier work this paper cites.
Interaction screening: Efficient and sample-optimal learning of ising models
Marc Vuffray, Sidhant Misra, Andrey Lokhov, and Michael Chertkov · 2016
Earlier work this paper cites.
Findings of the 2017 conference on machine translation (WMT17)
Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Shujian Huang, Matthias Huck, Philipp Koehn, Qun Liu, Varvara Logacheva, Christof Monz, Matteo Negri, Matt Post, Raphael Rubino, Lucia Specia, and Marco Turchi · 2017
Earlier work this paper cites.
Maximum-likelihood augmented discrete generative adversarial networks, 2017
Tong Che, Yanran Li, Ruixiang Zhang, R Devon Hjelm, Wenjie Li, Yangqiu Song, and Yoshua Bengio · 2017
Earlier work this paper cites.
Adversarial ranking for language generation
Kevin Lin, Dianqi Li, Xiaodong He, Zhengyou Zhang, and Ming-Ting Sun · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Seqgan: Sequence generative adversarial nets with policy gradient
Lantao Yu, Weinan Zhang, Jun Wang, and Yong Yu · 2017
Earlier work this paper cites.
Non-autoregressive neural machine translation
Jiatao Gu, James Bradbury, Caiming Xiong, Victor O.K. Li, and Richard Socher · 2018
Earlier work this paper cites.
Long text generation via adversarial training with leaked information
Jiaxian Guo, Sidi Lu, Han Cai, Weinan Zhang, Yong Yu, and Jun Wang · 2018
Earlier work this paper cites.
Deterministic non-autoregressive neural sequence modeling by iterative refinement
Jason Lee, Elman Mansimov, and Kyunghyun Cho · 2018
Earlier work this paper cites.
A call for clarity in reporting BLEU scores
Matt Post · 2018
Earlier work this paper cites.
Adafactor: Adaptive learning rates with sublinear memory cost
Noam Shazeer and Mitchell Stern · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Mask-predict: Parallel decoding of conditional masked language models
Marjan Ghazvininejad, Omer Levy, Yinhan Liu, and Luke Zettlemoyer · 2019
Earlier work this paper cites.
FlowSeq: Non-autoregressive conditional sequence generation with generative flow
Xuezhe Ma, Chunting Zhou, Xian Li, Graham Neubig, and Eduard Hovy · 2019
Cited alongside, same era.
Insertion transformer: Flexible sequence generation via insertion operations
Mitchell Stern, William Chan, Jamie Kiros, and Jakob Uszkoreit · 2019
Cited alongside, same era.
BERT has a mouth, and it must speak: BERT as a Markov random field language model
Alex Wang and Kyunghyun Cho · 2019
Cited alongside, same era.
Latent normalizing flows for discrete sequences
Zachary Ziegler and Alexander Rush · 2019
Cited alongside, same era.
Do sequence-to-sequence VAEs learn global features of sentences?
Tom Bosc and Pascal Vincent · 2020
Cited alongside, same era.
Language models are few-shot learners
Colin Wei, Yining Chen, and Tengyu Ma · 2021
Later among the works it cites.
Self-attention networks can process bounded hierarchical languages
Shunyu Yao, Binghui Peng, Christos Papadimitriou, and Karthik Narasimhan · 2021
Later among the works it cites.
Inductive biases and variable creation in self-attention mechanisms
Benjamin L Edelman, Surbhi Goel, Sham Kakade, and Cyril Zhang · 2022
Later among the works it cites.
Exposing the implicit energy networks behind masked language models via metropolis–hastings
Kartik Goyal, Chris Dyer, and Taylor Berg-Kirkpatrick · 2022
Later among the works it cites.
Vision transformers provably learn spatial structure
Samy Jelassi, Michael Eli Sander, and Yuanzhi Li · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Cited alongside, same era.
Imputer: Sequence modelling via imputation and dynamic programming
William Chan, Chitwan Saharia, Geoffrey Hinton, Mohammad Norouzi, and Navdeep Jaitly · 2020
Cited alongside, same era.
Residual energy-based models for text generation
Yuntian Deng, Anton Bakhtin, Myle Ott, Arthur Szlam, and Marc’Aurelio Ranzato · 2020
Cited alongside, same era.
Semi-autoregressive training improves mask-predict decoding, 2020
Marjan Ghazvininejad, Omer Levy, and Luke Zettlemoyer · 2020
Cited alongside, same era.
Jointly masked sequence-to-sequence model for non-autoregressive neural machine translation
Junliang Guo, Linli Xu, and Enhong Chen · 2020
Cited alongside, same era.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi · 2020
Cited alongside, same era.
Non-autoregressive machine translation with disentangled context transformer
Jungo Kasai, James Cross, Marjan Ghazvininejad, and Jiatao Gu · 2020
Cited alongside, same era.
Diffusion-LM improves controllable text generation
Xiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang, and Tatsunori Hashimoto · 2022
Later among the works it cites.
Masked prediction: A parameter identifiability view
Bingbin Liu, Daniel Hsu, Pradeep Kumar Ravikumar, and Andrej Risteski · 2022
Later among the works it cites.
COLD decoding: Energy-based constrained text generation with langevin dynamics
Lianhui Qin, Sean Welleck, Daniel Khashabi, and Yejin Choi · 2022
Later among the works it cites.
Diffuser: Diffusion via edit-based reconstruction
Machel Reid, Vincent Josua Hellendoorn, and Graham Neubig · 2022
Later among the works it cites.
Scaling up models and data with t5x
Adam Roberts, Hyung Won Chung, Anselm Levskaya, Gaurav Mishra, James Bradbury, Daniel Andor, Sharan Narang, Brian Lester, Colin Gaffney, Afroz Mohiuddin, Curtis Hawthorne, Aitor Lewkowycz, Alex Salcianu, Marc van Zee, Jacob Austin, Sebastian Goodman, Livio Baldini Soares, Haitang Hu, Sasha Tsvyashchenko, Aakanksha Chowdhery, Jasmijn Bastings, Jannis Bulian, Xavier Garcia, Jianmo Ni, Andrew Chen, Kathleen Kenealy, Jonathan H. Clark, Stephan Lee, Dan Garrette, James Lee-Thorp, Colin Raffel, Noam Shazeer, Marvin Ritter, Maarten Bosma, Alexandre Passos, Jeremy Maitin-Shepard, Noah Fiedel, Mark Omernick, Brennan Saeta, Ryan Sepassi, Alexander Spiridonov, Joshua Newlan, and Andrea Gesmundo · 2022
Later among the works it cites.
Step-unrolled denoising autoencoders for text generation
Nikolay Savinov, Junyoung Chung, Mikolaj Binkowski, Erich Elsen, and Aaron van den Oord · 2022
Later among the works it cites.
Non-autoregressive neural machine translation: A call for clarity, 2022
Robin M. Schmidt, Telmo Pires, Stephan Peitz, and Jonas Lööf · 2022
Later among the works it cites.
On the inconsistencies of conditionals learned by masked language models
Tom Young and Yang You · 2022
Later among the works it cites.
Parallel discrete sampling via continuous walks
Nima Anari, Yizhi Huang, Tianyu Liu, Thuy-Duong Vuong, Brian Xu, and Katherine Yu · 2023
Later among the works it cites.
Diffuseq: Sequence to sequence text generation with diffusion models
Shansan Gong, Mukai Li, Jiangtao Feng, Zhiyong Wu, and Lingpeng Kong · 2023
Later among the works it cites.
Statistical efficiency of score matching: The view from isoperimetry
Frederic Koehler, Alexander Heckett, and Andrej Risteski · 2023
Later among the works it cites.
Parallelising glauber dynamics, 2023
Holden Lee · 2023
Later among the works it cites.
How do transformers learn topic structure: Towards a mechanistic understanding
Yuchen Li, Yuanzhi Li, and Andrej Risteski · 2023
Later among the works it cites.
Transformers learn shortcuts to automata
Bingbin Liu, Jordan T. Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang · 2023
Later among the works it cites.
Discrete diffusion language modeling by estimating the ratios of the data distribution, 2023
Aaron Lou, Chenlin Meng, and Stefano Ermon · 2023
Later among the works it cites.
Representation deficiency in masked language modeling
Yu Meng, Jitin Krishnan, Sinong Wang, Qifan Wang, Yuning Mao, Han Fang, Marjan Ghazvininejad, Jiawei Han, and Luke Zettlemoyer · 2023
Later among the works it cites.
Provable benefits of score matching
Chirag Pabbaraju, Dhruv Rohatgi, Anish Sevekari, Holden Lee, Ankur Moitra, and Andrej Risteski · 2023
Later among the works it cites.
Fit like you sample: Sample-efficient generalized score matching from fast mixing diffusions, 2023
Yilong Qin and Andrej Risteski · 2023
Later among the works it cites.
Deriving language models from masked language models
Lucas Torroba Hennigen and Yoon Kim · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
Transformers are uninterpretable with myopic methods: a case study with bounded dyck grammars
Kaiyue Wen, Yuchen Li, Bingbin Liu, and Andrej Risteski · 2023
Later among the works it cites.
Do transformers parse while predicting the masked word?
Haoyu Zhao, Abhishek Panigrahi, Rong Ge, and Sanjeev Arora · 2023
Later among the works it cites.
A reparameterized discrete diffusion model for text generation, 2023
Lin Zheng, Jianbo Yuan, Lei Yu, and Lingpeng Kong · 2023
Later among the works it cites.
The pitfalls of next-token prediction, 2024
Gregor Bachmann and Vaishnavh Nagarajan · 2024
Closest in time.
Medusa: Simple llm inference acceleration framework with multiple decoding heads, 2024
Tianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng, Jason D. Lee, Deming Chen, and Tri Dao · 2024
Closest in time.