Fetching the paper…
Reading the bibliography…
Despite their remarkable success in language modeling, transformers trained to predict the next token in a sequence struggle with long-term planning.
Value iteration networks
Aviv Tamar, S. Levine, P. Abbeel, Yi Wu, and G. Thomas · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
Structured denoising diffusion models in discrete state-spaces
Jacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow, and Rianne van den Berg · 2021
Earlier work this paper cites.
Deep reinforcement learning based robot navigation in dynamic environments using occupancy values of motion primitives
Neset Unver Akmandor, Hongyu Li, Gary Lvov, Eric Dusel, and T. Padır · 2022
Earlier work this paper cites.
Diffusionbert: Improving generative masked language models with diffusion models
Zhengfu He, Tianxiang Sun, Kuan Wang, Xuanjing Huang, and Xipeng Qiu · 2022
Earlier work this paper cites.
Planning with diffusion for flexible behavior synthesis
Michael Janner, Yilun Du, Joshua B. Tenenbaum, and Sergey Levine · 2022
Earlier work this paper cites.
Diffusion-lm improves controllable text generation
Xiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang, and Tatsunori Hashimoto · 2022
Cited alongside, same era.
Investigating navigation strategies in the morris water maze through deep reinforcement learning
A. Liu and A. Borisyuk · 2023
Cited alongside, same era.
Roformer: Enhanced transformer with rotary position embedding, 2023
Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu · 2023
Cited alongside, same era.
Attention is all you need, 2023
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2023
Cited alongside, same era.
The pitfalls of next-token prediction
Gregor Bachmann and Vaishnavh Nagarajan · 2024
Cited alongside, same era.
Faith and fate: Limits of transformers on compositionality
Nouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li, Liwei Jiang, Bill Yuchen Lin, Sean Welleck, Peter West, Chandra Bhagavatula, Ronan Le Bras, et al · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Gemma · 2024
Closest in time.
Better & faster large language models via multi-token prediction
Fabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière, David Lopez-Paz, and Gabriel Synnaeve · 2024
Closest in time.
Latent plan transformer: Planning as latent variable inference
Deqian Kong, Dehong Xu, Minglu Zhao, Bo Pang, Jianwen Xie, Andrew Lizarraga, Yuhao Huang, Sirui Xie, and Ying Nian Wu · 2024
Closest in time.
Beyond a*: Better planning with transformers via search dynamics bootstrapping
Lucas Lehnert, Sainbayar Sukhbaatar, Paul Mcvay, Michael Rabbat, and Yuandong Tian · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lukas Berglund, Meg Tong, Max Kaufmann, Mikita Balesni, Asa Cooper Stickland, Tomasz Korbak, and Owain Evans · 2024
Cited alongside, same era.
The llama 3 herd of models
Abhimanyu Dubey et al · 2024
Cited alongside, same era.
A configurable library for generating and manipulating maze datasets
Michael Igorevich Ivanitskiy, Rusheb Shah, Alex F Spies, Tilman Räuker, Dan Valentine, Can Rager, Lucia Quirke, Chris Mathwin, Guillaume Corlouer, Cecilia Diniz Behn, et al
Cited in the paper.
Structured world representations in maze-solving transformers
Michael Igorevich Ivanitskiy, Alex F. Spies, Tilman Räuker, Guillaume Corlouer, Chris Mathwin, Lucia Quirke, Can Rager, Rusheb Shah, Dan Valentine, Cecilia Diniz Behn, Katsumi Inoue, and Samy Wu Fung
Cited in the paper.
The factorization curse: Which tokens you predict underlie the reversal curse and more, 2024a
Ouail Kitouni, Niklas Nolte, Diane Bouchacourt, Adina Williams, Mike Rabbat, and Mark Ibrahim
Cited in the paper.
Disk: A diffusion model for structured knowledge, 2024b
Ouail Kitouni, Niklas Nolte, James Hensman, and Bhaskar Mitra
Cited in the paper.
Haitong Wang, Aaron Hao Tan, and Goldie Nejat
Cited in the paper.
Simple and effective masked diffusion language models
Subham Sekhar Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan, Edgar Marroquin, Justin T Chiu, Alexander Rush, and Volodymyr Kuleshov · 2024
Closest in time.
Simplified and generalized masked diffusion for discrete data, 2024
Jiaxin Shi, Kehang Han, Zhe Wang, Arnaud Doucet, and Michalis K. Titsias · 2024
Closest in time.