Fetching the paper…
Reading the bibliography…
We introduce the "Belief State Transformer", a next-token predictor that takes both a prefix and suffix as inputs, with a novel objective of predicting both the next token for the prefix and the previous token for the suffix.
Liquid time-constant networks, 2020
Ramin Hasani, Mathias Lechner, Alexander Amini, Daniela Rus, and Radu Grosu · 2006
Earlier work this paper cites.
Sequence transduction with recurrent neural networks
Alex Graves · 2012
Earlier work this paper cites.
Backbone language modeling for constrained natural language generation
Lili Mou, Rui Yan, Ge Li, Lu Zhang, and Zhi Jin · 2015
Earlier work this paper cites.
Agreement on target-bidirectional neural machine translation
Lemao Liu, Masao Utiyama, Andrew Finch, and Eiichiro Sumita · 2016
Earlier work this paper cites.
Failures of gradient-based deep learning
Shai Shalev-Shwartz, Ohad Shamir, and Shaked Shammah · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Deep contextualized word representations. corr
ME Peters, M Neumann, M Iyyer, M Gardner, C Clark, K Lee, and L Zettlemoyer · 2018
Earlier work this paper cites.
Twin networks: Matching the future for sequence generation, 2018
Dmitriy Serdyuk, Nan Rosemary Ke, Alessandro Sordoni, Adam Trischler, Chris Pal, and Yoshua Bengio · 2018
Earlier work this paper cites.
Regularizing neural machine translation by target-bidirectional agreement, 2018
Zhirui Zhang, Shuangzhi Wu, Shujie Liu, Mu Li, Ming Zhou, and Tong Xu · 2018
Earlier work this paper cites.
Insertion-based decoding with automatically inferred generation order
Jiatao Gu, Qi Liu, and Kyunghyun Cho · 2019
Earlier work this paper cites.
Lstm vs. gru vs. bidirectional rnn for script generation
Sanidhya Mangal, Poorva Joshi, and Rahul Modak · 2019
Cited alongside, same era.
Non-monotonic sequential text generation
Sean Welleck, Kianté Brantley, Hal Daumé Iii, and Kyunghyun Cho · 2019
Cited alongside, same era.
On the non-universality of deep learning: quantifying the cost of symmetry
Emmanuel Abbe and Enric Boix-Adsera · 2022
Cited alongside, same era.
Efficient training of language models to fill in the middle
Mohammad Bavarian, Heewoo Jun, Nikolas Tezak, John Schulman, Christine McLeavey, Jerry Tworek, and Mark Chen · 2022
Cited alongside, same era.
Polynomial-time universality and limitations of deep learning
Emmanuel Abbe and Colin Sandon · 2023
Cited alongside, same era.
Medusa: Simple llm inference acceleration framework with multiple decoding heads
Tianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng, Jason D Lee, Deming Chen, and Tri Dao · 2024
Closest in time.
The mystery of the pathological path-star task for language models
Arvid Frydenlund · 2024
Closest in time.
Better & faster large language models via multi-token prediction
Fabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière, David Lopez-Paz, and Gabriel Synnaeve · 2024
Closest in time.
Teaching arithmetic to small transformers
Nayoung Lee, Kartik Sreenivasan, Jason D. Lee, Kangwook Lee, and Dimitris Papailiopoulos · 2024
Closest in time.
Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al · 2023
Cited alongside, same era.
Tinystories: How small can language models be and still speak coherent english?
Ronen Eldan and Yuanzhi Li · 2023
Cited alongside, same era.
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao · 2023
Cited alongside, same era.
The pitfalls of next-token prediction
Gregor Bachmann and Vaishnavh Nagarajan · 2024
Cited alongside, same era.
Closest in time.
Evaluating cognitive maps and planning in large language models with cogeval
Ida Momennejad, Hosein Hasanbeig, Felipe Vieira Frujeri, Hiteshi Sharma, Nebojsa Jojic, Hamid Palangi, Robert Ness, and Jonathan Larson · 2024
Closest in time.
Meet in the middle: A new pre-training paradigm
Anh Nguyen, Nikos Karampatziakis, and Weizhu Chen · 2024
Closest in time.
Arrows of time for large language models
Vassilis Papadopoulos, Jérémie Wenger, and Clément Hongler · 2024
Closest in time.
Transformers, parallel computation, and logarithmic depth
Clayton Sanford, Daniel Hsu, and Matus Telgarsky · 2024
Closest in time.
Transformers represent belief state geometry in their residual stream, 2024
Adam S. Shai, Sarah E. Marzen, Lucas Teixeira, Alexander Gietelink Oldenziel, and Paul M. Riechers · 2024
Closest in time.