Fetching the paper…
Reading the bibliography…
Can a mere next-token predictor faithfully model human intelligence? We crystallize this emerging concern and correct popular misconceptions surrounding it, and advocate a simple multi-token objective.
Clever Hans (the horse of Mr. Von Osten) a contribution to experimental animal and human psychology
Pfungst, O. and Rahn, C. L · 1911
Earlier work this paper cites.
A mathematical theory of communication
Shannon, C. E · 1948
Earlier work this paper cites.
Prediction and entropy of printed english
Shannon, C. E · 1951
Earlier work this paper cites.
A learning algorithm for continually running fully recurrent neural networks
Williams, R. J. and Zipser, D · 1989
Earlier work this paper cites.
Lower bounds for reductions
Kääriäinen, M · 2006
Earlier work this paper cites.
Search-based structured prediction
Daumé III, H., Langford, J., and Marcu, D · 2009
Earlier work this paper cites.
Efficient reductions for imitation learning
Ross, S. and Bagnell, D · 2010
Earlier work this paper cites.
Thinking, fast and slow
Kahneman, D · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G. J., and Bagnell, D · 2011
Earlier work this paper cites.
Reinforcement and imitation learning via interactive no-regret learning
Ross, S. and Bagnell, J. A · 2014
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
Bengio, S., Vinyals, O., Jaitly, N., and Shazeer, N · 2015
Earlier work this paper cites.
Learning to search better than your teacher
Chang, K., Krishnamurthy, A., Agarwal, A., III, H. D., and Langford, J · 2015
Earlier work this paper cites.
Professor forcing: A new algorithm for training recurrent networks
Goyal, A., Lamb, A., Zhang, Y., Zhang, S., Courville, A. C., and Bengio, Y · 2016
Earlier work this paper cites.
Knowledge matters: Importance of prior information for optimization
Gülçehre, Ç. and Bengio, Y · 2016
Earlier work this paper cites.
Sequence level training with recurrent neural networks
Ranzato, M., Chopra, S., Auli, M., and Zaremba, W · 2016
Earlier work this paper cites.
On the sample complexity of end-to-end training vs. semantic abstraction training
Shalev-Shwartz, S. and Shashua, A · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Wu, Y., Schuster, M., Chen, Z., Le, Q. V., Norouzi, M., Macherey, W., Krikun, M., Cao, Y., Gao, Q., Macherey, K., et al · 2016
Earlier work this paper cites.
An actor-critic algorithm for sequence prediction
Bahdanau, D., Brakel, P., Xu, K., Goyal, A., Lowe, R., Pineau, J., Courville, A. C., and Bengio, Y · 2017
Earlier work this paper cites.
Limits of end-to-end learning
Glasmachers, T · 2017
Earlier work this paper cites.
Program induction by rationale generation: Learning to solve and explain algebraic word problems
Ling, W., Yogatama, D., Dyer, C., and Blunsom, P · 2017
Earlier work this paper cites.
Failures of gradient-based deep learning
Shalev-Shwartz, S., Shamir, O., and Shammah, S · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Non-autoregressive neural machine translation
Gu, J., Bradbury, J., Xiong, C., Li, V. O. K., and Socher, R · 2018
Earlier work this paper cites.
A deep reinforced model for abstractive summarization
Paulus, R., Xiong, C., and Socher, R · 2018
Earlier work this paper cites.
Blockwise parallel decoding for deep autoregressive models
Stern, M., Shazeer, N., and Uszkoreit, J · 2018
Earlier work this paper cites.
Openwebtext corpus
Gokaslan, A. and Cohen, V · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 2019
Earlier work this paper cites.
Burtsev, M. S., Kuratov, Y., Peganov, A., and Sapunov, G. V · 2020
Earlier work this paper cites.
Learning to summarize from human feedback
Liu, F. et al · 2020
Earlier work this paper cites.
Prophetnet: Predicting future n-gram for sequence-to-sequence pre-training
Qi, W., Yan, Y., Gong, Y., Liu, D., Duan, N., Chen, J., Zhang, R., and Zhou, M · 2020
Earlier work this paper cites.
The pitfalls of simplicity bias in neural networks
Shah, H., Tamuly, K., Raghunathan, A., Jain, P., and Netrapalli, P · 2020
Earlier work this paper cites.
Consistency of a recurrent language model with respect to incomplete decoding
Welleck, S., Kulikov, I., Kim, J., Pang, R. Y., and Cho, K · 2020
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?
Bender, E. M., Gebru, T., McMillan-Major, A., and Shmitchell, S · 2021
Cited alongside, same era.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Cited alongside, same era.
Why machine reading comprehension models learn shortcuts?
Lai, Y., Zhang, C., Feng, Y., Huang, Q., and Zhao, D · 2021
Cited alongside, same era.
Discovering non-monotonic autoregressive orderings with variational inference
Li, X., Trabucco, B., Park, D. H., Luo, M., Shen, S., Darrell, T., and Gao, Y · 2021
Cited alongside, same era.
Limitations of autoregressive models and their alternatives
Lin, C., Jaech, A., Li, X., Gormley, M. R., and Eisner, J · 2021
Cited alongside, same era.
Transformers learn shortcuts to automata
Liu, B., Ash, J. T., Goel, S., Krishnamurthy, A., and Zhang, C · 2023
Later among the works it cites.
Are we falling in a middle-intelligence trap? an analysis and mitigation of the reversal curse
Lv, A., Zhang, K., Xie, S., Tu, Q., Chen, Y., Wen, J.-R., and Yan, R · 2023
Later among the works it cites.
Auto-regressive next-token predictors are universal learners
Malach, E · 2023
Later among the works it cites.
McCoy, R. T., Yao, S., Friedman, D., Hardy, M., and Griffiths, T. L · 2023
Later among the works it cites.
The parallelism tradeoff: Limitations of log-precision transformers, 2023
Merrill, W. and Sabharwal, A · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nye, M. I., Andreassen, A. J., Gur-Ari, G., Michalewski, H., Austin, J., Bieber, D., Dohan, D., Lewkowycz, A., Bosma, M., Luan, D., Sutton, C., and Odena, A · 2021
Cited alongside, same era.
Gradient starvation: A learning proclivity in neural networks
Pezeshki, M., Kaba, S., Bengio, Y., Courville, A. C., Precup, D., and Lajoie, G · 2021
Cited alongside, same era.
Measuring and improving bert’s mathematical abilities by predicting the order of reasoning
Piekos, P., Malinowski, M., and Michalewski, H · 2021
Cited alongside, same era.
Teaching autoregressive language models complex tasks by demonstration
Recchia, G · 2021
Cited alongside, same era.
Prompt programming for large language models: Beyond the few-shot paradigm
Reynolds, L. and McDonell, K · 2021
Cited alongside, same era.
Exploring length generalization in large language models
Anil, C., Wu, Y., Andreassen, A., Lewkowycz, A., Misra, V., Ramasesh, V. V., Slone, A., Gur-Ari, G., Dyer, E., and Neyshabur, B · 2022
Cited alongside, same era.
On the role of bidirectionality in language model pre-training
Artetxe, M., Du, J., Goyal, N., Zettlemoyer, L., and Stoyanov, V · 2022
Cited alongside, same era.
Later among the works it cites.
Evaluating cognitive maps and planning in large language models with cogeval
Momennejad, I., Hasanbeig, H., Frujeri, F. V., Sharma, H., Ness, R. O., Jojic, N., Palangi, H., and Larson, J · 2023
Later among the works it cites.
Pass: Parallel speculative sampling
Monea, G., Joulin, A., and Grave, E · 2023
Later among the works it cites.
Future lens: Anticipating subsequent tokens from a single hidden state
Pal, K., Sun, J., Yuan, A., Wallace, B. C., and Bau, D · 2023
Later among the works it cites.
Eliciting language model behaviors using reverse language models
Pfau, J., Infanger, A., Sheshadri, A., Panda, A., Michael, J., and Huebner, C · 2023
Later among the works it cites.
Hans, are you clever? clever hans effect analysis of neural systems, 2023
Ranaldi, L. and Zanzotto, F. M · 2023
Later among the works it cites.
Positional description matters for transformers arithmetic
Shen, R., Bubeck, S., Eldan, R., Lee, Y. T., Li, Y., and Zhang, Y · 2023
Later among the works it cites.
Reflexion: Language agents with verbal reinforcement learning, 2023
Shinn, N., Cassano, F., Berman, E., Gopinath, A., Narasimhan, K., and Yao, S · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Image captioners are scalable vision learners too
Tschannen, M., Kumar, M., Steiner, A., Zhai, X., Houlsby, N., and Beyer, L · 2023
Later among the works it cites.
On the planning abilities of large language models - A critical investigation
Valmeekam, K., Marquez, M., Sreedharan, S., and Kambhampati, S · 2023
Later among the works it cites.
Sub-task decomposition enables learning in sequence to sequence tasks
Wies, N., Levine, Y., and Shashua, A · 2023
Later among the works it cites.
Adaptive computation with elastic input sequence
Xue, F., Likhosherstov, V., Arnab, A., Houlsby, N., Dehghani, M., and You, Y · 2023
Later among the works it cites.
React: Synergizing reasoning and acting in language models
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K. R., and Cao, Y · 2023
Later among the works it cites.
On the paradox of learning to reason from data
Zhang, H., Li, L. H., Meng, T., Chang, K., and den Broeck, G. V · 2023
Later among the works it cites.
Learning fine-grained bimanual manipulation with low-cost hardware
Zhao, T. Z., Kumar, V., Levine, S., and Finn, C · 2023
Later among the works it cites.
Fractal patterns may unravel the intelligence in next-token prediction, 2024
Alabdulmohsin, I., Tran, V. Q., and Dehghani, M · 2024
Closest in time.
Graph of thoughts: Solving elaborate problems with large language models
Besta, M., Blach, N., Kubicek, A., Gerstenberger, R., Podstawski, M., Gianinazzi, L., Gajda, J., Lehmann, T., Niewiadomski, H., Nyczyk, P., et al · 2024
Closest in time.
Faith and fate: Limits of transformers on compositionality
Dziri, N., Lu, X., Sclar, M., Li, X. L., Jiang, L., Lin, B. Y., Welleck, S., West, P., Bhagavatula, C., Le Bras, R., et al · 2024
Closest in time.
Think before you speak: Training language models with pause tokens
Goyal, S., Ji, Z., Rawat, A. S., Menon, A. K., Kumar, S., and Nagarajan, V · 2024
Closest in time.
Teaching large language models to reason with reinforcement learning, 2024
Havrilla, A., Du, Y., Raparthy, S. C., Nalmpantis, C., Dwivedi-Yu, J., Zhuravinskyi, M., Hambro, E., Sukhbaatar, S., and Raileanu, R · 2024
Closest in time.
Do large language models need sensory grounding for meaning and understanding?
LeCun, Y · 2024
Closest in time.
Mechanics of next token prediction with self-attention
Li, Y., Huang, Y., Ildiz, M. E., Rawat, A. S., and Oymak, S · 2024
Closest in time.
Self-refine: Iterative refinement with self-feedback
Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., et al · 2024
Closest in time.
The expressive power of transformers with chain of thought
Merrill, W. and Sabharwal, A · 2024
Closest in time.
Arrows of time for large language models, 2024
Papadopoulos, V., Wenger, J., and Hongler, C · 2024
Closest in time.
Transformers, parallel computation, and logarithmic depth
Sanford, C., Hsu, D., and Telgarsky, M · 2024
Closest in time.
Repetition improves language model embeddings, 2024
Springer, J. M., Kotha, S., Fried, D., Neubig, G., and Raghunathan, A · 2024
Closest in time.
Implicit bias of next-token prediction, 2024
Thrampoulidis, C · 2024
Closest in time.