Fetching the paper…
Reading the bibliography…
The recently introduced path-star task is a minimal task designed to exemplify limitations to the abilities of language models (Bachmann and Nagarajan, 2024).
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 1904
Earlier work this paper cites.
A learning algorithm for continually running fully recurrent neural networks
Ronald J Williams and David Zipser. 1989 · 1989
Earlier work this paper cites.
Glancing transformer for non-autoregressive neural machine translation
Lihua Qian, Hao Zhou, Yu Bao, Mingxuan Wang, Lin Qiu, Weinan Zhang, Yong Yu, and Lei Li. 2021 · 2003
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer. 2015 · 2015
Earlier work this paper cites.
Sequence level training with recurrent neural networks
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba. 2016 · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Non-autoregressive neural machine translation
Jiatao Gu, James Bradbury, Caiming Xiong, Victor OK Li, and Richard Socher. 2018 · 2018
Earlier work this paper cites.
Deterministic non-autoregressive neural sequence modeling by iterative refinement
Jason Lee, Elman Mansimov, and Kyunghyun Cho. 2018 · 2018
Earlier work this paper cites.
Semi-autoregressive neural machine translation
Chunqi Wang, Ji Zhang, and Haiqing Chen. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Mask-predict: Parallel decoding of conditional masked language models
Marjan Ghazvininejad, Omer Levy, Yinhan Liu, and Luke Zettlemoyer. 2019 · 2019
Earlier work this paper cites.
Transformer dissection: An unified understanding for transformer’s attention via the lens of kernel
Yao-Hung Hubert Tsai, Shaojie Bai, Makoto Yamada, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2019 · 2019
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. 2019 · 2019
Earlier work this paper cites.
Probabilistically masked language model capable of autoregressive generation in arbitrary word order
Yi Liao, Xin Jiang, and Qun Liu. 2020 · 2020
Cited alongside, same era.
Fully non-autoregressive neural machine translation: Tricks of the trade
Jiatao Gu and Xiang Kong. 2021 · 2021
Cited alongside, same era.
Sensitivity as a complexity measure for sequence classification tasks
Michael Hahn, Dan Jurafsky, and Richard Futrell. 2021 · 2021
Cited alongside, same era.
Thinking like transformers
Gail Weiss, Yoav Goldberg, and Eran Yahav. 2021 · 2021
Cited alongside, same era.
Why exposure bias matters: An imitation learning perspective of error accumulation in language generation
Kushal Arora, Layla El Asri, Hareesh Bahuleyan, and Jackie Cheung. 2022 · 2022
Cited alongside, same era.
Transformer language models without positional encodings still learn positional information
The impact of positional encoding on length generalization in transformers
Amirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Payel Das, and Siva Reddy. 2023 · 2023
Later among the works it cites.
Pass: Parallel speculative sampling
Giovanni Monea, Armand Joulin, and Edouard Grave. 2023 · 2023
Later among the works it cites.
On the planning abilities of large language models - a critical investigation
Karthik Valmeekam, Matthew Marquez, Sarath Sreedharan, and Subbarao Kambhampati. 2023 · 2023
Later among the works it cites.
Sub-task decomposition enables learning in sequence to sequence tasks
Noam Wies, Yoav Levine, and Amnon Shashua. 2023 · 2023
Later among the works it cites.
The pitfalls of next-token prediction
Gregor Bachmann and Vaishnavh Nagarajan. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Adi Haviv, Ori Ram, Ofir Press, Peter Izsak, and Omer Levy. 2022 · 2022
Cited alongside, same era.
Emergent abilities of large language models
Barret Zoph, Colin Raffel, Dale Schuurmans, Dani Yogatama, Denny Zhou, Don Metzler, Ed H. Chi, Jason Wei, Jeff Dean, Liam B. Fedus, Maarten Paul Bosma, Oriol Vinyals, Percy Liang, Sebastian Borgeaud, Tatsunori B. Hashimoto, and Yi Tay. 2022 · 2022
Cited alongside, same era.
Simplicity bias in transformers and their ability to learn sparse Boolean functions
Satwik Bhattamishra, Arkil Patel, Varun Kanade, and Phil Blunsom. 2023 · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al. 2023 · 2023
Cited alongside, same era.
Zero-shot approach to overcome perturbation sensitivity of prompts
Mohna Chakraborty, Adithya Kulkarni, and Qi Li. 2023 · 2023
Cited alongside, same era.
On the relation between sensitivity and accuracy in in-context learning
Yanda Chen, Chen Zhao, Zhou Yu, Kathleen McKeown, and He He. 2023 · 2023
Cited alongside, same era.
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao. 2023 · 2023
Cited alongside, same era.
Lukas Berglund, Meg Tong, Maximilian Kaufmann, Mikita Balesni, Asa Cooper Stickland, Tomasz Korbak, and Owain Evans. 2024 · 2024
Closest in time.
Premise order matters in reasoning with large language models
Xinyun Chen, Ryan Andrew Chi, Xuezhi Wang, and Denny Zhou. 2024 · 2024
Closest in time.
Reverse training to nurse the reversal curse
Olga Golovneva, Zeyuan Allen-Zhu, Jason E Weston, and Sainbayar Sukhbaatar. 2024 · 2024
Closest in time.
Think before you speak: Training language models with pause tokens
Sachin Goyal, Ziwei Ji, Ankit Singh Rawat, Aditya Krishna Menon, Sanjiv Kumar, and Vaishnavh Nagarajan. 2024 · 2024
Closest in time.
Why are sensitive functions hard for transformers?
Michael Hahn and Mark Rofin. 2024 · 2024
Closest in time.
Marianna Nezhurina, Lucia Cipolina-Kun, Mehdi Cherti, and Jenia Jitsev. 2024 · 2024
Closest in time.
Transformers, parallel computation, and logarithmic depth
Clayton Sanford, Daniel Hsu, and Matus Telgarsky. 2024 · 2024
Closest in time.
What algorithms can transformers learn? a study in length generalization
Hattie Zhou, Arwen Bradley, Etai Littwin, Noam Razin, Omid Saremi, Joshua M. Susskind, Samy Bengio, and Preetum Nakkiran. 2024 · 2024
Closest in time.