Fetching the paper…
Reading the bibliography…
Why do large language models sometimes output factual inaccuracies and exhibit erroneous reasoning? The brittleness of these models, particularly when executing long chains of reasoning, currently seems to be an inevitable price to pay for their advanced capabilities of coherently synthesizing knowledge, pragmatics, and abstract thought.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020) · 1901
Earlier work this paper cites.
Sticking to the facts: Confident decoding for faithful data-to-text generation
Tian, R., Narayan, S., Sellam, T., and Parikh, A. P. (2019) · 1910
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al. (2019) · 1910
Earlier work this paper cites.
Improvements in ionic relays
Eccles, W. H. and Jordan, F. W. (1918) · 1918
Earlier work this paper cites.
A trigger relay utilizing three-electrode thermionic vacuum tubes
Eccles, W. and Jordan, F. (1919) · 1919
Earlier work this paper cites.
The algebraic theory of context-free languages
Chomsky, N. and Schützenberger, M. P. (1959) · 1959
Earlier work this paper cites.
Algebraic theory of machines, I: Prime decomposition theorem for finite semigroups and machines
Krohn, K. and Rhodes, J. (1965) · 1965
Earlier work this paper cites.
Cascade synthesis of finite-state machines
Zeiger, H. P. (1967) · 1967
Earlier work this paper cites.
Automata, languages, and machines
Eilenberg, S. (1974) · 1974
Earlier work this paper cites.
Finite monoids and the fine structure of 𝖭𝖢 1 \mathsf{NC}^{1}
Barrington, D. A. M. and Thérien, D. (1988) · 1988
Earlier work this paper cites.
Finding structure in time
Elman, J. L. (1990) · 1990
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
Glu variants improve transformer
Shazeer, N. (2020) · 2002
Earlier work this paper cites.
Transferring inductive biases through knowledge distillation
Abnar, S., Dehghani, M., and Zuidema, W. (2020) · 2006
Earlier work this paper cites.
Long range dependence
Samorodnitsky, G. et al. (2007) · 2007
Earlier work this paper cites.
Applications of automata theory and algebra: via the mathematical theory of complexity to biology, physics, psychology, philosophy, and games
Rhodes, J., Nehaniv, C. L., and Hirsch, M. W. (2010) · 2010
Earlier work this paper cites.
Long range arena: A benchmark for efficient transformers
Tay, Y., Dehghani, M., Abnar, S., Shen, Y., Bahri, D., Pham, P., Rao, J., Yang, L., Ruder, S., and Metzler, D. (2020) · 2011
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Sutskever, I., Martens, J., Dahl, G., and Hinton, G. (2013) · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
Graves, A., Wayne, G., and Danihelka, I. (2014) · 2014
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Luong, M.-T., Pham, H., and Manning, C. D. (2015) · 2015
Earlier work this paper cites.
Assessing the ability of lstms to learn syntax-sensitive dependencies
Linzen, T., Dupoux, E., and Goldberg, Y. (2016) · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Wu, Y., Schuster, M., Chen, Z., Le, Q. V., Norouzi, M., Macherey, W., Krikun, M., Cao, Y., Gao, Q., Macherey, K., et al. (2016) · 2016
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F. (2017) · 2017
Earlier work this paper cites.
Automatic differentiation in pytorch
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A. (2017) · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017) · 2017
Earlier work this paper cites.
Dehghani, M., Gouws, S., Vinyals, O., Uszkoreit, J., and Kaiser, Ł. (2018) · 2018
Earlier work this paper cites.
Sharp nearby, fuzzy far away: How neural language models use context
Khandelwal, U., He, H., Qi, P., and Jurafsky, D. (2018) · 2018
Earlier work this paper cites.
Attention with sparsity regularization for neural machine translation and summarization
Zhang, J., Zhao, Y., Li, H., and Zong, C. (2018) · 2018
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q. V., and Salakhutdinov, R. (2019) · 2019
Earlier work this paper cites.
Handling divergent reference texts when evaluating table-to-text generation
Dhingra, B., Faruqui, M., Parikh, A., Chang, M.-W., Das, D., and Cohen, W. (2019) · 2019
Earlier work this paper cites.
Exposure bias versus self-recovery: Are distortions really incremental for autoregressive text generation?
He, T., Zhang, J., Zhou, Z., and Glass, J. R. (2019) · 2019
Earlier work this paper cites.
Attention is not explanation
Jain, S. and Wallace, B. C. (2019) · 2019
Cited alongside, same era.
Generalization through memorization: Nearest neighbor language models
Khandelwal, U., Levy, O., Jurafsky, D., Zettlemoyer, L., and Lewis, M. (2019) · 2019
Cited alongside, same era.
Language models as knowledge bases?
Petroni, F., Rocktäschel, T., Lewis, P., Bakhtin, A., Wu, Y., Miller, A. H., and Riedel, S. (2019) · 2019
Cited alongside, same era.
On the ability and limitations of transformers to recognize formal languages
Bhattamishra, S., Ahuja, K., and Goyal, N. (2020) · 2020
Cited alongside, same era.
Theoretical limitations of self-attention in neural sequence models
Hahn, M. (2020) · 2020
Cited alongside, same era.
Constructing a multi-hop QA dataset for comprehensive evaluation of reasoning steps
Ho, X., Duong Nguyen, A.-K., Sugawara, S., and Aizawa, A. (2020) · 2020
Vision transformers provably learn spatial structure
Jelassi, S., Sander, M. E., and Li, Y. (2022) · 2022
Later among the works it cites.
Coherence boosting: When your pretrained language model is not paying enough attention
Malkin, N., Wang, Z., and Jojic, N. (2022) · 2022
Later among the works it cites.
A mechanistic interpretability analysis of grokking
Nanda, N. and Lieberum, T. (2022) · 2022
Later among the works it cites.
Formal algorithms for transformers
Phuong, M. and Hutter, M. (2022) · 2022
Later among the works it cites.
Language models are greedy reasoners: A systematic formal analysis of chain-of-thought
Saparov, A. and He, H. (2022) · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Transformers are rnns: Fast autoregressive transformers with linear attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F. (2020) · 2020
Cited alongside, same era.
Totto: A controlled table-to-text generation dataset
Parikh, A. P., Wang, X., Gehrmann, S., Faruqui, M., Dhingra, B., Yang, D., and Das, D. (2020) · 2020
Cited alongside, same era.
Proofwriter: Generating implications, proofs, and abductive statements over natural language
Tafjord, O., Dalvi, B., and Clark, P. (2020) · 2020
Cited alongside, same era.
An interpretability illusion for BERT
Bolukbasi, T., Pearce, A., Yuan, A., Coenen, A., Reif, E., Viégas, F., and Wattenberg, M. (2021) · 2021
Cited alongside, same era.
Improving faithfulness in abstractive summarization with contrast candidate generation and selection
Chen, S., Zhang, F., Sone, K., and Roth, D. (2021) · 2021
Cited alongside, same era.
Neural path hunter: Reducing hallucination in dialogue systems via path grounding
Dziri, N., Madotto, A., Zaiane, O., and Bose, A. (2021) · 2021
Cited alongside, same era.
Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., et al. (2022) · 2022
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al. (2022) · 2022
Later among the works it cites.
Self-consistency improves chain of thought reasoning in language models
Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdhery, A., and Zhou, D. (2022) · 2022
Later among the works it cites.
Memorizing transformers
Wu, Y., Rabe, M. N., Hutchins, D. S., and Szegedy, C. (2022) · 2022
Later among the works it cites.
Unveiling transformers with lego: a synthetic reasoning task
Zhang, Y., Backurs, A., Bubeck, S., Eldan, R., Gunasekar, S., and Wagner, T. (2022) · 2022
Later among the works it cites.
Teaching algorithmic reasoning via in-context learning
Zhou, H., Nova, A., Larochelle, H., Courville, A., Neyshabur, B., and Sedghi, H. (2022) · 2022
Later among the works it cites.
Linear attention is (maybe) all you need (to understand transformer optimization)
Ahn, K., Cheng, X., Song, M., Yun, C., Jadbabaie, A., and Sra, S. (2023) · 2023
Closest in time.
Mamba: Linear-time sequence modeling with selective state spaces
Anonymous (2023) · 2023
Closest in time.
Unlimiformer: Long-range transformers with unlimited length input
Bertsch, A., Alon, U., Neubig, G., and Gormley, M. R. (2023) · 2023
Closest in time.
Selection-inference: Exploiting large language models for interpretable logical reasoning
Creswell, A., Shanahan, M., and Higgins, I. (2023) · 2023
Closest in time.
TinyStories: How small can language models be and still speak coherent English?
Eldan, R. and Li, Y. (2023) · 2023
Closest in time.
Survey of hallucination in natural language generation
Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., and Fung, P. (2023) · 2023
Closest in time.
Learning to reason and memorize with self-notes
Lanchantin, J., Toshniwal, S., Weston, J., Szlam, A., and Sukhbaatar, S. (2023) · 2023
Closest in time.
Teaching arithmetic to small transformers
Lee, N., Sreenivasan, K., Lee, J. D., Lee, K., and Papailiopoulos, D. (2023) · 2023
Closest in time.
How do transformers learn topic structure: Towards a mechanistic understanding
Li, Y., Li, Y.-F., and Risteski, A. (2023) · 2023
Closest in time.
Transformers learn shortcuts to automata
Liu, B., Ash, J. T., Goel, S., Krishnamurthy, A., and Zhang, C. (2023) · 2023
Closest in time.
The larger they are, the harder they fail: Language models do not recognize identifier swaps in python
Miceli-Barone, A. V., Barez, F., Konstas, I., and Cohen, S. B. (2023) · 2023
Closest in time.
OpenAI (2023) · 2023
Closest in time.
Resurrecting recurrent neural networks for long sequences
Orvieto, A., Smith, S. L., Gu, A., Fernando, A., Gulcehre, C., Pascanu, R., and De, S. (2023) · 2023
Closest in time.
RWKV: Reinventing RNNs for the Transformer era
Peng, B., Alcaide, E., Anthony, Q., Albalak, A., Arcadinho, S., Cao, H., Cheng, X., Chung, M., Grella, M., GV, K. K., He, X., Hou, H., Kazienko, P., Kocon, J., Kong, J., Koptyra, B., Lau, H., Mantri, K. S. I., Mom, F., Saito, A., Tang, X., Wang, B., Wind, J. S., Wozniak, S., Zhang, R., Zhang, Z., Zhao, Q., Zhou, P., Zhu, J., and Zhu, R.-J. (2023) · 2023
Closest in time.
Synthetic prompting: Generating chain-of-thought demonstrations for large language models
Shao, Z., Gong, Y., Shen, Y., Huang, M., Duan, N., and Chen, W. (2023) · 2023
Closest in time.
Large language models can be easily distracted by irrelevant context
Shi, F., Chen, X., Misra, K., Scales, N., Dohan, D., Chi, E., Schärli, N., and Zhou, D. (2023) · 2023
Closest in time.
Mlregtest: A benchmark for the machine learning of regular languages
van der Poel, S., Lambert, D., Kostyszyn, K., Gao, T., Verma, R., Andersen, D., Chau, J., Peterson, E., Clair, C. S., Fodor, P., Shibata, C., and Heinz, J. (2023) · 2023
Closest in time.
Transformers are uninterpretable with myopic methods: a case study with bounded Dyck grammars
Wen, K., Li, Y., Liu, B., and Risteski, A. (2023) · 2023
Closest in time.
How language model hallucinations can snowball
Zhang, M., Press, O., Merrill, W., Liu, A., and Smith, N. A. (2023) · 2023
Closest in time.
Do transformers parse while predicting the masked word?
Zhao, H., Panigrahi, A., Ge, R., and Arora, S. (2023) · 2023
Closest in time.
Progressive-hint prompting improves reasoning in large language models
Zheng, C., Liu, Z., Xie, E., Li, Z., and Li, Y. (2023) · 2023
Closest in time.