Fetching the paper…
Reading the bibliography…
Transformers trained on huge text corpora exhibit a remarkable set of capabilities, e.g., performing basic arithmetic.
The language of thought , volume 5
Fodor, J. A · 1975
Earlier work this paper cites.
Connectionism and cognitive architecture: A critical analysis
Fodor, J. A. and Pylyshyn, Z. W · 1988
Earlier work this paper cites.
The compositionality papers
Fodor, J. A. and Lepore, E · 2002
Earlier work this paper cites.
Some new aspects of the coupon collector’s problem
Myers, A. N. and Wilf, H. S · 2006
Earlier work this paper cites.
A generalized coupon collector problem
Xu, W. and Tang, A. K · 2011
Earlier work this paper cites.
Convergent learning: Do different neural networks learn the same representations?
Li, Y., Yosinski, J., Clune, J., Lipson, H., and Hopcroft, J · 2015
Earlier work this paper cites.
Gaussian error linear units (gelus)
Hendrycks, D. and Gimpel, K · 2016
Earlier work this paper cites.
Using the output embedding to improve language models
Press, O. and Wolf, L · 2016
Earlier work this paper cites.
Probing the compositionality of intuitive functions
Schulz, E., Tenenbaum, J., Duvenaud, D. K., Speekenbrink, M., and Gershman, S. J · 2016
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, 𝖫 \mathsf{L} ., and Polosukhin, I · 2017
Earlier work this paper cites.
Learning compositionally through attentive guidance
Hupkes, D., Singh, A., Korrel, K., Kruszewski, G., and Bruni, E · 2018
Earlier work this paper cites.
Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks
Lake, B. and Baroni, M · 2018
Earlier work this paper cites.
Memorize or generalize? searching for a compositional rnn in a haystack
Liška, A., Kruszewski, G., and Baroni, M · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al · 2018
Earlier work this paper cites.
Reducing sentiment bias in language models via counterfactual evaluation
Huang, P.-S., Zhang, H., Jiang, R., Stanforth, R., Welbl, J., Rae, J., Maini, V., Yogatama, D., and Kohli, P · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
The woman worked as a babysitter: On biases in language generation
Sheng, E., Chang, K.-W., Natarajan, P., and Peng, N · 2019
Earlier work this paper cites.
Bert rediscovers the classical nlp pipeline
Tenney, I., Das, D., and Pavlick, E · 2019
Earlier work this paper cites.
On the ability and limitations of transformers to recognize formal languages
Bhattamishra, S., Ahuja, K., and Goyal, N · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Gehman, S., Gururangan, S., Sap, M., Choi, Y., and Smith, N. A · 2020
Earlier work this paper cites.
Scaling laws for autoregressive generative modeling
Henighan, T., Kaplan, J., Katz, M., Chen, M., Hesse, C., Jackson, J., Jun, H., Brown, T. B., Dhariwal, P., Gray, S., et al · 2020
Earlier work this paper cites.
Compositionality decomposed: How do neural networks generalise?
Hupkes, D., Dankers, V., Mul, M., and Bruni, E · 2020
Earlier work this paper cites.
The radicalization risks of gpt-3 and advanced neural language models
McGuffie, K. and Newhouse, A · 2020
Earlier work this paper cites.
Zoom in: An introduction to circuits
Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., and Carter, S · 2020
Earlier work this paper cites.
A neural scaling law from the dimension of the data manifold
Sharma, U. and Kaplan, J · 2020
Earlier work this paper cites.
Recipes for safety in open-domain chatbots
Xu, J., Ju, D., Li, M., Boureau, Y.-L., Weston, J., and Dinan, E · 2020
Earlier work this paper cites.
Persistent anti-muslim bias in large language models
Abid, A., Farooqi, M., and Zou, J · 2021
Earlier work this paper cites.
Program synthesis with large language models
Austin, J., Odena, A., Nye, M., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C., Terry, M., Le, Q., et al · 2021
Earlier work this paper cites.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al · 2021
Cited alongside, same era.
A survey on bias in deep nlp
Garrido-Muñoz, I., Montejo-Ráez, A., Martínez-Santiago, F., and Ureña-López, L. A · 2021
Cited alongside, same era.
Hernandez, D., Kaplan, J., Henighan, T., and McCandlish, S · 2021
Cited alongside, same era.
Can machines learn morality? the delphi experiment
Jiang, L., Hwang, J. D., Bhagavatula, C., Le Bras, R., Liang, J., Dodge, J., Sakaguchi, K., Forbes, M., Borchardt, J., Gabriel, S., et al · 2021
Cited alongside, same era.
Scaling scaling laws with board games
Jones, A. L · 2021
Cited alongside, same era.
Transformers learn shortcuts to automata
Liu, B., Ash, J. T., Goel, S., Krishnamurthy, A., and Zhang, C · 2022
Later among the works it cites.
In-context learning and induction heads
Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Johnston, S., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2022
Later among the works it cites.
Impact of pretraining term frequencies on few-shot reasoning
Razeghi, Y., Logan IV, R. L., Gardner, M., and Singh, S · 2022
Later among the works it cites.
Language models are greedy reasoners: A systematic formal analysis of chain-of-thought
Saparov, A. and He, H · 2022
Later among the works it cites.
Goal misgeneralization: Why correct specifications aren’t enough for correct goals
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Truthfulqa: Measuring how models mimic human falsehoods
Lin, S., Hilton, J., and Evans, O · 2021
Cited alongside, same era.
A mathematical framework for transformer circuits
Nelson, E., Neel, N., Catherine, O., Tom, H., Nicholas, J., Ben, M., Amanda, A., Yuntao, B., Anna, C., Tom, C., et al · 2021
Cited alongside, same era.
Show your work: Scratchpads for intermediate computation with language models
Nye, M., Andreassen, A. J., Gur-Ari, G., Michalewski, H., Austin, J., Bieber, D., Dohan, D., Lewkowycz, A., Bosma, M., Luan, D., et al · 2021
Cited alongside, same era.
Making transformers solve compositional tasks
Ontanón, S., Ainslie, J., Cvicek, V., and Fisher, Z · 2021
Cited alongside, same era.
Bbq: A hand-built bias benchmark for question answering
Parrish, A., Chen, A., Nangia, N., Padmakumar, V., Phang, J., Thompson, J., Htut, P. M., and Bowman, S. R · 2021
Cited alongside, same era.
Scaling language models: Methods, analysis & insights from training gopher
Rae, J. W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, F., Aslanides, J., Henderson, S., Ring, R., Young, S., et al · 2021
Cited alongside, same era.
Multitask prompted training enables zero-shot task generalization
Sanh, V., Webson, A., Raffel, C., Bach, S. H., Sutawika, L., Alyafeai, Z., Chaffin, A., Stiegler, A., Scao, T. L., Raja, A., et al · 2021
Cited alongside, same era.
Shah, R., Varma, V., Kumar, R., Phuong, M., Krakovna, V., Uesato, J., and Kenton, Z · 2022
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al · 2022
Later among the works it cites.
Challenging big-bench tasks and whether chain-of-thought can solve them
Suzgun, M., Scales, N., Schärli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q. V., Chi, E. H., Zhou, D., et al · 2022
Later among the works it cites.
Transcending scaling laws with 0.1% extra compute
Tay, Y., Wei, J., Chung, H. W., Tran, V. Q., So, D. R., Shakeri, S., Garcia, X., Zheng, H. S., Rao, J., Chowdhery, A., et al · 2022
Later among the works it cites.
Lamda: Language models for dialog applications
Thoppilan, R., De Freitas, D., Hall, J., Shazeer, N., Kulshreshtha, A., Cheng, H.-T., Jin, A., Bos, T., Baker, L., Du, Y., et al · 2022
Later among the works it cites.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small
Wang, K., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J · 2022
Later among the works it cites.
Do vision-language pretrained models learn composable primitive concepts?
Yun, T., Bhalla, U., Pavlick, E., and Sun, C · 2022
Later among the works it cites.
Transformers learn to implement preconditioned gradient descent for in-context learning
Ahn, K., Cheng, X., Daneshmand, H., and Sra, S · 2023
Closest in time.
A theory for emergence of complex skills in language models
Arora, S. and Goyal, A · 2023
Closest in time.
Sparks of artificial general intelligence: Early experiments with gpt-4
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al · 2023
Closest in time.
Harms from increasingly agentic algorithmic systems
Chan, A., Salganik, R., Markelius, A., Pang, C., Rajkumar, N., Krasheninnikov, D., Langosco, L., He, Z., Duan, Y., Carroll, M., et al · 2023
Closest in time.
A toy model of universality: Reverse engineering how networks learn group operations, may 2023
Chughtai, B., Chan, L., and Nanda, N · 2023
Closest in time.
Holistic evaluation of text-to-image models
Lee, T., Yasunaga, M., Meng, C., Mai, Y., Park, J. S., Gupta, A., Zhang, Y., Narayanan, D., Teufel, H. B., Bellagente, M., et al · 2023
Closest in time.
Break it down: Evidence for structural compositionality in neural networks
Lepori, M. A., Serre, T., and Pavlick, E · 2023
Closest in time.
Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task, 2023a
Li, K., Hopkins, A. K., Bau, D., Viégas, F., Pfister, H., and Wattenberg, M · 2023
Closest in time.
Liu, H., Li, C., Wu, Q., and Lee, Y. J · 2023
Closest in time.
Mechanistic mode connectivity
Lubana, E. S., Bigelow, E. J., Dick, R. P., Krueger, D., and Tanaka, H · 2023
Closest in time.
The expresssive power of transformers with chain of thought
Merrill, W. and Sabharwal, A · 2023
Closest in time.
Compositional abilities emerge multiplicatively: Exploring diffusion models on a synthetic task
Okawa, M., Lubana, E. S., Dick, R. P., and Tanaka, H · 2023
Closest in time.
Toolllm: Facilitating large language models to master 16000+ real-world apis
Qin, Y., Liang, S., Ye, Y., Zhu, K., Yan, L., Lu, Y., Lin, Y., Cong, X., Tang, X., Qian, B., et al · 2023
Closest in time.
Model evaluation for extreme risks
Shevlane, T., Farquhar, S., Garfinkel, B., Phuong, M., Whittlestone, J., Leung, J., Kokotajlo, D., Marchal, N., Anderljung, M., Kolt, N., et al · 2023
Closest in time.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Closest in time.
Transformers learn in-context by gradient descent
Von Oswald, J., Niklasson, E., Randazzo, E., Sacramento, J., Mordvintsev, A., Zhmoginov, A., and Vladymyrov, M · 2023
Closest in time.
Skill-mix: A flexible and expandable family of evaluations for ai models
Yu, D., Kaur, S., Gupta, A., Brown-Cohen, J., Goyal, A., and Arora, S · 2023
Closest in time.
What algorithms can transformers learn? a study in length generalization
Zhou, H., Bradley, A., Littwin, E., Razin, N., Saremi, O., Susskind, J., Bengio, S., and Nakkiran, P · 2023
Closest in time.