Fetching the paper…
Reading the bibliography…
Large language models like GPT-4 exhibit emergent capabilities across general-purpose tasks, such as basic arithmetic, when trained on extensive text data, even though these tasks are not explicitly encoded by the unsupervised, next-token prediction objective.
A simpler approach to matrix completion
Recht, B · 2011
Earlier work this paper cites.
Can recursive neural tensor networks learn logical reasoning?
Bowman, S. R · 2013
Earlier work this paper cites.
Recursive neural networks for learning logical semantics
Bowman, S. R., Potts, C., and Manning, C. D · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., and Le, Q. V · 2014
Earlier work this paper cites.
Zaremba, W. and Sutskever, I · 2014
Earlier work this paper cites.
Learning to discover efficient mathematical identities
Zaremba, W., Kurach, K., and Fergus, R · 2014
Earlier work this paper cites.
Kaiser, Ł. and Sutskever, I · 2015
Earlier work this paper cites.
char-rnn
Karpathy, A · 2015
Earlier work this paper cites.
The algebraic combinatorial approach for low-rank matrix completion
Király, F. J., Theran, L., and Tomioka, R · 2015
Earlier work this paper cites.
Neural programmer-interpreters
Reed, S. and De Freitas, N · 2015
Earlier work this paper cites.
Solving general arithmetic word problems
Roy, S. and Roth, D · 2016
Earlier work this paper cites.
Making neural programming architectures generalize via recursion
Cai, J., Shin, R., and Song, D · 2017
Earlier work this paper cites.
Towards synthesizing complex programs from input-output examples
Chen, X., Liu, C., and Song, D · 2017
Earlier work this paper cites.
Program induction by rationale generation: Learning to solve and explain algebraic word problems
Ling, W., Yogatama, D., Dyer, C., and Blunsom, P · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Dehghani, M., Gouws, S., Vinyals, O., Uszkoreit, J., and Kaiser, Ł · 2018
Earlier work this paper cites.
Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks
Lake, B. and Baroni, M · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A. and Narasimhan, K · 2018
Earlier work this paper cites.
Self-attention with relative position representations
Shaw, P., Uszkoreit, J., and Vaswani, A · 2018
Earlier work this paper cites.
Open clone of openai’s unreleased webtext dataset scraper
Peterson, J., Meylan, S., and Bourgin, D · 2019
Earlier work this paper cites.
Explain yourself! leveraging language models for commonsense reasoning
Rajani, N. F., McCann, B., Xiong, C., and Socher, R · 2019
Earlier work this paper cites.
Do nlp models know numbers? probing numeracy in embeddings
Wallace, E., Wang, Y., Li, S., Singh, S., and Gardner, M · 2019
Earlier work this paper cites.
Are transformers universal approximators of sequence-to-sequence functions?
Yun, C., Bhojanapalli, S., Rawat, A. S., Reddi, S. J., and Kumar, S · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
Leap-of-thought: Teaching pre-trained models to systematically reason over implicit knowledge
Talmor, A., Tafjord, O., Clark, P., Goldberg, Y., and Berant, J · 2020
Cited alongside, same era.
Linear algebra with transformers
Charton, F · 2021
Cited alongside, same era.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Limitations of language models in arithmetic and symbolic induction
Qian, J., Wang, H., Li, Z., Li, S., and Yan, X · 2022
Later among the works it cites.
Impact of pretraining term frequencies on few-shot reasoning
Razeghi, Y., Logan IV, R. L., Gardner, M., and Singh, S · 2022
Later among the works it cites.
Language models are multilingual chain-of-thought reasoners
Shi, F., Suzgun, M., Freitag, M., Wang, X., Srivats, S., Vosoughi, S., Chung, H. W., Tay, Y., Ruder, S., Zhou, D., et al · 2022
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Data-centric ai requires rethinking data notion
Hajij, M., Zamzmi, G., Ramamurthy, K. N., and Saenz, A. G · 2021
Cited alongside, same era.
Have you seen that number? investigating extrapolation in question answering models
Kim, J., Hong, G., Kim, K.-m., Kang, J., and Myaeng, S.-H · 2021
Cited alongside, same era.
A data-centric approach for training deep neural networks with less data
Motamedi, M., Sakharnykh, N., and Kaldewey, T · 2021
Cited alongside, same era.
Investigating the limitations of transformers with simple arithmetic tasks
Nogueira, R., Jiang, Z., and Lin, J · 2021
Cited alongside, same era.
Show your work: Scratchpads for intermediate computation with language models
Nye, M., Andreassen, A. J., Gur-Ari, G., Michalewski, H., Austin, J., Bieber, D., Dohan, D., Lewkowycz, A., Bosma, M., Luan, D., et al · 2021
Cited alongside, same era.
Making transformers solve compositional tasks
Ontanón, S., Ainslie, J., Cvicek, V., and Fisher, Z · 2021
Cited alongside, same era.
Attention is turing complete
Pérez, J., Barceló, P., and Marinkovic, J · 2021
Cited alongside, same era.
Sun, Y., Dong, L., Patra, B., Ma, S., Huang, S., Benhaim, A., Chaudhary, V., Song, X., and Wei, F · 2022
Later among the works it cites.
Transcending scaling laws with 0.1% extra compute
Tay, Y., Wei, J., Chung, H. W., Tran, V. Q., So, D. R., Shakeri, S., Garcia, X., Zheng, H. S., Rao, J., Chowdhery, A., et al · 2022
Later among the works it cites.
Lamda: Language models for dialog applications
Thoppilan, R., De Freitas, D., Hall, J., Shazeer, N., Kulshreshtha, A., Cheng, H.-T., Jin, A., Bos, T., Baker, L., Du, Y., et al · 2022
Later among the works it cites.
Solving math word problems with process-and outcome-based feedback
Uesato, J., Kushman, N., Kumar, R., Song, F., Siegel, N., Wang, L., Creswell, A., Irving, G., and Higgins, I · 2022
Later among the works it cites.
Self-consistency improves chain of thought reasoning in language models
Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., and Zhou, D · 2022
Later among the works it cites.
Emergent analogical reasoning in large language models
Webb, T., Holyoak, K. J., and Lu, H · 2022
Later among the works it cites.
Star: Self-taught reasoner bootstrapping reasoning with reasoning
Zelikman, E., Mu, J., Goodman, N. D., and Wu, Y. T · 2022
Later among the works it cites.
Sparks of artificial general intelligence: Early experiments with gpt-4
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al · 2023
Closest in time.
Teaching large language models to self-debug
Chen, X., Lin, M., Schärli, N., and Zhou, D · 2023
Closest in time.
Datacomp: In search of the next generation of multimodal datasets
Gadre, S. Y., Ilharco, G., Fang, A., Hayase, J., Smyrnis, G., Nguyen, T., Marten, R., Wortsman, M., Ghosh, D., Zhang, J., et al · 2023
Closest in time.
Looped transformers as programmable computers
Giannou, A., Rajput, S., Sohn, J.-y., Lee, K., Lee, J. D., and Papailiopoulos, D · 2023
Closest in time.
Hanna, M., Liu, O., and Variengien, A · 2023
Closest in time.
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K · 2023
Closest in time.
Exposing attention glitches with flip-flop language modeling
Liu, B., Ash, J. T., Goel, S., Krishnamurthy, A., and Zhang, C · 2023
Closest in time.
Introducing mpt-7b: A new standard for open source, commercially usable llms, 2023
MosaicML · 2023
Closest in time.
Efficiently scaling transformer inference
Pope, R., Douglas, S., Chowdhery, A., Devlin, J., Bradbury, J., Heek, J., Xiao, K., Agrawal, S., and Dean, J · 2023
Closest in time.
Llama: Open and efficient foundation language models, 2023
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G · 2023
Closest in time.
How well do large language models perform in arithmetic tasks?
Yuan, Z., Yuan, H., Tan, C., Wang, W., and Huang, S · 2023
Closest in time.