Fetching the paper…
Reading the bibliography…
Recent work suggests that large language models may implicitly learn world models.
Finite automata and the representation of events
Myhill, J · 1957
Earlier work this paper cites.
Linear automaton transformations
Nerode, A · 1958
Earlier work this paper cites.
Introduction to the Theory of Computation, Third Edition
Sipser, M · 2013
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Hierarchical neural story generation
Fan, A., Lewis, M., and Dauphin, Y · 2018
Earlier work this paper cites.
On evaluating the generalization of LSTM models in formal languages
Suzgun, M., Belinkov, Y., and Shieber, S. M · 2018
Earlier work this paper cites.
Designing and interpreting probes with control tasks
Hewitt, J. and Liang, P · 2019
Earlier work this paper cites.
The curious case of neural text degeneration
Holtzman, A., Buys, J., Du, L., Forbes, M., and Choi, Y · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
On the ability and limitations of transformers to recognize formal languages
Bhattamishra, S., Ahuja, K., and Goyal, N · 2020
Earlier work this paper cites.
Can language models encode perceptual structure without grounding? A case study in color
Abdou, M., Kulmizev, A., Hershcovich, D., Frank, S., Pavlick, E., and Søgaard, A · 2021
Earlier work this paper cites.
Implicit representations of meaning in neural language models
Li, B. Z., Nye, M., and Andreas, J · 2021
Earlier work this paper cites.
Mapping language models to grounded conceptual spaces
Patel, R. and Pavlick, E · 2021
Earlier work this paper cites.
Generating landmark navigation instructions from maps as a graph-to-text problem
Schumann, R. and Riezler, S · 2021
Cited alongside, same era.
Single-sequence protein structure prediction using a language model and deep learning
Chowdhury, R., Bouatta, N., Biswas, S., Floristean, C., Kharkar, A., Roy, K., Rochereau, C., Ahdritz, G., Zhang, J., Church, G. M., Sorger, P. K., and AlQuraishi, M · 2022
Cited alongside, same era.
Truncation sampling as language model desmoothing
Hewitt, J., Manning, C. D., and Liang, P · 2022
Cited alongside, same era.
Transformers learn shortcuts to automata
Liu, B., Ash, J. T., Goel, S., Krishnamurthy, A., and Zhang, C · 2022
Cited alongside, same era.
Analyzing generalization of vision and language navigation to unseen outdoor areas
Schumann, R. and Riezler, S · 2022
Cited alongside, same era.
Evidence of meaning in language models trained on programs
Jin, C. and Rinard, M · 2023
Later among the works it cites.
Causal reasoning and large language models: Opening a new frontier for causality
Kıcıman, E., Ness, R., Sharma, A., and Tan, C · 2023
Later among the works it cites.
Kuo, M.-T., Hsueh, C.-C., and Tsai, R. T.-H · 2023
Later among the works it cites.
Emergent world representations: Exploring a sequence model trained on a synthetic task
Li, K., Hopkins, A. K., Bau, D., Viégas, F., Pfister, H., and Wattenberg, M · 2023
Later among the works it cites.
Evolutionary-scale prediction of atomic-level protein structure with a language model
Lin, Z., Akin, H., Rao, R., Hie, B., Zhu, Z., Lu, W., Smetanin, N., Verkuil, R., Kabeli, O., Shmueli, Y., dos Santos Costa, A., Fazel-Zarandi, M., Sercu, T., Candido, S., and Rives, A · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Challenging BIG-bench tasks and whether chain-of-thought can solve them
Suzgun, M., Scales, N., Schärli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q. V., Chi, E. H., Zhou, D., et al · 2022
Cited alongside, same era.
Chess as a testbed for language model state tracking
Toshniwal, S., Wiseman, S., Livescu, K., and Gimpel, K · 2022
Cited alongside, same era.
Bai, J., Bai, S., Chu, Y., Cui, Z., Dang, K., Deng, X., Fan, Y., Ge, W., Han, Y., Huang, F., et al · 2023
Cited alongside, same era.
DNA language models are powerful predictors of genome-wide variant effects
Benegas, G., Batra, S. S., and Song, Y. S · 2023
Cited alongside, same era.
Autonomous chemical research with large language models
Boiko, D. A., MacKnight, R., Kline, B., and Gomes, G · 2023
Cited alongside, same era.
Leveraging pre-trained large language models to construct and utilize world models for model-based task planning
Guan, L., Valmeekam, K., Sreedharan, S., and Kambhampati, S · 2023
Cited alongside, same era.
Linear latent world models in simple transformers: A case study on Othello-GPT
Hazineh, D. S., Zhang, Z., and Chiu, J · 2023
Cited alongside, same era.
Later among the works it cites.
The parallelism tradeoff: Limitations of log-precision transformers
Merrill, W. and Sabharwal, A · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Modeling and analyzing urban networks and amenities with OSMnx
Boeing, G · 2024
Closest in time.
Leveraging large language models for predictive chemistry
Jablonka, K. M., Schwaller, P., Ortega-Guerrero, A., and Smit, B · 2024
Closest in time.
The illusion of state in state-space models
Merrill, W., Petty, J., and Sabharwal, A · 2024
Closest in time.
2014 New York City taxi trips
Murray, K. W · 2024
Closest in time.
VELMA: Verbalization embodiment of LLM agents for vision and language navigation in street view
Schumann, R., Zhu, W., Feng, W., Fu, T.-J., Riezler, S., and Wang, W. Y · 2024
Closest in time.