Fetching the paper…
Reading the bibliography…
Autoregressive language models have demonstrated a remarkable ability to extract latent structure from text.
A new approach to linear filtering and prediction problems
Rudolph Emil Kalman · 1960
Earlier work this paper cites.
Statistical inference for probabilistic functions of finite state markov chains
Leonard E Baum and Ted Petrie · 1966
Earlier work this paper cites.
A tutorial on hidden markov models and selected applications in speech recognition
L.R. Rabiner · 1989
Earlier work this paper cites.
Bayesian modeling of human concept learning
Joshua Tenenbaum · 1998
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, and Pascal Vincent · 2000
Earlier work this paper cites.
Bayesian Theory
Jose M Bernardo and Adrian Smith · 2000
Earlier work this paper cites.
Bayesian Data Analysis
Andrew Gelman, John B. Carlin, Hal S. Stern, and Donald B. Rubin · 2004
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2014
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Distributional vectors encode referential attributes
Abhijeet Gupta, Gemma Boleda, Marco Baroni, and Sebastian Padó · 2015
Earlier work this paper cites.
What’s in an embedding? analyzing word embeddings through multilingual evaluation
Arne Köhn · 2015
Earlier work this paper cites.
Probing for semantic evidence of composition by means of simple classification tasks
Allyson Ettinger, Ahmed Elgohary, and Philip Resnik · 2016
Earlier work this paper cites.
Does string-based neural MT learn source syntax?
Xing Shi, Inkit Padhi, and Kevin Knight · 2016
Earlier work this paper cites.
Fine-grained analysis of sentence embeddings using auxiliary prediction tasks
Yossi Adi, Einat Kermany, Yonatan Belinkov, Ofer Lavi, and Yoav Goldberg · 2017
Earlier work this paper cites.
What do neural machine translation models learn about morphology?
Yonatan Belinkov, Nadir Durrani, Fahim Dalvi, Hassan Sajjad, and James Glass · 2017
Earlier work this paper cites.
Learning to generate reviews and discovering sentiment
Alec Radford, Rafal Jozefowicz, and Ilya Sutskever · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Visualisation and’diagnostic classifiers’ reveal how recurrent and recursive neural networks process hierarchical structure
Dieuwke Hupkes, Sara Veldhoen, and Willem Zuidema · 2018
Cited alongside, same era.
Under the hood: Using diagnostic classifiers to investigate and improve how language models track agreement information
Mario Giulianelli, Jack Harding, Florian Mohnert, Dieuwke Hupkes, and Willem Zuidema · 2018
Cited alongside, same era.
What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties
Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni · 2018
Language models as agent models
Jacob Andreas · 2022
Later among the works it cites.
Topic discovery via latent space clustering of pretrained language model representations
Yu Meng, Yunyi Zhang, Jiaxin Huang, Yu Zhang, and Jiawei Han · 2022
Later among the works it cites.
Probing classifiers: Promises, shortcomings, and advances
Yonatan Belinkov · 2022
Later among the works it cites.
Semantic projection recovers rich human knowledge of multiple object features from word embeddings
Gabriel Grand, Idan Asher Blank, Francisco Pereira, and Evelina Fedorenko · 2022
Later among the works it cites.
What can transformers learn in-context? a case study of simple function classes
Shivam Garg, Dimitris Tsipras, Percy Liang, and Gregory Valiant · 2022
Later among the works it cites.
Transformers can do bayesian inference
Samuel Müller, Noah Hollmann, Sebastian Pineda Arango, Josif Grabocka, and Frank Hutter · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D. Manning · 2019
Cited alongside, same era.
What do you learn from context? probing for sentence structure in contextualized word representations
Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R Thomas McCoy, Najoung Kim, Benjamin Van Durme, Sam Bowman, Dipanjan Das, and Ellie Pavlick · 2019
Cited alongside, same era.
Linguistic knowledge and transferability of contextual representations
Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, and Noah A. Smith · 2019
Cited alongside, same era.
Open sesame: Getting inside BERT’s linguistic knowledge
Yongjie Lin, Yi Chern Tan, and Robert Frank · 2019
Cited alongside, same era.
A primer in BERTology: What we know about how BERT works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky · 2020
Cited alongside, same era.
Meta-trained agents implement bayes-optimal agents
Vladimir Mikulik, Grégoire Delétang, Tom McGrath, Tim Genewein, Miljan Martic, Shane Legg, and Pedro A. Ortega · 2020
Cited alongside, same era.
Later among the works it cites.
Sparse autoencoders find highly interpretable features in language models
Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey · 2023
Later among the works it cites.
How do transformers learn topic structure: Towards a mechanistic understanding
Yuchen Li, Yuan-Fang Li, and Andrej Risteski · 2023
Later among the works it cites.
Revisiting topic-guided language models, 2023
Carolina Zheng, Keyon Vafa, and David M. Blei · 2023
Later among the works it cites.
R Thomas McCoy, Shunyu Yao, Dan Friedman, Matthew Hardy, and Thomas L Griffiths · 2023
Later among the works it cites.
Deep de finetti: Recovering topic distributions from large language models, 2023
Liyi Zhang, R. Thomas McCoy, Theodore R. Sumers, Jian-Qiao Zhu, and Thomas L. Griffiths · 2023
Later among the works it cites.
What learning algorithm is in-context learning? investigations with linear models
Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou · 2023
Later among the works it cites.
Large language models are implicitly latent variable models: Explaining and finding good demonstrations for in-context learning, 2024
Xinyi Wang, Wanrong Zhu, Michael Saxon, Mark Steyvers, and William Yang Wang · 2024
Closest in time.
In-context learning through the bayesian prism
Madhur Panwar, Kabir Ahuja, and Navin Goyal · 2024
Closest in time.