Fetching the paper…
Reading the bibliography…
This paper systematically compares different methods of deriving item-level predictions of language models for multiple-choice tasks.
“Language Models are Few-Shot Learners”
Tom Brown et al · 1901
Earlier work this paper cites.
“Why Most Published Research Findings Are False”
John.. Ioannidis · 2005
Earlier work this paper cites.
“On the Predictive Power of Neural Language Models for Human Real-Time Comprehension Behavior”
Ethan Wilcox et al · 2006
Earlier work this paper cites.
“Embedded Implicatures?!?”
Bart Geurts and Nausicaa Pouscoulous · 2009
Earlier work this paper cites.
“Experimental Evidence for Embedded Scalar Implicatures”
Emmanuel Chemla and Benjamin Spector · 2011
Earlier work this paper cites.
“Lost your Marbles? The Puzzle of Dependent Measures in Experimental Pragmatics”
Judith Degen and Noah. Goodman · 2014
Earlier work this paper cites.
“Latent semantic analysis cosines as a cognitive similarity measure: Evidence from priming studies” PMID: 25952009
Fritz Günther, Carolin Dudschig and Barbara Kaup · 2015
Earlier work this paper cites.
“Task types, link functions & probabilistic modeling in experimental pragmatics”
Michael Franke · 2016
Earlier work this paper cites.
“The Seven Deadly Sins of Psychology”
Chris Chambers · 2017
Earlier work this paper cites.
“Semantic vector evaluation and human performance on a new vocabulary MCQ test”
Joseph Levy, John Bullinaria and Samantha McCormick · 2017
Earlier work this paper cites.
“A manifesto for reproducible science”
Marcus. Munafò et al · 2017
Earlier work this paper cites.
“Reproducibility in Computational Linguistics: Are We Willing to Share?”
Martijn Wieling, Josine Rawee and Gertjan van Noord · 2018
Earlier work this paper cites.
“What do RNN Language Models Learn about Filler-Gap Dependencies?”
Ethan Wilcox, Roger Levy, Takashi Morita and Richard Futrell · 2018
Earlier work this paper cites.
“Neural language models as psycholinguistic subjects: Representations of syntactic state”
Richard Futrell et al · 2019
Earlier work this paper cites.
“Linking Hypothesis and Number of Response Options Modulate Inferred Scalar Implicature Rate”
Masoud Jasbi, Brandon Waldon and Judith Degen · 2019
Earlier work this paper cites.
“Superglue: A stickier benchmark for general-purpose language understanding systems”
Alex Wang et al · 2019
Earlier work this paper cites.
“Incrementality and efficiency shape pragmatics across languages”
Paula Rubio-Fernandez and Julian Jara-Ettinger · 2020
Earlier work this paper cites.
“Masked Language Model Scoring”
Julian Salazar, Davis Liang, Toan. Nguyen and Katrin Kirchhoff · 2020
Earlier work this paper cites.
“Polite speech emerges from competing social goals”
Erica Yoon, Michael Tessler, Noah Goodman and Michael Frank · 2020
Cited alongside, same era.
“On the opportunities and risks of foundation models”
Rishi Bommasani et al · 2021
Cited alongside, same era.
“Surface Form Competition: Why the Highest Probability Answer Isn’t Always Right”
Ari Holtzman et al · 2021
Cited alongside, same era.
“Language Model Evaluation Beyond Perplexity”
Clara Meister and Ryan Cotterell · 2021
Cited alongside, same era.
“Show your work: Scratchpads for intermediate computation with language models”
Maxwell Nye et al · 2021
Cited alongside, same era.
“Sparks of artificial general intelligence: Early experiments with gpt-4”
Sébastien Bubeck et al · 2023
Later among the works it cites.
“How to handle the truth: A model of politeness as strategic truth-stretching”
Fausto Carcassi and Michael Franke · 2023
Later among the works it cites.
Mario Giulianelli et al · 2023
Later among the works it cites.
“Machine Psychology: Investigating Emergent Capabilities and Behavior in Large Language Models Using Psychological Methods”, 2023
Thilo Hagendorff · 2023
Later among the works it cites.
“A fine-grained comparison of pragmatic language understanding in humans and language models”
Jennifer Hu et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Suhas Arehalli, Brian Dillon and Tal Linzen · 2022
Cited alongside, same era.
“Palm: Scaling language modeling with pathways”
Aakanksha Chowdhery et al · 2022
Cited alongside, same era.
“Scaling instruction-finetuned language models”
Hyung Chung et al · 2022
Cited alongside, same era.
“New and improved embedding model”
Ryan Greene, Ted Sanders, Lilian Weng and Arvind Neelakantan · 2022
Cited alongside, same era.
“Language Models (Mostly) Know What They Know”, 2022
Saurav Kadavath et al · 2022
Cited alongside, same era.
“Event knowledge in large language models: the gap between the impossible and the unlikely”
Carina Kauf et al · 2022
Cited alongside, same era.
“Reducing Conversational Agents’ Overconfidence Through Linguistic Calibration”
Sabrina. Mielke, Arthur Szlam, Emily Dinan and Y-Lan Boureau · 2022
Cited alongside, same era.
“Prompt-based methods may underestimate large language models’ linguistic generalizations”
Jennifer Hu and Roger Levy · 2023
Later among the works it cites.
“Why large language models are poor theories of human linguistic cognition. A reply to Piantadosi (2023).”
Roni Katzir · 2023
Later among the works it cites.
“A Better Way to Do Masked Language Model Scoring”
Carina Kauf and Anna Ivanova · 2023
Later among the works it cites.
“Holistic Evaluation of Language Models”
Percy Liang et al · 2023
Later among the works it cites.
“GPT-4 Technical Report”, 2023
OpenAI · 2023
Later among the works it cites.
“Modern language models refute Chomsky’s approach to language”
Steven Piantadosi · 2023
Later among the works it cites.
“Hypothesis Only Baselines in Natural Language Inference”
Adam Poliak et al · 2023
Later among the works it cites.
“The pragmatic profile of ChatGPT: assessing the pragmatic skills of a conversational agent”
Chiara Sandi, Federico Frau, Veronica Mangiaterra and Valentina Bambini · 2023
Later among the works it cites.
“Probing the psychology of AI models”
Richard Shiffrin and Melanie Mitchell · 2023
Later among the works it cites.
“What’s the Meaning of Superhuman Performance in Today’s NLU?”
Simone Tedeschi et al · 2023
Later among the works it cites.
“Llama 2: Open foundation and fine-tuned chat models”
Hugo Touvron et al · 2023
Later among the works it cites.
“Overinformative Question Answering by Humans and Machines”
Polina Tsvilodub, Michael Franke, Robert Hawkins and Noah. Goodman · 2023
Later among the works it cites.
“Tree of thoughts: Deliberate problem solving with large language models”
Shunyu Yao et al · 2023
Later among the works it cites.