Fetching the paper…
Reading the bibliography…
A characteristic feature of human semantic cognition is its ability to not only store and retrieve the properties of concepts observed through experience, but to also facilitate the inheritance of properties (can breathe) from superordinate concepts (animal) to their subordinates (dog) -- i.e.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
The child’s learning of english morphology
Jean Berko. 1958 · 1958
Earlier work this paper cites.
Word concepts: A theory and simulation of some basic semantic capabilities
M Ross Quillian. 1967 · 1967
Earlier work this paper cites.
Features of similarity
Amos Tversky. 1977 · 1977
Earlier work this paper cites.
Theories of semantic memory
Edward E Smith and William K Estes. 1978 · 1978
Earlier work this paper cites.
Feature-based induction
Steven A Sloman. 1993 · 1993
Earlier work this paper cites.
Verb Semantics and Lexical Selection
Zhibiao Wu and Martha Palmer. 1994 · 1994
Earlier work this paper cites.
WordNet: a lexical database for English
George A Miller. 1995 · 1995
Earlier work this paper cites.
Categorical inference is not a tree: The myth of inheritance hierarchies
Steven A Sloman. 1998 · 1998
Earlier work this paper cites.
Detecting blickets: How young children use information about novel causal powers in categorization and induction
Alison Gopnik and David M Sobel. 2000 · 2000
Earlier work this paper cites.
The Big Book of Concepts
Gregory L Murphy. 2002 · 2002
Earlier work this paper cites.
English gigaword
David Graff, Junbo Kong, Ke Chen, and Kazuaki Maeda. 2003 · 2003
Earlier work this paper cites.
Semantic relations and the lexicon: Antonymy, synonymy and other paradigms
M Lynne Murphy. 2003 · 2003
Earlier work this paper cites.
Semantic cognition: A parallel distributed processing approach
Timothy T Rogers and James L McClelland. 2004 · 2004
Earlier work this paper cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 2005
Earlier work this paper cites.
Clueweb09 data set
Jamie Callan, Mark Hoy, Changkuk Yoo, and Le Zhao. 2009 · 2009
Earlier work this paper cites.
Concepts and categories: Memory, meaning, and metaphysics
Lance J Rips, Edward E Smith, and Douglas L Medin. 2012 · 2012
Earlier work this paper cites.
Reporting bias and knowledge acquisition
Jonathan Gordon and Benjamin Van Durme. 2013 · 2013
Earlier work this paper cites.
The Centre for Speech, Language and the Brain (CSLB) concept property norms
Barry J Devereux, Lorraine K Tyler, Jeroen Geertzen, and Billi Randall. 2014 · 2014
Earlier work this paper cites.
Empowering Faculty: A Campus Cyberinfrastructure Strategy for Research Communities
Gerry McCartney, Thomas Hacker, and Baijian Yang. 2014 · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014 · 2014
Earlier work this paper cites.
How well do distributional models capture different types of semantic knowledge?
Dana Rubinstein, Effi Levi, Roy Schwartz, and Ari Rappoport. 2015 · 2015
Earlier work this paper cites.
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Earlier work this paper cites.
CC-News
Sebastian Nagel. 2016 · 2016
Earlier work this paper cites.
Are distributional representations ready for the real world? evaluating word vectors for grounded perceptual meaning
Li Lucy and Jon Gauthier. 2017 · 2017
Earlier work this paper cites.
Targeted syntactic evaluation of language models
Rebecca Marvin and Tal Linzen. 2018 · 2018
Earlier work this paper cites.
Firearms and tigers are dangerous, kitchen knives and zebras are not: Testing whether word embeddings can tell
Pia Sommerauer and Antske Fokkens. 2018 · 2018
Earlier work this paper cites.
A simple method for commonsense reasoning
Trieu H Trinh and Quoc V Le. 2018 · 2018
Earlier work this paper cites.
Cracking the contextual commonsense code: Understanding commonsense reasoning aptitude of deep contextual representations
Jeff Da and Jungo Kasai. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Do Neural Language Representations Learn Physical Commonsense?
Maxwell Forbes, Ari Holtzman, and Yejin Choi. 2019 · 2019
Cited alongside, same era.
Neural language models as psycholinguistic subjects: Representations of syntactic state
Richard Futrell, Ethan Wilcox, Takashi Morita, Peng Qian, Miguel Ballesteros, and Roger Levy. 2019 · 2019
Cited alongside, same era.
Openwebtext corpus
Aaron Gokaslan and Vanya Cohen. 2019 · 2019
Cited alongside, same era.
Designing and interpreting probes with control tasks
Surface form competition: Why the highest probability answer isn’t always right
Ari Holtzman, Peter West, Vered Shwartz, Yejin Choi, and Luke Zettlemoyer. 2021 · 2021
Later among the works it cites.
How can we know when language models know? on the calibration of language models for question answering
Zhengbao Jiang, Jun Araki, Haibo Ding, and Graham Neubig. 2021 · 2021
Later among the works it cites.
Word meaning in minds and machines
Brenden M Lake and Gregory L Murphy. 2021 · 2021
Later among the works it cites.
Do language models learn typicality judgments from text?
Kanishka Misra, Allyson Ettinger, and Julia Rayz. 2021 · 2021
Later among the works it cites.
Show your work: Scratchpads for intermediate computation with language models
Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, et al. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
John Hewitt and Percy Liang. 2019 · 2019
Cited alongside, same era.
Inherent disagreements in human textual inferences
Ellie Pavlick and Tom Kwiatkowski. 2019 · 2019
Cited alongside, same era.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020 · 2020
Cited alongside, same era.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al. 2020 · 2020
Cited alongside, same era.
Neural natural language inference models partially embed theories of lexical entailment and negation
Atticus Geiger, Kyle Richardson, and Christopher Potts. 2020 · 2020
Cited alongside, same era.
Sorting through the noise: Testing robustness of information processing in pre-trained language models
Lalchand Pandia and Allyson Ettinger. 2021 · 2021
Later among the works it cites.
Probing the probing paradigm: Does probing accuracy entail task relevance?
Abhilasha Ravichander, Yonatan Belinkov, and Eduard Hovy. 2021 · 2021
Later among the works it cites.
Relational World Knowledge Representation in Contextual Language Models: A Review
Tara Safavi and Danai Koutra. 2021 · 2021
Later among the works it cites.
The multiberts: Bert reproductions for robustness analysis
Thibault Sellam, Steve Yadlowsky, Jason Wei, Naomi Saphra, Alexander D’Amour, Tal Linzen, Jasmijn Bastings, Iulia Turc, Jacob Eisenstein, Dipanjan Das, et al. 2021 · 2021
Later among the works it cites.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Ben Wang and Aran Komatsuzaki. 2021 · 2021
Later among the works it cites.
When do you need billions of words of pretraining data?
Yian Zhang, Alex Warstadt, Xiaocheng Li, and Samuel R. Bowman. 2021 · 2021
Later among the works it cites.
Probing classifiers: Promises, shortcomings, and advances
Yonatan Belinkov. 2022 · 2022
Closest in time.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022 · 2022
Closest in time.
Language models show human-like content effects on reasoning
Ishita Dasgupta, Andrew K Lampinen, Stephanie CY Chan, Antonia Creswell, Dharshan Kumaran, James L McClelland, and Felix Hill. 2022 · 2022
Closest in time.
Najoung Kim, Tal Linzen, and Paul Smolensky. 2022 · 2022
Closest in time.
Can language models learn from explanations in context?
Andrew K Lampinen, Ishita Dasgupta, Stephanie CY Chan, Kory Matthewson, Michael Henry Tessler, Antonia Creswell, James L McClelland, Jane X Wang, and Felix Hill. 2022 · 2022
Closest in time.
Andrew Kyle Lampinen. 2022 · 2022
Closest in time.
minicons: Enabling flexible behavioral and representational analyses of transformer language models
Kanishka Misra. 2022 · 2022
Closest in time.
A property induction framework for neural language models
Kanishka Misra, Julia Taylor Rayz, and Allyson Ettinger. 2022 · 2022
Closest in time.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Closest in time.
Mapping language models to grounded conceptual spaces
Roma Patel and Ellie Pavlick. 2022 · 2022
Closest in time.
Meaning without reference in large language models
Steven T Piantadosi and Felix Hill. 2022 · 2022
Closest in time.
Language model acceptability judgements are not always robust to context
Koustuv Sinha, Jon Gauthier, Aaron Mueller, Kanishka Misra, Keren Fuentes, Roger Levy, and Adina Williams. 2022 · 2022
Closest in time.
Diagnosing Semantic Properties in Distributional Representations of Word Meaning
Pia Johanna Maria Sommerauer. 2022 · 2022
Closest in time.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, and Denny Zhou. 2022 · 2022
Closest in time.
Least-to-most prompting enables complex reasoning in large language models
Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Olivier Bousquet, Quoc Le, and Ed Chi. 2022 · 2022
Closest in time.
Large language models can be easily distracted by irrelevant context
Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed Chi, Nathanael Schärli, and Denny Zhou. 2023 · 2023
Closest in time.
Are language models worse than humans at following prompts? it’s complicated
Albert Webson, Alyssa Marie Loo, Qinan Yu, and Ellie Pavlick. 2023 · 2023
Closest in time.