Fetching the paper…
Reading the bibliography…
With the advent of large language models (LLMs), the trend in NLP has been to train LLMs on vast amounts of data to solve diverse language understanding and generation tasks.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020) · 1901
Earlier work this paper cites.
Negated and misprimed probes for pretrained language models: Birds can talk, but cannot fly
Kassner, N. and Schütze, H. (2019) · 1911
Earlier work this paper cites.
The semantic conception of truth: and the foundations of semantics
Tarski, A. (1944) · 1944
Earlier work this paper cites.
The concept of truth in formalized languages
Tarski, A. (1956) · 1956
Earlier work this paper cites.
Conjectures and refutations: The growth of scientific knowledge
Popper, K. (1963) · 1963
Earlier work this paper cites.
Automatic methods of inductive inference
Plotkin, G. (1972) · 1972
Earlier work this paper cites.
Model theory
Chang, C. C. and Keisler, H. J. (1973) · 1973
Earlier work this paper cites.
Formal Philosophy
Montague, R. (1974) · 1974
Earlier work this paper cites.
On the relation between direct and continuation semantics
Reynolds, J. C. (1974) · 1974
Earlier work this paper cites.
Sometime is sometimes not never: On the temporal logic of programs
Lamport, L. (1980) · 1980
Earlier work this paper cites.
Generalized quantifiers in natural language
Barwise, J. and Cooper, R. (1981) · 1981
Earlier work this paper cites.
Introduction to Montague semantics
Dowty, D. R., Wall, R., and Peters, S. (1981) · 1981
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Hornik, K., Stinchcombe, M., and White, H. (1989) · 1989
Earlier work this paper cites.
Efficient induction of logic programs
Muggleton, S., Feng, C., et al. (1992) · 1992
Earlier work this paper cites.
On the computational power of neural nets
Siegelmann, H. T. and Sontag, E. D. (1992) · 1992
Earlier work this paper cites.
Reference to Abstract Objects in Discourse
Asher, N. (1993) · 1993
Earlier work this paper cites.
From Discourse to Logic: Introduction to Modeltheoretic Semantics of Natural Language, Formal Logic and Discourse Representation Theory
Kamp, H. and Reyle, U. (1993) · 1993
Earlier work this paper cites.
Classical descriptive set theory
Kechris, A. (1995) · 1995
Earlier work this paper cites.
No free lunch theorems for search
Wolpert, D. H., Macready, W. G., et al. (1995) · 1995
Cited alongside, same era.
Neural network learning: Theoretical foundations
Anthony, M., Bartlett, P. L., Bartlett, P. L., et al. (1999) · 1999
Cited alongside, same era.
A finite-state approach to events in natural language semantics
Fernando, T. (2004) · 2004
Cited alongside, same era.
Towards a montagovian account of dynamics
De Groote, P. (2006) · 2006
Cited alongside, same era.
Sdrt and continuation semantics
Asher, N. and Pogodalla, S. (2011) · 2010
Cited alongside, same era.
Learnability, stability and uniform convergence
Shalev-Shwartz, S., Shamir, O., Srebro, N., and Sridharan, K. (2010) · 2010
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2019) · 2019
Later among the works it cites.
A subregular bound on the complexity of lexical quantifiers
Graf, T. (2019) · 2019
Later among the works it cites.
The meaning of “most” for visual question answering models
Kuhnle, A. and Copestake, A. (2019) · 2019
Later among the works it cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. (2019) · 2019
Later among the works it cites.
CoQA: A Conversational Question Answering Challenge
Reddy, S., Chen, D., and Manning, C. D. (2019) · 2019
Later among the works it cites.
Learnability and semantic universals
Steinert-Threlkeld, S. and Szymanik, J. (2019) · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Siegelmann, H. T. (2012) · 2012
Cited alongside, same era.
On learnability, complexity and stability
Villa, S., Rosasco, L., and Poggio, T. (2013) · 2013
Cited alongside, same era.
Message exchange games in strategic conversations
Asher, N., Paul, S., and Venant, A. (2017) · 2017
Cited alongside, same era.
Generalization in deep learning
Kawaguchi, K., Kaelbling, L. P., and Bengio, Y. (2017) · 2017
Cited alongside, same era.
Overcoming catastrophic forgetting in neural networks
Kirkpatrick, J., Pascanu, R., Rabinowitz, N., Veness, J., Desjardins, G., Rusu, A. A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., et al. (2017) · 2017
Cited alongside, same era.
Exploring generalization in deep learning
Neyshabur, B., Bhojanapalli, S., McAllester, D., and Srebro, N. (2017) · 2017
Cited alongside, same era.
Climbing towards NLU: On meaning, form, and understanding in the age of data
Bender, E. M. and Koller, A. (2020) · 2020
Later among the works it cites.
An analysis of natural language inference benchmarks through the lens of negation
Hossain, M. M., Kovatchev, V., Dutta, P., Kao, T., Wei, E., and Blanco, E. (2020) · 2020
Later among the works it cites.
Sinha, K., Parthasarathi, P., Pineau, J., and Williams, A. (2020) · 2020
Later among the works it cites.
Understanding by understanding not: Modeling negation in language models
Hosseini, A., Reddy, S., Bahdanau, D., Hjelm, R. D., Sordoni, A., and Courville, A. (2021) · 2021
Later among the works it cites.
Chaturvedi, A., Bhar, S., Saha, S., Garain, U., and Asher, N. (2022) · 2022
Later among the works it cites.
The difficulty of computing stable and accurate neural networks: On the barriers of deep learning and smale’s 18th problem
Colbrook, M. J., Antun, V., and Hansen, A. C. (2022) · 2022
Later among the works it cites.
Strings from neurons to language
Fernando, T. (2022) · 2022
Later among the works it cites.
Negation, coordination, and quantifiers in contextualized language models
Kalouli, A.-L., Sevastjanova, R., Beck, C., and Romero, M. (2022) · 2022
Later among the works it cites.
When and why vision-language models behave like bag-of-words models, and what to do about it?
Yuksekgonul, M., Bianchi, F., Kalluri, P., Jurafsky, D., and Zou, J. (2022) · 2022
Later among the works it cites.
Dissociating language and thought in large language models: a cognitive perspective
Mahowald, K., Ivanova, A. A., Blank, I. A., Kanwisher, N., Tenenbaum, J. B., and Fedorenko, E. (2023) · 2023
Closest in time.
Augmented language models: a survey
Mialon, G., Dessì, R., Lomeli, M., Nalmpantis, C., Pasunuru, R., Raileanu, R., Rozière, B., Schick, T., Dwivedi-Yu, J., Celikyilmaz, A., et al. (2023) · 2023
Closest in time.