Fetching the paper…
Reading the bibliography…
Recent research has revealed that neural language models at scale suffer from poor temporal generalization capability, i.e., the language model pre-trained on static data from past years performs worse over time on emerging data.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019 · 1909
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2019 · 1910
Earlier work this paper cites.
Basic properties of the generalized boltzmann-gibbs-shannon entropy
W Ochs. 1976 · 1976
Earlier work this paper cites.
Silhouettes: a graphical aid to the interpretation and validation of cluster analysis
Peter J Rousseeuw. 1987 · 1987
Earlier work this paper cites.
Divergence measures based on the shannon entropy
Jianhua Lin. 1991 · 1991
Earlier work this paper cites.
Ibd (isolation by distance): a program for analyses of isolation by distance
AJ Bohonak. 2002 · 2002
Earlier work this paper cites.
Using tf-idf to determine word relevance in document queries
Juan Ramos et al. 2003 · 2003
Earlier work this paper cites.
Analysing lexical semantic change with contextualised word representations
Mario Giulianelli, Marco Del Tredici, and Raquel Fernández. 2020 · 2004
Earlier work this paper cites.
Kea: Practical automated keyphrase extraction
Ian H Witten, Gordon W Paynter, Eibe Frank, Carl Gutwin, and Craig G Nevill-Manning. 2005 · 2005
Earlier work this paper cites.
Domain adaptation with structural correspondence learning
John Blitzer, Ryan McDonald, and Fernando Pereira. 2006 · 2006
Earlier work this paper cites.
k-means++: The advantages of careful seeding
Sergei Vassilvitskii and David Arthur. 2006 · 2006
Earlier work this paper cites.
Biographies, bollywood, boom-boxes and blenders: Domain adaptation for sentiment classification
John Blitzer, Mark Dredze, and Fernando Pereira. 2007 · 2007
Earlier work this paper cites.
Automatically identifying changes in the semantic orientation of words
Paul Cook and Suzanne Stevenson. 2010 · 2010
Earlier work this paper cites.
Stream-based translation models for statistical machine translation
Abby Levenberg, Chris Callison-Burch, and Miles Osborne. 2010 · 2010
Earlier work this paper cites.
Automatic keyword extraction from individual documents
Stuart Rose, Dave Engel, Nick Cramer, and Wendy Cowley. 2010 · 2010
Earlier work this paper cites.
From frequency to meaning: Vector space models of semantics
Peter D Turney and Patrick Pantel. 2010 · 2010
Earlier work this paper cites.
A distributional similarity approach to the detection of semantic change in the google books ngram corpus
Kristina Gulordava and Marco Baroni. 2011 · 2011
Cited alongside, same era.
Temporal analysis of language through neural language models
Yoon Kim, Yi-I Chiu, Kentaro Hanaki, Darshan Hegde, and Slav Petrov. 2014 · 2014
Cited alongside, same era.
Pasquale Nardone. 2014 · 2014
Cited alongside, same era.
A bottom up approach to category mapping and meaning change
Haim Dubossarsky, Yulia Tsvetkov, Chris Dyer, and Eitan Grossman. 2015 · 2015
Cited alongside, same era.
Statistically significant detection of linguistic change
Vivek Kulkarni, Rami Al-Rfou, Bryan Perozzi, and Steven Skiena. 2015 · 2015
Cited alongside, same era.
Diachronic word embeddings reveal statistical laws of semantic change
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Later among the works it cites.
Mining the uk web archive for semantic change detection
Adam Tsakalidis, Marya Bazzi, Mihai Cucuringu, Pierpaolo Basile, and Barbara McGillivray. 2019 · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. 2019 · 2019
Later among the works it cites.
Perl: Pivot-based domain adaptation for pre-trained deep contextualized embedding models
Eyal Ben-David, Carmel Rabinovitz, and Roi Reichart. 2020 · 2020
Later among the works it cites.
Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
William L Hamilton, Jure Leskovec, and Dan Jurafsky. 2016 · 2016
Cited alongside, same era.
Dynamic word embeddings
Robert Bamler and Stephan Mandt. 2017 · 2017
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2017 · 2017
Cited alongside, same era.
Exploiting the web for semantic change detection
Pierpaolo Basile and Barbara McGillivray. 2018 · 2018
Cited alongside, same era.
Yake! collection-independent automatic keyword extractor
Ricardo Campos, Vítor Mangaravite, Arian Pasquali, Alípio Mário Jorge, Célia Nunes, and Adam Jatowt. 2018 · 2018
Cited alongside, same era.
Short-term meaning shift: A distributional exploration
Marco Del Tredici, Raquel Fernández, and Gemma Boleda. 2018 · 2018
Cited alongside, same era.
Examining temporality in document classification
Xiaolei Huang and Michael J Paul. 2018 · 2018
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al. 2020 · 2020
Later among the works it cites.
Temporally-informed analysis of named entity recognition
Shruti Rijhwani and Daniel Preoţiuc-Pietro. 2020 · 2020
Later among the works it cites.
Dynamic language models for continuously evolving content
Spurthi Amba Hombaiah, Tao Chen, Mingyang Zhang, Michael Bendersky, and Marc Najork. 2021 · 2021
Later among the works it cites.
A data-driven approach to studying changing vocabularies in historical newspaper collections
Simon Hengchen, Ruben Ros, Jani Marjanen, and Mikko Tolonen. 2021 · 2021
Later among the works it cites.
Computational approaches to lexical semantic change: Visualization systems and novel applications
Adam Jatowta, Nina Tahmasebib, and Lars Borinb. 2021 · 2021
Later among the works it cites.
Lexical semantic change discovery
Sinan Kurtyigit, Maike Park, Dominik Schlechtweg, Jonas Kuhn, and Sabine Schulte im Walde. 2021 · 2021
Later among the works it cites.
Mind the gap: Assessing temporal generalization in neural language models
Angeliki Lazaridou, Adhi Kuncoro, Elena Gribovskaya, Devang Agrawal, Adam Liska, Tayfun Terzi, Mai Gimenez, Cyprien de Masson d’Autume, Tomas Kocisky, Sebastian Ruder, et al. 2021 · 2021
Later among the works it cites.
Dilbert: Customized pre-training for domain adaptation with category shift, with an application to aspect extraction
Entony Lekhtman, Yftah Ziser, and Roi Reichart. 2021 · 2021
Later among the works it cites.
Time waits for no one! analysis and challenges of temporal misalignment
Kelvin Luu, Daniel Khashabi, Suchin Gururangan, Karishma Mandyam, and Noah A Smith. 2021 · 2021
Later among the works it cites.
Temporal adaptation of bert and performance on downstream document classification: Insights from social media
Paul Röttger and Janet Pierrehumbert. 2021 · 2021
Later among the works it cites.
A commentary of gpt-3 in mit technology review 2021
Min Zhang and Juntao Li. 2021 · 2021
Later among the works it cites.
Unifying language learning paradigms
Yi Tay, Mostafa Dehghani, Vinh Q Tran, Xavier Garcia, Dara Bahri, Tal Schuster, Huaixiu Steven Zheng, Neil Houlsby, and Donald Metzler. 2022 · 2022
Closest in time.