Fetching the paper…
Reading the bibliography…
Our world is open-ended, non-stationary, and constantly evolving; thus what we talk about and how we talk about it change over time.
Catastrophic interference in connectionist networks: The sequential learning problem
Michael Mccloskey and Neil J. Cohen · 1989
Earlier work this paper cites.
A dynamic language model for speech recognition
F. Jelinek, B. Merialdo, S. Roukos, and M. Strauss · 1991
Earlier work this paper cites.
Lifelong robot learning
Sebastian Thrun and Tom M. Mitchell · 1995
Earlier work this paper cites.
Learning in the presence of concept drift and hidden contexts
Gerhard Widmer and Miroslav Kubat · 1996
Earlier work this paper cites.
Catastrophic forgetting in connectionist networks
Robert M. French · 1999
Earlier work this paper cites.
Latent dirichlet allocation
David M. Blei, Andrew Y. Ng, and Michael. I Jordan · 2003
Earlier work this paper cites.
Detecting change in data streams
Daniel Kifer, Shai Ben-David, and Johannes Gehrke · 2004
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandara Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al · 2005
Earlier work this paper cites.
Early drift detection method
Manuel Baena-Garcıa, José del Campo-Ávila, Raúl Fidalgo, Albert Bifet, R Gavalda, and R Morales-Bueno · 2006
Earlier work this paper cites.
Domain adaptation with structural correspondence learning
John Blitzer, Ryan McDonald, and Fernando Pereira · 2006
Earlier work this paper cites.
Frustratingly easy domain adaptation
Hal Daumé III · 2007
Earlier work this paper cites.
Question and answer test-train overlap in open-domain question answering datasets
Patrick Lewis, Pontus Stenetorp, and Sebastian Riedel · 2008
Earlier work this paper cites.
Continuous time dynamic topic models
Chong Wang, David Blei, and David Heckerman · 2008
Earlier work this paper cites.
Adaptive concept drift detection
Anton Dries and Ulrich Rückert · 2009
Earlier work this paper cites.
Streaming for large scale NLP: Language modeling
Amit Goyal, Hal Daumé III, and Suresh Venkatasubramanian · 2009
Earlier work this paper cites.
Stream-based translation models for statistical machine translation
Abby Levenberg, Chris Callison-Burch, and Miles Osborne · 2010
Earlier work this paper cites.
Recurrent neural network based language model
Tomas Mikolov, Martin Karafiát, Lukás Burget, Jan Cernocký, and Sanjeev Khudanpur · 2010
Earlier work this paper cites.
Domain adaptation via pseudo in-domain data selection
Amittai Axelrod, Xiaodong He, and Jianfeng Gao · 2011
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson · 2013
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Alex Graves · 2013
Earlier work this paper cites.
Crowdsourcing and annotating NER for Twitter #drift
Hege Fromreide, Dirk Hovy, and Anders Søgaard · 2014
Earlier work this paper cites.
Exponential reservoir sampling for streaming language models
Miles Osborne, Ashwin Lall, and Benjamin Van Durme · 2014
Earlier work this paper cites.
Dynamic language models for streaming text
Dani Yogatama, Chong Wang, Bryan R. Routledge, Noah A. Smith, and Eric P. Xing · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Deep learning for event-driven stock prediction
Xiao Ding, Yue Zhang, Ting Liu, and Junwen Duan · 2015
Earlier work this paper cites.
Never-ending learning
T. Mitchell, W. Cohen, E. Hruschka, P. Talukdar, J. Betteridge, A. Carlson, B. Dalvi, M. Gardner, B. Kisiel, J. Krishnamurthy, N. Lao, K. Mazaitis, T. Mohamed, N. Nakashole, E. Platanios, A. Ritter, M. Samadi, B. Settles, R. Wang, D. Wijaya, A. Gupta, X. Chen, A. Saparov, M. Greaves, and J. Welling · 2015
Earlier work this paper cites.
Pointer networks
Oriol Vinyals, Meire Fortunato, and Navdeep Jaitly · 2015
Earlier work this paper cites.
Diachronic Word Embeddings Reveal Statistical Laws of Semantic Change
William L Hamilton, Jure Leskovec, and Dan Jurafsky · 2016
Earlier work this paper cites.
The LAMBADA dataset: Word prediction requiring a broad discourse context
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Ngoc Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández · 2016
Earlier work this paper cites.
Andrei A. Rusu, Neil C. Rabinowitz, Guillaume Desjardins, Hubert Soyer, James Kirkpatrick, Koray Kavukcuoglu, Razvan Pascanu, and Raia Hadsell · 2016
Earlier work this paper cites.
Searchqa: A new q&a dataset augmented with context from a search engine
Matthew Dunn, Levent Sagun, Mike Higgins, V. U. Güney, Volkan Cirik, and Kyunghyun Cho · 2017
Cited alongside, same era.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A. Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell · 2017
Cited alongside, same era.
Gradient episodic memory for continual learning
David Lopez-Paz and Marc’Aurelio Ranzato · 2017
Cited alongside, same era.
Temporal Word Analogies: Identifying Lexical Replacement with Diachronic Word Embeddings
Terrence Szymanski · 2017
Cited alongside, same era.
NewsQA: A machine comprehension dataset
Adam Trischler, Tong Wang, Xingdi Yuan, Justin Harris, Alessandro Sordoni, Philip Bachman, and Kaheer Suleman · 2017
Cited alongside, same era.
Defending against neural fake news
Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi · 2019
Later among the works it cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E. Peters, and Arman Cohan · 2020
Later among the works it cites.
Back to the future - temporal adaptation of text representations
Johannes Bjerva, Wouter Kouw, and Isabelle Augenstein · 2020
Later among the works it cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Later among the works it cites.
How public opinion has moved on black lives matter
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Continuous adaptation via meta-learning in nonstationary and competitive environments
Maruan Al-Shedivat, Trapit Bansal, Yuri Burda, Ilya Sutskever, Igor Mordatch, and Pieter Abbeel · 2018
Cited alongside, same era.
Dynamic evaluation of neural sequence models
Ben Krause, Emmanuel Kahembwe, Iain Murray, and Steve Renals · 2018
Cited alongside, same era.
Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson · 2018
Cited alongside, same era.
Sentiment analysis under temporal shift
Jan Lukes and Anders Søgaard · 2018
Cited alongside, same era.
Deep neural models of semantic shift
Alex Rosenfeld and Katrin Erk · 2018
Cited alongside, same era.
Automated fact checking: Task formulations, methods and future directions
James Thorne and Andreas Vlachos · 2018
Cited alongside, same era.
Nate Cohn and Kevin Quealy · 2020
Later among the works it cites.
Conformer: Convolution-augmented transformer for speech recognition
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, and Ruoming Pang · 2020
Later among the works it cites.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith · 2020
Later among the works it cites.
Retrieval-augmented language model pre-training
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang · 2020
Later among the works it cites.
Embracing change: Continual learning in deep neural networks
Raia Hadsell, Dushyant Rao, Andrei A. Rusu, and Razvan Pascanu · 2020
Later among the works it cites.
Pretrained transformers improve out-of-distribution robustness
Dan Hendrycks, Xiaoyuan Liu, Eric Wallace, Adam Dziedzic, Rishabh Krishnan, and Dawn Song · 2020
Later among the works it cites.
Scaling laws for neural language models, 2020
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Later among the works it cites.
Dense passage retrieval for open-domain question answering
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih · 2020
Later among the works it cites.
Generalization through memorization: Nearest neighbor language models
Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis · 2020
Later among the works it cites.
Reformer: The efficient transformer
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya · 2020
Later among the works it cites.
Evaluating online continual learning with calm, 2020
Germán Kruszewski, Ionut-Teodor Sorodoc, and Tomas Mikolov · 2020
Later among the works it cites.
Adversarial NLI: A new benchmark for natural language understanding
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela · 2020
Later among the works it cites.
Dynasent: A dynamic benchmark for sentiment analysis
Christopher Potts, Zhengxuan Wu, Atticus Geiger, and Douwe Kiela · 2020
Later among the works it cites.
Temporally-informed analysis of named entity recognition
Shruti Rijhwani and Daniel Preotiuc-Pietro · 2020
Later among the works it cites.
Anton Sinitsin, Vsevolod Plokhotnyuk, Dmitriy Pyrkin, Sergei Popov, and Artem Babenko · 2020
Later among the works it cites.
LAMOL: LAnguage MOdeling for Lifelong Language Learning
Fan-Keng Sun, Cheng-Hao Ho, and Hung-Yi Lee · 2020
Later among the works it cites.
Modifying memories in transformer models
Chen Zhu, Ankit Singh Rawat, Manzil Zaheer, Srinadh Bhojanapalli, Daliang Li, Felix Yu, and Sanjiv Kumar · 2020
Later among the works it cites.
Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models, 2021
Elad Ben-Zaken, Shauli Ravfogel, and Yoav Goldberg · 2021
Closest in time.
On the dangers of stochastic parrots: Can language models be too big?
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell · 2021
Closest in time.
Editing factual knowledge in language models
Nicola De Cao, Wilker Aziz, and Ivan Titov · 2021
Closest in time.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2021
Closest in time.
The pile: An 800gb dataset of diverse text for language modeling, 2021
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy · 2021
Closest in time.
Carbon emissions and large neural network training
David A. Patterson, Joseph Gonzalez, Quoc V. Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David R. So, Maud Texier, and Jeff Dean · 2021
Closest in time.
We need to talk about random splits
Anders Søgaard, Sebastian Ebert, Jasmijn Bastings, and Katja Filippova · 2021
Closest in time.
Adaptive semiparametric language models
Dani Yogatama, Cyprien de Masson d’Autume, and Lingpeng Kong · 2021
Closest in time.