Fetching the paper…
Reading the bibliography…
Even the largest neural networks make errors, and once-correct predictions can become invalid as the world changes.
Roberta: A robustly optimized BERT pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 1907
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Reimers, N. and Gurevych, I · 1908
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., and Brew, J · 1910
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
A novel connectionist system for unconstrained handwriting recognition
Graves, A., Liwicki, M., Fernández, S., Bertolami, R., Bunke, H., and Schmidhuber, J · 2008
Earlier work this paper cites.
Modifying memories in transformer models, 2020
Zhu, C., Rawat, A. S., Zaheer, M., Bhojanapalli, S., Li, D., Yu, F., and Kumar, S · 2012
Earlier work this paper cites.
Graves, A., Wayne, G., and Danihelka, I · 2014
Earlier work this paper cites.
Siamese neural networks for one-shot image recognition
Koch, G., Zemel, R., Salakhutdinov, R., et al · 2015
Earlier work this paper cites.
Control of memory, active perception, and action in minecraft
Oh, J., Chockalingam, V., Lee, H., et al · 2016
Earlier work this paper cites.
Meta-learning with memory-augmented neural networks
Santoro, A., Bartunov, S., Botvinick, M., Wierstra, D., and Lillicrap, T · 2016
Earlier work this paper cites.
Matching networks for one shot learning
Vinyals, O., Blundell, C., Lillicrap, T., Wierstra, D., et al · 2016
Earlier work this paper cites.
Reading wikipedia to answer open-domain questions
Chen, D., Fisch, A., Weston, J., and Bordes, A · 2017
Earlier work this paper cites.
Zero-shot relation extraction via reading comprehension
Levy, O., Seo, M., Choi, E., and Zettlemoyer, L · 2017
Earlier work this paper cites.
Gradient episodic memory for continual learning
Lopez-Paz, D. and Ranzato, M. A · 2017
Earlier work this paper cites.
Neural episodic control
Pritzel, A., Uria, B., Srinivasan, S., Badia, A. P., Vinyals, O., Hassabis, D., Wierstra, D., and Blundell, C · 2017
Earlier work this paper cites.
Prototypical networks for few-shot learning
Snell, J., Swersky, K., and Zemel, R. S · 2017
Earlier work this paper cites.
Transforming question answering datasets into natural language inference datasets
Demszky, D., Guu, K., and Liang, P · 2018
Earlier work this paper cites.
FEVER: a large-scale dataset for fact extraction and VERification
Thorne, J., Vlachos, A., Christodoulopoulos, C., and Mittal, A · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Natural questions: a benchmark for question answering research
Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Kelcey, M., Devlin, J., Lee, K., Toutanova, K. N., Jones, L., Chang, M.-W., Dai, A., Uszkoreit, J., Le, Q., and Petrov, S · 2019
Cited alongside, same era.
Latent retrieval for weakly supervised open domain question answering
Lee, K., Chang, M.-W., and Toutanova, K · 2019
Cited alongside, same era.
Are red roses red? evaluating consistency of question-answering models
Ribeiro, M. T., Guestrin, C., and Singh, S · 2019
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Later among the works it cites.
How much knowledge can you pack into the parameters of a language model?, 2020
Roberts, A., Raffel, C., and Shazeer, N · 2020
Later among the works it cites.
Meta-neighborhoods
Shan, S., Li, Y., and Oliva, J. B · 2020
Later among the works it cites.
Editable neural networks
Sinitsin, A., Plokhotnyuk, V., Pyrkin, D., Popov, S., and Babenko, A · 2020
Later among the works it cites.
Are neural nets modular? Inspecting functional modularity through differentiable weight masks
Csordás, R., van Steenkiste, S., and Schmidhuber, J · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rolnick, D., Ahuja, A., Schwarz, J., Lillicrap, T., and Wayne, G · 2019
Cited alongside, same era.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 2019
Cited alongside, same era.
Correcting deep neural networks with small, generalizing patches
Sotoudeh, M. and Thakur, A. V · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
Dark experience for general continual learning: a strong, simple baseline
Buzzega, P., Boschini, M., Porrello, A., Abati, D., and Calderara, S · 2020
Cited alongside, same era.
More than a feeling: Benchmarks for sentiment analysis accuracy
Heitmann, M., Siebert, C., Hartmann, J., and Schamp, C · 2020
Cited alongside, same era.
Dense passage retrieval for open-domain question answering
Karpukhin, V., Oguz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., and Yih, W.-t · 2020
Cited alongside, same era.
Dai, D., Dong, L., Hao, Y., Sui, Z., and Wei, F · 2021
Later among the works it cites.
Editing factual knowledge in language models
De Cao, N., Aziz, W., and Titov, I · 2021
Later among the works it cites.
Hase, P., Diab, M., Celikyilmaz, A., Li, X., Kozareva, Z., Stoyanov, V., Bansal, M., and Iyer, S · 2021
Later among the works it cites.
BeliefBank: Adding memory to a pre-trained language model for a systematic notion of belief
Kassner, N., Tafjord, O., Schütze, H., and Clark, P · 2021
Later among the works it cites.
Mind the gap: Assessing temporal generalization in neural language models
Lazaridou, A., Kuncoro, A., Gribovskaya, E., Agrawal, D., Liska, A., Terzi, T., Gimenez, M., de Masson d’Autume, C., Ruder, S., Yogatama, D., Cao, K., Kociský, T., Young, S., and Blunsom, P · 2021
Later among the works it cites.
Mitchell, E., Lin, C., Bosselut, A., Finn, C., and Manning, C. D · 2021
Later among the works it cites.
Hindsight: Posterior-guided training of retrievers for improved open-ended generation
Paranjape, A., Khattab, O., Potts, C., Zaharia, M. A., and Manning, C. D · 2021
Later among the works it cites.
Recipes for building an open-domain chatbot
Roller, S., Dinan, E., Goyal, N., Ju, D., Williamson, M., Liu, Y., Xu, J., Ott, M., Smith, E. M., Boureau, Y.-L., and Weston, J · 2021
Later among the works it cites.
Colbertv2: Effective and efficient retrieval via lightweight late interaction
Santhanam, K., Khattab, O., Saad-Falcon, J., Potts, C., and Zaharia, M · 2021
Later among the works it cites.
Get your vitamin C! robust fact verification with contrastive evidence
Schuster, T., Fisch, A., and Barzilay, R · 2021
Later among the works it cites.
Locating and editing factual associations in GPT, 2022
Meng, K., Bau, D., Andonian, A., and Belinkov, Y · 2022
Closest in time.