Fetching the paper…
Reading the bibliography…
We enhance auto-regressive language models by conditioning on document chunks retrieved from a large corpus, based on local similarity with preceding tokens.
Latent Retrieval for Weakly Supervised Open Domain Question Answering
K. Lee, M.-W. Chang, and K. Toutanova · 1906
Earlier work this paper cites.
Accelerating large-scale inference with anisotropic vector quantization
R. Guo, P. Sun, E. Lindgren, Q. Geng, D. Simcha, F. Chern, and S. Kumar · 1908
Earlier work this paper cites.
Megatron-LM: Training multi-billion parameter language models using model parallelism
M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro · 1909
Earlier work this paper cites.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2001
Earlier work this paper cites.
Latent Dirichlet Allocation
D. M. Blei, A. Y. Ng, and M. I. Jordan · 2003
Earlier work this paper cites.
LDA-based document models for ad-hoc retrieval
X. Wei and W. B. Croft · 2006
Earlier work this paper cites.
Large Language models in machine translation
T. Brants, A. C. Popat, P. Xu, F. J. Och, and J. Dean · 2007
Earlier work this paper cites.
The probabilistic relevance framework: BM25 and beyond
S. Robertson and H. Zaragoza · 2009
Earlier work this paper cites.
Recurrent neural network based language model
T. Mikolov, M. Karafiát, L. Burget, J. Cernockỳ, and S. Khudanpur · 2010
Earlier work this paper cites.
Generating sequences with recurrent neural networks
A. Graves · 2013
Earlier work this paper cites.
Deep learning with differential privacy
M. Abadi, A. Chu, I. Goodfellow, H. B. McMahan, I. Mironov, K. Talwar, and L. Zhang · 2016
Earlier work this paper cites.
A neural knowledge language model
S. Ahn, H. Choi, T. Pärnamaa, and Y. Bengio · 2016
Earlier work this paper cites.
Exploring the limits of language modeling
R. Jozefowicz, O. Vinyals, M. Schuster, N. Shazeer, and Y. Wu · 2016
Earlier work this paper cites.
The LAMBADA dataset: Word prediction requiring a broad discourse context
D. Paperno, G. Kruszewski, A. Lazaridou, N. Q. Pham, R. Bernardi, S. Pezzelle, M. Baroni, G. Boleda, and R. Fernández · 2016
Earlier work this paper cites.
Improving neural language models with a continuous cache
E. Grave, A. Joulin, and N. Usunier · 2017
Earlier work this paper cites.
Pointer sentinel mixture models
S. Merity, C. Xiong, J. Bradbury, and R. Socher · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
J. Bradbury, R. Frostig, P. Hawkins, M. J. Johnson, C. Leary, D. Maclaurin, G. Necula, A. Paszke, J. V. der Plas, S. Wanderman-Milne, and Q. Zhang · 2018
Earlier work this paper cites.
Search engine guided neural machine translation
J. Gu, Y. Wang, K. Cho, and V. O. Li · 2018
Earlier work this paper cites.
T. Kudo and J. Richardson · 2018
Earlier work this paper cites.
Guiding neural machine translation with retrieved translation pieces
J. Zhang, M. Utiyama, E. Sumita, G. Neubig, and S. Nakamura · 2018
Cited alongside, same era.
Adaptive input representations for neural language modeling
A. Baevski and M. Auli · 2019
Cited alongside, same era.
Wikilinkgraphs: a complete, longitudinal and multi-language dataset of the wikipedia link networks
C. Consonni, D. Laniado, and A. Montresor · 2019
Cited alongside, same era.
Transformer-XL: Attentive language models beyond a fixed-length context
Z. Dai, Z. Yang, Y. Yang, J. Carbonell, Q. Le, and R. Salakhutdinov · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Cited alongside, same era.
Attention is not Explanation
S. Jain and B. C. Wallace · 2019
Lessons from archives: Strategies for collecting sociocultural data in machine learning
E. S. Jo and T. Gebru · 2020
Later among the works it cites.
Dense passage retrieval for open-domain question answering
V. Karpukhin, B. Oguz, S. Min, P. Lewis, L. Wu, S. Edunov, D. Chen, and W.-t. Yih · 2020
Later among the works it cites.
Generalization through memorization: Nearest neighbor language models
U. Khandelwal, O. Levy, D. Jurafsky, L. Zettlemoyer, and M. Lewis · 2020
Later among the works it cites.
Retrieval-augmented generation for knowledge-intensive NLP tasks
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel, S. Riedel, and D. Kiela · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Natural Questions: a benchmark for question answering research
T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, M. Kelcey, J. Devlin, K. Lee, K. N. Toutanova, L. Jones, M.-W. Chang, A. Dai, J. Uszkoreit, Q. Le, and S. Petrov · 2019
Cited alongside, same era.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever · 2019
Cited alongside, same era.
Energy and policy considerations for deep learning in NLP
E. Strubell, A. Ganesh, and A. McCallum · 2019
Cited alongside, same era.
Root mean square layer normalization
B. Zhang and R. Sennrich · 2019
Cited alongside, same era.
Interpretability and analysis in neural NLP
Y. Belinkov, S. Gehrmann, and E. Pavlick · 2020
Cited alongside, same era.
Zero: Memory optimizations toward training trillion parameter models
S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He · 2020
Later among the works it cites.
Green AI
R. Schwartz, J. Dodge, N. A. Smith, and O. Etzioni · 2020
Later among the works it cites.
On the dangers of stochastic parrots: Can language models be too big?
E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell · 2021
Closest in time.
Extracting training data from large language models
N. Carlini, F. Tramer, E. Wallace, M. Jagielski, A. Herbert-Voss, K. Lee, A. Roberts, T. Brown, D. Song, U. Erlingsson, A. Oprea, and C. Raffel · 2021
Closest in time.
Leveraging passage retrieval with generative models for open domain question answering
G. Izacard and E. Grave · 2021
Closest in time.
Internet-augmented dialogue generation
M. Komeili, K. Shuster, and J. Weston · 2021
Closest in time.
Pitfalls of static language modelling
A. Lazaridou, A. Kuncoro, E. Gribovskaya, D. Agrawal, A. Liska, T. Terzi, M. Gimenez, C. de Masson d’Autume, S. Ruder, D. Yogatama, K. Cao, T. Kociský, S. Young, and P. Blunsom · 2021
Closest in time.
Deduplicating training data makes language models better
K. Lee, D. Ippolito, A. Nystrom, C. Zhang, D. Eck, C. Callison-Burch, and N. Carlini · 2021
Closest in time.
Question and answer test-train overlap in open-domain question answering datasets
P. Lewis, P. Stenetorp, and S. Riedel · 2021
Closest in time.
Jurassic-1: Technical details and evaluation
O. Lieber, O. Sharir, B. Lenz, and Y. Shoham · 2021
Closest in time.
Scaling language models: Methods, analysis & insights from training Gopher
J. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, F. Song, J. Aslanides, S. Henderson, R. Ring, S. Young, E. Rutherford, T. Hennigan, J. Menick, A. Cassirer, R. Powell, G. van den Driessche, L. A. Hendricks, M. Rauh, P.-S. Huang, A. Glaese, J. Welbl, S. Dathathri, S. Huang, J. Uesato, J. Mellor, I. Higgins, A. Creswell, N. McAleese, A. Wu, E. Elsen, S. Jayakumar, E. Buchatskaya, D. Budden, E. Sutherland, K. Simonyan, M. Paganini, L. Sifre, L. Martens, X. L. Li, A. Kuncoro, A. Nematzadeh, E. Gribovskaya, D. Donato, A. Lazaridou, A. Mensch, J.-B. Lespiau, M. Tsimpoukelli, N. Grigorev, D. Fritz, T. Sottiaux, M. Pajarskas, T. Pohlen, Z. Gong, D. Toyama, C. de Masson d’Autume, Y. Li, T. Terzi, V. Mikulik, I. Babuschkin, A. Clark, D. de Las Casas, A. Guy, J. Bradbury, M. Johnson, B. Hechtman, L. Weidinger, I. Gabriel, W. Isaac, E. Lockhart, S. Osindero, L. Rimell, C. Dyer, O. Vinyals, K. Ayoub, J. Stanway, L. Bennett, D. Hassabis, K. Kavukcuoglu, and G. Irving · 2021
Closest in time.
End-to-end training of multi-document reader and retriever for open-domain question answering
D. S. Sachan, S. Reddy, W. Hamilton, C. Dyer, and D. Yogatama · 2021
Closest in time.
Retrieval augmentation reduces hallucination in conversation
K. Shuster, S. Poff, M. Chen, D. Kiela, and J. Weston · 2021
Closest in time.
Ethical and social risks of harm from language models
L. Weidinger, I. Gabriel, C. Griffin, M. Rauh, J. Uesato, J. Mellor, W. Isaac, P.-S. Huang, L. A. Hendricks, M. Cheng, B. Balle, J. Haas, C. Biles, L. Rimell, W. Hawkins, M. Glaese, A. Kasirzadeh, Z. Kenton, S. Brown, A. Birhane, T. Stepleton, G. Irving, and S. Legassick · 2021
Closest in time.
Adaptive semiparametric language models
D. Yogatama, C. de Masson d’Autume, and L. Kong · 2021
Closest in time.