Fetching the paper…
Reading the bibliography…
Pretraining Neural Language Models (NLMs) over a large corpus involves chunking the text into training examples, which are contiguous text segments of sizes processable by the neural architecture.
How small is a unit ball?
David J. Smith and Mavina K. Vamanamurthy · 1989
Earlier work this paper cites.
Numerical operator calculus in higher dimensions
Gregory Beylkin and Martin J Mohlenkamp · 2002
Earlier work this paper cites.
Multiresolution quantum chemistry in multiwavelet bases
Robert J Harrison, George I Fann, Takeshi Yanai, and Gregory Beylkin · 2003
Earlier work this paper cites.
On the efficient evaluation of coalescence integrals in population balance models
Wolfgang Hackbusch · 2006
Earlier work this paper cites.
Multivariate regression and machine learning with sums of separable functions
Gregory Beylkin, Jochen Garcke, and Martin J Mohlenkamp · 2009
Earlier work this paper cites.
Recognizing textual entailment: Rational, evaluation and approaches–erratum
Ido Dagan, Bill Dolan, Bernardo Magnini, and Dan Roth · 2010
Earlier work this paper cites.
Tensor spaces and numerical tensor calculus , volume 42
Wolfgang Hackbusch · 2012
Earlier work this paper cites.
The winograd schema challenge
Hector Levesque, Ernest Davis, and Leora Morgenstern · 2012
Earlier work this paper cites.
The approximate rank of a matrix and its algorithmic applications: approximate rank
Noga Alon, Troy Lee, Adi Shraibman, and Santosh Vempala · 2013
Earlier work this paper cites.
The volume of n-balls
Jake Gipple · 2014
Earlier work this paper cites.
Inductive bias of deep convolutional networks through pooling geometry
Nadav Cohen and Amnon Shashua · 2017
Earlier work this paper cites.
Analysis and design of convolutional networks via hierarchical tensor decompositions
Nadav Cohen, Or Sharir, Yoav Levine, Ronen Tamari, David Yakira, and Amnon Shashua · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel R Bowman · 2017
Earlier work this paper cites.
Senteval: An evaluation toolkit for universal sentence representations
Alexis Conneau and Douwe Kiela · 2018
Cited alongside, same era.
Benefits of depth for long-term memory of recurrent networks
Yoav Levine, Or Sharir, Alon Ziv, and Amnon Shashua · 2018
Cited alongside, same era.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman · 2018
Cited alongside, same era.
Learning to retrieve reasoning paths over wikipedia graph for question answering
Akari Asai, Kazuma Hashimoto, Hannaneh Hajishirzi, Richard Socher, and Caiming Xiong · 2019
Cited alongside, same era.
Billion-scale similarity search with gpus
Jeff Johnson, Matthijs Douze, and Hervé Jégou · 2019
Cited alongside, same era.
Realm: Retrieval-augmented language model pre-training
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang · 2020
Later among the works it cites.
Poly-encoders: Architectures and pre-training strategies for fast and accurate multi-sentence scoring
Samuel Humeau, Kurt Shuster, Marie-Anne Lachaux, and Jason Weston · 2020
Later among the works it cites.
Leveraging passage retrieval with generative models for open domain question answering
Gautier Izacard and Edouard Grave · 2020
Later among the works it cites.
The depth-to-width interplay in self-attention
Yoav Levine, Noam Wies, Or Sharir, Hofit Bata, and Amnon Shashua · 2020
Later among the works it cites.
Mike Lewis, Marjan Ghazvininejad, Gargi Ghosh, Armen Aghajanyan, Sida Wang, and Luke Zettlemoyer · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Generalization through memorization: Nearest neighbor language models
Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis · 2019
Cited alongside, same era.
Natural questions: a benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al · 2019
Cited alongside, same era.
Latent retrieval for weakly supervised open domain question answering
Kenton Lee, Ming-Wei Chang, and Kristina Toutanova · 2019
Cited alongside, same era.
Quantum entanglement in deep learning architectures
Yoav Levine, Or Sharir, Nadav Cohen, and Amnon Shashua · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Cited alongside, same era.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Later among the works it cites.
How much knowledge can you pack into the parameters of a language model?
Adam Roberts, Colin Raffel, and Noam Shazeer · 2020
Later among the works it cites.
It’s not just size that matters: Small language models are also few-shot learners
Timo Schick and Hinrich Schütze · 2020
Later among the works it cites.
Deep autoregressive models for the efficient variational simulation of many-body quantum systems
Or Sharir, Yoav Levine, Noam Wies, Giuseppe Carleo, and Amnon Shashua · 2020
Later among the works it cites.
Long range arena: A benchmark for efficient transformers
Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler · 2020
Later among the works it cites.
Nandan Thakur, Nils Reimers, Johannes Daxenberger, and Iryna Gurevych · 2020
Later among the works it cites.
Improving language models by retrieving from trillions of tokens
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George van den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, et al · 2021
Closest in time.
Msa transformer
Roshan Rao, Jason Liu, Robert Verkuil, Joshua Meier, John F Canny, Pieter Abbeel, Tom Sercu, and Alexander Rives · 2021
Closest in time.
Scale efficiently: Insights from pre-training and fine-tuning transformers
Yi Tay, Mostafa Dehghani, Jinfeng Rao, William Fedus, Samira Abnar, Hyung Won Chung, Sharan Narang, Dani Yogatama, Ashish Vaswani, and Donald Metzler · 2021
Closest in time.
Which transformer architecture fits my data? a vocabulary bottleneck in self-attention
Noam Wies, Yoav Levine, Daniel Jannai, and Amnon Shashua · 2021
Closest in time.