Fetching the paper…
Reading the bibliography…
Existing language models (LMs) predict tokens with a softmax over a finite vocabulary, which can make it difficult to predict rare tokens or phrases.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Bert is not a knowledge base (yet): Factual knowledge vs. name-based reasoning in unsupervised qa
Nina Poerner, Ulli Waltinger, and Hinrich Schütze. 2019 · 1911
Earlier work this paper cites.
Nonparametric statistics
Sidney Siegel. 1957 · 1957
Earlier work this paper cites.
Example-based super-resolution
William T Freeman, Thouis R Jones, and Egon C Pasztor. 2002 · 2002
Earlier work this paper cites.
Improving machine translation quality with automatic named entity recognition
Bogdan Babych and Anthony F. Hartley. 2003 · 2003
Earlier work this paper cites.
Learning translations of named-entity phrases from parallel corpora
Robert C Moore. 2003 · 2003
Earlier work this paper cites.
Mining and summarizing customer reviews
Minqing Hu and Bing Liu. 2004 · 2004
Earlier work this paper cites.
A sentimental education: Sentiment analysis using subjectivity summarization based on minimum cuts
Bo Pang and Lillian Lee. 2004 · 2004
Earlier work this paper cites.
The pascal recognising textual entailment challenge
Ido Dagan, Oren Glickman, and Bernardo Magnini. 2005 · 2005
Earlier work this paper cites.
Hierarchical probabilistic neural network language model
Frederic Morin and Yoshua Bengio. 2005 · 2005
Earlier work this paper cites.
Improving named entity translation by exploiting comparable and parallel corpora
Ahmed Hassan, Haytham Fahmy, and Hany Hassan. 2007 · 2007
Earlier work this paper cites.
The probabilistic relevance framework: Bm25 and beyond
Stephen Robertson, Hugo Zaragoza, et al. 2009 · 2009
Earlier work this paper cites.
Product quantization for nearest neighbor search
Herve Jegou, Matthijs Douze, and Cordelia Schmid. 2010 · 2010
Earlier work this paper cites.
A memory efficient baseline for open domain question answering
Gautier Izacard, Fabio Petroni, Lucas Hosseini, Nicola De Cao, Sebastian Riedel, and Edouard Grave. 2020 · 2012
Earlier work this paper cites.
Fast search in hamming space with multi-index hashing
Mohammad Norouzi, Ali Punjani, and David J Fleet. 2012 · 2012
Earlier work this paper cites.
Nonparametric statistical methods
Myles Hollander, Douglas A Wolfe, and Eric Chicken. 2013 · 2013
Earlier work this paper cites.
Hidden factors and hidden topics: understanding rating dimensions with review text
Julian McAuley and Jure Leskovec. 2013 · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Optimized product quantization
Tiezheng Ge, Kaiming He, Qifa Ke, and Jian Sun. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015 · 2015
Earlier work this paper cites.
Strategies for training large vocabulary neural language models
Wenlin Chen, David Grangier, and Michael Auli. 2016 · 2016
Earlier work this paper cites.
Efficient softmax approximation for GPUs
Édouard Grave, Armand Joulin, Moustapha Cissé, David Grangier, and Hervé Jégou. 2017 · 2017
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017 · 2017
Earlier work this paper cites.
Cross-lingual entity alignment via joint attribute-preserving embedding
Zequn Sun, Wei Hu, and Chengkai Li. 2017 · 2017
Earlier work this paper cites.
T-rex: A large scale alignment of natural language with knowledge base triples
Hady Elsahar, Pavlos Vougiouklis, Arslen Remaci, Christophe Gravier, Jonathon Hare, Frederique Laforest, and Elena Simperl. 2018 · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. 2018 · 2018
Earlier work this paper cites.
Phrase-indexed question answering: A new challenge for scalable document comprehension
Minjoon Seo, Tom Kwiatkowski, Ankur P Parikh, Ali Farhadi, and Hannaneh Hajishirzi. 2018 · 2018
Earlier work this paper cites.
The impact of named entity translation for neural machine translation
Jinghui Yan, Jiajun Zhang, JinAn Xu, and Chengqing Zong. 2018 · 2018
Cited alongside, same era.
Breaking the softmax bottleneck: A high-rank rnn language model
Zhilin Yang, Zihang Dai, Ruslan Salakhutdinov, and William W Cohen. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Billion-scale similarity search with gpus
Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2019 · 2019
Cited alongside, same era.
Natural questions: a benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. 2019 · 2019
Cited alongside, same era.
Latent retrieval for weakly supervised open domain question answering
Surface form competition: Why the highest probability answer isn’t always right
Ari Holtzman, Peter West, Vered Schwartz, Yejin Choi, and Luke Zettlemoyer. 2021 · 2021
Later among the works it cites.
Learning dense representations of phrases at scale
Jinhyuk Lee, Mujeen Sung, Jaewoo Kang, and Danqi Chen. 2021 · 2021
Later among the works it cites.
Scaling language models: Methods, analysis & insights from training gopher
Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al. 2021 · 2021
Later among the works it cites.
Few-shot question answering by pretraining span selection
Ori Ram, Yuval Kirstain, Jonathan Berant, Amir Globerson, and Omer Levy. 2021 · 2021
Later among the works it cites.
mt5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kenton Lee, Ming-Wei Chang, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019 · 2019
Cited alongside, same era.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Real-time open-domain question answering with dense-sparse phrase index
Minjoon Seo, Jinhyuk Lee, Tom Kwiatkowski, Ankur P Parikh, Ali Farhadi, and Hannaneh Hajishirzi. 2019 · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Cited alongside, same era.
Tydi qa: A benchmark for information-seeking question answering in typologically diverse languages
Jonathan H. Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki. 2020 · 2020
Cited alongside, same era.
Efficient passage retrieval with hashing for open-domain question answering
Ikuya Yamada, Akari Asai, and Hannaneh Hajishirzi. 2021 · 2021
Later among the works it cites.
Adaptive semiparametric language models
Dani Yogatama, Cyprien de Masson d’Autume, and Lingpeng Kong. 2021 · 2021
Later among the works it cites.
Calibrate before use: Improving few-shot performance of language models
Tony Z Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021 · 2021
Later among the works it cites.
Factual probing is [mask]: Learning vs. learning to recall
Zexuan Zhong, Dan Friedman, and Danqi Chen. 2021 · 2021
Later among the works it cites.
Neuro-symbolic language modeling with automaton-augmented retrieval
Uri Alon, Frank Xu, Junxian He, Sudipta Sengupta, Dan Roth, and Graham Neubig. 2022 · 2022
Closest in time.
On the role of bidirectionality in language model pre-training
Mikel Artetxe, Jingfei Du, Naman Goyal, Luke Zettlemoyer, and Ves Stoyanov. 2022 · 2022
Closest in time.
Improving language models by retrieving from trillions of tokens
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, et al. 2022 · 2022
Closest in time.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022 · 2022
Closest in time.
Time-aware language models as temporal knowledge bases
Bhuwan Dhingra, Jeremy R Cole, Julian Martin Eisenschlos, Daniel Gillick, Jacob Eisenstein, and William W Cohen. 2022 · 2022
Closest in time.
Attributed text generation via post-hoc research and revision
Luyu Gao, Zhuyun Dai, Panupong Pasupat, Anthony Chen, Arun Tejasvi Chaganty, Yicheng Fan, Vincent Y Zhao, Ni Lao, Hongrae Lee, Da-Cheng Juan, et al. 2022 · 2022
Closest in time.
Few-shot learning with retrieval augmented language models
Gautier Izacard, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave. 2022 · 2022
Closest in time.
Temporalwiki: A lifelong benchmark for training and evaluating ever-evolving language models
Joel Jang, Seonghyeon Ye, Changho Lee, Sohee Yang, Joongbo Shin, Janghoon Han, Gyeonghun Kim, and Minjoon Seo. 2022 · 2022
Closest in time.
Kamel: Knowledge analysis with multitoken entities in language models
Jan-Christoph Kalo and Leandra Fichtel. 2022 · 2022
Closest in time.
Large language models struggle to learn long-tail knowledge
Nikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace, and Colin Raffel. 2022 · 2022
Closest in time.
Learning rich representation of keyphrases from text
Mayank Kulkarni, Debanjan Mahata, Ravneet Arora, and Rajarshi Bhowmik. 2022 · 2022
Closest in time.
Bidirectional language models are also few-shot learners
Ajay Patel, Bryan Li, Mohammad Sadegh Rasooli, Noah Constant, Colin Raffel, and Chris Callison-Burch. 2022 · 2022
Closest in time.
Bloom: A 176b-parameter open-access multilingual language model
Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2022 · 2022
Closest in time.
Peer: A collaborative language model
Timo Schick, Jane Dwivedi-Yu, Zhengbao Jiang, Fabio Petroni, Patrick Lewis, Gautier Izacard, Qingfei You, Christoforos Nalmpantis, Edouard Grave, and Sebastian Riedel. 2022 · 2022
Closest in time.
Nearest neighbor zero-shot inference
Weijia Shi, Julian Michael, Suchin Gururangan, and Luke Zettlemoyer. 2022 · 2022
Closest in time.
Capturing structural locality in non-parametric language models
Frank F. Xu, Junxian He, Graham Neubig, and Vincent Josua Hellendoorn. 2022 · 2022
Closest in time.
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022 · 2022
Closest in time.
Copy is all you need
Tian Lan, Deng Cai, Yan Wang, Heyan Huang, and Xian-Ling Mao. 2023 · 2023
Closest in time.