Fetching the paper…
Reading the bibliography…
Language model pre-training has been shown to capture a surprising amount of world knowledge, crucial for NLP tasks such as question answering.
A discrete hard em approach for weakly supervised question answering
Min, S., Chen, D., Hajishirzi, H., and Zettlemoyer, L · 1909
Earlier work this paper cites.
Knowledge guided text retrieval and reading for open domain question answering
Min, S., Chen, D., Zettlemoyer, L., and Hajishirzi, H · 1911
Earlier work this paper cites.
An analysis of the askmsr question-answering system
Brill, E., Dumais, S., and Banko, M · 2002
Earlier work this paper cites.
Introduction to the conll-2003 shared task: Language-independent named entity recognition
Sang, E. T. K. and De Meulder, F · 2003
Earlier work this paper cites.
The probabilistic relevance framework: Bm25 and beyond
Robertson, S., Zaragoza, H., et al · 2009
Earlier work this paper cites.
Maximum inner-product search using cone trees
Ram, P. and Gray, A. G · 2012
Earlier work this paper cites.
Semantic parsing on freebase from question-answer pairs
Berant, J., Chou, A., Frostig, R., and Liang, P · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2014
Earlier work this paper cites.
Graves, A., Wayne, G., and Danihelka, I · 2014
Earlier work this paper cites.
Asymmetric lsh (alsh) for sublinear time maximum inner product search (mips)
Shrivastava, A. and Li, P · 2014
Earlier work this paper cites.
Weston, J., Chopra, S., and Bordes, A · 2014
Earlier work this paper cites.
Semi-supervised sequence learning
Dai, A. M. and Le, Q. V · 2015
Earlier work this paper cites.
Skip-thought vectors
Kiros, R., Zhu, Y., Salakhutdinov, R. R., Zemel, R., Urtasun, R., Torralba, A., and Fidler, S · 2015
Earlier work this paper cites.
Learning binary codes for maximum inner product search
Shen, F., Liu, W., Zhang, S., Yang, Y., and Tao Shen, H · 2015
Earlier work this paper cites.
End-to-end memory networks
Sukhbaatar, S., Weston, J., Fergus, R., et al · 2015
Cited alongside, same era.
Learning recurrent span representations for extractive question answering
Lee, K., Salant, S., Kwiatkowski, T., Parikh, A., Das, D., and Berant, J · 2016
Cited alongside, same era.
Key-value memory networks for directly reading documents
Miller, A., Fisch, A., Dodge, J., Karimi, A.-H., Bordes, A., and Weston, J · 2016
Cited alongside, same era.
Squad: 100,000+ questions for machine comprehension of text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P · 2016
Cited alongside, same era.
Bidirectional attention flow for machine comprehension
Seo, M., Kembhavi, A., Farhadi, A., and Hajishirzi, H · 2016
Cited alongside, same era.
Learning to retrieve reasoning paths over wikipedia graph for question answering
Asai, A., Hashimoto, K., Hajishirzi, H., Socher, R., and Xiong, C · 2019
Later among the works it cites.
SpanBERT: Improving pre-training by representing and predicting spans
Joshi, M., Chen, D., Liu, Y., Weld, D. S., Zettlemoyer, L., and Levy, O · 2019
Later among the works it cites.
Generalization through memorization: Nearest neighbor language models
Khandelwal, U., Levy, O., Jurafsky, D., Zettlemoyer, L., and Lewis, M · 2019
Later among the works it cites.
Natural questions: a benchmark for question answering research
Kwiatkowski, T., Palomaki, J., Rhinehart, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Kelcey, M., Devlin, J., et al · 2019
Later among the works it cites.
Large memory layers with product keys
Lample, G., Sablayrolles, A., Ranzato, M., Denoyer, L., and Jégou, H · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chen, D., Fisch, A., Weston, J., and Bordes, A · 2017
Cited alongside, same era.
Simple and effective multi-paragraph reading comprehension
Clark, C. and Gardner, M · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Cited alongside, same era.
Generating sentences by editing prototypes
Guu, K., Hashimoto, T. B., Oren, Y., and Liang, P · 2018
Cited alongside, same era.
A retrieve-and-edit framework for predicting structured outputs
Hashimoto, T. B., Guu, K., Oren, Y., and Liang, P. S · 2018
Cited alongside, same era.
Deep contextualized word representations
Peters, M. E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., and Zettlemoyer, L · 2018
Cited alongside, same era.
Improving language understanding with unsupervised learning
Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I · 2018
Cited alongside, same era.
Later among the works it cites.
Latent retrieval for weakly supervised open domain question answering
Lee, K., Chang, M.-W., and Toutanova, K · 2019
Later among the works it cites.
Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., and Zettlemoyer, L · 2019
Later among the works it cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Later among the works it cites.
Knowledge enhanced contextual word representations, 2019
Peters, M. E., Neumann, M., IV, R. L. L., Schwartz, R., Joshi, V., Singh, S., and Smith, N. A · 2019
Later among the works it cites.
Language models as knowledge bases?
Petroni, F., Rocktäschel, T., Lewis, P., Bakhtin, A., Wu, Y., Miller, A. H., and Riedel, S · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2019
Later among the works it cites.
How much knowledge can you pack into the parameters of a language model?
Roberts, A., Raffel, C., and Shazeer, N · 2020
Closest in time.