Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning · 2018
Earlier work this paper cites.
Cognitive graph for multi-hop reading comprehension at scale
Original
Ming Ding, Chang Zhou, Qibin Chen, Hongxia Yang, and Jie Tang · 2019
Earlier work this paper cites.
Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs
Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner · 2019
Earlier work this paper cites.
Multi-hop reading comprehension through question decomposition and rescoring
Sewon Min, Victor Zhong, Luke Zettlemoyer, and Hannaneh Hajishirzi · 2019
Earlier work this paper cites.
Dynamically fused graph network for multi-hop reasoning
Original
Yunxuan Xiao, Yanru Qu, Lin Qiu, Hao Zhou, Lei Li, Weinan Zhang, and Yong Yu · 2019
Earlier work this paper cites.
Constructing a multi-hop QA dataset for comprehensive evaluation of reasoning steps
Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa · 2020
Earlier work this paper cites.
Models in the loop: Aiding crowdworkers with generative annotation assistants
Max Bartolo, Tristan Thrush, Sebastian Riedel, Pontus Stenetorp, Robin Jia, and Douwe Kiela · 2021
Earlier work this paper cites.
Improving language models by retrieving from trillions of tokens
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George van den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, Diego de Las Casas, Aurelia Guy, Jacob Menick, Roman Ring, T. W. Hennigan, Saffron Huang, Lorenzo Maggiore, Chris Jones, Albin Cassirer, Andy Brock, Michela Paganini, Geoffrey Irving, Oriol Vinyals, Simon Osindero, Karen Simonyan, Jack W. Rae, Erich Elsen, and L. Sifre · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Original
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman · 2021
Earlier work this paper cites.
Baleen: Robust multi-hop reasoning at scale via condensed retrieval
Original
O. Khattab, Christopher Potts, and Matei A. Zaharia · 2021
Earlier work this paper cites.
Do multi-hop question answering systems know how to answer the single-hop sub-questions?
Yixuan Tang, Hwee Tou Ng, and Anthony Tung · 2021
Earlier work this paper cites.
Musique: Multihop questions via single-hop question composition
H. Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal · 2021
Earlier work this paper cites.
Ask me anything: A simple strategy for prompting language models
Original
Simran Arora, Avanika Narayan, Mayee F Chen, Laurel J Orr, Neel Guha, Kush Bhatia, Ines Chami, Frederic Sala, and Christopher Ré · 2022
Earlier work this paper cites.
A survey on measuring and mitigating reasoning shortcuts in machine reading comprehension
Original
Xanh Ho, Johannes Mario Meissner, Saku Sugawara, and Akiko Aizawa · 2022
Earlier work this paper cites.