Fetching the paper…
Reading the bibliography…
Most existing multi-hop datasets are extractive answer datasets, where the answers to the questions can be extracted directly from the provided context.
The value of semantic parse labeling for knowledge base question answering
Wen-tau Yih, Matthew Richardson, Chris Meek, Ming-Wei Chang, and Jina Suh. 2016 · 2016
Earlier work this paper cites.
What makes reading comprehension questions easier?
Saku Sugawara, Kentaro Inui, Satoshi Sekine, and Akiko Aizawa. 2018 · 2018
Earlier work this paper cites.
The web as a knowledge-base for answering complex questions
Alon Talmor and Jonathan Berant. 2018 · 2018
Earlier work this paper cites.
Constructing datasets for multi-hop reading comprehension across documents
Johannes Welbl, Pontus Stenetorp, and Sebastian Riedel. 2018 · 2018
Earlier work this paper cites.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018 · 2018
Earlier work this paper cites.
Understanding dataset design choices for multi-hop reasoning
Jifan Chen and Greg Durrett. 2019 · 2019
Earlier work this paper cites.
Avoiding reasoning shortcuts: Adversarial evaluation, training, and model development for multi-hop QA
Yichen Jiang and Mohit Bansal. 2019 · 2019
Earlier work this paper cites.
Shortcut learning in deep neural networks
R. Geirhos, J.-H. Jacobsen, C. Michaelis, R. Zemel, W. Brendel, M. Bethge, and F. A. Wichmann. 2020 · 2020
Earlier work this paper cites.
Constructing a multi-hop QA dataset for comprehensive evaluation of reasoning steps
Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa. 2020 · 2020
Earlier work this paper cites.
R4C: A benchmark for evaluating RC systems to get the right answer for the right reason
Naoya Inoue, Pontus Stenetorp, and Kentaro Inui. 2020 · 2020
Earlier work this paper cites.
Is multihop QA in DiRe condition? measuring and reducing disconnected reasoning
Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2020 · 2020
Earlier work this paper cites.
Decomposing complex questions makes multi-hop QA easier and more interpretable
Ruiliu Fu, Han Wang, Xuejun Zhang, Jun Zhou, and Yonghong Yan. 2021 · 2021
Cited alongside, same era.
Do multi-hop question answering systems know how to answer the single-hop sub-questions?
Yixuan Tang, Hwee Tou Ng, and Anthony Tung. 2021 · 2021
Cited alongside, same era.
Successive prompting for decomposing complex questions
Dheeru Dua, Shivanshu Gupta, Sameer Singh, and Matt Gardner. 2022 · 2022
Cited alongside, same era.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang (Shane) Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022 · 2022
Cited alongside, same era.
Is a question decomposition unit all we need?
Pruthvi Patel, Swaroop Mishra, Mihir Parmar, and Chitta Baral. 2022 · 2022
Cited alongside, same era.
MuSiQue: Multihop questions via single-hop question composition
Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2022 · 2022
Measuring and narrowing the compositionality gap in language models
Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah Smith, and Mike Lewis. 2023 · 2023
Later among the works it cites.
Reasoning with language model prompting: A survey
Shuofei Qiao, Yixin Ou, Ningyu Zhang, Xiang Chen, Yunzhi Yao, Shumin Deng, Chuanqi Tan, Fei Huang, and Huajun Chen. 2023 · 2023
Later among the works it cites.
When do decompositions help for machine reading?
Kangda Wei, Dawn Lawrie, Benjamin Van Durme, Yunmo Chen, and Orion Weller. 2023 · 2023
Later among the works it cites.
MQuAKE: Assessing knowledge editing in language models via multi-hop questions
Zexuan Zhong, Zhengxuan Wu, Christopher Manning, Christopher Potts, and Danqi Chen. 2023 · 2023
Later among the works it cites.
Least-to-most prompting enables complex reasoning in large language models
Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc V Le, and Ed H. Chi. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou. 2022 · 2022
Cited alongside, same era.
A survey on measuring and mitigating reasoning shortcuts in machine reading comprehension
Xanh Ho, Johannes Mario Meissner, Saku Sugawara, and Akiko Aizawa. 2023 · 2023
Cited alongside, same era.
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023 · 2023
Cited alongside, same era.
Decomposed prompting: A modular approach for solving complex tasks
Tushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu, Kyle Richardson, Peter Clark, and Ashish Sabharwal. 2023 · 2023
Cited alongside, same era.
Compositional questions do not necessitate multi-hop reasoning
Sewon Min, Eric Wallace, Sameer Singh, Matt Gardner, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2019a
Cited in the paper.
Multi-hop reading comprehension through question decomposition and rescoring
Sewon Min, Victor Zhong, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2019b
Cited in the paper.
Llama 3 model card
AI@Meta. 2024 · 2024
Closest in time.
OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, and et al. 2024 · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, and et al. 2024 · 2024
Closest in time.
MRKE: The multi-hop reasoning evaluation of llms by knowledge edition
Jian Wu, Linyi Yang, Manabu Okumura, and Yue Zhang. 2024 · 2024
Closest in time.
FanOutQA: Multi-hop, multi-document question answering for large language models
Andrew Zhu, Alyssa Hwang, Liam Dugan, and Chris Callison-Burch. 2024 · 2024
Closest in time.