Fetching the paper…
Reading the bibliography…
As large language models (LLMs) grow larger and more sophisticated, assessing their "reasoning" capabilities in natural language grows more challenging.
Eli5: Long form question answering, 2019
Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli · 1907
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach, 2019
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 1907
Earlier work this paper cites.
Learning to retrieve reasoning paths over wikipedia graph for question answering, 2019
Akari Asai, Kazuma Hashimoto, Hannaneh Hajishirzi, Richard Socher, and Caiming Xiong · 1911
Earlier work this paper cites.
Recall and recognition of self-performed acts
Gilbert Mohr, Johannes Engelkamp, and Hubert D. Zimmer · 1989
Earlier work this paper cites.
Dense passage retrieval for open-domain question answering, 2020
Vladimir Karpukhin, Barlas Oğuz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih · 2004
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Leveraging passage retrieval with generative models for open domain question answering, 2020
Gautier Izacard and Edouard Grave · 2007
Earlier work this paper cites.
Causality
Judea Pearl · 2009
Earlier work this paper cites.
The probabilistic relevance framework: Bm25 and beyond
Stephen Robertson, Hugo Zaragoza, et al · 2009
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Vqa: Visual question answering, 2015
Aishwarya Agrawal, Jiasen Lu, Stanislaw Antol, Margaret Mitchell, C. Lawrence Zitnick, Dhruv Batra, and Devi Parikh · 2015
Earlier work this paper cites.
How question types reveal student thinking: An experimental comparison of multiple-true-false and free-response formats
Joanna K. Hubbard, Macy A. Potts, and Brian A. Couch · 2017
Earlier work this paper cites.
e-snli: Natural language inference with natural language explanations, 2018
Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom · 2018
Earlier work this paper cites.
A call for clarity in reporting BLEU scores
Matt Post · 2018
Earlier work this paper cites.
FEVER: a large-scale dataset for fact extraction and VERification
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal · 2018
Earlier work this paper cites.
Towards inference-oriented reading comprehension: Parallelqa
Soumya Wadhwa, Varsha Embar, Matthias Grabmair, and Eric Nyberg · 2018
Cited alongside, same era.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning · 2018
Cited alongside, same era.
Sentence mover’s similarity: Automatic evaluation for multi-sentence texts
Elizabeth Clark, Asli Celikyilmaz, and Noah A. Smith · 2019
Cited alongside, same era.
Natural questions: A benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Deberta: Decoding-enhanced bert with disentangled attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen · 2021
Later among the works it cites.
Tellmewhy: A dataset for answering why-questions in narratives
Yash Kumar Lal, Nathanael Chambers, Raymond Mooney, and Niranjan Balasubramanian · 2021
Later among the works it cites.
Pyserini: A python toolkit for reproducible information retrieval research with sparse and dense representations
Jimmy Lin, Xueguang Ma, Sheng-Chieh Lin, Jheng-Hong Yang, Ronak Pradeep, and Rodrigo Nogueira · 2021
Later among the works it cites.
Re-evaluating word mover’s distance, 2021
Ryoma Sato, Makoto Yamada, and Hisashi Kashima · 2021
Later among the works it cites.
News summarization and evaluation in the era of gpt-3, 2022
Tanya Goyal, Junyi Jessy Li, and Greg Durrett · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Explain yourself! leveraging language models for commonsense reasoning
Nazneen Fatema Rajani, Bryan McCann, Caiming Xiong, and Richard Socher · 2019
Cited alongside, same era.
CommonsenseQA: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
HybridQA: A dataset of multi-hop question answering over tabular and textual data
Wenhu Chen, Hanwen Zha, Zhiyu Chen, Wenhan Xiong, Hong Wang, and William Yang Wang · 2020
Cited alongside, same era.
R4c: A benchmark for evaluating rc systems to get the right answer for the right reason
Naoya Inoue, Pontus Stenetorp, and Kentaro Inui · 2020
Cited alongside, same era.
Learning to explain: Datasets and models for identifying valid reasoning chains in multihop question-answering
Harsh Jhamtani and Peter Clark · 2020
Cited alongside, same era.
Bleurt: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur P. Parikh · 2020
Cited alongside, same era.
Closest in time.
The abduction of sherlock holmes: A dataset for visual abductive reasoning
Jack Hessel, Jena D. Hwang, Jae Sung Park, Rowan Zellers, Chandra Bhagavatula, Anna Rohrbach, Kate Saenko, and Yejin Choi · 2022
Closest in time.
Learn to explain: Multimodal reasoning via thought chains for science question answering, 2022
Pan Lu, Swaroop Mishra, Tony Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan · 2022
Closest in time.
Entailment tree explanations via iterative retrieval-generation reasoner
Danilo Neves Ribeiro, Shen Wang, Xiaofei Ma, Rui Dong, Xiaokai Wei, Henry Zhu, Xinchi Chen, Zhiheng Huang, Peng Xu, Andrew O. Arnold, and Dan Roth · 2022
Closest in time.
Am i me or you? state-of-the-art dialogue models cannot maintain an identity
Kurt Shuster, Jack Urbanek, Arthur D. Szlam, and Jason Weston · 2022
Closest in time.
Causality-aware enhanced model for multi-hop question answering over knowledge graphs
Yuan Sui, Shanshan Feng, Huaxiang Zhang, Jian Cao, Liang Hu, and Nengjun Zhu · 2022
Closest in time.
What language model architecture and pretraining objective work best for zero-shot generalization?
Thomas Wang, Adam Roberts, Daniel Hesslow, Teven Le Scao, Hyung Won Chung, Iz Beltagy, Julien Launay, and Colin Raffel · 2022
Closest in time.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou · 2022
Closest in time.
Towards fine-grained causal reasoning and qa
Linyi Yang, Zhen Wang, Yuxiang Wu, Jie Yang, and Yue Zhang · 2022
Closest in time.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi · 2022
Closest in time.