Fetching the paper…
Reading the bibliography…
Predicting future events based on news on the Web stands as one of the ultimate aspirations of artificial intelligence.
Verification of forecasts expressed in terms of probability
Glenn W Brier. 1950 · 1950
Earlier work this paper cites.
An introduction to causal inference
Judea Pearl. 2010 · 2010
Earlier work this paper cites.
The Integrated Crisis Early Warning System (ICEWS)
Philip A Schrodt, David J Gerner, Peter W Foltz, Moon-Soo Cho, and Young Joon Park. 2012 · 2012
Earlier work this paper cites.
GDELT: Global Data on Events, Location, and Tone, 1979-2012
Kale Leetaru and Philip A Schrodt. 2013 · 2013
Earlier work this paper cites.
What happens next? event prediction using a compositional neural network model. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 30
Granroth-Wilding. 2016 · 2016
Earlier work this paper cites.
Lsdsem 2017 shared task: The story cloze test. In Proceedings of the 2nd Workshop on Linking Models of Lexical, Sentential and Discourse-level Semantics . 46–51
Nasrin Mostafazadeh, Michael Roth, Annie Louis, Nathanael Chambers, and James Allen. 2017 · 2017
Earlier work this paper cites.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W Cohen, Ruslan Salakhutdinov, and Christopher D Manning. 2018 · 2018
Earlier work this paper cites.
Constructing a multi-hop QA dataset for comprehensive evaluation of reasoning steps
Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa. 2020 · 2020
Earlier work this paper cites.
2WikiMultiHopQA: A Dataset for Multi-Hop Question Answering on Wikipedia. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics . Association for Computational Linguistics, Online, 7380–7391
Chin-Yew Lin, Xi Victoria Lin, and Jimmy Lin. 2020 · 2020
Earlier work this paper cites.
StrategyQA: A Question Answering Benchmark Requiring Strategy and Planning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, Online, 5835–5847
Mor Geva, Daniel Khashabi, Tushar Khot, Ashish Sabharwal, and Dan Roth. 2021 · 2021
Earlier work this paper cites.
Measuring Massive Multitask Language Understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. 2021 · 2021
Earlier work this paper cites.
e-CARE: a New Dataset for Exploring Explainable Causal Reasoning. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 432–446
Li Du, Xiao Ding, Kai Xiong, Ting Liu, and Bing Qin. 2022 · 2022
Earlier work this paper cites.
Event-level prediction of urban crime reveals a signature of enforcement bias in US cities
Victor Rotaru, Yi Huang, Timmy Li, James Evans, and Ishanu Chattopadhyay. 2022 · 2022
Earlier work this paper cites.
ASQA: Factoid Questions Meet Long-Form Answers
Ivan Stelmakh, Yi Luan, Bhuwan Dhingra, and Ming-Wei Chang. 2022 · 2022
Cited alongside, same era.
A Causal Framework to Quantify the Robustness of Mathematical Reasoning with Language Models. In The 61st Annual Meeting Of The Association For Computational Linguistics
Alessandro Stolfo, Zhijing Jin, Kumar Shridhar, Bernhard Schoelkopf, and Mrinmaya Sachan. 2023 · 2023
Cited alongside, same era.
React: Synergizing reasoning and acting in language models. In International Conference on Learning Representations (ICLR)
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023 · 2023
Cited alongside, same era.
Distilling Script Knowledge from Large Language Models for Constrained Language Planning. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 4303–4325
Siyu Yuan, Jiangjie Chen, Ziquan Fu, Xuyang Ge, Soham Shah, Charles Jankowski, Yanghua Xiao, and Deqing Yang. 2023 · 2023
Cited alongside, same era.
Benchmarking Large Language Models in Retrieval-Augmented Generation. In Proceedings of the Thirty-Eighth AAAI Conference on Artificial Intelligence and Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence and Fourteenth Symposium on Educational Advances in Artificial Intelligence . AAAI Press, Washington, DC, USA, 17754–17762
Nianzu Liu, Tianyi Zhang, and Percy Liang. 2024 · 2024
Later among the works it cites.
WeQA: A Benchmark for Retrieval Augmented Generation in Wind Energy Domain
Rounak Meyur, Hung Phan, Sridevi Wagle, Jan Strube, Mahantesh Halappanavar, Sameera Horawalavithana, Anurag Acharya, and Sai Munikoti. 2024 · 2024
Later among the works it cites.
Can Language Models Use Forecasting Strategies?
Sarah Pratt, Seth Blumberg, Pietro Kreitlon Carolino, and Meredith Ringel Morris. 2024 · 2024
Later among the works it cites.
MTRAG: A Multi-Turn Conversational Benchmark for Evaluating Retrieval-Augmented Generation Systems
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
PubHealth: A Benchmark for Public Health Question Answering. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . Association for Computational Linguistics, Toronto, Canada, 9802–9822
Yuxuan Zhang, Zhiyuan Zhang, Yicheng Wang, Yuxuan Su, Yixuan Su, Yixuan Su, Yixuan Su, Yixuan Su, Yixuan Su, Yixuan Su, Yixuan Su, Yixuan Su, Yixuan Su, and Yixuan Su. 2023 · 2023
Cited alongside, same era.
VERA: Validation and Enhancement for Retrieval Augmented systems
Nitin Aravind Birur, Tanay Baswa, Divyanshu Kumar, Jatan Loya, Sahil Agarwal, and Prashanth Harshangi. 2024 · 2024
Cited alongside, same era.
Language Models as Causal Effect Generators
Lucius EJ Bynum and Kyunghyun Cho. 2024 · 2024
Cited alongside, same era.
Ragbench: Explainable benchmark for retrieval-augmented generation systems
Robert Friel, Masha Belyi, and Atindriyo Sanyal. 2024 · 2024
Cited alongside, same era.
OpenEP: Open-Ended Future Event Prediction
Yong Guan, Hao Peng, Xiaozhi Wang, Lei Hou, and Juanzi Li. 2024 · 2024
Cited alongside, same era.
Approaching Human-Level Forecasting with Language Models
Danny Halawi, Fred Zhang, Chen Yueh-Han, and Jacob Steinhardt. 2024 · 2024
Cited alongside, same era.
Reasoning and tools for human-level forecasting
Elvis Hsieh, Preston Fu, and Jonathan Chen. 2024 · 2024
Cited alongside, same era.
Forecastbench: A dynamic benchmark of ai forecasting capabilities
Ezra Karger, Houtan Bastani, Chen Yueh-Han, Zachary Jacobs, Danny Halawi, Fred Zhang, and Philip E Tetlock. 2024 · 2024
Cited alongside, same era.
Yixuan Tang and Yi Yang. 2024 · 2024
Later among the works it cites.
A comprehensive evaluation on event reasoning of large language models
Zhengwei Tao, Zhi Jin, Yifan Zhang, Xiancai Chen, Haiyan Zhao, Jia Li, Bing Liang, Chongyang Tao, Qun Liu, and Kam-Fai Wong. 2024 · 2024
Later among the works it cites.
CRAG: Corrective Retrieval-Augmented Generation for Robust Knowledge Grounding
Steven H. Wang, Antoine Scardigli, Leonard Tang, Wei Chen, Dimitry Levkin, Anya Chen, Spencer Ball, Thomas Woodside, Oliver Zhang, and Dan Hendrycks. 2024a · 2024
Later among the works it cites.
LegalBench-RAG: A Domain-Specific Benchmark for Evaluating Retrieval in Legal RAG Systems
Steven H. Wang, Antoine Scardigli, Leonard Tang, Wei Chen, Dimitry Levkin, Anya Chen, Spencer Ball, Thomas Woodside, Oliver Zhang, and Dan Hendrycks. 2024b · 2024
Later among the works it cites.
Exploring Large Language Models for Climate Forecasting
Yang Wang and Hassan A Karimi. 2024 · 2024
Later among the works it cites.
Siyun Zhao, Yuqing Yang, Zilong Wang, Zhiyuan He, Luna K Qiu, and Lili Qiu. 2024 · 2024
Later among the works it cites.
Chang Zong, Yuchen Yan, Weiming Lu, Jian Shao, Eliot Huang, Heng Chang, and Yueting Zhuang. 2024 · 2024
Later among the works it cites.
Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Mei Li, Mingfeng Xue, Pei Zhang, Qin Zhu, Rui Men, Runji Lin, Tianhao Li, Tianyi Tang, Tingyu Xia, Xingzhang Ren, Xuancheng Ren, Yang Fan, Yang Su, Yichang Zhang, Yu Wan, Yuqiong Liu, Zeyu Cui, Zhenru Zhang, and Zihan Qiu. 2025 · 2025
Closest in time.
FutureX: An Advanced Live Benchmark for LLM Agents in Future Prediction
Zhiyuan Zeng, Jiashuo Liu, Siyuan Chen, Tianci He, Yali Liao, Yixiao Tian, Jinpeng Wang, Zaiyuan Wang, Yang Yang, Lingyue Yin, et al · 2025
Closest in time.