Fetching the paper…
Reading the bibliography…
As question answering (QA) systems advance alongside the rapid evolution of foundation models, the need for robust, adaptable, and large-scale evaluation benchmarks becomes increasingly critical.
A faster approximation algorithm for the Steiner problem in graphs
Kurt Mehlhorn. 1988 · 1988
Earlier work this paper cites.
Pattern recognition and machine learning . Vol. 4
Christopher M Bishop and Nasser M Nasrabadi. 2006 · 2006
Earlier work this paper cites.
Freebase: a collaboratively created graph database for structuring human knowledge. In Proceedings of the 2008 ACM SIGMOD international conference on Management of data . 1247–1250
Kurt Bollacker, Colin Evans, Praveen Paritosh, Tim Sturge, and Jamie Taylor. 2008 · 2008
Earlier work this paper cites.
Semantic parsing on freebase from question-answer pairs. In Proceedings of the 2013 conference on empirical methods in natural language processing . 1533–1544
Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang. 2013 · 2013
Earlier work this paper cites.
Semantic parsing via paraphrasing. In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 1415–1425
Jonathan Berant and Percy Liang. 2014 · 2014
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. 2014 · 2014
Earlier work this paper cites.
Wikidata: a free collaborative knowledgebase
Denny Vrandečić and Markus Krötzsch. 2014 · 2014
Earlier work this paper cites.
Large-scale simple question answering with memory networks
Antoine Bordes, Nicolas Usunier, Sumit Chopra, and Jason Weston. 2015 · 2015
Earlier work this paper cites.
Dbpedia–a large-scale, multilingual knowledge base extracted from wikipedia
Jens Lehmann, Robert Isele, Max Jakob, Anja Jentzsch, Dimitris Kontokostas, Pablo N Mendes, Sebastian Hellmann, Mohamed Morsey, Patrick Van Kleef, Sören Auer, et al · 2015
Earlier work this paper cites.
Constraint-based question answering with knowledge graph. In Proceedings of COLING 2016, the 26th international conference on computational linguistics: technical papers . 2503–2514
Junwei Bao, Nan Duan, Zhao Yan, Ming Zhou, and Tiejun Zhao. 2016 · 2016
Earlier work this paper cites.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville. 2016 · 2016
Earlier work this paper cites.
Key-value memory networks for directly reading documents
Alexander Miller, Adam Fisch, Jesse Dodge, Amir-Hossein Karimi, Antoine Bordes, and Jason Weston. 2016 · 2016
Earlier work this paper cites.
From freebase to wikidata: The great migration. In Proceedings of the 25th international conference on world wide web . 1419–1428
Thomas Pellissier Tanon, Denny Vrandečić, Sebastian Schaffert, Thomas Steiner, and Lydia Pintscher. 2016 · 2016
Earlier work this paper cites.
The value of semantic parse labeling for knowledge base question answering. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) . 201–206
Wen-tau Yih, Matthew Richardson, Christopher Meek, Ming-Wei Chang, and Jina Suh. 2016 · 2016
Earlier work this paper cites.
Lc-quad: A corpus for complex question answering over knowledge graphs. In The Semantic Web–ISWC 2017: 16th International Semantic Web Conference, Vienna, Austria, October 21-25, 2017, Proceedings, Part II 16 . Springer, 210–218
Priyansh Trivedi, Gaurav Maheshwari, Mohnish Dubey, and Jens Lehmann. 2017 · 2017
Earlier work this paper cites.
9th challenge on question answering over linked data (QALD-9)
Ngonga Ngomo. 2018 · 2018
Earlier work this paper cites.
Open domain question answering using early fusion of knowledge bases and text
Haitian Sun, Bhuwan Dhingra, Manzil Zaheer, Kathryn Mazaitis, Ruslan Salakhutdinov, and William W Cohen. 2018 · 2018
Earlier work this paper cites.
The web as a knowledge-base for answering complex questions
Alon Talmor and Jonathan Berant. 2018 · 2018
Earlier work this paper cites.
Variational reasoning for question answering with knowledge graph. In Proceedings of the AAAI conference on artificial intelligence , Vol. 32
Yuyu Zhang, Hanjun Dai, Zornitsa Kozareva, Alexander Smola, and Le Song. 2018 · 2018
Earlier work this paper cites.
Lc-quad 2.0: A large dataset for complex question answering over wikidata and dbpedia. In The Semantic Web–ISWC 2019: 18th International Semantic Web Conference, Auckland, New Zealand, October 26–30, 2019, Proceedings, Part II 18 . Springer, 69–78
Mohnish Dubey, Debayan Banerjee, Abdelrahman Abdelkawi, and Jens Lehmann. 2019 · 2019
Earlier work this paper cites.
Mike Lewis. 2019 · 2019
Earlier work this paper cites.
Pullnet: Open domain question answering with iterative retrieval on knowledge bases and text
Haitian Sun, Tania Bedrax-Weiss, and William W Cohen. 2019 · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom B Brown. 2020 · 2020
Earlier work this paper cites.
Model patching: Closing the subgroup performance gap with data augmentation
Karan Goel, Albert Gu, Yixuan Li, and Christopher Ré. 2020 · 2020
Earlier work this paper cites.
Wiki-40b: Multilingual language model dataset. In Proceedings of the Twelfth Language Resources and Evaluation Conference . 2440–2452
Mandy Guo, Zihang Dai, Denny Vrandečić, and Rami Al-Rfou. 2020 · 2020
Earlier work this paper cites.
Query graph generation for answering multi-hop complex questions from knowledge bases. Association for Computational Linguistics
Yunshi Lan and Jing Jiang. 2020 · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al · 2020
Earlier work this paper cites.
Textattack: A framework for adversarial attacks, data augmentation, and adversarial training in nlp
John X Morris, Eli Lifland, Jin Yong Yoo, Jake Grigsby, Di Jin, and Yanjun Qi. 2020 · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Cited alongside, same era.
Improving multi-hop question answering over knowledge graphs using knowledge base embeddings. In Proceedings of the 58th annual meeting of the association for computational linguistics . 4498–4507
Apoorv Saxena, Aditay Tripathi, and Partha Talukdar. 2020 · 2020
Cited alongside, same era.
Case-based reasoning for natural language queries over knowledge bases
Rajarshi Das, Manzil Zaheer, Dung Thai, Ameya Godbole, Ethan Perez, Jay-Yoon Lee, Lizhen Tan, Lazaros Polymenakos, and Andrew McCallum. 2021 · 2021
Cited alongside, same era.
Rethinking benchmark and contamination for language models with rephrased samples
Shuo Yang, Wei-Lin Chiang, Lianmin Zheng, Joseph E Gonzalez, and Ion Stoica. 2023 · 2023
Later among the works it cites.
Siren’s Song in the AI Ocean: A Survey on Hallucination in Large Language Models
Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, et al · 2023
Later among the works it cites.
Don’t make your llm an evaluation benchmark cheater
Kun Zhou, Yutao Zhu, Zhipeng Chen, Wentong Chen, Wayne Xin Zhao, Xu Chen, Yankai Lin, Ji-Rong Wen, and Jiawei Han. 2023 · 2023
Later among the works it cites.
DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks. In International Conference on Learning Representations
Kaijie Zhu, Jiaao Chen, Jindong Wang, Neil Zhenqiang Gong, Diyi Yang, and Xing Xie. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Beyond IID: three levels of generalization for question answering on knowledge bases. In Proceedings of the Web Conference 2021 . 3477–3488
Yu Gu, Sue Kase, Michelle Vanni, Brian Sadler, Percy Liang, Xifeng Yan, and Yu Su. 2021 · 2021
Cited alongside, same era.
What disease does this patient have? a large-scale open domain question answering dataset from medical exams
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. 2021 · 2021
Cited alongside, same era.
Dynabench: Rethinking benchmarking in NLP
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, et al · 2021
Cited alongside, same era.
Dynaboard: An evaluation-as-a-service platform for holistic next-generation benchmarking
Zhiyi Ma, Kawin Ethayarajh, Tristan Thrush, Somya Jain, Ledell Wu, Robin Jia, Christopher Potts, Adina Williams, and Douwe Kiela. 2021 · 2021
Cited alongside, same era.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Cited alongside, same era.
Jinhao Jiang, Kun Zhou, Wayne Xin Zhao, and Ji-Rong Wen. 2022 · 2022
Cited alongside, same era.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022 · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Cited alongside, same era.
Yongshuo Zong, Tingyang Yu, Ruchika Chavhan, Bingchen Zhao, and Timothy Hospedales. 2023 · 2023
Later among the works it cites.
Scaling Laws for Predicting Downstream Performance in LLMs
Yangyi Chen, Binxuan Huang, Yifan Gao, Zhengyang Wang, Jingfeng Yang, and Heng Ji. 2024a · 2024
Later among the works it cites.
Enhancing complex question answering over knowledge graphs through evidence pattern retrieval. In Proceedings of the ACM on Web Conference 2024 . 2106–2115
Wentao Ding, Jinmao Li, Liangchuan Luo, and Yuzhong Qu. 2024 · 2024
Later among the works it cites.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Later among the works it cites.
Jiawei Gu, Xuhui Jiang, Zhichao Shi, Hexiang Tan, Xuehao Zhai, Chengjin Xu, Wei Li, Yinghan Shen, Shengjie Ma, Honghao Liu, et al · 2024
Later among the works it cites.
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al · 2024
Later among the works it cites.
The Amazon Nova family of models: Technical report and model card
Amazon Artificial General Intelligence. 2024 · 2024
Later among the works it cites.
A survey of generative search and recommendation in the era of large language models
Yongqi Li, Xinyu Lin, Wenjie Wang, Fuli Feng, Liang Pang, Wenjie Li, Liqiang Nie, Xiangnan He, and Tat-Seng Chua. 2024 · 2024
Later among the works it cites.
Reasoning on Graphs: Faithful and Interpretable Large Language Model Reasoning. In The Twelfth International Conference on Learning Representations
LINHAO LUO, Yuan-Fang Li, Reza Haf, and Shirui Pan. 2024 · 2024
Later among the works it cites.
AndroidWorld: A dynamic benchmarking environment for autonomous agents
Christopher Rawles, Sarah Clinckemaillie, Yifan Chang, Jonathan Waltz, Gabrielle Lau, Marybeth Fair, Alice Li, William Bishop, Wei Li, Folawiyo Campbell-Ajala, et al · 2024
Later among the works it cites.
Wisdom of the silicon crowd: Llm ensemble prediction capabilities match human crowd accuracy
Philipp Schoenegger, Indre Tuminauskaite, Peter S Park, and Philip E Tetlock. 2024 · 2024
Later among the works it cites.
Yago 4.5: A large and clean knowledge base with a rich taxonomy. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval . 131–140
Fabian M Suchanek, Mehwish Alam, Thomas Bonald, Lihu Chen, Pierre-Henri Paris, and Jules Soria. 2024 · 2024
Later among the works it cites.
Think-on-Graph: Deep and Responsible Reasoning of Large Language Model on Knowledge Graph. In The Twelfth International Conference on Learning Representations
Jiashuo Sun, Chengjin Xu, Lumingyuan Tang, Saizhuo Wang, Chen Lin, Yeyun Gong, Lionel Ni, Heung-Yeung Shum, and Jian Guo. 2024 · 2024
Later among the works it cites.
QALD-10–The 10th challenge on question answering over linked data: Shifting from DBpedia to Wikidata as a KG for KGQA
Ricardo Usbeck, Xi Yan, Aleksandr Perevalov, Longquan Jiang, Julius Schulz, Angelie Kraft, Cedric Möller, Junbo Huang, Jan Reineke, Axel-Cyrille Ngonga Ngomo, et al · 2024
Later among the works it cites.
Hallucination is Inevitable: An Innate Limitation of Large Language Models
Ziwei Xu, Sanjay Jain, and Mohan Kankanhalli. 2024 · 2024
Later among the works it cites.
t a u tau -bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Shunyu Yao, Noah Shinn, Pedram Razavi, and Karthik Narasimhan. 2024 · 2024
Later among the works it cites.
Darg: Dynamic evaluation of large language models via adaptive reasoning graph
Zhehao Zhang, Jiaao Chen, and Diyi Yang. 2024 · 2024
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2024
Later among the works it cites.
DyVal 2: Dynamic Evaluation of Large Language Models by Meta Probing Agents
Kaijie Zhu, Jindong Wang, Qinlin Zhao, Ruochen Xu, and Xing Xie. 2024 · 2024
Later among the works it cites.
The Claude 3 Model Family: Opus, Sonnet, Haiku
Anthropic. 2024 · 2025
Closest in time.
Command R: Details and Application
Cohere. 2024 · 2025
Closest in time.
Mistral Small 3
Mistral AI Team. 2025 · 2025
Closest in time.