Fetching the paper…
Reading the bibliography…
Despite its real-world significance, model performance on tabular data remains underexplored, leaving uncertainty about which model to rely on and which prompt configuration to adopt.
The problem of m rankings
Maurice G Kendall and B Babington Smith. 1939 · 1939
Earlier work this paper cites.
The control of the false discovery rate in multiple testing under dependency
Yoav Benjamini and Daniel Yekutieli. 2001 · 2001
Earlier work this paper cites.
Turl: Table understanding through representation learning
Xiang Deng, Huan Sun, Alyssa Lees, You Wu, and Cong Yu. 2020 · 2006
Earlier work this paper cites.
Calculating and synthesizing effect sizes
III Turner, Herbert M and Robert M Bernard. 2006 · 2006
Earlier work this paper cites.
Compositional semantic parsing on semi-structured tables
Panupong Pasupat and Percy Liang. 2015 · 2015
Earlier work this paper cites.
The harmonic mean p-value for combining dependent tests
Daniel J Wilson. 2019 · 2019
Earlier work this paper cites.
Tabfact: A large-scale dataset for table-based fact verification
Wenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang, Hong Wang, Shiyang Li, Xiyou Zhou, and William Yang Wang. 2020 · 2020
Earlier work this paper cites.
SciGen: a dataset for reasoning-aware text generation from scientific tables
Nafise Sadat Moosavi, Andreas Rücklé, Dan Roth, and Iryna Gurevych. 2021 · 2021
Earlier work this paper cites.
Towards table-to-text generation with numerical reasoning
Lya Hulliyyatus Suadaa, Hidetaka Kamigaito, Kotaro Funakoshi, Manabu Okumura, and Hiroya Takamura. 2021 · 2021
Earlier work this paper cites.
Large language models are few (1)-shot table reasoners
Wenhu Chen. 2022 · 2022
Earlier work this paper cites.
FinQA: A dataset of numerical reasoning over financial data
Zhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan Routledge, and William Yang Wang. 2022 · 2022
Earlier work this paper cites.
PASTA: Table-operations aware fact verification via sentence-table cloze pre-training
Zihui Gu, Ju Fan, Nan Tang, Preslav Nakov, Xiaoman Zhao, and Xiaoyong Du. 2022 · 2022
Cited alongside, same era.
Holistic evaluation of language models
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Manning, Christopher Ré, Diana Acosta-Navas, Drew A. Hudson, and 31 others. 2023 · 2023
Cited alongside, same era.
Rethinking tabular data understanding with large language models
Tianyang Liu, Fei Wang, and Muhao Chen. 2023 · 2023
Cited alongside, same era.
Tabular representation, noisy operators, and impacts on table structure understanding tasks in LLMs
Ananya Singha, José Cambronero, Sumit Gulwani, Vu Le, and Chris Parnin. 2023 · 2023
Cited alongside, same era.
When benchmarks are targets: Revealing the sensitivity of large language model leaderboards
Large language model for table processing: A survey
Weizheng Lu, Jiaming Zhang, Jing Zhang, and Yueguo Chen. 2024 · 2024
Later among the works it cites.
State of what art? a call for multi-prompt LLM evaluation
Moran Mizrahi, Guy Kaplan, Dan Malkin, Rotem Dror, Dafna Shahaf, and Gabriel Stanovsky. 2024 · 2024
Later among the works it cites.
Uncovering limitations of large language models in information seeking from tables
Chaoxu Pang, Yixuan Cao, Chunhao Yang, and Ping Luo. 2024 · 2024
Later among the works it cites.
Efficient benchmarking of language models
Yotam Perlitz, Elron Bandel, Ariel Gera, Ofir Arviv, Liat Ein-Dor, Eyal Shnarch, Noam Slonim, Michal Shmueli-Scheuer, and Leshem Choshen. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Norah Alzahrani, Hisham Abdullah Alyahya, Yazeed Alnumay, Sultan Alrashed, Shaykhah Alsubaie, Yusef Almushaykeh, Faisal Mirza, Nouf Alotaibi, Nora Altwairesh, Areeb Alowisheq, M Saiful Bari, and Haidar Khan. 2024 · 2024
Cited alongside, same era.
Unitxt: Flexible, shareable and reusable data preparation and evaluation for generative AI
Elron Bandel, Yotam Perlitz, Elad Venezian, Roni Friedman, Ofir Arviv, Matan Orbach, Shachar Don-Yehiya, Dafna Sheinwald, Ariel Gera, Leshem Choshen, Michal Shmueli-Scheuer, and Yoav Katz. 2024 · 2024
Cited alongside, same era.
On the robustness of language models for tabular question answering
Kushal Raj Bhandari, Sixue Xing, Soham Dan, and Jianxi Gao. 2024 · 2024
Cited alongside, same era.
Large language models on tabular data–a survey
Xi Fang, Weijie Xu, Fiona Anting Tan, Jiani Zhang, Ziqing Hu, Yanjun Qi, Scott Nickleach, Diego Socolinsky, Srinivasan Sengamedu, and Christos Faloutsos. 2024 · 2024
Cited alongside, same era.
Question answering over tabular data with databench: A large-scale empirical evaluation of LLMs
Jorge Osés Grijalba, L Alfonso Urena Lopez, Eugenio Martínez-Cámara, and Jose Camacho-Collados. 2024 · 2024
Cited alongside, same era.
From crowdsourced data to high-quality benchmarks: Arena-hard and benchbuilder pipeline
Tianle Li, Wei-Lin Chiang, Evan Frick, Lisa Dunlap, Tianhao Wu, Banghua Zhu, Joseph E. Gonzalez, and Ion Stoica. 2024 · 2024
Cited alongside, same era.
Qtsumm: Query-focused summarization over tabular data
Yilun Zhao, Zhenting Qi, Linyong Nan, Boyu Mi, Yixin Liu, Weijin Zou, Simeng Han, Ruizhe Chen, Xiangru Tang, Yumo Xu, Dragomir Radev, and Arman Cohan. 2023a
Cited in the paper.
Yilun Zhao, Haowei Zhang, Shengyun Si, Linyong Nan, Xiangru Tang, and Arman Cohan. 2023b
Cited in the paper.
Zipeng Qiu, You Peng, Guangxin He, Binhang Yuan, and Chen Wang. 2024 · 2024
Later among the works it cites.
Betterbench: Assessing ai benchmarks, uncovering issues, and establishing best practices
Anka Reuel, Amelia Hardy, Chandler Smith, Max Lamparth, Malcolm Hardy, and Mykel J. Kochenderfer. 2024 · 2024
Later among the works it cites.
Language modeling on tabular data: A survey of foundations, techniques and evolution
Yucheng Ruan, Xiang Lan, Jingying Ma, Yizhi Dong, Kai He, and Mengling Feng. 2024 · 2024
Later among the works it cites.
Table meets LLM: Can large language models understand structured table data? a benchmark and empirical study
Yuan Sui, Mengyu Zhou, Mingjie Zhou, Shi Han, and Dongmei Zhang. 2024 · 2024
Later among the works it cites.
Tablebench: A comprehensive and complex benchmark for table question answering
Xianjie Wu, Jian Yang, Linzheng Chai, Ge Zhang, Jiaheng Liu, Xinrun Du, Di Liang, Daixin Shu, Xianfu Cheng, Tianzhen Sun, Guanglin Niu, Tongliang Li, and Zhoujun Li. 2024 · 2024
Later among the works it cites.
Statistical multi-metric evaluation and visualization of llm system predictive performance
Samuel Ackerman, Eitan Farchi, Orna Raz, and Assaf Toledo. 2025 · 2025
Closest in time.