Fetching the paper…
Reading the bibliography…
Spreadsheets are characterized by their extensive two-dimensional grids, flexible layouts, and varied formatting options, which pose significant challenges for large language models (LLMs).
Detecting and recognizing tables in spreadsheets
Iyad Abu Doush and Enrico Pontelli. 2010 · 2010
Earlier work this paper cites.
Automatic web spreadsheet data extraction
Zhe Chen and Michael Cafarella. 2013 · 2013
Earlier work this paper cites.
Integrating spreadsheet data via accurate and low-effort extraction
Zhe Chen and Michael Cafarella. 2014 · 2014
Earlier work this paper cites.
Cacheck: detecting and repairing cell arrays in spreadsheets
Wensheng Dou, Chang Xu, Shing-Chi Cheung, and Jun Wei. 2016 · 2016
Earlier work this paper cites.
Understanding the semantic structures of tables with a hybrid deep neural network architecture
Kyosuke Nishida, Kugatsu Sadamitsu, Ryuichiro Higashinaka, and Yoshihiro Matsuo. 2017 · 2017
Earlier work this paper cites.
Expandable group identification in spreadsheets
Wensheng Dou, Shi Han, Liang Xu, Dongmei Zhang, and Jun Wei. 2018 · 2018
Earlier work this paper cites.
Semantic structure extraction for spreadsheet tables with a multi-task learning architecture
Haoyu Dong, Shijie Liu, Zhouyu Fu, Shi Han, and Dongmei Zhang. 2019a · 2019
Earlier work this paper cites.
Tabular cell classification using pre-trained cell embeddings
Majid Ghasemi Gol, Jay Pujara, and Pedro Szekely. 2019 · 2019
Earlier work this paper cites.
Sherlock: A deep learning approach to semantic data type detection
Madelon Hulsebos, Kevin Hu, Michiel Bakker, Emanuel Zgraggen, Arvind Satyanarayan, Tim Kraska, Çagatay Demiralp, and César Hidalgo. 2019 · 2019
Earlier work this paper cites.
Uni-detect: A unified approach to automated error detection in tables
Pei Wang and Yeye He. 2019 · 2019
Earlier work this paper cites.
Spreadsheetcoder: Formula prediction from semi-structured context
Xinyun Chen, Petros Maniatis, Rishabh Singh, Charles Sutton, Hanjun Dai, Max Lin, and Denny Zhou. 2021 · 2021
Earlier work this paper cites.
TUTA: Tree-based transformers for generally structured table pre-training
Zhiruo Wang, Haoyu Dong, Ran Jia, Jia Li, Zhiyi Fu, Shi Han, and Dongmei Zhang. 2021 · 2021
Earlier work this paper cites.
Binding language models in symbolic languages
Zhoujun Cheng, Tianbao Xie, Peng Shi, Chengzu Li, Rahul Nadkarni, Yushi Hu, Caiming Xiong, Dragomir Radev, Mari Ostendorf, Luke Zettlemoyer, et al. 2022 · 2022
Cited alongside, same era.
Turl: Table understanding through representation learning
Xiang Deng, Huan Sun, Alyssa Lees, You Wu, and Cong Yu. 2022 · 2022
Cited alongside, same era.
Table Pretraining: A survey on model architectures, pretraining objectives, and downstream tasks
Haoyu Dong, Zhoujun Cheng, Xinyi He, Mengyu Zhou, Anda Zhou, Fan Zhou, Ao Liu, Shi Han, and Dongmei Zhang. 2022 · 2022
Cited alongside, same era.
Omnitab: Pretraining with natural and synthetic data for few-shot table-based question answering
Zhengbao Jiang, Yi Mao, Pengcheng He, Graham Neubig, and Weizhu Chen. 2022 · 2022
Cited alongside, same era.
Tapex: Table pre-training via learning a neural sql executor
Sibei Chen, Yeye He, Weiwei Cui, Ju Fan, Song Ge, Haidong Zhang, Dongmei Zhang, and Surajit Chaudhuri. 2024 · 2024
Closest in time.
Naihao Deng, Zhenjie Sun, Ruiqi He, Aman Sikka, Yulong Chen, Lin Ma, Yue Zhang, and Rada Mihalcea. 2024 · 2024
Closest in time.
Large language models for tabular data: Progresses and future directions
Haoyu Dong and Zhiruo Wang. 2024 · 2024
Closest in time.
Text2analysis: A benchmark of table question answering with advanced data analysis and unclear queries
Xinyi He, Mengyu Zhou, Xinrun Xu, Xiaojun Ma, Rui Ding, Lun Du, Yan Gao, Ran Jia, Xu Chen, Shi Han, et al. 2024 · 2024
Closest in time.
Flame: A small language model for spreadsheet formulas
Harshit Joshi, Abishai Ebenezer, José Cambronero Sanchez, Sumit Gulwani, Aditya Kanade, Vu Le, Ivan Radiček, and Gust Verbruggen. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Qian Liu, Bei Chen, Jiaqi Guo, Morteza Ziyadi, Zeqi Lin, Weizhu Chen, and Jian-Guang Lou. 2022 · 2022
Cited alongside, same era.
Mondrian: Spreadsheet layout detection
Gerardo Vitagliano, Lucas Reisener, Lan Jiang, Mazhar Hameed, and Felix Naumann. 2022 · 2022
Cited alongside, same era.
Yucheng Li. 2023 · 2023
Cited alongside, same era.
Yuan Sui, Mengyu Zhou, Mingjie Zhou, Shi Han, and Dongmei Zhang. 2023 · 2023
Cited alongside, same era.
Dbcopilot: Scaling natural language querying to massive databases
Tianshu Wang, Hongyu Lin, Xianpei Han, Le Sun, Xiaoyang Chen, Hao Wang, and Zhenyu Zeng. 2023 · 2023
Cited alongside, same era.
Retrieval meets long context large language models
Peng Xu, Wei Ping, Xianchao Wu, Lawrence McAfee, Chen Zhu, Zihan Liu, Sandeep Subramanian, Evelina Bakhturina, Mohammad Shoeybi, and Bryan Catanzaro. 2023 · 2023
Cited alongside, same era.
Tablellama: Towards open large generalist models for tables
Tianshu Zhang, Xiang Yue, Yifei Li, and Huan Sun. 2023 · 2023
Cited alongside, same era.
Chain-of-thought reasoning in tabular language models
Mingyu Zheng, Hao Yang, Wenbin Jiang, Zheng Lin, Yajuan Lyu, Qiaoqiao She, and Weiping Wang. 2023 · 2023
Cited alongside, same era.
Closest in time.
Auto-tables: Relationalize tables without using examples
Peng Li, Yeye He, Cong Yan, Yue Wang, and Surajit Chaudhuri. 2024 · 2024
Closest in time.
Lost in the middle: How language models use long contexts
Nelson F Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024 · 2024
Closest in time.
Llmlingua-2: Data distillation for efficient and faithful task-agnostic prompt compression
Zhuoshi Pan, Qianhui Wu, Huiqiang Jiang, Menglin Xia, Xufang Luo, Jue Zhang, Qingwei Lin, Victor Rühle, Yuqing Yang, Chin-Yew Lin, et al. 2024 · 2024
Closest in time.
Table meets llm: Can large language models understand structured table data? a benchmark and empirical study
Yuan Sui, Mengyu Zhou, Mingjie Zhou, Shi Han, and Dongmei Zhang. 2024 · 2024
Closest in time.
Vision language models for spreadsheet understanding: Challenges and opportunities
Shiyu Xia, Junyu Xiong, Haoyu Dong, Jianbo Zhao, Yuzhang Tian, Mengyu Zhou, Yeye He, Shi Han, and Dongmei Zhang. 2024 · 2024
Closest in time.
Tablellm: Enabling tabular data manipulation by llms in real office usage scenarios
Xiaokang Zhang, Jing Zhang, Zeyao Ma, Yang Li, Bohan Zhang, Guanlin Li, Zijun Yao, Kangli Xu, Jinchang Zhou, Daniel Zhang-Li, et al. 2024 · 2024
Closest in time.
Pytheas: pattern-based table discovery in csv files
Christina Christodoulakis, Eric B Munson, Moshe Gabel, Angela Demke Brown, and Renée J Miller. 2020 · 2089
Closest in time.