Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) promise to automate data engineering on tabular data, offering enterprises a valuable opportunity to cut the high costs of manual data handling.
On schema matching with opaque column names and data values. In Proceedings of the 2003 ACM SIGMOD international conference on Management of data (SIGMOD ’03) . Association for Computing Machinery, New York, NY, USA, 205–216
Jaewoo Kang and Jeffrey F. Naughton. 2003 · 2003
Earlier work this paper cites.
Progressive Duplicate Detection
Thorsten Papenbrock, Arvid Heise, and Felix Naumann. 2015 · 2014
Earlier work this paper cites.
TabEL: Entity Linking in Web Tables. In The Semantic Web - ISWC 2015 . Springer International Publishing, Cham, 425–441
Chandra Sekhar Bhagavatula, Thanapon Noraset, and Doug Downey. 2015 · 2015
Earlier work this paper cites.
Magellan: toward building entity matching management systems over data science stacks
Pradap Konda, Sanjib Das, Paul Suganthan G. C., et al · 2016
Earlier work this paper cites.
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer. In International Conference on Learning Representations
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, et al · 2017
Earlier work this paper cites.
Deep learning for entity matching: A design space exploration. In Proceedings of the 2018 international conference on management of data . 19–34
Sidharth Mudgal, Han Li, Theodoros Rekatsinas, et al · 2018
Earlier work this paper cites.
Get Real: How Benchmarks Fail to Represent the Real World. In Proceedings of the Workshop on Testing Database Systems . ACM, Houston TX USA, 1–6
Adrian Vogelsgesang, Michael Haubenschild, Jan Finis, et al · 2018
Earlier work this paper cites.
Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018 , Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun’ichi Tsujii (Eds.). Association for Computational Linguistics, 3911–3921
Tao Yu, Rui Zhang, Kai Yang, et al · 2018
Earlier work this paper cites.
Sherlock: A Deep Learning Approach to Semantic Data Type Detection. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (KDD ’19) . Association for Computing Machinery, New York, NY, USA, 1500–1508
Madelon Hulsebos, Kevin Hu, Michiel Bakker, et al · 2019
Earlier work this paper cites.
JOSIE: Overlap Set Similarity Search for Finding Joinable Tables in Data Lakes. In Proceedings of the 2019 International Conference on Management of Data (SIGMOD ’19) . Association for Computing Machinery, New York, NY, USA, 847–864
Erkang Zhu, Dong Deng, Fatemeh Nargesian, and Renée J. Miller. 2019 · 2019
Earlier work this paper cites.
Language Models are Few-Shot Learners. In Advances in Neural Information Processing Systems , H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin (Eds.), Vol. 33. Curran Associates, Inc., 1877–1901
Tom Brown, Benjamin Mann, Nick Ryder, et al · 2020
Earlier work this paper cites.
TURL: table understanding through representation learning
Xiang Deng, Huan Sun, Alyssa Lees, You Wu, and Cong Yu. 2020 · 2020
Earlier work this paper cites.
TaPas: Weakly Supervised Table Parsing via Pre-training. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics . Association for Computational Linguistics, Online, 4320–4333
Jonathan Herzig, Pawel Krzysztof Nowak, Thomas Müller, Francesco Piccinno, and Julian Eisenschlos. 2020 · 2020
Earlier work this paper cites.
TaBERT: Pretraining for Joint Understanding of Textual and Tabular Data. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics . Association for Computational Linguistics, Online, 8413–8426
Pengcheng Yin, Graham Neubig, Wen-tau Yih, and Sebastian Riedel. 2020 · 2020
Earlier work this paper cites.
Sato: contextual semantic type detection in tables
Dan Zhang, Madelon Hulsebos, Yoshihiko Suhara, Çağatay Demiralp, Jinfeng Li, and Wang-Chiew Tan. 2020 · 2020
Earlier work this paper cites.
MATE: Multi-view Attention for Table Transformer Efficiency. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, Online and Punta Cana, Dominican Republic, 7606–7619
Julian Eisenschlos, Maharshi Gor, Thomas Müller, and William Cohen. 2021 · 2021
Earlier work this paper cites.
GitTables benchmark - column type detection
Madelon Hulsebos, Cağatay Demiralp, and Paul Demiralp. 2021 · 2021
Earlier work this paper cites.
Capturing Semantics for Imputation with Pre-trained Language Models. In 2021 IEEE 37th International Conference on Data Engineering (ICDE) . IEEE, Chania, Greece, 61–72
Yinan Mei, Shaoxu Song, Chenguang Fang, Haifeng Yang, Jingyun Fang, and Jiang Long. 2021 · 2021
Earlier work this paper cites.
How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) , Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli (Eds.). Association for Computational Linguistics, Online, 3118–3135
Phillip Rust, Jonas Pfeiffer, Ivan Vulić, Sebastian Ruder, and Iryna Gurevych. 2021 · 2021
Earlier work this paper cites.
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness. In Advances in Neural Information Processing Systems , S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Vol. 35. Curran Associates, Inc., 16344–16359
Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. 2022 · 2022
Earlier work this paper cites.
SOTAB: The WDC Schema.org Table Annotation Benchmark
Keti Korini, Ralph Peeters, and Christian Bizer. 2022 · 2022
Earlier work this paper cites.
Can Foundation Models Wrangle Your Data?
Avanika Narayan, Ines Chami, Laurel Orr, and Christopher Ré. 2022 · 2022
Earlier work this paper cites.
The Need for Tabular Representation Learning: An Industry Perspective
Alexandra Savelieva, Andreas Mueller, Avrilia Floratou, et al · 2022
Earlier work this paper cites.
Towards Foundation Models for Relational Databases [Vision Paper]. In NeurIPS 2022 First Table Representation Workshop
Liane Vogel, Benjamin Hilprecht, and Carsten Binnig. 2022 · 2022
Cited alongside, same era.
Finetuned Language Models are Zero-Shot Learners. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022 . OpenReview.net
Jason Wei, Maarten Bosma, Vincent Y. Zhao, et al · 2022
Cited alongside, same era.
TableFormer: Robust Transformer Modeling for Table-Text Encoding. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . Association for Computational Linguistics, Dublin, Ireland, 528–537
Jingfeng Yang, Aditya Gupta, Shyam Upadhyay, Luheng He, Rahul Goel, and Shachi Paul. 2022 · 2022
Cited alongside, same era.
Mathematical Capabilities of ChatGPT. In Advances in Neural Information Processing Systems , A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Eds.), Vol. 36. Curran Associates, Inc., 27699–27744
Simon Frieder, Luca Pinchetti, Chevalier Chevalier, et al · 2023
Lost in the Middle: How Language Models Use Long Contexts
Nelson F. Liu, Kevin Lin, John Hewitt, et al · 2024
Later among the works it cites.
Llama 3.1 Model Card
Meta. 2024 · 2024
Later among the works it cites.
OpenAI, Josh Achiam, Steven Adler, et al · 2024
Later among the works it cites.
Schema Matching with Large Language Models: an Experimental Study. In Proceedings of Workshops at the 50th International Conference on Very Large Data Bases, VLDB 2024, Guangzhou, China, August 26-30, 2024 . VLDB.org
Marcel Parciak, Brecht Vandevoort, Frank Neven, Liesbet M. Peeters, and Stijn Vansummeren. 2024 · 2024
Later among the works it cites.
WDC Products: A Multi-Dimensional Entity Matching Benchmark. In Proceedings 27th International Conference on Extending Database Technology, EDBT 2024, Paestum, Italy, March 25 - March 28 , Letizia Tanca, Qiong Luo, Giuseppe Polese, Loredana Caruccio, Xavier Oriol, and Donatella Firmani (Eds.). OpenProceedings.org, 22–33
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
GitTables: A Large-Scale Corpus of Relational Tables
Madelon Hulsebos, Cagatay Demiralp, and Paul Groth. 2023 · 2023
Cited alongside, same era.
SIGNAL – The SAP Signavio Analytics Query Language
Timotheus Kampik, Andre Lücke, Jörn Horstmann, Mark Wheeler, and David Eickhoff. 2023 · 2023
Cited alongside, same era.
Column Type Annotation using ChatGPT. In Joint Proceedings of Workshops at the 49th International Conference on Very Large Data Bases (VLDB 2023), Vancouver, Canada, August 28 - September 1, 2023 (CEUR Workshop Proceedings) , Vol. 3462. CEUR-WS.org
Keti Korini and Christian Bizer. 2023 · 2023
Cited alongside, same era.
SportsTables: A new Corpus for Semantic Type Detection
Sven Langenecker, Christoph Sturm, Christian Schalles, and Carsten Binnig. 2023 · 2023
Cited alongside, same era.
Can LLM Already Serve as A Database Interface? A BIg Bench for Large-Scale Database Grounded Text-to-SQLs. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023 , Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (Eds.)
Jinyang Li, Binyuan Hui, Ge Qu, et al · 2023
Cited alongside, same era.
Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2023 · 2023
Cited alongside, same era.
Using ChatGPT for Entity Matching. In New Trends in Database and Information Systems - ADBIS 2023 Short Papers, Doctoral Consortium and Workshops: AIDMA, DOING, K-Gals, MADEISD, PeRS, Barcelona, Spain, September 4-7, 2023, Proceedings (Communications in Computer and Information Science) , Vol. 1850. Springer, 221–230
Ralph Peeters and Christian Bizer. 2023 · 2023
Cited alongside, same era.
Synthetic Data Generation for Enterprise DBMS. In 2023 IEEE 39th International Conference on Data Engineering (ICDE) . 3585–3588
Anupam Sanghi and Jayant R. Haritsa. 2023 · 2023
Cited alongside, same era.
Ralph Peeters, Reng Chiz Der, and Christian Bizer. 2024 · 2024
Later among the works it cites.
Table Meets LLM: Can Large Language Models Understand Structured Table Data? A Benchmark and Empirical Study. In Proceedings of the 17th ACM International Conference on Web Search and Data Mining . ACM, Merida Mexico, 645–654
Yuan Sui, Mengyu Zhou, Mingjie Zhou, Shi Han, and Dongmei Zhang. 2024 · 2024
Later among the works it cites.
Generating Succinct Descriptions of Database Schemata for Cost-Efficient Prompting of Large Language Models
Immanuel Trummer. 2024 · 2024
Later among the works it cites.
CAESURA: Language Models as Multi-Modal Query Planners. In 14th Conference on Innovative Data Systems Research, CIDR 2024, Chaminade, HI, USA, January 14-17, 2024 . www.cidrdb.org
Matthias Urban and Carsten Binnig. 2024 · 2024
Later among the works it cites.
Automating the Enterprise with Foundation Models
Michael Wornow, Avanika Narayan, Krista Opsahl-Ong, Quinn McIntyre, Nigam Shah, and Christopher Ré. 2024 · 2024
Later among the works it cites.
Large Language Models as Data Preprocessors. In Proceedings of Workshops at the 50th International Conference on Very Large Data Bases, VLDB 2024, Guangzhou, China, August 26-30, 2024 . VLDB.org
Haochen Zhang, Yuyang Dong, Chuan Xiao, and Masafumi Oyamada. 2024a · 2024
Later among the works it cites.
Directions Towards Efficient and Automated Data Wrangling with Large Language Models. In 2024 IEEE 40th International Conference on Data Engineering Workshops (ICDEW) . 301–304
Zeyu Zhang, Paul Groth, Iacer Calixto, and Sebastian Schelter. 2024b · 2024
Later among the works it cites.
The Mighty ToRR: A Benchmark for Table Reasoning and Robustness
Shir Ashury-Tahan, Yifan Mai, Rajmohan C, et al · 2025
Closest in time.
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
DeepSeek-AI, Daya Guo, Dejian Yang, et al · 2025
Closest in time.
A Preview of XiYan-SQL: A Multi-Generator Ensemble Framework for Text-to-SQL
Yingqi Gao, Yifu Liu, Xiaoxia Li, et al · 2025
Closest in time.
TableLoRA: Low-rank Adaptation on Table Structure Understanding for Large Language Models. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025 , Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (Eds.). Association for Computational Linguistics, 22376–22391
Xinyi He, Yihao Liu, Mengyu Zhou, et al · 2025
Closest in time.
Mind the Data Gap: Bridging LLMs to Enterprise Data Integration. In 15th Annual Conference on Innovative Data Systems Research
Moe Kayali, Fabian Wenz, Nesime Tatbul, and Çağatay Demiralp. 2025 · 2025
Closest in time.
Magneto: Combining Small and Large Language Models for Schema Matching
Yurong Liu, Eduardo Peña, Aécio S. R. Santos, Eden Wu, and Juliana Freire. 2025a · 2025
Closest in time.
Learning to Reason with LLMs
OpenAI. 2024b · 2025
Closest in time.
QATCH: Automatic Evaluation of SQL-Centric Tasks on Proprietary Data
Simone Papicchio, Paolo Papotti, and Luca Cagliero. 2025 · 2025
Closest in time.
TabICL: A Tabular Foundation Model for In-Context Learning on Large Data
Jingang Qu, David Holzmüller, Gaël Varoquaux, and Marine Le Morvan. 2025 · 2025
Closest in time.
Match, Compare, or Select? An Investigation of Large Language Models for Entity Matching. In Proceedings of the 31st International Conference on Computational Linguistics , Owen Rambow, Leo Wanner, Marianna Apidianaki, Hend Al-Khalifa, Barbara Di Eugenio, and Steven Schockaert (Eds.). Association for Computational Linguistics, Abu Dhabi, UAE, 96–109
Tianshu Wang, Xiaoyang Chen, Hongyu Lin, et al · 2025
Closest in time.
Learning Relational Tabular Data without Shared Features
Zhaomin Wu, Shida Wang, Ziyang Wang, and Bingsheng He. 2025 · 2025
Closest in time.
Can language models automate data wrangling?
Gonzalo Jaimovitch-López, Cèsar Ferri, José Hernández-Orallo, Fernando Martínez-Plumed, and María José Ramírez-Quintana. 2022 · 2082
Closest in time.