Fetching the paper…
Reading the bibliography…
Language models, such as GPT-3.5 and ChatGPT, demonstrate remarkable abilities to follow diverse human instructions and perform a wide range of tasks.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Data cleaning: Problems and current approaches
Erhard Rahm, Hong Hai Do, et al · 2000
Earlier work this paper cites.
Generic schema matching with cupid. In vldb , Vol. 1. 49–58
Jayant Madhavan, Philip A Bernstein, and Erhard Rahm. 2001 · 2001
Earlier work this paper cites.
A survey of approaches to automatic schema matching
Erhard Rahm and Philip A Bernstein. 2001 · 2001
Earlier work this paper cites.
Uncovering the Relational Web.. In WebDB . Citeseer, 1–6
Michael J Cafarella, Alon Y Halevy, Yang Zhang, Daisy Zhe Wang, and Eugene Wu. 2008 · 2008
Earlier work this paper cites.
Harvesting relational tables from lists on the web
Hazem Elmeleegy, Jayant Madhavan, and Alon Halevy. 2009 · 2009
Earlier work this paper cites.
ERACER: a database approach for statistical inference and data cleaning. In Proceedings of the 2010 ACM SIGMOD International Conference on Management of data . 75–86
Chris Mayfield, Jennifer Neville, and Sunil Prabhakar. 2010 · 2010
Earlier work this paper cites.
Spreadsheet table transformations from examples
William R Harris and Sumit Gulwani. 2011 · 2011
Earlier work this paper cites.
Wrangler: Interactive visual specification of data transformation scripts. In Proceedings of the sigchi conference on human factors in computing systems . 3363–3372
Sean Kandel, Andreas Paepcke, Joseph Hellerstein, and Jeffrey Heer. 2011 · 2011
Earlier work this paper cites.
Tegra: Table extraction by global record alignment. In Proceedings of the 2015 ACM SIGMOD international conference on management of data . 1713–1728
Xu Chu, Yeye He, Kaushik Chakrabarti, and Kris Ganjam. 2015 · 2015
Earlier work this paper cites.
Compositional semantic parsing on semi-structured tables
Panupong Pasupat and Percy Liang. 2015 · 2015
Earlier work this paper cites.
Data services leveraging Bing’s data assets
Kaushik Chakrabarti, Surajit Chaudhuri, Zhimin Chen, Kris Ganjam, Yeye He, and W Redmond. 2016 · 2016
Earlier work this paper cites.
Data cleaning: Overview and emerging challenges. In Proceedings of the 2016 international conference on management of data . 2201–2206
Xu Chu, Ihab F Ilyas, Sanjay Krishnan, and Jiannan Wang. 2016 · 2016
Earlier work this paper cites.
Table cell search for question answering. In Proceedings of the 25th International Conference on World Wide Web . 771–782
Huan Sun, Hao Ma, Xiaodong He, Wen-tau Yih, Yu Su, and Xifeng Yan. 2016 · 2016
Earlier work this paper cites.
Multi-hypothesis CSV parsing. In Proceedings of the 29th International Conference on Scientific and Statistical Database Management . 1–12
Till Döhmen, Hannes Mühleisen, and Peter Boncz. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Sqlnet: Generating structured queries from natural language without reinforcement learning
Xiaojun Xu, Chang Liu, and Dawn Song. 2017 · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Transform-data-by-example (TDE) an extensible search engine for data transformations
Yeye He, Xu Chu, Kris Ganjam, Yudian Zheng, Vivek Narasayya, and Surajit Chaudhuri. 2018 · 2018
Earlier work this paper cites.
Deep learning for entity matching: A design space exploration. In Proceedings of the 2018 International Conference on Management of Data . 19–34
Sidharth Mudgal, Han Li, Theodoros Rekatsinas, AnHai Doan, Youngchoon Park, Ganesh Krishnan, Rohit Deep, Esteban Arcaute, and Vijay Raghavendra. 2018 · 2018
Earlier work this paper cites.
Synthesizing type-detection logic for rich semantic data types using open-source code. In Proceedings of the 2018 International Conference on Management of Data . 35–50
Cong Yan and Yeye He. 2018 · 2018
Earlier work this paper cites.
Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, Brussels, Belgium
Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, Zilin Zhang, and Dragomir Radev. 2018 · 2018
Cited alongside, same era.
DataWig: Missing Value Imputation for Tables
Felix Biessmann, Tammo Rukat, Philipp Schmidt, Prathik Naidu, Sebastian Schelter, Andrey Taptunov, Dustin Lange, and David Salinas. 2019 · 2019
Cited alongside, same era.
Tabfact: A large-scale dataset for table-based fact verification
Wenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang, Hong Wang, Shiyang Li, Xiyou Zhou, and William Yang Wang. 2019 · 2019
Cited alongside, same era.
Sherlock: A deep learning approach to semantic data type detection. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining . 1500–1508
Madelon Hulsebos, Kevin Hu, Michiel Bakker, Emanuel Zgraggen, Arvind Satyanarayan, Tim Kraska, Çagatay Demiralp, and César Hidalgo. 2019 · 2019
Can foundation models wrangle your data?
Avanika Narayan, Ines Chami, Laurel Orr, Simran Arora, and Christopher Ré. 2022 · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
Annotating columns with pre-trained language models. In Proceedings of the 2022 International Conference on Management of Data . 1493–1503
Yoshihiko Suhara, Jinfeng Li, Yuliang Li, Dan Zhang, Çağatay Demiralp, Chen Chen, and Wang-Chiew Tan. 2022 · 2022
Later among the works it cites.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2022c · 2022
Later among the works it cites.
Self-instruct: Aligning language model with self generated instructions
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Cited alongside, same era.
Wrangling messy CSV files by detecting row and type patterns
Gerrit JJ van den Burg, Alfredo Nazábal, and Charles Sutton. 2019 · 2019
Cited alongside, same era.
Auto-em: End-to-end fuzzy entity-matching using pre-trained deep models and transfer learning. In The World Wide Web Conference . 2413–2424
Chen Zhao and Yeye He. 2019 · 2019
Cited alongside, same era.
Making pre-trained language models better few-shot learners
Tianyu Gao, Adam Fisch, and Danqi Chen. 2020 · 2020
Cited alongside, same era.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A Smith. 2020 · 2020
Cited alongside, same era.
Auto-transform: learning-to-transform by patterns
Zhongjun Jin, Yeye He, and Surajit Chauduri. 2020 · 2020
Cited alongside, same era.
Deep entity matching with pre-trained language models
Yuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan, and Wang-Chiew Tan. 2020 · 2020
Cited alongside, same era.
TaBERT: Pretraining for joint understanding of textual and tabular data
Pengcheng Yin, Graham Neubig, Wen-tau Yih, and Sebastian Riedel. 2020 · 2020
Cited alongside, same era.
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2022a · 2022
Later among the works it cites.
Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks
Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Anjana Arunkumar, Arjun Ashok, Arut Selvan Dhanasekaran, Atharva Naik, David Stap, et al · 2022
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Later among the works it cites.
Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al · 2023
Closest in time.
How Large Language Models Will Disrupt Data Management
Raul Castro Fernandez, Aaron J Elmore, Michael J Franklin, Sanjay Krishnan, and Chenhao Tan. 2023 · 2023
Closest in time.
Large language models are state-of-the-art evaluators of translation quality
Tom Kocmi and Christian Federmann. 2023 · 2023
Closest in time.
Column Type Annotation using ChatGPT
Keti Korini and Christian Bizer. 2023 · 2023
Closest in time.
Self-Alignment with Instruction Backtranslation
Xian Li, Ping Yu, Chunting Zhou, Timo Schick, Luke Zettlemoyer, Omer Levy, Jason Weston, and Mike Lewis. 2023 · 2023
Closest in time.
Auto-BI: Automatically Build BI-Models Leveraging Local Join Prediction and Global Schema Graph
Yiming Lin, Yeye He, and Surajit Chaudhuri. 2023 · 2023
Closest in time.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2023 · 2023
Closest in time.
Using ChatGPT for Entity Matching
Ralph Peeters and Christian Bizer. 2023 · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.
Unicorn: A unified multi-tasking model for supporting matching tasks in data integration
Jianhong Tu, Ju Fan, Nan Tang, Peng Wang, Guoliang Li, Xiaoyong Du, Xiaofeng Jia, and Song Gao. 2023 · 2023
Closest in time.
Pollock: A Data Loading Benchmark
Gerardo Vitagliano, Mazhar Hameed, Lan Jiang, Lucas Reisener, Eugene Wu, and Felix Naumann. 2023 · 2023
Closest in time.
How Far Can Camels Go? Exploring the State of Instruction Tuning on Open Resources
Yizhong Wang, Hamish Ivison, Pradeep Dasigi, Jack Hessel, Tushar Khot, Khyathi Raghavi Chandu, David Wadden, Kelsey MacMillan, Noah A Smith, Iz Beltagy, et al · 2023
Closest in time.
A prompt pattern catalog to enhance prompt engineering with chatgpt
Jules White, Quchen Fu, Sam Hays, Michael Sandborn, Carlos Olea, Henry Gilbert, Ashraf Elnashar, Jesse Spencer-Smith, and Douglas C Schmidt. 2023 · 2023
Closest in time.
Lima: Less is more for alignment
Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, et al · 2023
Closest in time.