Fetching the paper…
Reading the bibliography…
Entity matching (EM) is the problem of determining whether two records refer to same real-world entity, which is crucial in data integration, e.g., for product catalogs or address databases.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Finding related tables in data lakes for interactive data science. In Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data . 1951–1966
Yi Zhang and Zachary G Ives. 2020 · 1966
Earlier work this paper cites.
Detecting data errors: Where are we and what needs to be done?
Ziawasch Abedjan, Xu Chu, Dong Deng, Raul Castro Fernandez, Ihab F Ilyas, Mourad Ouzzani, Paolo Papotti, Michael Stonebraker, and Nan Tang. 2016 · 2016
Earlier work this paper cites.
The Data Linter: Lightweight Automated Sanity Checking for ML Data Sets
Nick Hynes, D. Sculley, and Michael Terry. 2017 · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Navigating the data lake with datamaran: Automatically extracting structure from log datasets. In Proceedings of the 2018 International Conference on Management of Data . 943–958
Yihan Gao, Silu Huang, and Aditya Parameswaran. 2018 · 2018
Earlier work this paper cites.
Deep learning for entity matching: A design space exploration. In Proceedings of the 2018 international conference on management of data . 19–34
Sidharth Mudgal, Han Li, Theodoros Rekatsinas, AnHai Doan, Youngchoon Park, Ganesh Krishnan, Rohit Deep, Esteban Arcaute, and Vijay Raghavendra. 2018 · 2018
Earlier work this paper cites.
Data Integration: The Current Status and the Way Forward
Michael Stonebraker, Ihab F Ilyas, et al · 2018
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Earlier work this paper cites.
Magellan: toward building ecosystems of entity matching solutions
AnHai Doan, Pradap Konda, Paul Suganthan GC, Yash Govind, Derek Paulsen, Kaushik Chandrasekhar, Philip Martinkus, and Matthew Christie. 2020 · 2020
Earlier work this paper cites.
Autogluon-tabular: Robust and accurate automl for structured data
Nick Erickson, Jonas Mueller, Alexander Shirkov, Hang Zhang, Pedro Larroy, Mu Li, and Alexander Smola. 2020 · 2020
Earlier work this paper cites.
Deep entity matching with pre-trained language models
Yuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan, and Wang-Chiew Tan. 2020 · 2020
Earlier work this paper cites.
Blocking and filtering techniques for entity resolution: A survey
George Papadakis, Dimitrios Skoutas, Emmanouil Thanos, and Themis Palpanas. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Cited alongside, same era.
Zeroer: Entity resolution using zero labeled examples. In Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data . 1149–1164
Renzhi Wu, Sanya Chaba, Saurabh Sawlani, Xu Chu, and Saravanan Thirumuruganathan. 2020 · 2020
Cited alongside, same era.
Multi-context attention for entity matching. In Proceedings of The Web Conference 2020 . 2634–2640
Dongxiang Zhang, Yuyang Nie, Sai Wu, Yanyan Shen, and Kian-Lee Tan. 2020 · 2020
Cited alongside, same era.
GNEM: a generic one-to-set neural entity matching framework. In Proceedings of the Web Conference 2021 . 1686–1694
Runjin Chen, Yanyan Shen, and Dongxiang Zhang. 2021 · 2021
Cited alongside, same era.
Hierarchical matching network for heterogeneous entity resolution. In Proceedings of the Twenty-Ninth International Conference on International Joint Conferences on Artificial Intelligence . 3665–3671
Entity matching using large language models
Ralph Peeters and Christian Bizer. 2023 · 2023
Later among the works it cites.
Through the Fairness Lens: Experimental Analysis and Evaluation of Entity Matching
Nima Shahbazi, Nikola Danevski, Fatemeh Nargesian, Abolfazl Asudeh, and Divesh Srivastava. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
Jellyfish: A Large Language Model for Data Preprocessing
Haochen Zhang, Yuyang Dong, Chuan Xiao, and Masafumi Oyamada. 2023 · 2023
Later among the works it cites.
Distributed inference and fine-tuning of large language models over the internet
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cheng Fu, Xianpei Han, Jiaming He, and Le Sun. 2021 · 2021
Cited alongside, same era.
CleanML: A study for evaluating the impact of data cleaning on ml classification tasks. In 2021 IEEE 37th International Conference on Data Engineering (ICDE) . IEEE, 13–24
Peng Li, Xi Rao, Jennifer Blase, Yue Zhang, Xu Chu, and Ce Zhang. 2021 · 2021
Cited alongside, same era.
Production machine learning pipelines: Empirical analysis and optimization opportunities. In Proceedings of the 2021 International Conference on Management of Data . 2639–2652
Doris Xin, Hui Miao, Aditya Parameswaran, and Neoklis Polyzotis. 2021 · 2021
Cited alongside, same era.
A critical re-evaluation of neural methods for entity alignment
Manuel Leone, Stefano Huber, Akhil Arora, Alberto García-Durán, and Robert West. 2022 · 2022
Cited alongside, same era.
Can Foundation Models Wrangle Your Data?
Avanika Narayan et al · 2022
Cited alongside, same era.
Towards parameter-efficient automation of data wrangling tasks with prefix-tuning. In NeurIPS 2022 First Table Representation Workshop
David Vos, Till Döhmen, and Sebastian Schelter. 2022 · 2022
Cited alongside, same era.
Machop: an end-to-end generalized entity matching framework. In Proceedings of the Fifth International Workshop on Exploiting Artificial Intelligence Techniques for Data Management . 1–10
Jin Wang, Yuliang Li, Wataru Hirota, and Eser Kandogan. 2022 · 2022
Cited alongside, same era.
REIN: A Comprehensive Benchmark Framework for Data Cleaning Methods in ML Pipelines
Mohamed Abdelaal, Christian Hammacher, and Harald Schoening. 2023 · 2023
Cited alongside, same era.
Alexander Borzunov, Max Ryabinin, Artem Chumachenko, Dmitry Baranchuk, Tim Dettmers, Younes Belkada, Pavel Samygin, and Colin A Raffel. 2024 · 2024
Closest in time.
Gen-T: Table Reclamation in Data Lakes
Grace Fan, Roee Shraga, and Renée J Miller. 2024 · 2024
Closest in time.
Relationalizing Tables with Large Language Models: The Promise and Challenges. In 2024 IEEE 40th International Conference on Data Engineering Workshops (ICDEW) . IEEE, 305–309
Zezhou Huang and Eugene Wu. 2024 · 2024
Closest in time.
Table-GPT: Table Fine-tuned GPT for Diverse Table Tasks
Peng Li, Yeye He, Dror Yashar, Weiwei Cui, Song Ge, Haidong Zhang, Danielle Rifinski Fainman, Dongmei Zhang, and Surajit Chaudhuri. 2024 · 2024
Closest in time.
A Declarative System for Optimizing AI Workloads
Chunwei Liu, Matthew Russo, Michael Cafarella, Lei Cao, Peter Baille Chen, Zui Chen, Michael Franklin, Tim Kraska, Samuel Madden, and Gerardo Vitagliano. 2024 · 2024
Closest in time.
A critical re-evaluation of benchmark datasets for (deep) learning-based matching algorithms
George Papadakis, Nishadi Kirielle, Peter Christen, and Themis Palpanas. 2024 · 2024
Closest in time.
Mixture-of-Agents Enhances Large Language Model Capabilities
Junlin Wang, Jue Wang, Ben Athiwaratkun, Ce Zhang, and James Zou. 2024 · 2024
Closest in time.
Deepmatcher: a deep transformer-based network for robust and accurate local feature matching
Tao Xie, Kun Dai, Ke Wang, Ruifeng Li, and Lijun Zhao. 2024 · 2024
Closest in time.
Directions Towards Efficient and Automated Data Wrangling with Large Language Models
Zeyu Zhang, Paul Groth, Iacer Calixto, and Sebastian Schelter. 2024 · 2024
Closest in time.