Fetching the paper…
Reading the bibliography…
This paper explores the utilization of LLMs for data preprocessing (DP), a crucial step in the data mining pipeline that transforms raw data into a clean format conducive to easy processing.
No free lunch theorems for optimization
D. H. Wolpert and W. G. Macready · 1997
Earlier work this paper cites.
Holistic data cleaning: Putting violations into context
X. Chu, I. F. Ilyas, and P. Papotti · 2013
Earlier work this paper cites.
Schema matching prediction with applications to data source discovery and dynamic ensembling
T. Sagi and A. Gal · 2013
Earlier work this paper cites.
Katara: A data cleaning system powered by knowledge bases and crowdsourcing
X. Chu, J. Morcos, I. F. Ilyas, M. Ouzzani, P. Papotti, N. Tang, and Y. Ye · 2015
Earlier work this paper cites.
Truth finding on the deep web: Is the problem solved?
X. Li, X. L. Dong, K. Lyons, W. Meng, and D. Srivastava · 2015
Earlier work this paper cites.
Combining quantitative and logical data cleaning
N. Prokoshyna, J. Szlichta, F. Chiang, R. J. Miller, and D. Srivastava · 2015
Earlier work this paper cites.
Magellan: toward building entity matching management systems over data science stacks
P. Konda, S. Das, A. Doan, A. Ardalan, J. R. Ballard, H. Li, F. Panahi, H. Zhang, J. Naughton, S. Prasad, et al · 2016
Earlier work this paper cites.
Recognizing salient entities in shopping queries
Z. Kozareva, Q. Li, K. Zhai, and W. Guo · 2016
Earlier work this paper cites.
Profiling the potential of web tables for augmenting cross-domain knowledge bases
D. Ritze, O. Lehmberg, Y. Oulabi, and C. Bizer · 2016
Earlier work this paper cites.
HoloClean: Holistic data repairs with probabilistic inference
T. Rekatsinas, X. Chu, I. F. Ilyas, and C. Ré · 2017
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Transform-data-by-example (TDE) an extensible search engine for data transformations
Y. He, X. Chu, K. Ganjam, Y. Zheng, V. Narasayya, and S. Chaudhuri · 2018
Earlier work this paper cites.
Auto-detect: Data-driven error detection in tables
Z. Huang and Y. He · 2018
Earlier work this paper cites.
Deep learning for entity matching: A design space exploration
S. Mudgal, H. Li, T. Rekatsinas, A. Doan, Y. Park, G. Krishnan, R. Deep, E. Arcaute, and V. Raghavendra · 2018
Earlier work this paper cites.
Enriching data imputation under similarity rule constraints
S. Song, Y. Sun, A. Zhang, L. Chen, and J. Wang · 2018
Earlier work this paper cites.
GAIN: Missing data imputation using generative adversarial nets
J. Yoon, J. Jordon, and M. Schaar · 2018
Earlier work this paper cites.
Opentag: Open attribute value extraction from product profiles
G. Zheng, S. Mukherjee, X. L. Dong, and F. Li · 2018
Earlier work this paper cites.
On the measure of intelligence
F. Chollet · 2019
Earlier work this paper cites.
Learning to rerank schema matches
A. Gal, H. Roitman, and R. Shraga · 2019
Earlier work this paper cites.
HoloDetect: Few-shot learning for error detection
A. Heidari, J. McGrath, I. F. Ilyas, and T. Rekatsinas · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for NLP
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. De Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly · 2019
Earlier work this paper cites.
TinyBERT: Distilling BERT for natural language understanding
X. Jiao, Y. Yin, L. Shang, X. Jiang, X. Chen, L. Li, F. Wang, and Q. Liu · 2019
Earlier work this paper cites.
RoBERTa: A robustly optimized BERT pretraining approach
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov · 2019
Earlier work this paper cites.
Raha: A configuration-free error detection system
M. Mahdavi, Z. Abedjan, R. Castro Fernandez, S. Madden, M. Ouzzani, M. Stonebraker, and N. Tang · 2019
Earlier work this paper cites.
Uni-detect: A unified approach to automated error detection in tables
P. Wang and Y. He · 2019
Earlier work this paper cites.
Scaling up open tagging from tens to thousands: Comprehension empowered attribute value extraction from product title
H. Xu, W. Wang, X. Mao, X. Jiang, and M. Lan · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi · 2019
Earlier work this paper cites.
Dataset discovery in data lakes
A. Bogatu, A. A. Fernandes, N. W. Paton, and N. Konstantinou · 2020
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Dataset search: a survey
A. Chapman, E. Simperl, L. Koesten, G. Konstantinidis, L.-D. Ibáñez, E. Kacprzak, and P. Groth · 2020
Earlier work this paper cites.
ARDA: automatic relational data augmentation for machine learning
N. Chepurko, R. Marcus, E. Zgraggen, R. C. Fernandez, T. Kraska, and D. Karger · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2020
Earlier work this paper cites.
Auto-transform: learning-to-transform by patterns
Z. Jin, Y. He, and S. Chauduri · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel, et al · 2020
Earlier work this paper cites.
Deep entity matching with pre-trained language models
Y. Li, J. Li, Y. Suhara, A. Doan, and W.-C. Tan · 2020
Earlier work this paper cites.
Baran: Effective error correction via a unified context representation and transfer learning
M. Mahdavi and Z. Abedjan · 2020
Earlier work this paper cites.
Handling incomplete heterogeneous data using VAEs
A. Nazabal, P. M. Olmos, Z. Ghahramani, and I. Valera · 2020
Earlier work this paper cites.
Blocking and filtering techniques for entity resolution: A survey
G. Papadakis, D. Skoutas, E. Thanos, and T. Palpanas · 2020
Earlier work this paper cites.
Adnev: Cross-domain schema matching using deep similarity matrix adjustment and evaluation
R. Shraga, A. Gal, and H. Roitman · 2020
Cited alongside, same era.
Learning to extract attribute value from product via question answering: A multi-task approach
Q. Wang, L. Yang, B. Kanagal, S. Sanghai, D. Sivakumar, B. Shu, Z. Yu, and J. Elsas · 2020
Cited alongside, same era.
Attention-based learning for missing data imputation in HoloClean
R. Wu, A. Zhang, I. Ilyas, and T. Rekatsinas · 2020
Cited alongside, same era.
Finding related tables in data lakes for interactive data science
Y. Zhang and Z. G. Ives · 2020
Cited alongside, same era.
Multimodal joint attribute prediction and value extraction for e-commerce product
T. Zhu, Y. Wang, H. Li, Y. Wu, X. He, and B. Zhou · 2020
Cited alongside, same era.
Large language models empowered agent-based modeling and simulation: A survey and perspectives
C. Gao, X. Lan, N. Li, Y. Yuan, J. Ding, Z. Zhou, F. Xu, and Y. Li · 2023
Closest in time.
Data lakes: A survey of functions and systems
R. Hai, C. Koutras, C. Quix, and M. Jarke · 2023
Closest in time.
Record fusion via inference and data augmentation
A. Heidari, G. Michalopoulos, I. F. Ilyas, and T. Rekatsinas · 2023
Closest in time.
MetaGPT: Meta programming for multi-agent collaborative framework
S. Hong, X. Zheng, J. Chen, Y. Cheng, C. Zhang, Z. Wang, S. K. S. Yau, Z. Lin, L. Zhou, C. Ran, et al · 2023
Closest in time.
Llama and llama 2 variants
Hugging Face · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A survey on data fusion: what for? in what form? what is next?
G. K. Canalle, A. C. Salgado, and B. F. Loscio · 2021
Cited alongside, same era.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al · 2021
Cited alongside, same era.
LoRA: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2021
Cited alongside, same era.
Tabbie: Pretrained representations of tabular data
H. Iida, D. Thai, V. Manjunatha, and M. Iyyer · 2021
Cited alongside, same era.
Valentine: Evaluating matching techniques for dataset discovery
C. Koutras, G. Siachamis, A. Ionescu, K. Psarakis, J. Brons, M. Fragkoulis, C. Lofi, A. Bonifati, and A. Katsifodimos · 2021
Cited alongside, same era.
PClean: Bayesian data cleaning at scale with domain-specific probabilistic programming
A. Lew, M. Agrawal, D. Sontag, and V. Mansinghka · 2021
Cited alongside, same era.
Prefix-tuning: Optimizing continuous prompts for generation
X. L. Li and P. Liang · 2021
Cited alongside, same era.
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. d. l. Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, et al · 2023
Closest in time.
SANTOS: Relationship-based semantic table union search
A. Khatiwada, G. Fan, R. Shraga, Z. Chen, W. Gatterbauer, R. J. Miller, and M. Riedewald · 2023
Closest in time.
Column type annotation using ChatGPT
K. Korini and C. Bizer · 2023
Closest in time.
Efficient memory management for large language model serving with pagedattention
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, and I. Stoica · 2023
Closest in time.
OpenOrcaPlatypus: Llama2-13B model instruct-tuned on filtered OpenOrcaV1 GPT-4 dataset and merged with divergent STEM and logic dataset model
A. N. Lee, C. J. Hunter, N. Ruiz, B. Goodson, W. Lian, G. Wang, E. Pentland, A. Cook, C. Vong, and ”Teknium” · 2023
Closest in time.
Table-GPT: Table-tuned GPT for diverse table tasks
P. Li, Y. He, D. Yashar, W. Cui, S. Ge, H. Zhang, D. R. Fainman, D. Zhang, and S. Chaudhuri · 2023
Closest in time.
Lost in the middle: How language models use long contexts
N. F. Liu, K. Lin, J. Hewitt, A. Paranjape, M. Bevilacqua, F. Petroni, and P. Liang · 2023
Closest in time.
LLM-FP4: 4-bit floating-point quantized transformers
S.-y. Liu, Z. Liu, X. Huang, P. Dong, and K.-T. Cheng · 2023
Closest in time.
Stable beluga 2
D. Mahan, R. Carlow, L. Castricato, N. Cooper, and C. Laforte · 2023
Closest in time.
March 20 ChatGPT outage: Here’s what happened, 2023
OpenAI · 2023
Closest in time.
Entity matching using large language models
R. Peeters and C. Bizer · 2023
Closest in time.
Communicative agents for software development
C. Qian, X. Cong, C. Yang, W. Chen, Y. Su, J. Xu, Z. Liu, and M. Sun · 2023
Closest in time.
BClean: A bayesian data cleaning system
J. Qin, S. Huang, Y. Wang, J. Zhu, Y. Zhang, Y. Miao, R. Mao, M. Onizuka, and C. Xiao · 2023
Closest in time.
Can gpt-4 support analysis of textual data in tasks requiring highly specialized domain expertise?
J. Savelka, K. D. Ashley, M. A. Gray, H. Westermann, and H. Xu · 2023
Closest in time.
LLaMA: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Closest in time.
Unicorn: A unified multi-tasking model for supporting matching tasks in data integration
J. Tu, J. Fan, N. Tang, P. Wang, G. Li, X. Du, X. Jia, and S. Gao · 2023
Closest in time.
Solar-0-70b-16bit
Upstage · 2023
Closest in time.
A survey on large language model based autonomous agents
L. Wang, C. Ma, X. Feng, Z. Zhang, H. Yang, J. Zhang, Z. Chen, J. Tang, X. Chen, Y. Lin, et al · 2023
Closest in time.
Prompt engineering
L. Weng · 2023
Closest in time.
Smart agent-based modeling: On the use of large language models in computer simulations
Z. Wu, R. Peng, X. Han, S. Zheng, Y. Zhang, and C. Xiao · 2023
Closest in time.
The rise and potential of large language model based agents: A survey
Z. Xi, W. Chen, X. Guo, W. He, Y. Ding, B. Hong, M. Zhang, J. Wang, S. Jin, E. Zhou, et al · 2023
Closest in time.
Large language models as data preprocessors
H. Zhang, Y. Dong, C. Xiao, and M. Oyamada · 2023
Closest in time.
Instruction tuning for large language models: A survey
S. Zhang, L. Dong, X. Li, S. Zhang, X. Sun, S. Wang, J. Li, R. Hu, T. Zhang, F. Wu, et al · 2023
Closest in time.
Tablellama: Towards open large generalist models for tables
T. Zhang, X. Yue, Y. Li, and H. Sun · 2023
Closest in time.
Siren’s song in the ai ocean: A survey on hallucination in large language models
Y. Zhang, Y. Li, L. Cui, D. Cai, L. Liu, T. Fu, X. Huang, E. Zhao, Y. Zhang, Y. Chen, et al · 2023
Closest in time.
A survey of large language models
W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, Y. Min, B. Zhang, J. Zhang, Z. Dong, et al · 2023
Closest in time.
Open llm leaderboard
H. Face · 2024
Closest in time.
Awq: Activation-aware weight quantization for on-device llm compression and acceleration
J. Lin, J. Tang, H. Tang, S. Yang, W.-M. Chen, W.-C. Wang, G. Xiao, X. Dang, C. Gan, and S. Han · 2024
Closest in time.
Large language model for table processing: A survey
W. Lu, J. Zhang, J. Zhang, and Y. Chen · 2024
Closest in time.
Introducing Meta Llama 3: The most capable openly available LLM to date
Meta AI · 2024
Closest in time.
Rematch: Retrieval enhanced schema matching with llms
E. Sheetrit, M. Brief, M. Mishaeli, and O. Elisha · 2024
Closest in time.
Collaborative agents for software engineering
D. Tang, Z. Chen, K. Kim, Y. Song, H. Tian, S. Ezzini, Y. Huang, and J. K. T. F. Bissyande · 2024
Closest in time.
vLLM: Easy, fast, and cheap LLM serving with PagedAttention
vLLM Team · 2024
Closest in time.