Fetching the paper…
Reading the bibliography…
Data curation tasks that prepare data for analytics are critical for turning data into actionable insights.
Data Civilizer 2.0: A Holistic Framework for Data Preparation and Analytics
El Kindi Rezig et al. 2019 · 1957
Earlier work this paper cites.
Multiple Imputation for Nonresponse in Surveys
D. B. Rubin. 1987 · 1987
Earlier work this paper cites.
Information Extraction
James R. Cowie and Wendy G. Lehnert. 1996 · 1996
Earlier work this paper cites.
The Probabilistic Relevance Framework: BM25 and Beyond
Stephen Robertson and Hugo Zaragoza. 2009 · 2009
Earlier work this paper cites.
Evaluation of entity resolution approaches on real-world match problems
Hanna Köpcke, Andreas Thor, and Erhard Rahm. 2010a · 2010
Earlier work this paper cites.
Evaluation of entity resolution approaches on real-world match problems
Hanna Köpcke, Andreas Thor, and Erhard Rahm. 2010b · 2010
Earlier work this paper cites.
Data Curation at Scale: The Data Tamer System. In CIDR
Michael Stonebraker et al. 2013 · 2013
Earlier work this paper cites.
Sequence to Sequence Learning with Neural Networks. In Advances in Neural Information Processing Systems 27: Annual Conference on Neural Information Processing Systems 2014, December 8-13 2014, Montreal, Quebec, Canada , Zoubin Ghahramani et al. (Ed.). 3104–3112
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014 · 2014
Earlier work this paper cites.
Big data curation
André Freitas and Edward Curry. 2016 · 2016
Earlier work this paper cites.
Reading Wikipedia to Answer Open-Domain Questions. In Association for Computational Linguistics (ACL)
Danqi Chen, Adam Fisch, Jason Weston, and Antoine Bordes. 2017 · 2017
Earlier work this paper cites.
Automated Hate Speech Detection and the Problem of Offensive Language. In ICWSM 2017 . AAAI Press, 512–515
Thomas Davidson, Dana Warmsley, Michael W. Macy, and Ingmar Weber. 2017 · 2017
Earlier work this paper cites.
HoloClean: Holistic Data Repairs with Probabilistic Inference
Theodoros Rekatsinas et al. 2017 · 2017
Earlier work this paper cites.
Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018 , Ellen Riloff et al. (Ed.). Association for Computational Linguistics, 3911–3921
Tao Yu et al. 2018 · 2018
Earlier work this paper cites.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Colin Raffel et al. 2020a · 2020
Earlier work this paper cites.
BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020 , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel R. Tetreault (Eds.). Association for Computational Linguistics, 7871–7880
Mike Lewis et al. 2020b · 2020
Earlier work this paper cites.
Transformers: State-of-the-Art Natural Language Processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations . Association for Computational Linguistics, Online
Thomas Wolf et al. 2020d · 2020
Cited alongside, same era.
Deep Entity Matching with Pre-Trained Language Models
Yuliang Li et al. 2020e · 2020
Cited alongside, same era.
Open Question Answering over Tables and Text
Wenhu Chen, Ming wei Chang, Eva Schlinger, William Wang, and William Cohen. 2021 · 2021
Cited alongside, same era.
Want To Reduce Labeling Cost? GPT-3 Can Help. In Findings of the Association for Computational Linguistics: EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 16-20 November, 2021 , Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (Eds.). Association for Computational Linguistics, 4195–4205
Shuohang Wang et al. 2021a · 2021
Cited alongside, same era.
PromptNER: Prompting For Named Entity Recognition
Dhananjay Ashok and Zachary C. Lipton. 2023 · 2023
Closest in time.
CodeT: Code Generation with Generated Tests. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023 . OpenReview.net
Bei Chen et al. 2023a · 2023
Closest in time.
ChatDB: Augmenting LLMs with Databases as Their Symbolic Memory
Chenxu Hu et al. 2023b · 2023
Closest in time.
StructGPT: A General Framework for Large Language Model to Reason over Structured Data
Jinhao Jiang et al. 2023d · 2023
Closest in time.
Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yichao Zhou et al. 2021c · 2021
Cited alongside, same era.
KaggleDBQA: Realistic Evaluation of Text-to-SQL Parsers. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) . Association for Computational Linguistics, Online, 2261–2273
Chia-Hsuan Lee, Oleksandr Polozov, and Matthew Richardson. 2021 · 2021
Cited alongside, same era.
Entailment as Few-Shot Learner
Sinong Wang, Han Fang, Madian Khabsa, Hanzi Mao, and Hao Ma. 2021 · 2021
Cited alongside, same era.
LangChain
Harrison Chase. 2022 · 2022
Cited alongside, same era.
CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022
Hung Le et al. 2022b · 2022
Cited alongside, same era.
DOM-LM: Learning Generalizable Representations for HTML Documents
Xiang Deng et al. 2022c · 2022
Cited alongside, same era.
GLM: General Language Model Pretraining with Autoregressive Blank Infilling. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022 , Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.). Association for Computational Linguistics, 320–335
Zhengxiao Du et al. 2022d · 2022
Cited alongside, same era.
LlamaIndex
Jerry Liu. 2022 · 2022
Cited alongside, same era.
Pan Lu et al. 2023e · 2023
Closest in time.
Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes
Simran Arora et al. 2023f · 2023
Closest in time.
ReCode: Robustness Evaluation of Code Generation Models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023 . Association for Computational Linguistics, 13818–13843
Shiqi Wang et al. 2023h · 2023
Closest in time.
HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in HuggingFace
Yongliang Shen et al. 2023j · 2023
Closest in time.
Interleaving Pre-Trained Language Models and Large Language Models for Zero-Shot NL2SQL Generation
Zihui Gu et al. 2023k · 2023
Closest in time.
GPTCache : A Library for Creating Semantic Cache for LLM Queries
Frank Liu. 2023 · 2023
Closest in time.
AgentGPT
reworkd. 2023 · 2023
Closest in time.
Auto-GPT: An Autonomous GPT-4 Experiment
Significant-Gravitas. 2023 · 2023
Closest in time.
AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation
Dong Huang et al. 2024 · 2024
Closest in time.
Entity Matching using Large Language Models
Ralph Peeters and Christian Bizer. 2024 · 2024
Closest in time.