Fetching the paper…
Reading the bibliography…
Data curation is a wide-ranging area which contains many critical but time-consuming data processing tasks.
Unsupervised Word Sense Disambiguation Rivaling Supervised Methods. In ACL
David Yarowsky. 1995 · 1995
Earlier work this paper cites.
CrowdDB: answering queries with crowdsourcing. In SIGMOD
Michael J. Franklin et al. 2011 · 2011
Earlier work this paper cites.
ZenCrowd: leveraging probabilistic reasoning and crowdsourcing techniques for large-scale entity linking. In WWW
Gianluca Demartini et al. 2012 · 2012
Earlier work this paper cites.
Data Curation at Scale: The Data Tamer System. In CIDR
Michael Stonebraker et al. 2013 · 2013
Earlier work this paper cites.
Big data curation
André Freitas and Edward Curry. 2016 · 2016
Earlier work this paper cites.
HoloClean: Holistic Data Repairs with Probabilistic Inference
Theodoros Rekatsinas et al. 2017 · 2017
Cited alongside, same era.
Deep Entity Matching with Pre-Trained Language Models
Yuliang Li et al. 2020b · 2020
Cited alongside, same era.
Evaluating Large Language Models Trained on Code
Mark Chen et al. 2021a · 2021
Cited alongside, same era.
Distilling Knowledge from Reader to Retriever for Question Answering. In ICLR
Gautier Izacard and Edouard Grave. 2021 · 2021
Cited alongside, same era.
Exploiting Cloze-Questions for Few-Shot Text Classification and Natural Language Inference. In EACL
Timo Schick and Hinrich Schütze. 2021 · 2021
Cited alongside, same era.
The Magellan Data Repository
Sanjib et al. Das. [n.d.]
Cited in the paper.
Awesome ChatGPT Prompts
Fatih Kadir Akın et al. 2023a
Cited in the paper.
Language Models are Few-Shot Learners. In NeurIPS
Tom B. Brown et al. 2020a
Cited in the paper.
Capturing Semantics for Imputation with Pre-trained Language Models. In ICDE
Yinan Mei et al. 2021b
Cited in the paper.
Can Foundation Models Wrangle Your Data?
Avanika Narayan et al. 2022 · 2022
Later among the works it cites.
A Survey on Pretrained Language Models for Neural Code Intelligence
Yichen Xu and Yanqiao Zhu. 2022 · 2022
Later among the works it cites.
LLaMA: Open and Efficient Foundation Language Models
Hugo Touvron et al. 2023b · 2023
Closest in time.
Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing
Pengfei Liu et al. 2023c · 2023
Closest in time.
Toolformer: Language Models Can Teach Themselves to Use Tools
Timo Schick et al. 2023d · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…