Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Parameter-Efficient Transfer Learning for NLP
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019 · 1902
Earlier work this paper cites.
On Faithfulness and Factuality in Abstractive Summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL) . 1906–1919
Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020 · 1919
Earlier work this paper cites.
Jukebox: A Generative Model for Music
Original
Prafulla Dhariwal, Heewoo Jun, Christine Payne, Jong Wook Kim, Alec Radford, and Ilya Sutskever. 2020 · 2005
Earlier work this paper cites.
MIDAS regressions: Further results and new directions
Eric Ghysels, Arthur Sinko, and Rossen Valkanov. 2007 · 2007
Earlier work this paper cites.
ERACER: a database approach for statistical inference and data cleaning. In Proceedings of the 2010 ACM SIGMOD International Conference on Management of data . 75–86
Chris Mayfield, Jennifer Neville, and Sunil Prabhakar. 2010 · 2010
Earlier work this paper cites.
Wrangler: Interactive visual specification of data transformation scripts. In Proceedings of the sigchi conference on human factors in computing systems . 3363–3372
Sean Kandel, Andreas Paepcke, Joseph Hellerstein, and Jeffrey Heer. 2011 · 2011
Earlier work this paper cites.
Statistical Distortion: Consequences of Data Cleaning
Tamraparni Dasu and Ji Meng Loh. 2012 · 2012
Earlier work this paper cites.
CrowdER: Crowdsourcing Entity Resolution
Jiannan Wang, Tim Kraska, Michael J Franklin, and Jianhua Feng. 2012 · 2012
Earlier work this paper cites.
Holistic data cleaning: Putting violations into context. In 2013 IEEE 29th International Conference on Data Engineering (ICDE) . IEEE, 458–469
Xu Chu, Ihab F Ilyas, and Paolo Papotti. 2013 · 2013
Earlier work this paper cites.
NADEEF: a commodity data cleaning system. In Proceedings of the 2013 ACM SIGMOD International Conference on Management of Data . 541–552
Michele Dallachiesa, Amr Ebaid, Ahmed Eldawy, Ahmed Elmagarmid, Ihab F Ilyas, Mourad Ouzzani, and Nan Tang. 2013 · 2013
Earlier work this paper cites.
Optimal hashing schemes for entity matching. In Proceedings of the 22nd international conference on world wide web . 295–306
Nilesh Dalvi, Vibhor Rastogi, Anirban Dasgupta, Anish Das Sarma, and Tamás Sarlós. 2013 · 2013
Earlier work this paper cites.
Data Curation at Scale: The Data Tamer System.. In Cidr , Vol. 2013. Citeseer
Michael Stonebraker, Daniel Bruckner, Ihab F Ilyas, George Beskales, Mitch Cherniack, Stanley B Zdonik, Alexander Pagan, and Shan Xu. 2013 · 2013
Earlier work this paper cites.
Corleone: Hands-off crowdsourcing for entity matching. In Proceedings of the 2014 ACM SIGMOD international conference on Management of data . 601–612
Chaitanya Gokhale, Sanjib Das, AnHai Doan, Jeffrey F Naughton, Narasimhan Rampalli, Jude Shavlik, and Xiaojin Zhu. 2014 · 2014
Earlier work this paper cites.
Katara: A data cleaning system powered by knowledge bases and crowdsourcing. In Proceedings of the 2015 ACM SIGMOD international conference on management of data . 1247–1261
Xu Chu, John Morcos, Ihab F Ilyas, Mourad Ouzzani, Paolo Papotti, Nan Tang, and Yin Ye. 2015 · 2015
Earlier work this paper cites.
Detecting data errors: Where are we and what needs to be done?
Ziawasch Abedjan, Xu Chu, Dong Deng, Raul Castro Fernandez, Ihab F Ilyas, Mourad Ouzzani, Paolo Papotti, Michael Stonebraker, and Nan Tang. 2016 · 2016
Earlier work this paper cites.
Magellan: toward building entity matching management systems over data science stacks
Pradap Konda, Sanjib Das, AnHai Doan, Adel Ardalan, Jeffrey R Ballard, Han Li, Fatemah Panahi, Haojun Zhang, Jeff Naughton, Shishir Prasad, et al · 2016
Earlier work this paper cites.
Survey: Models and Prototypes of Schema Matching
Edhy Sutanta, Retantyo Wardoyo, Khabib Mustofa, and Edi Winarko. 2016 · 2016
Earlier work this paper cites.
Decoupled weight decay regularization
Original
Ilya Loshchilov and Frank Hutter. 2017 · 2017
Earlier work this paper cites.
HoloClean: Holistic Data Repairs with Probabilistic Inference
Theodoros Rekatsinas, Xu Chu, Ihab F Ilyas, and Christopher Ré. 2017 · 2017
Earlier work this paper cites.
The EU General Data Protection Regulation (GDPR). In Springer International Publishing
P. Voigt and A. Von dem Bussche. 2017 · 2017
Earlier work this paper cites.
How to turn data exhaust into a competitive edge
2018 · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Original
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Transform-Data-by-Example (TDE): An Extensible Search Engine for Data Transformations
Yeye He, Xu Chu, Kris Ganjam, Yudian Zheng, Vivek Narasayya, and Surajit Chaudhuri. 2018 · 2018
Earlier work this paper cites.
Deep learning for entity matching: A design space exploration. In Proceedings of the 2018 International Conference on Management of Data . 19–34
Sidharth Mudgal, Han Li, Theodoros Rekatsinas, AnHai Doan, Youngchoon Park, Ganesh Krishnan, Rohit Deep, Esteban Arcaute, and Vijay Raghavendra. 2018 · 2018
Earlier work this paper cites.
Deep Contextualized Word Representations. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers) . Association for Computational Linguistics, New Orleans, Louisiana, 2227–2237
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Earlier work this paper cites.
DataWig: Missing Value Imputation for Tables
Felix Biessmann, Tammo Rukat, Phillipp Schmidt, Prathik Naidu, Sebastian Schelter, Andrey Taptunov, Dustin Lange, and David Salinas. 2019a · 2019
Earlier work this paper cites.
DataWig: Missing Value Imputation for Tables
Felix Biessmann, Tammo Rukat, Philipp Schmidt, Prathik Naidu, Sebastian Schelter, Andrey Taptunov, Dustin Lange, and David Salinas. 2019b · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) . Association for Computational Linguistics, Minneapolis, Minnesota, 4171–4186
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Holodetect: Few-shot learning for error detection. In Proceedings of the 2019 International Conference on Management of Data . 829–846
Alireza Heidari, Joshua McGrath, Ihab F Ilyas, and Theodoros Rekatsinas. 2019 · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Original
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Earlier work this paper cites.
What is unstructured data and why is it so important to businesses?
Bernard Marr. 2021 · 2019
Earlier work this paper cites.
Language Models as Knowledge Bases?. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) . 2463–2473
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019 · 2019
Earlier work this paper cites.
Federal Judicial Caseload Statistics 2020
2020 · 2020
Earlier work this paper cites.