Fetching the paper…
Reading the bibliography…
We apply foundation models to data discovery and exploration tasks.
Binary codes capable of correcting deletions, insertions and reversals
Vladimir Levenshtein · 1966
Earlier work this paper cites.
The mindlessness of ostensibly thoughtful action: The role of" placebic" information in interpersonal interaction
Ellen J Langer, Arthur Blank, and Benzion Chanowitz · 1978
Earlier work this paper cites.
Wrapper induction for information extraction
Nicholas Kushmerick, Daniel S. Weld, and Robert B. Doorenbos · 1997
Earlier work this paper cites.
Comp.basilisk faq
David Langford · 1999
Earlier work this paper cites.
Mining database structure; or, how to build a data quality browser
Tamraparni Dasu, Theodore Johnson, S. Muthukrishnan, and Vladislav Shkapenyuk · 2002
Earlier work this paper cites.
Provenance semirings
Todd J. Green, Grigoris Karvounarakis, and Val Tannen · 2007
Earlier work this paper cites.
Webtables: exploring the power of tables on the web
Michael J. Cafarella, Alon Y. Halevy, Daisy Zhe Wang, Eugene Wu, and Yang Zhang · 2008
Earlier work this paper cites.
Shedding light on the dark data in the long tail of science
P. Bryan Heidorn · 2008
Earlier work this paper cites.
Common crawl, 2011
Common Crawl Foundation · 2011
Earlier work this paper cites.
Wrangler: interactive visual specification of data transformation scripts
Sean Kandel, Andreas Paepcke, Joseph M. Hellerstein, and Jeffrey Heer · 2011
Earlier work this paper cites.
Dbpedia: A multilingual cross-domain knowledge base
Pablo N. Mendes, Max Jakob, and Christian Bizer · 2012
Earlier work this paper cites.
Fast foreign-key detection in microsoft SQL server powerpivot for excel
Zhimin Chen, Vivek R. Narasayya, and Surajit Chaudhuri · 2014
Earlier work this paper cites.
The Forrester Wave: Big Data Hadoop Distributions, Q1 2016
Mike Gualtieri and Noel Yuhanna · 2016
Earlier work this paper cites.
Extracting databases from dark data with deepdive
Ce Zhang, Jaeho Shin, Christopher Ré, Michael J. Cafarella, and Feng Niu · 2016
Earlier work this paper cites.
Matching web tables to dbpedia - A feature utility study
Dominique Ritze and Christian Bizer · 2017
Earlier work this paper cites.
Ten years of webtables
Michael J. Cafarella, Alon Y. Halevy, Hongrae Lee, Jayant Madhavan, Cong Yu, Daisy Zhe Wang, and Eugene Wu · 2018
Earlier work this paper cites.
Google dataset search: Building a search engine for datasets in an open web ecosystem
Dan Brickley, Matthew Burgess, and Natasha F. Noy · 2019
Earlier work this paper cites.
Viznet: Towards a large-scale visualization learning and benchmarking repository
Kevin Hu, Neil Gaikwad, Michiel Bakker, Madelon Hulsebos, Emanuel Zgraggen, César Hidalgo, Tim Kraska, Guoliang Li, Arvind Satyanarayan, and Çağatay Demiralp · 2019
Earlier work this paper cites.
Sherlock: A deep learning approach to semantic data type detection
Madelon Hulsebos, Kevin Zeng Hu, Michiel A. Bakker, Emanuel Zgraggen, Arvind Satyanarayan, Tim Kraska, Çagatay Demiralp, and César A. Hidalgo · 2019
Earlier work this paper cites.
Data lake management: Challenges and opportunities
Fatemeh Nargesian, Erkang Zhu, Renée J. Miller, Ken Q. Pu, and Patricia C. Arocena · 2019
Earlier work this paper cites.
JOSIE: overlap set similarity search for finding joinable tables in data lakes
Erkang Zhu, Dong Deng, Fatemeh Nargesian, and Renée J. Miller · 2019
Earlier work this paper cites.
Dataset search: a survey
Adriane Chapman, Elena Simperl, Laura Koesten, George Konstantinidis, Luis-Daniel Ibáñez, Emilia Kacprzak, and Paul Groth · 2020
Earlier work this paper cites.
TURL: table understanding through representation learning
Xiang Deng, Huan Sun, Alyssa Lees, You Wu, and Cong Yu · 2020
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith · 2020
Earlier work this paper cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi · 2020
Earlier work this paper cites.
Auto-suggest: Learning-to-recommend data preparation steps using data science notebooks
Cong Yan and Yeye He · 2020
Cited alongside, same era.
Tabert: Pretraining for joint understanding of textual and tabular data
Pengcheng Yin, Graham Neubig, Wen-tau Yih, and Sebastian Riedel · 2020
Cited alongside, same era.
Sato: Contextual semantic type detection in tables
Dan Zhang, Yoshihiko Suhara, Jinfeng Li, Madelon Hulsebos, Çagatay Demiralp, and Wang-Chiew Tan · 2020
Cited alongside, same era.
State of data science
Inc. Anaconda · 2021
Cited alongside, same era.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ B. Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, Erik Brynjolfsson, Shyamal Buch, Dallas Card, Rodrigo Castellon, Niladri S. Chatterji, Annie S. Chen, Kathleen Creel, Jared Quincy Davis, Dorottya Demszky, Chris Donahue, Moussa Doumbouya, Esin Durmus, Stefano Ermon, John Etchemendy, Kawin Ethayarajh, Li Fei-Fei, Chelsea Finn, Trevor Gale, Lauren Gillespie, Karan Goel, Noah D. Goodman, Shelby Grossman, Neel Guha, Tatsunori Hashimoto, Peter Henderson, John Hewitt, Daniel E. Ho, Jenny Hong, Kyle Hsu, Jing Huang, Thomas Icard, Saahil Jain, Dan Jurafsky, Pratyusha Kalluri, Siddharth Karamcheti, Geoff Keeling, Fereshte Khani, Omar Khattab, Pang Wei Koh, Mark S. Krass, Ranjay Krishna, Rohith Kuditipudi, and et al · 2021
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe · 2022
Later among the works it cites.
A sketch-based index for correlated dataset search
Aécio S. R. Santos, Aline Bessa, Christopher Musco, and Juliana Freire · 2022
Later among the works it cites.
Large language models encode clinical knowledge
Karan Singhal, Shekoofeh Azizi, Tao Tu, S. Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Kumar Tanwani, Heather Cole-Lewis, Stephen Pfohl, Perry Payne, Martin Seneviratne, Paul Gamble, Chris Kelly, Nathaneal Schärli, Aakanksha Chowdhery, Philip Andrew Mansfield, Blaise Agüera y Arcas, Dale R. Webster, Gregory S. Corrado, Yossi Matias, Katherine Chou, Juraj Gottweis, Nenad Tomasev, Yun Liu, Alvin Rajkomar, Joelle K. Barral, Christopher Semturs, Alan Karthikesalingam, and Vivek Natarajan · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Auctus: A dataset search engine for data discovery and augmentation
Sonia Castelo, Rémi Rampin, Aécio S. R. Santos, Aline Bessa, Fernando Chirigati, and Juliana Freire · 2021
Cited alongside, same era.
Documenting large webtext corpora: A case study on the colossal clean crawled corpus
Jesse Dodge, Maarten Sap, Ana Marasovic, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner · 2021
Cited alongside, same era.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy · 2021
Cited alongside, same era.
Gittables: A large-scale corpus of relational tables
Madelon Hulsebos, Çagatay Demiralp, and Paul Groth · 2021
Cited alongside, same era.
TABBIE: pretrained representations of tabular data
Hiroshi Iida, Dung Thai, Varun Manjunatha, and Mohit Iyyer · 2021
Cited alongside, same era.
Generating table vector representations
Aneta Koleva, Martin Ringsquandl, Mitchell Joblin, and Volker Tresp · 2021
Cited alongside, same era.
Data lake organization
Fatemeh Nargesian, Ken Q. Pu, Bahar Ghadiri Bashardoost, Erkang Zhu, and Renée J. Miller · 2021
Cited alongside, same era.
Later among the works it cites.
Annotating columns with pre-trained language models
Yoshihiko Suhara, Jinfeng Li, Yuliang Li, Dan Zhang, Çağatay Demiralp, Chen Chen, and Wang-Chiew Tan · 2022
Later among the works it cites.
Galactica: A large language model for science
Ross Taylor, Marcin Kardas, Guillem Cucurull, Thomas Scialom, Anthony Hartshorn, Elvis Saravia, Andrew Poulton, Viktor Kerkez, and Robert Stojnic · 2022
Later among the works it cites.
Towards nlp-enhanced data profiling tools
Immanuel Trummer · 2022
Later among the works it cites.
Emergent abilities of large language models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus · 2022
Later among the works it cites.
Symbolic knowledge distillation: from general language models to commonsense models
Peter West, Chandra Bhagavatula, Jack Hessel, Jena D. Hwang, Liwei Jiang, Ronan Le Bras, Ximing Lu, Sean Welleck, and Yejin Choi · 2022
Later among the works it cites.
Language models enable simple systems for generating structured views of heterogeneous data lakes, 2023
Simran Arora, Brandon Yang, Sabri Eyuboglu, Avanika Narayan, Andrew Hojel, Immanuel Trummer, and Christopher Ré · 2023
Closest in time.
Sparks of artificial general intelligence: Early experiments with gpt-4, 2023
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang · 2023
Closest in time.
Mata v. avianca, inc. (1:22-cv-01461)
District Court · 2023
Closest in time.
Qlora: Efficient finetuning of quantized llms, 2023
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer · 2023
Closest in time.
Faith and fate: Limits of transformers on compositionality
Nouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li, Liwei Jiang, Bill Yuchen Lin, Peter West, Chandra Bhagavatula, Ronan Le Bras, Jena D. Hwang, Soumya Sanyal, Sean Welleck, Xiang Ren, Allyson Ettinger, Zaïd Harchaoui, and Yejin Choi · 2023
Closest in time.
Table discovery in data lakes: State-of-the-art and future directions
Grace Fan, Jin Wang, Yuliang Li, and Renée J. Miller · 2023
Closest in time.
The false promise of imitating proprietary llms, 2023
Arnav Gudibande, Eric Wallace, Charlie Snell, Xinyang Geng, Hao Liu, Pieter Abbeel, Sergey Levine, and Dawn Song · 2023
Closest in time.
CHORUS: foundation models for unified data discovery and exploration
Moe Kayali, Anton Lykov, Ilias Fountalis, Nikolaos Vasiloglou, Dan Olteanu, and Dan Suciu · 2023
Closest in time.
SANTOS: relationship-based semantic table union search
Aamod Khatiwada, Grace Fan, Roee Shraga, Zixuan Chen, Wolfgang Gatterbauer, Renée J. Miller, and Mirek Riedewald · 2023
Closest in time.
Electric vehicle population data electric vehicle population data, 04 2023
Washington State Department of Licensing · 2023
Closest in time.
Large language models can be easily distracted by irrelevant context
Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed H. Chi, Nathanael Schärli, and Denny Zhou · 2023
Closest in time.
Trifacta wrangler
Trifacta · 2023
Closest in time.
Chatgpt for robotics: Design principles and model abilities
Sai Vemprala, Rogerio Bonatti, Arthur Bucker, and Ashish Kapoor · 2023
Closest in time.
How language model hallucinations can snowball, 2023
Muru Zhang, Ofir Press, William Merrill, Alisa Liu, and Noah A. Smith · 2023
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica · 2023
Closest in time.