Fetching the paper…
Reading the bibliography…
It is well-established that large, diverse datasets play a pivotal role in the performance of modern AI systems for text and image modalities.
Adriane Chapman, Elena Simperl, Laura Koesten, George Konstantinidis, Luis-Daniel Ibáñez-Gonzalez, Emilia Kacprzak, and Paul Groth · 1901
Earlier work this paper cites.
VizNet: Towards A Large-Scale Visualization Learning and Benchmarking Repository, May 2019
Kevin Hu, Neil Gaikwad, Michiel Bakker, Madelon Hulsebos, Emanuel Zgraggen, César Hidalgo, Tim Kraska, Guoliang Li, Arvind Satyanarayan, and Çağatay Demiralp · 1905
Earlier work this paper cites.
Sherlock: A Deep Learning Approach to Semantic Data Type Detection, May 2019
Madelon Hulsebos, Kevin Hu, Michiel Bakker, Emanuel Zgraggen, Arvind Satyanarayan, Tim Kraska, Çağatay Demiralp, and César Hidalgo · 1905
Earlier work this paper cites.
Recommending Related Tables, July 2019
Shuo Zhang and Krisztian Balog · 1907
Earlier work this paper cites.
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks, August 2019
Nils Reimers and Iryna Gurevych · 1908
Earlier work this paper cites.
Powerlaw: a Python package for analysis of heavy-tailed distributions
Jeff Alstott, Ed Bullmore, and Dietmar Plenz · 1932
Earlier work this paper cites.
ToTTo: A Controlled Table-To-Text Generation Dataset, October 2020
Ankur P. Parikh, Xuezhi Wang, Sebastian Gehrmann, Manaal Faruqui, Bhuwan Dhingra, Diyi Yang, and Dipanjan Das · 2004
Earlier work this paper cites.
TAPAS: Weakly Supervised Table Parsing via Pre-training
Jonathan Herzig, Paweł Krzysztof Nowak, Thomas Müller, Francesco Piccinno, and Julian Martin Eisenschlos · 2004
Earlier work this paper cites.
TaBERT: Pretraining for Joint Understanding of Textual and Tabular Data, May 2020
Pengcheng Yin, Graham Neubig, Wen-tau Yih, and Sebastian Riedel · 2005
Earlier work this paper cites.
Power laws, Pareto distributions and Zipf’s law
M. E. J. Newman · 2005
Earlier work this paper cites.
Google Dataset Search by the Numbers, June 2020
Omar Benjelloun, Shiyu Chen, and Natasha Noy · 2006
Earlier work this paper cites.
WebTables: exploring the power of tables on the web
Michael J. Cafarella, Alon Halevy, Daisy Zhe Wang, Eugene Wu, and Yang Zhang · 2008
Earlier work this paper cites.
Yuyang Dong, Kunihiro Takeoka, Chuan Xiao, and Masafumi Oyamada · 2010
Earlier work this paper cites.
Data Structures for Statistical Computing in Python
Wes McKinney · 2010
Earlier work this paper cites.
Bridging Textual and Tabular Data for Cross-Domain Text-to-SQL Semantic Parsing, December 2020
Xi Victoria Lin, Richard Socher, and Caiming Xiong · 2012
Earlier work this paper cites.
Efficient Estimation of Word Representations in Vector Space, September 2013
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
Knowledge vault: a web-scale approach to probabilistic knowledge fusion
Xin Dong, Evgeniy Gabrilovich, Geremy Heitz, Wilko Horn, Ni Lao, Kevin Murphy, Thomas Strohmann, Shaohua Sun, and Wei Zhang · 2014
Earlier work this paper cites.
TabEL: Entity Linking in Web Tables
Chandra Sekhar Bhagavatula, Thanapon Noraset, and Doug Downey · 2015
Earlier work this paper cites.
A Large Public Corpus of Web Tables containing Time and Context Metadata
Oliver Lehmberg, Dominique Ritze, Robert Meusel, and Christian Bizer · 2016
Earlier work this paper cites.
Matching Web Tables with Knowledge Base Entities: From Entity Lookups to Entity Embeddings
Vasilis Efthymiou, Oktie Hassanzadeh, Mariano Rodriguez-Muro, and Vassilis Christophides · 2017
Cited alongside, same era.
Auto-join: joining tables by leveraging transformations
Erkang Zhu, Yeye He, and Surajit Chaudhuri · 2017
Cited alongside, same era.
Effective and efficient Semantic Table Interpretation using TableMiner+
Ziqi Zhang · 2017
Cited alongside, same era.
Ad Hoc Table Retrieval using Semantic Similarity
Shuo Zhang and Krisztian Balog · 2018
Cited alongside, same era.
Table union search on open data
Fatemeh Nargesian, Erkang Zhu, Ken Q. Pu, and Renée J. Miller · 2018
Cited alongside, same era.
Ray: A Distributed Framework for Emerging AI Applications, September 2018
RPT: relational pre-trained transformer is almost all you need towards democratizing data preparation
Nan Tang, Ju Fan, Fangyi Li, Jianhong Tu, Xiaoyong Du, Guoliang Li, Sam Madden, and Mourad Ouzzani · 2021
Later among the works it cites.
TABBIE: Pretrained Representations of Tabular Data, May 2021
Hiroshi Iida, Dung Thai, Varun Manjunatha, and Mohit Iyyer · 2021
Later among the works it cites.
Training Compute-Optimal Large Language Models, March 2022
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals, and Laurent Sifre · 2022
Later among the works it cites.
LAION-5B: An open large-scale dataset for training next generation image-text models, October 2022
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, Patrick Schramowski, Srivatsa Kundurthy, Katherine Crowson, Ludwig Schmidt, Robert Kaczmarczyk, and Jenia Jitsev · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Philipp Moritz, Robert Nishihara, Stephanie Wang, Alexey Tumanov, Richard Liaw, Eric Liang, Melih Elibol, Zongheng Yang, William Paul, Michael I. Jordan, and Ion Stoica · 2018
Cited alongside, same era.
Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, Zilin Zhang, and Dragomir Radev · 2019
Cited alongside, same era.
JOSIE: Overlap Set Similarity Search for Finding Joinable Tables in Data Lakes
Erkang Zhu, Dong Deng, Fatemeh Nargesian, and Renée J. Miller · 2019
Cited alongside, same era.
Survey on deep learning with class imbalance
Justin M. Johnson and Taghi M. Khoshgoftaar · 2019
Cited alongside, same era.
SemTab 2019: Resources to Benchmark Tabular Data to Knowledge Graph Matching Systems
Ernesto Jiménez-Ruiz, Oktie Hassanzadeh, Vasilis Efthymiou, Jiaoyan Chen, and Kavitha Srinivas · 2020
Cited alongside, same era.
TURL: table understanding through representation learning
Xiang Deng, Huan Sun, Alyssa Lees, You Wu, and Cong Yu · 2020
Cited alongside, same era.
The Pile: An 800GB Dataset of Diverse Text for Language Modeling, December 2020
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy · 2020
Cited alongside, same era.
Later among the works it cites.
High-Resolution Image Synthesis with Latent Diffusion Models, April 2022
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Later among the works it cites.
A Survey on Table Question Answering: Recent Advances, July 2022
Nengzheng Jin, Joanna Siebert, Dongfang Li, and Qingcai Chen · 2022
Later among the works it cites.
Haoyu Dong, Zhoujun Cheng, Xinyi He, Mengyu Zhou, Anda Zhou, Fan Zhou, Ao Liu, Shi Han, and Dongmei Zhang · 2022
Later among the works it cites.
From tabular data to knowledge graphs: A survey of semantic table interpretation tasks and methods
Jixiong Liu, Yoan Chabot, Raphaël Troncy, Viet-Phi Huynh, Thomas Labbé, and Pierre Monnin · 2022
Later among the works it cites.
Deduplicating Training Data Makes Language Models Better, March 2022
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini · 2022
Later among the works it cites.
Transformers for Tabular Data Representation: A Survey of Models and Applications
Gilbert Badaro, Mohammed Saeed, and Paolo Papotti · 2023
Closest in time.
GitTables: A Large-Scale Corpus of Relational Tables
Madelon Hulsebos, Çağatay Demiralp, and Paul Groth · 2023
Closest in time.
WikiDBs: A Corpus of Relational Databases From Wikidata
Liane Vogel and Carsten Binnig · 2023
Closest in time.
LakeBench: Benchmarks for Data Discovery over Data Lakes, July 2023
Kavitha Srinivas, Julian Dolby, Ibrahim Abdelaziz, Oktie Hassanzadeh, Harsha Kokel, Aamod Khatiwada, Tejaswini Pedapati, Subhajit Chaudhury, and Horst Samulowitz · 2023
Closest in time.
Is GPT-4 a Good Data Analyst?, May 2023
Liying Cheng, Xingxuan Li, and Lidong Bing · 2023
Closest in time.
Data-Copilot: Bridging Billions of Data and Humans with Autonomous Workflow, June 2023
Wenqi Zhang, Yongliang Shen, Weiming Lu, and Yueting Zhuang · 2023
Closest in time.
Jinyang Li, Binyuan Hui, Ge Qu, Binhua Li, Jiaxi Yang, Bowen Li, Bailin Wang, Bowen Qin, Rongyu Cao, Ruiying Geng, Nan Huo, Xuanhe Zhou, Chenhao Ma, Guoliang Li, Kevin C. C. Chang, Fei Huang, Reynold Cheng, and Yongbin Li · 2023
Closest in time.
DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-Correction, April 2023
Mohammadreza Pourreza and Davood Rafiei · 2023
Closest in time.
Column Type Annotation using ChatGPT, July 2023
Keti Korini and Christian Bizer · 2023
Closest in time.