Fetching the paper…
Reading the bibliography…
Synthetic data generation is integral to ML pipelines, e.g., to augment training data, replace sensitive information, and even to power advanced platforms like DeepSeek.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Database Systems: The Complete Book (2 ed.)
Hector Garcia-Molina, Jeffrey D. Ullman, and Jennifer Widom. 2008 · 2008
Earlier work this paper cites.
Probabilistic graphical models: principles and techniques
Daphne Koller and Nir Friedman. 2009 · 2009
Earlier work this paper cites.
Input-output analysis: foundations and extensions
R. E. Miller and P. D. Blair. 2009 · 2009
Earlier work this paper cites.
States Shapefile
Dominique Evans-Bye. 2015 · 2015
Earlier work this paper cites.
Deep feature synthesis: Towards automating data science endeavors. In 2015 IEEE international conference on data science and advanced analytics (DSAA) . IEEE, 1–10
James Max Kanter and Kalyan Veeramachaneni. 2015 · 2015
Earlier work this paper cites.
Functional dependency discovery: An experimental evaluation of seven algorithms
Thorsten Papenbrock, Jens Ehrlich, Jannik Marten, Tommy Neubert, Jan-Peer Rudolph, Martin Schönberg, Jakob Zwiener, and Felix Naumann. 2015 · 2015
Earlier work this paper cites.
A hybrid approach to functional dependency discovery. In Proceedings of the 2016 International Conference on Management of Data . 821–833
Thorsten Papenbrock and Felix Naumann. 2016 · 2016
Earlier work this paper cites.
The Synthetic data vault. In IEEE International Conference on Data Science and Advanced Analytics (DSAA) . 399–410
Neha Patki, Roy Wedge, and Kalyan Veeramachaneni. 2016 · 2016
Earlier work this paper cites.
Beijing PM2.5 Data
Song Chen. 2017 · 2017
Earlier work this paper cites.
California Housing Prices
Cam Nugent. 2017 · 2017
Earlier work this paper cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
GTR-LSTM: A triple encoder for sentence generation from RDF data. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 1627–1637
Bayu Distiawan, Jianzhong Qi, Rui Zhang, and Wei Wang. 2018 · 2018
Earlier work this paper cites.
illiad: Intelligent invariant and anomaly detection in cyber-physical systems
Nikhil Muralidhar, Chen Wang, Nathan Self, Marjan Momtazpour, Kiyoshi Nakayama, Ratnesh Sharma, and Naren Ramakrishnan. 2018 · 2018
Earlier work this paper cites.
Data synthesis based on generative adversarial networks
Noseong Park, Mahmoud Mohammadi, Kshitij Gorde, Sushil Jajodia, Hongkyu Park, and Youngmin Kim. 2018 · 2018
Earlier work this paper cites.
Synthesizing tabular data using generative adversarial networks
Lei Xu and Kalyan Veeramachaneni. 2018 · 2018
Earlier work this paper cites.
Graphical-model based estimation and inference for differential privacy. In International Conference on Machine Learning . PMLR, 4435–4444
Ryan McKenna, Daniel Sheldon, and Gerome Miklau. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Enhancing AMR-to-Text Generation with Dual Graph Representations. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) . 3183–3194
Leonardo FR Ribeiro, Claire Gardent, and Iryna Gurevych. 2019 · 2019
Earlier work this paper cites.
DataShot: Automatic Generation of Fact Sheets from Tabular Data
Yun Wang, Zhida Sun, Haidong Zhang, Weiwei Cui, Ke Xu, Xiaojuan Ma, and Dongmei Zhang. 2020 · 2019
Earlier work this paper cites.
Modeling tabular data using conditional gan
Lei Xu, Maria Skoularidou, Alfredo Cuesta-Infante, and Kalyan Veeramachaneni. 2019 · 2019
Cited alongside, same era.
Supervised learning on relational databases with graph neural networks
Milan Cvitkovic. 2020 · 2020
Cited alongside, same era.
Using gans for sharing networked time series data: Challenges, initial promise, and open questions. In Proceedings of the ACM Internet Measurement Conference . 464–483
Zinan Lin, Alankar Jain, Chen Wang, Giulia Fanti, and Vyas Sekar. 2020 · 2020
Cited alongside, same era.
Discovering functional dependencies from mixed-type data. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining . 1404–1414
Panagiotis Mandros, David Kaltenpoth, Mario Boley, and Jilles Vreeken. 2020 · 2020
Cited alongside, same era.
Discovering approximate functional dependencies using smoothed mutual information. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining . 1254–1264
Permutation-Invariant Tabular Data Synthesis. In 2022 IEEE International Conference on Big Data (Big Data) . IEEE, 5855–5864
Yujin Zhu, Zilong Zhao, Robert Birke, and Lydia Y Chen. 2022 · 2022
Later among the works it cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Later among the works it cites.
Graph Deep Factors for Probabilistic Time-series Forecasting
Hongjie Chen, Ryan A Rossi, Kanak Mahadik, Sungchul Kim, and Hoda Eldardiry. 2023 · 2023
Later among the works it cites.
Mathematical discoveries from program search with large language models
Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, Matej Balog, M Pawan Kumar, Emilien Dupont, Francisco JR Ruiz, Jordan S Ellenberg, Pengming Wang, Omar Fawzi, et al · 2023
Later among the works it cites.
Realtabformer: Generating realistic relational and tabular data using transformers
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Frédéric Pennerath, Panagiotis Mandros, and Jilles Vreeken. 2020 · 2020
Cited alongside, same era.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, Sharan S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu. 2020 · 2020
Cited alongside, same era.
A statistical perspective on discovering functional dependencies in noisy data. In Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data . 861–876
Yunjia Zhang, Zhihan Guo, and Theodoros Rekatsinas. 2020 · 2020
Cited alongside, same era.
TabularNet: A neural network architecture for understanding semantic structures of tabular data. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining . 322–331
Lun Du, Fei Gao, Xu Chen, Ran Jia, Junshan Wang, Jiang Zhang, Shi Han, and Dongmei Zhang. 2021 · 2021
Cited alongside, same era.
LoRA: Low-Rank Adaptation of Large Language Models
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen. 2021 · 2021
Cited alongside, same era.
Tour & Travels Customer Churn Prediction
Tejashvi. 2021 · 2021
Cited alongside, same era.
Tuta: Tree-based transformers for generally structured table pre-training. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining . 1780–1790
Zhiruo Wang, Haoyu Dong, Ran Jia, Jia Li, Zhiyi Fu, Shi Han, and Dongmei Zhang. 2021 · 2021
Cited alongside, same era.
Stan: Synthetic network traffic generation with generative neural models. In Deployable Machine Learning for Security Defense: Second International Workshop, MLHat 2021, Virtual Event, August 15, 2021, Proceedings 2 . Springer, 3–29
Shengzhe Xu, Manish Marwah, Martin Arlitt, and Naren Ramakrishnan. 2021 · 2021
Cited alongside, same era.
Aivin V Solatorio and Olivier Dupriez. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
Mixed-Type Tabular Data Synthesis with Score-based Diffusion in Latent Space
Hengrui Zhang, Jiani Zhang, Balasubramaniam Srinivasan, Zhengyuan Shen, Xiao Qin, Christos Faloutsos, Huzefa Rangwala, and George Karypis. 2023b · 2023
Later among the works it cites.
Generative Table Pre-training Empowers Models for Tabular Prediction. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . 14836–14854
Tianping Zhang, Shaowen Wang, Shuicheng Yan, Li Jian, and Qian Liu. 2023a · 2023
Later among the works it cites.
Tabula: Harnessing language models for tabular data synthesis
Zilong Zhao, Robert Birke, and Lydia Chen. 2023 · 2023
Later among the works it cites.
Dense Representation Learning and Retrieval for Tabular Data Prediction. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 3559–3569
Lei Zheng, Ning Li, Xianyu Chen, Quan Gan, and Weinan Zhang. 2023 · 2023
Later among the works it cites.
Premise Order Matters in Reasoning with Large Language Models
Xinyun Chen, Ryan A Chi, Xuezhi Wang, and Denny Zhou. 2024 · 2024
Closest in time.
Towards principled assessment of tabular data synthesis algorithms
Yuntao Du and Ninghui Li. 2024 · 2024
Closest in time.
Large language models (LLMs) on tabular data: Prediction, generation, and understanding — a survey
Xi Fang, Weijie Xu, Fiona Anting Tan, Jiani Zhang, Ziqing Hu, Yanjun (Jane) Qi, Scott Nickleach, Diego Socolinsky, "SHS" Srinivasan Sengamedu, and Christos Faloutsos. 2024 · 2024
Closest in time.
Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al · 2024
Closest in time.
Why think step by step? Reasoning emerges from the locality of experience
Ben Prystawski, Michael Li, and Noah Goodman. 2024 · 2024
Closest in time.
Curated LLM: Synergy of LLMs and Data Curation for tabular augmentation in low-data regimes. In Forty-first International Conference on Machine Learning
Nabeel Seedat, Nicolas Huynh, Boris van Breugel, and Mihaela van der Schaar. 2024 · 2024
Closest in time.
Neural Methods for Data-to-text Generation
Mandar Sharma, Ajay Kumar Gogineni, and Naren Ramakrishnan. 2024 · 2024
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al · 2025
Closest in time.
FakeTables: Using GANs to Generate Functional Dependency Preserving Tables with Bounded Real Data.. In IJCAI . 2074–2080
Haipeng Chen, Sushil Jajodia, Jing Liu, Noseong Park, Vadim Sokolov, and VS Subrahmanian. 2019 · 2080
Closest in time.