Fetching the paper…
Reading the bibliography…
Data serves as the fundamental foundation for advancing deep learning, particularly tabular data presented in a structured format, which is highly conducive to modeling.
Nearest neighbor pattern classification
Thomas Cover and Peter Hart · 1967
Earlier work this paper cites.
Algorithms for finding patterns in strings, handbook of theoretical computer science vol a
Alfred V Aho and AJ van Leeuwen · 1990
Earlier work this paper cites.
Statlog (German Credit Data)
Hans Hofmann · 1994
Earlier work this paper cites.
Scaling up the accuracy of naive-bayes classifiers: A decision-tree hybrid
Ron Kohavi et al · 1996
Earlier work this paper cites.
Smote: synthetic minority over-sampling technique
Nitesh V Chawla, Kevin W Bowyer, Lawrence O Hall, and W Philip Kegelmeyer · 2002
Earlier work this paper cites.
Airline passenger name record generation using generative adversarial networks
Alejandro Mottini, Alix Lheritier, and Rodrigo Acuna-Agost · 2018
Earlier work this paper cites.
Catboost: unbiased boosting with categorical features
Liudmila Prokhorenkova, Gleb Gusev, Aleksandr Vorobev, Anna Veronika Dorogush, and Andrey Gulin · 2018
Earlier work this paper cites.
Modeling tabular data using conditional gan
Lei Xu, Maria Skoularidou, Alfredo Cuesta-Infante, and Kalyan Veeramachaneni · 2019
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 2019
Earlier work this paper cites.
Synthetic data in machine learning for medicine and healthcare
Richard J Chen, Ming Y Lu, Tiffany Y Chen, Drew FK Williamson, and Faisal Mahmood · 2021
Earlier work this paper cites.
Ctab-gan: Effective table data synthesizing
Zilong Zhao, Aditya Kunar, Robert Birke, and Lydia Y Chen · 2021
Earlier work this paper cites.
Deep neural networks and tabular data: A survey. arxiv 2021
V Borisov, T Leemann, K Seßler, J Haug, M Pawelczyk, and G Kasneci · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Cited alongside, same era.
Customs import declaration datasets
Chaeyoon Jeong, Sundong Kim, Jaewoo Park, and Yeonsoo Choi · 2022
Cited alongside, same era.
Causal-tgan: Modeling tabular data using causally-aware gan
Bingyang Wen, Yupeng Cao, Fan Yang, Koduvayur Subbalakshmi, and Rajarathnam Chandramouli · 2022
Cited alongside, same era.
Language models are realistic tabular data generators
Vadim Borisov, Kathrin Seßler, Tobias Leemann, Martin Pawelczyk, and Gjergji Kasneci · 2022
Cited alongside, same era.
Tabula: Harnessing language models for tabular data synthesis
Skywork: A more open bilingual foundation model
Tianwen Wei, Liang Zhao, Lichang Zhang, Bo Zhu, Lijie Wang, Haihua Yang, Biye Li, Cheng Cheng, Weiwei Lü, Rui Hu, et al · 2023
Later among the works it cites.
Mistral 7b, 2023
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2023
Later among the works it cites.
Large language models on tabular data–a survey
Xi Fang, Weijie Xu, Fiona Anting Tan, Jiani Zhang, Ziqing Hu, Yanjun Qi, Scott Nickleach, Diego Socolinsky, Srinivasan Sengamedu, and Christos Faloutsos · 2024
Closest in time.
Large language model for table processing: A survey, 2024
Weizheng Lu, Jiaming Zhang, Jing Zhang, and Yueguo Chen · 2024
Closest in time.
Scaling while privacy preserving: A comprehensive synthetic tabular data generation and evaluation in learning analytics
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zilong Zhao, Robert Birke, and Lydia Chen · 2023
Cited alongside, same era.
Tabddpm: Modelling tabular data with diffusion models
Akim Kotelnikov, Dmitry Baranchuk, Ivan Rubachev, and Artem Babenko · 2023
Cited alongside, same era.
Mixed-type tabular data synthesis with score-based diffusion in latent space
Hengrui Zhang, Jiani Zhang, Balasubramaniam Srinivasan, Zhengyuan Shen, Xiao Qin, Christos Faloutsos, Huzefa Rangwala, and George Karypis · 2023
Cited alongside, same era.
Findiff: Diffusion models for financial tabular data generation
Timur Sattarov, Marco Schreyer, and Damian Borth · 2023
Cited alongside, same era.
Realtabformer: Generating realistic relational and tabular data using transformers
Aivin V Solatorio and Olivier Dupriez · 2023
Cited alongside, same era.
Empowering many, biasing a few: Generalist credit scoring through large language models
Duanyu Feng, Yongfu Dai, Jimin Huang, Yifang Zhang, Qianqian Xie, Weiguang Han, Alejandro Lopez-Lira, and Hao Wang · 2023
Cited alongside, same era.
Synthetic Data Metrics
DataCebo, Inc · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Cited alongside, same era.
Qinyi Liu, Mohammad Khalil, Jelena Jovanovic, and Ronas Shakya · 2024
Closest in time.
Dp-tabicl: In-context learning with differentially private tabular data
Alycia N Carey, Karuna Bhaila, Kennedy Edemacu, and Xintao Wu · 2024
Closest in time.
An improved tabular data generator with vae-gmm integration
Patricia A Apellániz, Juan Parras, and Santiago Zazo · 2024
Closest in time.
Controllable tabular data synthesis using diffusion models
Tongyu Liu, Ju Fan, Nan Tang, Guoliang Li, and Xiaoyong Du · 2024
Closest in time.
On protecting the data privacy of large language models (llms): A survey
Biwei Yan, Kun Li, Minghui Xu, Yueyan Dong, Yue Zhang, Zhaochun Ren, and Xiuzheng Cheng · 2024
Closest in time.
Initial exploration of zero-shot privacy utility tradeoffs in tabular data using gpt-4
Bishwas Mandal, George Amariucai, and Shuangqing Wei · 2024
Closest in time.
Pandora’s white-box: Increased training data leakage in open llms
Jeffrey G Wang, Jason Wang, Marvin Li, and Seth Neel · 2024
Closest in time.
Making pre-trained language models great on tabular prediction
Jiahuan Yan, Bo Zheng, Hongxia Xu, Yiheng Zhu, Danny Chen, Jimeng Sun, Jian Wu, and Jintai Chen · 2024
Closest in time.