Fetching the paper…
Reading the bibliography…
Tabular data is foundational to predictive modeling in various crucial industries, including healthcare, finance, retail, sustainability, etc.
Greedy function approximation: a gradient boosting machine
Jerome H Friedman. 2001 · 2001
Earlier work this paper cites.
Data Science for Business: What you need to know about data mining and data-analytic thinking
Foster Provost and Tom Fawcett. 2013 · 2013
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. 2015 · 2015
Earlier work this paper cites.
XGBoost: A scalable tree boosting system. In KDD
Tianqi Chen and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
LightGBM: A highly efficient gradient boosting decision tree. In Advances in neural information processing systems
Guolin Ke, Qi Meng, Thomas Finley, Taifeng Wang, Wei Chen, Weidong Ma, Qiwei Ye, and Tie-Yan Liu. 2017 · 2017
Earlier work this paper cites.
Decoupled Weight Decay Regularization. In International Conference on Learning Representations
Ilya Loshchilov and Frank Hutter. 2018 · 2018
Earlier work this paper cites.
CatBoost: unbiased boosting with categorical features. In NeurIPS
Liudmila Prokhorenkova, Gleb Gusev, Aleksandr Vorobev, Anna Veronika Dorogush, and Andrey Gulin. 2018 · 2018
Earlier work this paper cites.
Language Models are Few-Shot Learners. In NeurIPS
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Earlier work this paper cites.
TabTransformer: Tabular data modeling using contextual embeddings
Xin Huang, Ashish Khetan, Milan Cvitkovic, and Zohar Karnin. 2020 · 2020
Earlier work this paper cites.
Fundamentals of machine learning for predictive data analytics: algorithms, worked examples, and case studies
John D Kelleher, Brian Mac Namee, and Aoife D’arcy. 2020 · 2020
Earlier work this paper cites.
VIME: Extending the success of self- and semi-supervised learning to tabular domain. In NeurIPS
Jinsung Yoon, Yao Zhang, James Jordon, and Mihaela van der Schaar. 2020 · 2020
Earlier work this paper cites.
TabNet: Attentive interpretable tabular learning. In AAAI
Sercan Ö Arik and Tomas Pfister. 2021 · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
Revisiting deep learning models for tabular data. In NeurIPS
Yury Gorishniy, Ivan Rubachev, Valentin Khrulkov, and Artem Babenko. 2021 · 2021
Earlier work this paper cites.
Net-DNF: Effective deep modeling of tabular data. In ICLR
Liran Katzir, Gal Elidan, and Ran El-Yaniv. 2021 · 2021
Earlier work this paper cites.
Transformers can do bayesian inference
Samuel Müller, Noah Hollmann, Sebastian Pineda Arango, Josif Grabocka, and Frank Hutter. 2021 · 2021
Earlier work this paper cites.
SAINT: Improved neural networks for tabular data via row attention and contrastive pre-training
Gowthami Somepalli, Micah Goldblum, Avi Schwarzschild, C Bayan Bruss, and Tom Goldstein. 2021 · 2021
Cited alongside, same era.
SubTab: Subsetting features of tabular data for self-supervised representation learning. In NeurIPS
Talip Ucar, Ehsan Hajiramezanali, and Lindsay Edwards. 2021 · 2021
Cited alongside, same era.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Ben Wang and Aran Komatsuzaki. 2021 · 2021
Cited alongside, same era.
SCARF: Self-Supervised Contrastive Learning using Random Feature Corruption. In ICLR
Dara Bahri, Heinrich Jiang, Yi Tay, and Donald Metzler. 2022 · 2022
Cited alongside, same era.
LIFT: Language-interfaced fine-tuning for non-language machine learning tasks. In NeurIPS
Tuan Dinh, Yuchen Zeng, Ruisu Zhang, Ziqian Lin, Michael Gira, Shashank Rajput, Jy-yong Sohn, Dimitris Papailiopoulos, and Kangwook Lee. 2022 · 2022
Cited alongside, same era.
Language is not all you need: Aligning perception with language models
Shaohan Huang, Li Dong, Wenhui Wang, Yaru Hao, Saksham Singhal, Shuming Ma, Tengchao Lv, Lei Cui, Owais Khan Mohammed, Qiang Liu, et al · 2023
Closest in time.
Transfer Learning with Deep Tabular Models. In ICLR
Roman Levin, Valeriia Cherepanova, Avi Schwarzschild, Arpit Bansal, C. Bayan Bruss, Tom Goldstein, Andrew Gordon Wilson, and Micah Goldblum. 2023 · 2023
Closest in time.
Augmented Language Models: a Survey
Grégoire Mialon, Roberto Dessi, Maria Lomeli, Christoforos Nalmpantis, Ramakanth Pasunuru, Roberta Raileanu, Baptiste Roziere, Timo Schick, Jane Dwivedi-Yu, Asli Celikyilmaz, Edouard Grave, Yann LeCun, and Thomas Scialom. 2023 · 2023
Closest in time.
STUNT: Few-shot Tabular Learning with Self-generated Tasks from Unlabeled Tables. In ICLR
Jaehyun Nam, Jihoon Tack, Kyungmin Lee, Hankook Lee, and Jinwoo Shin. 2023 · 2023
Closest in time.
OpenAI. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Why do tree-based models still outperform deep learning on typical tabular data?. In NeurIPS
Léo Grinsztajn, Edouard Oyallon, and Gaël Varoquaux. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback. In NeurIPS
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
Multitask Prompted Training Enables Zero-Shot Task Generalization. In ICLR
Victor Sanh, Albert Webson, Colin Raffel, Stephen Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Teven Le Scao, Stella Biderman, Leo Gao, Thomas Wolf, and Alexander M Rush. 2022 · 2022
Cited alongside, same era.
Tabular data: Deep learning is not all you need
Ravid Shwartz-Ziv and Amitai Armon. 2022 · 2022
Cited alongside, same era.
TransTab: Learning transferable tabular transformers across tables. In NeurIPS
Zifeng Wang and Jimeng Sun. 2022 · 2022
Cited alongside, same era.
An Explanation of In-context Learning as Implicit Bayesian Inference. In ICLR
Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma. 2022 · 2022
Cited alongside, same era.
Elephants Never Forget: Testing Language Models for Memorization of Tabular Data. In NeurIPS 2023 Second Table Representation Learning Workshop
Sebastian Bordt, Harsha Nori, and Rich Caruana. 2023 · 2023
Cited alongside, same era.
Closest in time.
Toolformer: Language models can teach themselves to use tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023 · 2023
Closest in time.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Closest in time.
MediTab: Scaling Medical Tabular Data Predictors via Data Consolidation, Enrichment, and Refinement
Zifeng Wang, Chufan Gao, Cao Xiao, and Jimeng Sun. 2023 · 2023
Closest in time.
ReAct: Synergizing Reasoning and Acting in Language Models. In ICLR
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao. 2023 · 2023
Closest in time.
Generative Table Pre-training Empowers Models for Tabular Prediction. In EMNLP
Tianping Zhang, Shaowen Wang, Shuicheng Yan, Li Jian, and Qian Liu. 2023 · 2023
Closest in time.
XTab: Cross-table Pretraining for Tabular Transformers. In ICML
Bingzhao Zhu, Xingjian Shi, Nick Erickson, Mu Li, George Karypis, and Mahsa Shoaran. 2023 · 2023
Closest in time.
Making Pre-trained Language Models Great on Tabular Prediction. In ICLR
Anonymous. 2024 · 2024
Closest in time.
UniPredict: Large Language Models are Universal Tabular Classifiers
Ruiyu Wang, Zifeng Wang, and Jimeng Sun. 2024 · 2024
Closest in time.
UniTabE: A Universal Pretraining Protocol for Tabular Foundation Model in Data Science. In ICLR
Yazheng Yang, Yuqi Wang, Guang Liu, Ledell Wu, and Qi Liu. 2024 · 2024
Closest in time.