Fetching the paper…
Reading the bibliography…
Tabular data is a pervasive modality spanning a wide range of domains, and the inherent diversity poses a considerable challenge for deep learning.
Locally weighted regression: An approach to regression analysis by local fitting
William S Cleveland and Susan J Devlin · 1988
Earlier work this paper cites.
The Elements of Statistical Learning: Data Mining, Inference, and Prediction , volume 2
Trevor Hastie, Robert Tibshirani, and Jerome H Friedman · 2009
Earlier work this paper cites.
The Nature of Statistical Learning Theory
Vladimir Vapnik · 2013
Earlier work this paper cites.
Semi-supervised sequence learning
Andrew M Dai and Quoc V Le · 2015
Earlier work this paper cites.
XGBoost: A scalable tree boosting system
Tianqi Chen and Carlos Guestrin · 2016
Earlier work this paper cites.
Generalized additive models
Trevor Hastie · 2017
Earlier work this paper cites.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder · 2018
Earlier work this paper cites.
Coresets-methods and history: A theoreticians design pattern for approximation and streaming algorithms
Alexander Munteanu and Chris Schwiegelshohn · 2018
Earlier work this paper cites.
CatBoost: Unbiased boosting with categorical features
Liudmila Prokhorenkova, Gleb Gusev, Aleksandr Vorobev, Anna Veronika Dorogush, and Andrey Gulin · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training, 2018
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Earlier work this paper cites.
Billion-scale similarity search with GPUs
Jeff Johnson, Matthijs Douze, and Hervé Jégou · 2019
Earlier work this paper cites.
Generalization through memorization: Nearest neighbor language models
Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Earlier work this paper cites.
Google dataset search by the numbers
Omar Benjelloun, Shiyu Chen, and Natasha Noy · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Sequential deep learning for credit risk monitoring with tabular financial data
Jillian M Clements, Di Xu, Nooshin Yousefi, and Dmitry Efimov · 2020
Earlier work this paper cites.
Retrieval augmented language model pre-training
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive NLP tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela · 2020
Earlier work this paper cites.
A customer churn prediction model based on XGBoost and MLP
Qi Tang, Guoen Xia, Xianquan Zhang, and Feng Long · 2020
Cited alongside, same era.
Trust issues: Uncertainty estimation does not enable reliable OOD detection on medical tabular data
Dennis Ulmer, Lotta Meijerink, and Giovanni Cinà · 2020
Cited alongside, same era.
VIME: Extending the success of self-and semi-supervised learning to tabular domain
Jinsung Yoon, Yao Zhang, James Jordon, and Mihaela Van der Schaar · 2020
Cited alongside, same era.
Deep reinforcement learning at the edge of the statistical precipice
Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C Courville, and Marc Bellemare · 2021
Cited alongside, same era.
Tabnet: Attentive interpretable tabular learning
Sercan Ö Arik and Tomas Pfister · 2021
Cited alongside, same era.
OpenML benchmarking suites
Bernd Bischl, Giuseppe Casalicchio, Matthias Feurer, Pieter Gijsbers, Frank Hutter, Michel Lang, Rafael Gomes Mantovani, Jan van Rijn, and Joaquin Vanschoren · 2021
DNNR: Differential nearest neighbors regression
Youssef Nader, Leon Sixt, and Tim Landgraf · 2022
Later among the works it cites.
Tabular data: Deep learning is not all you need
Ravid Shwartz-Ziv and Amitai Armon · 2022
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou · 2022
Later among the works it cites.
An inductive bias for tabular deep learning
Ege Beyazit, Jonathan Kozaczuk, Bo Li, Vanessa Wallace, and Bilal Fadlallah · 2023
Later among the works it cites.
Fine-tuning the retrieval mechanism for tabular deep learning
Felix den Breejen, Sangmin Bae, Stephen Cha, Tae-Young Kim, Seoung Hyun Koh, and Se-Young Yun · 2023
Later among the works it cites.
TabLLM: Few-shot classification of tabular data with large language models
Stefan Hegselmann, Alejandro Buendia, Hunter Lang, Monica Agrawal, Xiaoyi Jiang, and David Sontag · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Cited alongside, same era.
Self-attention between datapoints: Going beyond individual input-output pairs in deep learning
Jannik Kossen, Neil Band, Clare Lyle, Aidan Gomez, Tom Rainforth, and Yarin Gal · 2021
Cited alongside, same era.
Show your work: Scratchpads for intermediate computation with language models
Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, Charles Sutton, and Augustus Odena · 2021
Cited alongside, same era.
Retrieval & interaction machine for tabular data prediction
Jiarui Qin, Weinan Zhang, Rong Su, Zhirong Liu, Weiwen Liu, Ruiming Tang, Xiuqiang He, and Yong Yu · 2021
Cited alongside, same era.
SAINT: Improved neural networks for tabular data via row attention and contrastive pre-training
Gowthami Somepalli, Micah Goldblum, Avi Schwarzschild, C Bayan Bruss, and Tom Goldstein · 2021
Cited alongside, same era.
Deep learning: A primer for psychologists
Christopher J Urban and Kathleen M Gates · 2021
Cited alongside, same era.
Later among the works it cites.
TabPFN: A transformer that solves small tabular classification problems in a second
Noah Hollmann, Samuel Müller, Katharina Eggensperger, and Frank Hutter · 2023
Later among the works it cites.
TabPFGen – Tabular data generation with TabPFN
Junwei Ma, Apoorv Dankar, George Stein, Guangwei Yu, and Anthony Caterini · 2023
Later among the works it cites.
When do neural nets outperform boosted trees on tabular data?
Duncan McElfresh, Sujay Khandagale, Jonathan Valverde, Vishak Prasad C, Ganesh Ramakrishnan, Micah Goldblum, and Colin White · 2023
Later among the works it cites.
Elephants never forget: Testing language models for memorization of tabular data
Sebastian Bordt, Harsha Nori, and Rich Caruana · 2024
Closest in time.
The Faiss library
Matthijs Douze, Alexandr Guzhva, Chengqi Deng, Jeff Johnson, Gergely Szilvasy, Pierre-Emmanuel Mazaré, Maria Lomeli, Lucas Hosseini, and Hervé Jégou · 2024
Closest in time.
Large language models (LLMs) on tabular data: Prediction, generation, and understanding – A survey
Xi Fang, Weijie Xu, Fiona Anting Tan, Jiani Zhang, Ziqing Hu, Yanjun Qi, Scott Nickleach, Diego Socolinsky, Srinivasan Sengamedu, and Christos Faloutsos · 2024
Closest in time.
TuneTables: Context optimization for scalable prior-data fitted networks
Benjamin Feuer, Robin Tibor Schirrmeister, Valeriia Cherepanova, Chinmay Hegde, Frank Hutter, Micah Goldblum, Niv Cohen, and Colin White · 2024
Closest in time.
TabR: Tabular deep learning meets nearest neighbors
Yury Gorishniy, Ivan Rubachev, Nikolay Kartashev, Daniil Shlenskii, Akim Kotelnikov, and Artem Babenko · 2024
Closest in time.
In-context data distillation with TabPFN
Junwei Ma, Valentin Thomas, Guangwei Yu, and Anthony Caterini · 2024
Closest in time.
Interpretable machine learning for TabPFN
David Rundel, Julius Kobialka, Constantin von Crailsheim, Matthias Feurer, Thomas Nagler, and David Rügamer · 2024
Closest in time.
Why tabular foundation models should be a research priority
Boris van Breugel and Mihaela van der Schaar · 2024
Closest in time.