Fetching the paper…
Reading the bibliography…
We present an Adversarially Pre-trained Transformer (APT) that is able to perform zero-shot meta-learning on tabular prediction tasks without pre-training on any real-world dataset, extending on the recent development of Prior-Data Fitted Networks (PFNs) and TabPFN.
Solution of incorrectly formulated problems and the regularization method
Tikhonov, A. N · 1963
Earlier work this paper cites.
Nearest neighbor pattern classification
Cover, T. and Hart, P · 1967
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Hornik, K., Stinchcombe, M., and White, H · 1989
Earlier work this paper cites.
Support-vector networks
Cortes, C · 1995
Earlier work this paper cites.
Random decision forests
Ho, T. K · 1995
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
Tibshirani, R · 1996
Earlier work this paper cites.
Importance of semantic representation: Dataless classification
Chang, M.-W., Ratinov, L.-A., Roth, D., and Srikumar, V · 2008
Earlier work this paper cites.
Zero-data learning of new tasks
Larochelle, H., Erhan, D., and Bengio, Y · 2008
Earlier work this paper cites.
Zero-shot learning with semantic output codes
Palatucci, M., Pomerleau, D., Hinton, G. E., and Mitchell, T. M · 2009
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2010
Earlier work this paper cites.
Tabtransformer: Tabular data modeling using contextual embeddings, 2020
Huang, X., Khetan, A., Cvitkovic, M., and Karnin, Z · 2012
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Earlier work this paper cites.
Efficient and robust automated machine learning
Feurer, M., Klein, A., Eggensperger, K., Springenberg, J., Blum, M., and Hutter, F · 2015
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Goodfellow, I., Shlens, J., and Szegedy, C · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S · 2015
Earlier work this paper cites.
Metalearning: a survey of trends and technologies
Lemke, C., Budka, M., and Gabrys, B · 2015
Earlier work this paper cites.
Xgboost: A scalable tree boosting system
Chen, T. and Guestrin, C · 2016
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Jang, E., Gu, S., and Poole, B · 2016
Earlier work this paper cites.
The concrete distribution: A continuous relaxation of discrete random variables
Maddison, C. J., Mnih, A., and Teh, Y. W · 2016
Earlier work this paper cites.
On the characterization of local nash equilibria in continuous games
Ratliff, L. J., Burden, S. A., and Sastry, S. S · 2016
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Earlier work this paper cites.
Lightgbm: A highly efficient gradient boosting decision tree
Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., Ye, Q., and Liu, T.-Y · 2017
Earlier work this paper cites.
Adversarial machine learning at scale
Kurakin, A., Goodfellow, I. J., and Bengio, S · 2017
Earlier work this paper cites.
Zero-shot learning - the good, the bad and the ugly
Xian, Y., Schiele, B., and Akata, Z · 2017
Earlier work this paper cites.
Notes from the ai frontier: Insights from hundreds of use cases
Chui, M., Manyika, J., Miremadi, M., Henke, N., Chung, R., Nel, P., and Malhotra, S · 2018
Earlier work this paper cites.
Neural architecture optimization
Luo, R., Tian, F., Qin, T., Chen, E., and Liu, T.-Y · 2018
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A · 2018
Earlier work this paper cites.
Reptile: a scalable metalearning algorithm
Nichol, A. and Schulman, J · 2018
Earlier work this paper cites.
Catboost: unbiased boosting with categorical features
Prokhorenkova, L., Gusev, G., Vorobev, A., Dorogush, A. V., and Gulin, A · 2018
Earlier work this paper cites.
Vanschoren, J · 2018
Cited alongside, same era.
Zero-shot learning—a comprehensive evaluation of the good, the bad and the ugly
Xian, Y., Lampert, C. H., Schiele, B., and Akata, Z · 2018
Cited alongside, same era.
Neural oblivious decision ensembles for deep learning on tabular data
Popov, S., Morozov, S., and Babenko, A · 2019
Cited alongside, same era.
Adversarial training for free!
Shafahi, A., Najibi, M., Ghiasi, M. A., Xu, Z., Dickerson, J., Studer, C., Davis, L. S., Taylor, G., and Goldstein, T · 2019
Cited alongside, same era.
You only propagate once: Accelerating adversarial training via maximal principle
Zhang, D., Zhang, T., Lu, Y., Zhu, Z., and Dong, B · 2019
Cited alongside, same era.
Auto-sklearn 2.0: Hands-free automl via meta-learning
Feurer, M., Eggensperger, K., Falkner, S., Lindauer, M., and Hutter, F · 2022
Later among the works it cites.
On embeddings for numerical features in tabular deep learning
Gorishniy, Y., Rubachev, I., and Babenko, A · 2022
Later among the works it cites.
Why do tree-based models still outperform deep learning on typical tabular data?
Grinsztajn, L., Oyallon, E., and Varoquaux, G · 2022
Later among the works it cites.
Tabpfn: A transformer that solves small tabular classification problems in a second
Hollmann, N., Müller, S., Eggensperger, K., and Hutter, F · 2022
Later among the works it cites.
Transfer learning with deep tabular models
Levin, R., Cherepanova, V., Schwarzschild, A., Bansal, A., Bruss, C. B., Goldstein, T., Wilson, A. G., and Goldblum, M · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Understanding and improving fast adversarial training
Andriushchenko, M. and Flammarion, N · 2020
Cited alongside, same era.
Language models are few-shot learners
Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., et al · 2020
Cited alongside, same era.
Meta-learning for generalized zero-shot learning
Verma, V. K., Brahma, D., and Rai, P · 2020
Cited alongside, same era.
Fast is better than free: Revisiting adversarial training
Wong, E., Rice, L., and Kolter, J. Z · 2020
Cited alongside, same era.
Tabnet: Attentive interpretable tabular learning
Arik, S. Ö. and Pfister, T · 2021
Cited alongside, same era.
Openml benchmarking suites
Bischl, B., Casalicchio, G., Feurer, M., Gijsbers, P., Hutter, F., Lang, M., Mantovani, R. G., van Rijn, J. N., and Vanschoren, J · 2021
Cited alongside, same era.
Population-based evolution optimizes a meta-learning objective
Frans, K. and Witkowski, O · 2021
Cited alongside, same era.
Revisiting pretraining objectives for tabular deep learning
Rubachev, I., Alekberov, A., Gorishniy, Y., and Babenko, A · 2022
Later among the works it cites.
Tabular data: Deep learning is not all you need
Shwartz-Ziv, R. and Armon, A · 2022
Later among the works it cites.
Memorizing transformers
Wu, Y., Rabe, M. N., Hutchins, D., and Szegedy, C · 2022
Later among the works it cites.
Language models are realistic tabular data generators
Borisov, V., Sessler, K., Leemann, T., Pawelczyk, M., and Kasneci, G · 2023
Later among the works it cites.
Scaling transformer to 1m tokens and beyond with RMT
Bulatov, A., Kuratov, Y., and Burtsev, M. S · 2023
Later among the works it cites.
Openml-ctr23–a curated tabular regression benchmarking suite
Fischer, S. F., Feurer, M., and Bischl, B · 2023
Later among the works it cites.
Tabllm: Few-shot classification of tabular data with large language models
Hegselmann, S., Buendia, A., Lang, H., Agrawal, M., Jiang, X., and Sontag, D · 2023
Later among the works it cites.
TabDDPM: Modelling tabular data with diffusion models, 2023
Kotelnikov, A., Baranchuk, D., Rubachev, I., and Babenko, A · 2023
Later among the works it cites.
Statistical foundations of prior-data fitted networks
Nagler, T · 2023
Later among the works it cites.
STUNT: Few-shot tabular learning with self-generated tasks from unlabeled tables
Nam, J., Tack, J., Lee, K., Lee, H., and Shin, J · 2023
Later among the works it cites.
Xtab: Cross-table pretraining for tabular transformers
Zhu, B., Shi, X., Erickson, N., Li, M., Karypis, G., and Shoaran, M · 2023
Later among the works it cites.
Llms are few-shot in-context low-resource language learners
Cahyawijaya, S., Lovenia, H., and Fung, P · 2024
Later among the works it cites.
Can a deep learning model be a sure bet for tabular prediction?
Chen, J., Yan, J., Chen, Q., Chen, D. Z., Wu, J., and Sun, J · 2024
Later among the works it cites.
Large scale transfer learning for tabular data via language modeling, 2024
Gardner, J., Perdomo, J. C., and Schmidt, L · 2024
Later among the works it cites.
Tabr: Tabular deep learning meets nearest neighbors
Gorishniy, Y., Rubachev, I., Kartashev, N., Shlenskii, D., Kotelnikov, A., and Babenko, A · 2024
Later among the works it cites.
Drift-resilient tabPFN: In-context learning distribution shifts on tabular data
Helli, K., Schnurr, D., Hollmann, N., Müller, S., and Hutter, F · 2024
Later among the works it cites.
CARTE: Pretraining and transfer for tabular learning
Kim, M. J., Grinsztajn, L., and Varoquaux, G · 2024
Later among the works it cites.
When do neural nets outperform boosted trees on tabular data?
McElfresh, D., Khandagale, S., Valverde, J., Prasad C, V., Ramakrishnan, G., Goldblum, M., and White, C · 2024
Later among the works it cites.
PORTAL: Scalable tabular foundation models via content-specific tokenization
Spinaci, M., Polewczyk, M., Hoffart, J., Kohler, M. C., Thelin, S., and Klein, T · 2024
Later among the works it cites.
Making pre-trained language models great on tabular prediction
Yan, J., Zheng, B., Xu, H., Zhu, Y., Chen, D., Sun, J., Wu, J., and Chen, J · 2024
Later among the works it cites.
Towards cross-table masked pretraining for web data mining
Ye, C., Lu, G., Wang, H., Li, L., Wu, S., Chen, G., and Zhao, J · 2024
Later among the works it cites.
Accurate predictions on small data with a tabular foundation model
Hollmann, N., Müller, S., Purucker, L., Krishnakumar, A., Körfer, M., Hoo, S. B., Schirrmeister, R. T., and Hutter, F · 2025
Closest in time.
Tabicl: A tabular foundation model for in-context learning on large data
Qu, J., Holzmüller, D., Varoquaux, G., and Morvan, M. L · 2025
Closest in time.