Fetching the paper…
Reading the bibliography…
The existing literature on deep learning for tabular data proposes a wide range of novel architectures and reports competitive results on various datasets.
Individual comparisons by ranking methods
F. Wilcoxon · 1945
Earlier work this paper cites.
Scaling up the accuracy of naive-bayes classifiers: a decision-tree hybrid
R. Kohavi · 1996
Earlier work this paper cites.
Sparse spatial autoregressions
R. Kelley Pace and R. Barry · 1997
Earlier work this paper cites.
Comparative accuracies of artificial neural networks and discriminant analysis in predicting forest cover types from cartographic variables
J. A. Blackard and D. J. Dean · 2000
Earlier work this paper cites.
Random forests
L. Breiman · 2001
Earlier work this paper cites.
Greedy function approximation: A gradient boosting machine
J. H. Friedman · 2001
Earlier work this paper cites.
The amsterdam library of object images
J. M. Geusebroek, G. J. Burghouts, , and A. W. M. Smeulders · 2005
Earlier work this paper cites.
Linformer: Self-attention with linear complexity
S. Wang, B. Z. Li, M. Khabsa, H. Fang, and H. Ma · 2006
Earlier work this paper cites.
Dcn v2: Improved deep & cross network and practical lessons for web-scale learning to rank systems
R. Wang, R. Shivanna, D. Z. Cheng, S. Jain, D. Lin, L. Hong, and E. H. Chi · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
The million song dataset
T. Bertin-Mahieux, D. P. Ellis, B. Whitman, and P. Lamere · 2011
Earlier work this paper cites.
Yahoo! learning to rank challenge overview
O. Chapelle and Y. Chang · 2011
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay · 2011
Earlier work this paper cites.
Tabtransformer: Tabular data modeling using contextual embeddings
X. Huang, A. Khetan, M. Cvitkovic, and Z. Karnin · 2012
Earlier work this paper cites.
Introducing LETOR 4.0 datasets
T. Qin and T. Liu · 2013
Earlier work this paper cites.
Searching for exotic particles in high-energy physics with deep learning
P. Baldi, P. Sadowski, and D. Whiteson · 2014
Earlier work this paper cites.
A data-driven approach to predict the success of bank telemarketing
S. Moro, P. Cortez, and P. Rita · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
N. Srivastava, G. E. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Earlier work this paper cites.
Openml: networked science in machine learning
J. Vanschoren, J. N. van Rijn, B. Bischl, and L. Torgo · 2014
Earlier work this paper cites.
Deep neural decision forests
P. Kontschieder, M. Fiterau, A. Criminisi, and S. Rota Bulo · 2015
Cited alongside, same era.
Layer normalization
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Cited alongside, same era.
Xgboost: A scalable tree boosting system
T. Chen and C. Guestrin · 2016
Cited alongside, same era.
Deep Learning
I. Goodfellow, Y. Bengio, and A. Courville · 2016
Cited alongside, same era.
Identity mappings in deep residual networks
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Lightgbm: A highly efficient gradient boosting decision tree
G. Ke, Q. Meng, T. Finley, T. Wang, W. Chen, W. Ma, Q. Ye, and T.-Y. Liu · 2017
Cited alongside, same era.
Adam: A method for stochastic optimization
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2019
Later among the works it cites.
Transformers without tears: Improving the normalization of self-attention
T. Q. Nguyen and J. Salazar · 2019
Later among the works it cites.
Autoint: Automatic feature interaction learning via self-attentive neural networks
W. Song, C. Shi, Z. Xiao, Z. Duan, Y. Xu, M. Zhang, and J. Tang · 2019
Later among the works it cites.
Tabnet: Attentive interpretable tabular learning
S. O. Arik and T. Pfister · 2020
Later among the works it cites.
Gradient boosting neural networks: Grownet
S. Badirli, X. Liu, Z. Xing, A. Bhowmik, K. Doan, and S. S. Keerthi · 2020
Later among the works it cites.
Deep ensembles: A loss landscape perspective
S. Fort, H. Hu, and B. Lakshminarayanan · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. P. Kingma and J. Ba · 2017
Cited alongside, same era.
Self-normalizing neural networks
G. Klambauer, T. Unterthiner, A. Mayr, and S. Hochreiter · 2017
Cited alongside, same era.
Bdt: Gradient boosted decision tables for high accuracy and scoring efficiency
Y. Lou and M. Obukhov · 2017
Cited alongside, same era.
Axiomatic attribution for deep networks
M. Sundararajan, A. Taly, and Q. Yan · 2017
Cited alongside, same era.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
Deep & cross network for ad click predictions
R. Wang, B. Fu, G. Fu, and M. Wang · 2017
Cited alongside, same era.
The tree ensemble layer: Differentiability meets conditional computation
H. Hazimeh, N. Ponomareva, P. Mol, Z. Tan, and R. Mazumder · 2020
Later among the works it cites.
Understanding the difficulty of training transformers
L. Liu, X. Liu, J. Gao, W. Chen, and J. Han · 2020
Later among the works it cites.
Neural oblivious decision ensembles for deep learning on tabular data
S. Popov, S. Morozov, and A. Babenko · 2020
Later among the works it cites.
Glu variants improve transformer
N. Shazeer · 2020
Later among the works it cites.
Efficient transformers: A survey
Y. Tay, M. Dehghani, D. Bahri, and D. Metzler · 2020
Later among the works it cites.
Huggingface’s transformers: State-of-the-art natural language processing
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, J. Davison, S. Shleifer, P. von Platen, C. Ma, Y. Jernite, J. Plu, C. Xu, T. L. Scao, S. Gugger, M. Drame, Q. Lhoest, and A. M. Rush · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2021
Closest in time.
Do transformer modifications transfer across implementations and applications?
S. Narang, H. W. Chung, Y. Tay, W. Fedus, T. Fevry, M. Matena, K. Malkan, N. Fiedel, N. Shazeer, Z. Lan, Y. Zhou, W. Li, N. Ding, J. Marcus, A. Roberts, and C. Raffel · 2021
Closest in time.
Are neural rankers still outperformed by gradient boosted decision trees?
Z. Qin, L. Yan, H. Zhuang, Y. Tay, R. K. Pasumarthi, X. Wang, M. Bendersky, and M. Najork · 2021
Closest in time.
Revisiting simple neural probabilistic language models
S. Sun and M. Iyyer · 2021
Closest in time.
R. Turner, D. Eriksson, M. McCourt, J. Kiili, E. Laaksonen, Z. Xu, and I. Guyon · 2021
Closest in time.
Which transformer architecture fits my data? a vocabulary bottleneck in self-attention
N. Wies, Y. Levine, D. Jannai, and A. Shashua · 2021
Closest in time.