Fetching the paper…
Reading the bibliography…
Pre-training is prevalent in deep learning for vision and text data, leveraging knowledge from other datasets to enhance downstream tasks.
Multidimensional scaling: I. theory and method
Torgerson, W. S · 1952
Earlier work this paper cites.
Variable kernel density estimation
Terrell, G. R. and Scott, D. W · 1992
Earlier work this paper cites.
Support-vector networks
Cortes, C. and Vapnik, V · 1995
Earlier work this paper cites.
A new definition of neighborhood of a point in multi-dimensional space
Chaudhuri, B · 1996
Earlier work this paper cites.
On the use of neighbourhood-based non-parametric classifiers
Sánchez, J. S., Pla, F., and Ferri, F. J · 1997
Earlier work this paper cites.
Multidimensional scaling
Carroll, J. D. and Arabie, P · 1998
Earlier work this paper cites.
Optimal kernel shapes for local linear regression
Ormoneit, D. and Hastie, T · 1999
Earlier work this paper cites.
A model of inductive bias learning
Baxter, J · 2000
Earlier work this paper cites.
Multidimensional scaling
Davison, M. L. and Sireci, S. G · 2000
Earlier work this paper cites.
Financial forecasting using support vector machines
Cao, L. and Tay, F. E. H · 2001
Earlier work this paper cites.
A generalized representer theorem
Schölkopf, B., Herbrich, R., and Smola, A. J · 2001
Earlier work this paper cites.
Stochastic neighbor embedding
Hinton, G. E. and Roweis, S · 2002
Earlier work this paper cites.
Learning with Kernels: support vector machines, regularization, optimization, and beyond
Schölkopf, B. and Smola, A. J · 2002
Earlier work this paper cites.
A perspective view and survey of meta-learning
Vilalta, R. and Drissi, Y · 2002
Earlier work this paper cites.
Distance metric learning with application to clustering with side-information
Xing, E. P., Ng, A. Y., Jordan, M. I., and Russell, S · 2002
Earlier work this paper cites.
Multiple kernel learning, conic duality, and the smo algorithm
Bach, F. R., Lanckriet, G. R., and Jordan, M. I · 2004
Earlier work this paper cites.
Neighbourhood components analysis
Goldberger, J., Hinton, G. E., Roweis, S., and Salakhutdinov, R. R · 2004
Earlier work this paper cites.
Learning a kernel matrix for nonlinear dimensionality reduction
Weinberger, K. Q., Sha, F., and Saul, L. K · 2004
Earlier work this paper cites.
A general and efficient multiple kernel learning algorithm
Sonnenburg, S., Rätsch, G., and Schäfer, C · 2005
Earlier work this paper cites.
Multi-task feature learning
Argyriou, A., Evgeniou, T., and Pontil, M · 2006
Earlier work this paper cites.
Pattern recognition and machine learning
Bishop, C · 2006
Earlier work this paper cites.
Gaussian processes for machine learning
Rasmussen, C. E. and Williams, C. K. I · 2006
Earlier work this paper cites.
Graph embedding and extensions: A general framework for dimensionality reduction
Yan, S., Xu, D., Zhang, B., Zhang, H.-J., Yang, Q., and Lin, S · 2006
Earlier work this paper cites.
Svm-knn: Discriminative nearest neighbor classification for visual category recognition
Zhang, H., Berg, A. C., Maire, M., and Malik, J · 2006
Earlier work this paper cites.
Predicting clicks: estimating the click-through rate for new ads
Richardson, M., Dominowska, E., and Ragno, R · 2007
Earlier work this paper cites.
In defense of nearest-neighbor based image classification
Boiman, O., Shechtman, E., and Irani, M · 2008
Earlier work this paper cites.
Visualizing data using t-sne
Van der Maaten, L. and Hinton, G · 2008
Earlier work this paper cites.
The Elements of Statistical Learning: Data Mining, Inference, and Prediction
Hastie, T., Tibshirani, R., and Friedman, J. H · 2009
Earlier work this paper cites.
Distance metric learning for large margin nearest neighbor classification
Weinberger, K. Q. and Saul, L. K · 2009
Earlier work this paper cites.
The comparisons of data mining techniques for the predictive accuracy of probability of default of credit card clients
Yeh, I.-C. and Lien, C.-h · 2009
Earlier work this paper cites.
Transfer metric learning by learning task relationships
Zhang, Y. and Yeung, D.-Y · 2010
Earlier work this paper cites.
Multiple kernel learning algorithms
Gönen, M. and Alpaydin, E · 2011
Earlier work this paper cites.
What you saw is not what you get: Domain adaptation using asymmetric kernel transforms
Kulis, B., Saenko, K., and Darrell, T · 2011
Earlier work this paper cites.
Multi-task low-rank metric learning based on common subspace
Yang, P., Huang, K., and Liu, C.-L · 2011
Earlier work this paper cites.
Conditional likelihood maximisation: A unifying framework for information theoretic feature selection
Brown, G., Pocock, A. C., Zhao, M.-J., and Luján, M · 2012
Earlier work this paper cites.
Foundations of Machine Learning
Mohri, M., Rostamizadeh, A., and Talwalkar, A · 2012
Earlier work this paper cites.
Heterogeneous ensemble for feature drifts in data streams
Nguyen, H.-L., Woon, Y.-K., Ng, W.-K., and Wan, L · 2012
Earlier work this paper cites.
Multilabel classification with meta-level features in a learning-to-rank framework
Yang, Y. and Gopal, S · 2012
Earlier work this paper cites.
Facial age estimation based on label-sensitive learning and age-oriented regression
Chao, W.-L., Liu, J.-Z., and Ding, J.-J · 2013
Earlier work this paper cites.
Adaptivity to local smoothness and dimension in kernel regression
Kpotufe, S. and Garg, V. K · 2013
Earlier work this paper cites.
Metric learning: A survey
Kulis, B · 2013
Earlier work this paper cites.
Metric learning: A survey
Kulis, B. et al · 2013
Earlier work this paper cites.
A survey on multi-view learning
Xu, C., Tao, D., and Xu, C · 2013
Earlier work this paper cites.
Geometry preserving multi-task metric learning
Yang, P., Huang, K., and Liu, C.-L · 2013
Earlier work this paper cites.
On efficient meta-level features for effective text classification
Canuto, S. D., Salles, T., Gonçalves, M. A., Rocha, L., Ramos, G. S., Gonçalves, L., Rosa, T. C., and Martins, W. S · 2014
Earlier work this paper cites.
Do we need hundreds of classifiers to solve real world classification problems?
Delgado, M. F., Cernadas, E., Barro, S., and Amorim, D. G · 2014
Earlier work this paper cites.
Learning categories from few examples with multi model knowledge transfer
Tommasi, T., Orabona, F., and Caputo, B · 2014
Earlier work this paper cites.
Openml: networked science in machine learning
Vanschoren, J., Van Rijn, J. N., Bischl, B., and Torgo, L · 2014
Earlier work this paper cites.
Metric Learning
Bellet, A., Habrard, A., and Sebban, M · 2015
Earlier work this paper cites.
Efficient and robust automated machine learning
Feurer, M., Klein, A., Eggensperger, K., Springenberg, J. T., Blum, M., and Hutter, F · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Rank consistency based multi-view learning: A privacy-preserving approach
Ye, H.-J., Zhan, D.-C., Miao, Y., Jiang, Y., and Zhou, Z.-H · 2015
Earlier work this paper cites.
Lift: Multi-label learning with label-specific features
Zhang, M.-L. and Wu, L · 2015
Earlier work this paper cites.
Xgboost: A scalable tree boosting system
Chen, T. and Guestrin, C · 2016
Earlier work this paper cites.
Wide & deep learning for recommender systems
Cheng, H.-T., Koc, L., Harmsen, J., Shaked, T., Chandra, T., Aradhye, H., Anderson, G., Corrado, G., Chai, W., Ispir, M., Anil, R., Haque, Z., Hong, L., Jain, V., Liu, X., and Shah, H · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Revisiting visual question answering baselines
Jabri, A., Joulin, A., and Van Der Maaten, L · 2016
Cited alongside, same era.
Matching networks for one shot learning
Vinyals, O., Blundell, C., Lillicrap, T., Kavukcuoglu, K., and Wierstra, D · 2016
Cited alongside, same era.
Deep learning over multi-field categorical data - - A case study on user response prediction
Zhang, W., Du, T., and Wang, J · 2016
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Self-attention between datapoints: Going beyond individual input-output pairs in deep learning
Kossen, J., Band, N., Lyle, C., Gomez, A. N., Rainforth, T., and Gal, Y · 2021
Later among the works it cites.
Learnable embedding sizes for recommender systems
Liu, S., Gao, C., Chen, Y., Jin, D., and Li, Y · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Later among the works it cites.
SAINT: improved neural networks for tabular data via row attention and contrastive pre-training
Somepalli, G., Goldblum, M., Schwarzschild, A., Bruss, C. B., and Goldstein, T · 2021
Later among the works it cites.
DCN V2: improved deep & cross network and practical lessons for web-scale learning to rank systems
Wang, R., Shivanna, R., Cheng, D. Z., Jain, S., Lin, D., Hong, L., and Chi, E. H · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deepfm: A factorization-machine based neural network for CTR prediction
Guo, H., Tang, R., Ye, Y., Li, Z., and He, X · 2017
Cited alongside, same era.
Learning with feature evolvable streams
Hou, B.-J., Zhang, L., and Zhou, Z.-H · 2017
Cited alongside, same era.
Lightgbm: A highly efficient gradient boosting decision tree
Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., Ye, Q., and Liu, T.-Y · 2017
Cited alongside, same era.
Self-normalizing neural networks
Klambauer, G., Unterthiner, T., Mayr, A., and Hochreiter, S · 2017
Cited alongside, same era.
Fast rates by transferring from auxiliary hypotheses
Kuzborskij, I. and Orabona, F · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Heterogeneous few-shot model rectification with semantic mapping
Ye, H.-J., Zhan, D.-C., Jiang, Y., and Zhou, Z.-H · 2021
Later among the works it cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., Ring, R., Rutherford, E., Cabi, S., Han, T., Gong, Z., Samangooei, S., Monteiro, M., Menick, J. L., Borgeaud, S., Brock, A., Nematzadeh, A., Sharifzadeh, S., Binkowski, M., Barreira, R., Vinyals, O., Zisserman, A., and Simonyan, K · 2022
Later among the works it cites.
You say factorization machine, I say neural network - it’s all in the activation
Almagor, C. and Hoshen, Y · 2022
Later among the works it cites.
Scarf: Self-supervised contrastive learning using random feature corruption
Bahri, D., Jiang, H., Tay, Y., and Metzler, D · 2022
Later among the works it cites.
Deep neural networks and tabular data: A survey
Borisov, V., Leemann, T., Seßler, K., Haug, J., Pawelczyk, M., and Kasneci, G · 2022
Later among the works it cites.
NODE-GAM: neural generalized additive model for interpretable deep learning
Chang, C.-H., Caruana, R., and Goldenberg, A · 2022
Later among the works it cites.
Danets: Deep abstract networks for tabular data classification and regression
Chen, J., Liao, K., Wan, Y., Chen, D. Z., and Wu, J · 2022
Later among the works it cites.
On embeddings for numerical features in tabular deep learning
Gorishniy, Y., Rubachev, I., and Babenko, A · 2022
Later among the works it cites.
Why do tree-based models still outperform deep learning on typical tabular data?
Grinsztajn, L., Oyallon, E., and Varoquaux, G · 2022
Later among the works it cites.
Prediction with unpredictable feature evolution
Hou, B.-J., Zhang, L., and Zhou, Z.-H · 2022
Later among the works it cites.
Turning the tables: Biased, imbalanced, dynamic tabular datasets for ML evaluation
Jesus, S. M., Pombal, J., Alves, D., Cruz, A. F., Saleiro, P., Ribeiro, R. P., Gama, J., and Bizarro, P · 2022
Later among the works it cites.
GATE: gated additive tree ensemble for tabular classification and regression
Joseph, M. and Raj, H · 2022
Later among the works it cites.
Few-shot learning for feature selection with hilbert-schmidt independence criterion
Kumagai, A., Iwata, T., Ida, Y., and Fujiwara, Y · 2022
Later among the works it cites.
Optembed: Learning optimal embedding table for click-through rate prediction
Lyu, F., Tang, X., Zhu, H., Guo, H., Zhang, Y., Tang, R., and Liu, X · 2022
Later among the works it cites.
Probabilistic machine learning: an introduction
Murphy, K. P · 2022
Later among the works it cites.
Can foundation models wrangle your data?
Narayan, A., Chami, I., Orr, L. J., and Ré, C · 2022
Later among the works it cites.
Revisiting pretraining objectives for tabular deep learning
Rubachev, I., Alekberov, A., Gorishniy, Y., and Babenko, A · 2022
Later among the works it cites.
Hopular: Modern hopfield networks for tabular data
Schäfl, B., Gruber, L., Bitto-Nemling, A., and Hochreiter, S · 2022
Later among the works it cites.
Tabular data: Deep learning is not all you need
Shwartz-Ziv, R. and Armon, A · 2022
Later among the works it cites.
Enhancing CTR prediction with context-aware feature representation learning
Wang, F., Wang, Y., Li, D., Gu, H., Lu, T., Zhang, P., and Gu, N · 2022
Later among the works it cites.
Transtab: Learning transferable tabular transformers across tables
Wang, Z. and Sun, J · 2022
Later among the works it cites.
Yan, J., Chen, J., Wu, Y., Chen, D. Z., and Wu, J · 2022
Later among the works it cites.
Tabnas: Rejection sampling for neural architecture search on tabular datasets
Yang, C., Bender, G., Liu, H., Kindermans, P.-J., Udell, M., Lu, Y., Le, Q. V., and Huang, D · 2022
Later among the works it cites.
A survey on multi-task learning
Zhang, Y. and Yang, Q · 2022
Later among the works it cites.
Tiger: Transferable interest graph embedding for domain-level zero-shot recommendation
Zhuo, J., Lian, J., Xu, L., Gong, M., Shou, L., Jiang, D., Xie, X., and Yue, Y · 2022
Later among the works it cites.
Efficient bayesian learning curve extrapolation using prior-data fitted networks
Adriaensen, S., Rakotoarison, H., Müller, S., and Hutter, F · 2023
Closest in time.
Hypothesis transfer learning with surrogate classification losses: Generalization bounds through algorithmic stability
Aghbalou, A. and Staerman, G · 2023
Closest in time.
Multi-layer attention-based explainability via transformers for tabular data
Gavito, A. T., Klabjan, D., and Utke, J · 2023
Closest in time.
Tabllm: few-shot classification of tabular data with large language models
Hegselmann, S., Buendia, A., Lang, H., Agrawal, M., Jiang, X., and Sontag, D · 2023
Closest in time.
Tabpfn: A transformer that solves small tabular classification problems in a second
Hollmann, N., Müller, S., Eggensperger, K., and Hutter, F · 2023
Closest in time.
TANGOS: regularizing tabular neural networks through gradient orthogonalization and specialization
Jeffares, A., Liu, T., Crabbé, J., Imrie, F., and van der Schaar, M · 2023
Closest in time.
Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., Dollár, P., and Girshick, R. B · 2023
Closest in time.
Tabddpm: Modelling tabular data with diffusion models
Kotelnikov, A., Baranchuk, D., Rubachev, I., and Babenko, A · 2023
Closest in time.
Transfer learning with deep tabular models
Levin, R., Cherepanova, V., Schwarzschild, A., Bansal, A., Bruss, C. B., Goldstein, T., Wilson, A. G., and Goldblum, M · 2023
Closest in time.
Luetto, S., Garuti, F., Sangineto, E., Forni, L., and Cucchiara, R · 2023
Closest in time.
When do neural nets outperform boosted trees on tabular data?
McElfresh, D. C., Khandagale, S., Valverde, J., C., V. P., Ramakrishnan, G., Goldblum, M., and White, C · 2023
Closest in time.
Tabret: Pre-training transformer-based tabular models for unseen columns
Onishi, S., Oono, K., and Hayashi, K · 2023
Closest in time.
Cross-modal fine-tuning: Align then refine
Shen, J., Li, L., Dery, L. M., Staten, C., Khodak, M., Neubig, G., and Talwalkar, A · 2023
Closest in time.
Graph neural network contextual embedding for deep learning on tabular data
Villaizán-Vallelado, M., Salvatori, M., Martínez, B. C., and Sánchez-Esguevillas, A. J · 2023
Closest in time.
Hypertab: Hypernetwork approach for deep learning on small tabular datasets
Wydmanski, W., Bulenok, O., and Smieja, M · 2023
Closest in time.
Xtab: Cross-table pretraining for tabular transformers
Zhu, B., Shi, X., Erickson, N., Li, M., Karypis, G., and Shoaran, M · 2023
Closest in time.
Tabr: Tabular deep learning meets nearest neighbors in 2023
Gorishniy, Y., Rubachev, I., Kartashev, N., Shlenskii, D., Kotelnikov, A., and Babenko, A · 2024
Closest in time.
The platonic representation hypothesis
Huh, M., Cheung, B., Wang, T., and Isola, P · 2024
Closest in time.
Dinov2: Learning robust visual features without supervision
Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., et al · 2024
Closest in time.
A closer look at deep learning on tabular data
Ye, H.-J., Liu, S.-Y., Cai, H.-R., Zhou, Q.-L., and Zhan, D.-C · 2024
Closest in time.
Tabpfn unleashed: A scalable and effective solution to tabular classification problems
Liu, S.-Y. and Ye, H.-J · 2025
Closest in time.
Tabred: A benchmark of tabular machine learning in-the-wild
Rubachev, I., Kartashev, N., Gorishniy, Y., and Babenko, A · 2025
Closest in time.
Revisiting nearest neighbor for tabular data: A deep tabular baseline two decades later
Ye, H.-J., Yin, H.-H., Zhan, D.-C., and Chao, W.-L · 2025
Closest in time.