Fetching the paper…
Reading the bibliography…
Tabular classification has traditionally relied on supervised algorithms, which estimate the parameters of a prediction model using its training data.
Coresets-Methods and History: A Theoreticians Design Pattern for Approximation and Streaming Algorithms
Alexander Munteanu and Chris Schwiegelshohn · 1987
Earlier work this paper cites.
Geometric approximation via coresets
Pankaj K Agarwal, Sariel Har-Peled, Kasturi R Varadarajan, et al · 2005
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay · 2011
Earlier work this paper cites.
A survey on feature selection methods
Girish Chandrashekar and Ferat Sahin · 2013
Earlier work this paper cites.
Feature selection: a literature review
Vipin Kumar and Sonajharia Minz · 2014
Earlier work this paper cites.
A review of feature selection methods based on mutual information
Jorge R Vergara and Pablo A Estévez · 2014
Earlier work this paper cites.
XGBoost: A scalable tree boosting system
Tianqi Chen and Carlos Guestrin · 2016
Earlier work this paper cites.
Feature selection: A data perspective
Jundong Li, Kewei Cheng, Suhang Wang, Fred Morstatter, Robert P Trevino, Jiliang Tang, and Huan Liu · 2017
Earlier work this paper cites.
Active learning for convolutional neural networks: A core-set approach
Ozan Sener and Silvio Savarese · 2017
Earlier work this paper cites.
CatBoost: unbiased boosting with categorical features
Liudmila Prokhorenkova, Gleb Gusev, Aleksandr Vorobev, Anna Veronika Dorogush, and Andrey Gulin · 2018
Earlier work this paper cites.
Towards efficient data valuation based on the shapley value
Ruoxi Jia, David Dao, Boxin Wang, Frances Ann Hubis, Nick Hynes, Nezihe Merve Gürel, Bo Li, Ce Zhang, Dawn Song, and Costas J Spanos · 2019
Earlier work this paper cites.
Billion-scale similarity search with GPUs
Jeff Johnson, Matthijs Douze, and Hervé Jégou · 2019
Cited alongside, same era.
Language Models are Few-Shot Learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Cited alongside, same era.
A comprehensive review of dimensionality reduction techniques for feature selection and feature extraction
Rizgar Zebari, Adnan Abdulazeez, Diyar Zeebaree, Dilovan Zebari, and Jwan Saeed · 2020
Cited alongside, same era.
Train short, test long: Attention with linear biases enables input length extrapolation
Ofir Press, Noah A Smith, and Mike Lewis · 2021
Cited alongside, same era.
Saint: Improved neural networks for tabular data via row attention and contrastive pre-training
Active learning through a covering lens
Ofer Yehuda, Avihu Dekel, Guy Hacohen, and Daphna Weinshall · 2022
Later among the works it cites.
Extending context window of large language models via positional interpolation
Shouyuan Chen, Sherman Wong, Liangjian Chen, and Yuandong Tian · 2023
Closest in time.
A Review of Data Valuation Approaches and Building and Scoring a Data Valuation Model
Mike Fleckenstein, Ali Obaidi, and Nektaria Tryfona · 2023
Closest in time.
HyperAttention: Long-context Attention in Near-Linear Time, October 2023
Insu Han, Rajesh Jayaram, Amin Karbasi, Vahab Mirrokni, David P. Woodruff, and Amir Zandieh · 2023
Closest in time.
TabPFN: A Transformer That Solves Small Tabular Classification Problems in a Second, May 2023
Noah Hollmann, Samuel Müller, Katharina Eggensperger, and Frank Hutter · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gowthami Somepalli, Micah Goldblum, Avi Schwarzschild, C Bayan Bruss, and Tom Goldstein · 2021
Cited alongside, same era.
Deep neural networks and tabular data: A survey
Vadim Borisov, Tobias Leemann, Kathrin Seßler, Johannes Haug, Martin Pawelczyk, and Gjergji Kasneci · 2022
Cited alongside, same era.
Active learning on a budget: Opposite strategies suit high and low budgets
Guy Hacohen, Avihu Dekel, and Daphna Weinshall · 2022
Cited alongside, same era.
Locating and editing factual associations in GPT
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2022
Cited alongside, same era.
Data valuation in machine learning:“ingredients”, strategies, and open challenges
Rachael Hwee Ling Sim, Xinyi Xu, and Bryan Kian Hsiang Low · 2022
Cited alongside, same era.
A length-extrapolatable transformer
Yutao Sun, Li Dong, Barun Patra, Shuming Ma, Shaohan Huang, Alon Benhaim, Vishrav Chaudhary, Xia Song, and Furu Wei · 2022
Cited alongside, same era.
An Explanation of In-context Learning as Implicit Bayesian Inference, July 2022
Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma · 2022
Cited alongside, same era.
Closest in time.
Emergent world representations: Exploring a sequence model trained on a synthetic task, 2023
Kenneth Li, Aspen K. Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg · 2023
Closest in time.
Goat: Fine-tuned LLaMA Outperforms GPT-4 on Arithmetic Tasks, May 2023
Tiedong Liu and Bryan Kian Hsiang Low · 2023
Closest in time.
When Do Neural Nets Outperform Boosted Trees on Tabular Data?, May 2023
Duncan McElfresh, Sujay Khandagale, Jonathan Valverde, Vishak Prasad C, Ganesh Ramakrishnan, Micah Goldblum, and Colin White · 2023
Closest in time.
Progress measures for grokking via mechanistic interpretability
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt · 2023
Closest in time.
Efficient Streaming Language Models with Attention Sinks, September 2023
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis · 2023
Closest in time.