Fetching the paper…
Reading the bibliography…
Recent efforts in fine-tuning language models often rely on automatic data selection, commonly using Nearest Neighbors retrieval from large datasets.
Adjustment of an inverse matrix corresponding to a change in one element of a given matrix
Jack Sherman and Winifred J Morrison · 1950
Earlier work this paper cites.
Discriminatory analysis: nonparametric discrimination, consistency properties , volume 1
Evelyn Fix and Joseph Lawson Hodges Jr · 1951
Earlier work this paper cites.
On estimating regression
Elizbar A Nadaraya · 1964
Earlier work this paper cites.
Smooth regression analysis
Geoffrey S Watson · 1964
Earlier work this paper cites.
Nearest neighbor pattern classification
Thomas Cover and Peter Hart · 1967
Earlier work this paper cites.
The sequential generation of d d -optimum experimental designs
Henry P Wynn · 1970
Earlier work this paper cites.
A statistical interpretation of term specificity and its application in retrieval
Karen Sparck Jones · 1972
Earlier work this paper cites.
Detection of influential observation in linear regression
R Dennis Cook · 1977
Earlier work this paper cites.
Accelerated greedy algorithms for maximizing submodular set functions
Michel Minoux · 1978
Earlier work this paper cites.
An analysis of approximations for maximizing submodular set functions—i
George L Nemhauser, Laurence A Wolsey, and Marshall L Fisher · 1978
Earlier work this paper cites.
Robust locally weighted regression and smoothing scatterplots
William S Cleveland · 1979
Earlier work this paper cites.
Locally weighted regression: an approach to regression analysis by local fitting
William S Cleveland and Susan J Devlin · 1988
Earlier work this paper cites.
Generalization and parameter estimation in feedforward nets: Some experiments
Nelson Morgan and Hervé Bourlard · 1989
Earlier work this paper cites.
Local learning algorithms
Léon Bottou and Vladimir Vapnik · 1992
Earlier work this paper cites.
Information-based objective functions for active data selection
David JC MacKay · 1992
Earlier work this paper cites.
Bayesian experimental design: A review
Kathryn Chaloner and Isabella Verdinelli · 1995
Earlier work this paper cites.
A sequential algorithm for training text classifiers: Corrigendum and additional data
David D Lewis · 1995
Earlier work this paper cites.
Locally weighted learning
Christopher G Atkeson, Andrew W Moore, and Stefan Schaal · 1997
Earlier work this paper cites.
A language modeling approach to information retrieval
Jay M. Ponte and W. Bruce Croft · 1998
Earlier work this paper cites.
Elements of information theory
Thomas M Cover · 1999
Earlier work this paper cites.
Gaussian process regression: Active data selection and test point rejection
Sambu Seo, Marko Wallat, Thore Graepel, and Klaus Obermayer · 2000
Earlier work this paper cites.
Gaussian processes for machine learning , volume 2
Christopher KI Williams and Carl Edward Rasmussen · 2006
Earlier work this paper cites.
Active learning via transductive experimental design
Kai Yu, Jinbo Bi, and Volker Tresp · 2006
Earlier work this paper cites.
On early stopping in gradient descent learning
Yuan Yao, Lorenzo Rosasco, and Andrea Caponnetto · 2007
Earlier work this paper cites.
The probabilistic relevance framework: Bm25 and beyond
Stephen Robertson, Hugo Zaragoza, et al · 2009
Earlier work this paper cites.
Active learning literature survey
Burr Settles · 2009
Earlier work this paper cites.
Gaussian process optimization in the bandit setting: No regret and experimental design
Niranjan Srinivas, Andreas Krause, Sham M Kakade, and Matthias Seeger · 2009
Earlier work this paper cites.
Online domain adaptation of a pre-trained cascade of classifiers
Vidit Jain and Erik Learned-Miller · 2011
Earlier work this paper cites.
Online learning for linearly parametrized control problems
Yasin Abbasi-Yadkori · 2013
Earlier work this paper cites.
Linguistic regularities in continuous space word representations
Tomáš Mikolov, Wen-tau Yih, and Geoffrey Zweig · 2013
Earlier work this paper cites.
The nature of statistical learning theory
Vladimir Vapnik · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy L Ba · 2014
Earlier work this paper cites.
Truncated variance reduction: A unified approach to bayesian optimization and level-set estimation
Ilija Bogunovic, Jonathan Scarlett, Andreas Krause, and Volkan Cevher · 2015
Earlier work this paper cites.
A latent variable model approach to pmi-based word embeddings
Sanjeev Arora, Yuanzhi Li, Yingyu Liang, Tengyu Ma, and Andrej Risteski · 2016
Earlier work this paper cites.
Training deep nets with sublinear memory cost
Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin · 2016
Earlier work this paper cites.
On kernelized multi-armed bandits
Sayak Ray Chowdhury and Aditya Gopalan · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2017
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang · 2017
Earlier work this paper cites.
Active learning for convolutional neural networks: A core-set approach
Ozan Sener and Silvio Savarese · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Batchbald: Efficient and diverse batch acquisition for deep bayesian active learning
Andreas Kirsch, Joost Van Amersfoort, and Yarin Gal · 2018
Cited alongside, same era.
Dynamic evaluation of neural sequence models
Ben Krause, Emmanuel Kahembwe, Iain Murray, and Steve Renals · 2018
Cited alongside, same era.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2018
Cited alongside, same era.
One sentence one model for neural machine translation
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2022
Later among the works it cites.
Datamodels: Predicting predictions from training data
Andrew Ilyas, Sung Min Park, Logan Engstrom, Guillaume Leclerc, and Aleksander Madry · 2022
Later among the works it cites.
Prism: A rich class of parameterized submodular information measures for guided data subset selection
Suraj Kothawade, Vishal Kaushal, Ganesh Ramakrishnan, Jeff Bilmes, and Rishabh Iyer · 2022
Later among the works it cites.
Glm-130b: An open bilingual pre-trained model
Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, et al · 2022
Later among the works it cites.
A statistical perspective on retrieval-based models
Soumya Basu, Ankit Singh Rawat, and Manzil Zaheer · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xiaoqing Li, Jiajun Zhang, and Chengqing Zong · 2018
Cited alongside, same era.
“zero-shot” super-resolution using deep internal learning
Assaf Shocher, Nadav Cohen, and Michal Irani · 2018
Cited alongside, same era.
A continuous-time view of early stopping for least squares regression
Alnur Ali, J Zico Kolter, and Ryan J Tibshirani · 2019
Cited alongside, same era.
On the measure of intelligence
François Chollet · 2019
Cited alongside, same era.
Billion-scale similarity search with gpus
Jeff Johnson, Matthijs Douze, and Hervé Jégou · 2019
Cited alongside, same era.
Dynamic evaluation of transformer language models
Ben Krause, Emmanuel Kahembwe, Iain Murray, and Steve Renals · 2019
Cited alongside, same era.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al · 2019
Cited alongside, same era.
Aman Bhargava, Cameron Witkowski, Manav Shah, and Matt Thomson · 2023
Later among the works it cites.
Prediction-oriented bayesian active learning
Freddie Bickford Smith, Andreas Kirsch, Sebastian Farquhar, Yarin Gal, Adam Foster, and Tom Rainforth · 2023
Later among the works it cites.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al · 2023
Later among the works it cites.
Simfluence: Modeling the influence of individual training examples by simulating training runs
Kelvin Guu, Albert Webson, Ellie Pavlick, Lucas Dixon, Ian Tenney, and Tolga Bolukbasi · 2023
Later among the works it cites.
A framework and benchmark for deep batch active learning for regression
David Holzmüller, Viktor Zaverkin, Johannes Kästner, and Ingo Steinwart · 2023
Later among the works it cites.
A kernel-based view of language model fine-tuning
Sadhika Malladi, Alexander Wettig, Dingli Yu, Danqi Chen, and Sanjeev Arora · 2023
Later among the works it cites.
Few-shot fine-tuning vs. in-context learning: A fair comparison and evaluation
Marius Mosbach, Tiago Pimentel, Shauli Ravfogel, Dietrich Klakow, and Yanai Elazar · 2023
Later among the works it cites.
Probabilistic machine learning: Advanced topics
Kevin P Murphy · 2023
Later among the works it cites.
Transformers learn in-context by gradient descent
Johannes Von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov · 2023
Later among the works it cites.
Large language models are latent variable models: Explaining and finding good demonstrations for in-context learning
Xinyi Wang, Wanrong Zhu, Michael Saxon, Mark Steyvers, and William Yang Wang · 2023
Later among the works it cites.
Compositional exemplars for in-context learning
Jiacheng Ye, Zhiyong Wu, Jiangtao Feng, Tao Yu, and Lingpeng Kong · 2023
Later among the works it cites.
Online (multinomial) logistic bandit: Improved regret and constant computation cost
Yu-Jie Zhang and Masashi Sugiyama · 2023
Later among the works it cites.
Phi-3 technical report: A highly capable language model locally on your phone
Marah Abdin, Sam Ade Jacobs, Ammar Ahmad Awan, Jyoti Aneja, Ahmed Awadallah, Hany Awadalla, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Harkirat Behl, et al · 2024
Closest in time.
The surprising effectiveness of test-time training for abstract reasoning
Ekin Akyürek, Mehul Damani, Linlu Qiu, Han Guo, Yoon Kim, and Jacob Andreas · 2024
Closest in time.
Active fine-tuning of generalist policies
Marco Bagatella, Jonas Hübotter, Georg Martius, and Andreas Krause · 2024
Closest in time.
Understanding in-context learning in transformers and llms by learning to learn discrete functions
Satwik Bhattamishra, Arkil Patel, Phil Blunsom, and Varun Kanade · 2024
Closest in time.
Large language monkeys: Scaling inference compute with repeated sampling
Bradley Brown, Jordan Juravsky, Ryan Ehrlich, Ronald Clark, Quoc V Le, Christopher Ré, and Azalia Mirhoseini · 2024
Closest in time.
Dataset-induced meta-learning (and other tricks): Improving model efficiency on arc
Jack Cole and Mohamed Osman · 2024
Closest in time.
Matthijs Douze, Alexandr Guzhva, Chengqi Deng, Jeff Johnson, Gergely Szilvasy, Pierre-Emmanuel Mazaré, Maria Lomeli, Lucas Hosseini, and Hervé Jégou · 2024
Closest in time.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Closest in time.
A framework for few-shot language model evaluation, 2024
Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, et al · 2024
Closest in time.
Towards flexible perception with visual memory
Robert Geirhos, Priyank Jaini, Austin Stone, Sourabh Medapati, Xi Yi, George Toderici, Abhijit Ogale, and Jonathon Shlens · 2024
Closest in time.
Test-time training on nearest neighbors for large language models
Moritz Hardt and Yu Sun · 2024
Closest in time.
Transductive active learning: Theory and applications
Jonas Hübotter, Bhavya Sukhija, Lenart Treven, Yarden As, and Andreas Krause · 2024
Closest in time.
Towards a statistical theory of data selection under weak supervision
Germain Kolossov, Andrea Montanari, and Pulkit Tandon · 2024
Closest in time.
In-context learning learns label relationships but is not conventional learning
Jannik Kossen, Yarin Gal, and Tom Rainforth · 2024
Closest in time.
Learning to reason with llms
OpenAI · 2024
Closest in time.
The linear representation hypothesis and the geometry of large language models
Kiho Park, Yo Joong Choe, and Victor Veitch · 2024
Closest in time.
Bandits with preference feedback: A stackelberg game perspective
Barna Pásztor, Parnian Kassraie, and Andreas Krause · 2024
Closest in time.
Learning to (learn at test time): Rnns with expressive hidden states
Yu Sun, Xinhao Li, Karan Dalal, Jiarui Xu, Arjun Vikram, Genghan Zhang, Yann Dubois, Xinlei Chen, Xiaolong Wang, Sanmi Koyejo, et al · 2024
Closest in time.
Gemma 2: Improving open language models at a practical size
Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, et al · 2024
Closest in time.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, et al · 2024
Closest in time.
Less: Selecting influential data for targeted instruction tuning
Mengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora, and Danqi Chen · 2024
Closest in time.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar · 2025
Closest in time.