Fetching the paper…
Reading the bibliography…
A major challenge in training large-scale machine learning models is configuring the training process to maximize model performance, i.e., finding the best training setup from a vast design space.
“Robust estimation of a location parameter”
Peter. Huber · 1964
Earlier work this paper cites.
“Register allocation via coloring”
Gregory Chaitin, Marc Auslander, Ashok Chandra, John Cocke, Martin Hopkins and Peter Markstein · 1981
Earlier work this paper cites.
“Optimization and nonsmooth analysis”
Frank Clarke · 1990
Earlier work this paper cites.
“Backpropagation through time: what it does and how to do it”
Paul Werbos · 1990
Earlier work this paper cites.
“Rematerialization”
Preston Briggs, Keith Cooper and Linda Torczon · 1992
Earlier work this paper cites.
“Learning long-term dependencies with gradient descent is difficult”
Yoshua Bengio, Patrice Simard and Paolo Frasconi · 1994
Earlier work this paper cites.
“An investigation of the gradient descent process in neural networks”
Barak Pearlmutter · 1996
Earlier work this paper cites.
“The MNIST database of handwritten digits”
Yann LeCun · 1998
Earlier work this paper cites.
“Iterative solution of nonlinear equations in several variables”
James Ortega and Werner Rheinboldt · 2000
Earlier work this paper cites.
“Exact alpha-beta computation in logarithmic space with application to MAP word graph construction”
Geoffrey Zweig and Mukund Padmanabhan · 2000
Earlier work this paper cites.
“Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories”
Li Fei-Fei, Rob Fergus and Pietro Perona · 2004
Earlier work this paper cites.
“The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution Generalization”, 2020
Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, Dawn Song, Jacob Steinhardt and Justin Gilmer · 2006
Earlier work this paper cites.
“Evaluating derivatives: principles and techniques of algorithmic differentiation”
Andreas Griewank and Andrea Walther · 2008
Earlier work this paper cites.
“Pareto-based multiobjective machine learning: An overview and case studies”
Yaochu Jin and Bernhard Sendhoff · 2008
Earlier work this paper cites.
“Optimal Jacobian accumulation is NP-complete”
Uwe Naumann · 2008
Earlier work this paper cites.
“Automated flower classification over a large number of classes”
Maria-Elena Nilsback and Andrew Zisserman · 2008
Earlier work this paper cites.
“Imagenet: A large-scale hierarchical image database”
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li and Li Fei-Fei · 2009
Earlier work this paper cites.
“Learning Multiple Layers of Features from Tiny Images”
Alex Krizhevsky · 2009
Earlier work this paper cites.
“The Pascal Visual Object Classes (VOC) Challenge”
M. Everingham, L. Van, C… Williams, J. Winn and A. Zisserman · 2010
Earlier work this paper cites.
“Sun database: Large-scale scene recognition from abbey to zoo”
Jianxiong Xiao, James Hays, Krista Ehinger, Aude Oliva and Antonio Torralba · 2010
Earlier work this paper cites.
“An analysis of single-layer networks in unsupervised feature learning”
Adam Coates, Andrew Ng and Honglak Lee · 2011
Earlier work this paper cites.
“Reading digits in natural images with unsupervised feature learning”
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Baolin Wu and Andrew Ng · 2011
Earlier work this paper cites.
“The German traffic sign recognition benchmark: a multi-class classification competition”
Johannes Stallkamp, Marc Schlipsing, Jan Salmen and Christian Igel · 2011
Earlier work this paper cites.
“Are we ready for autonomous driving? The KITTI vision benchmark suite”
Andreas Geiger, Philip Lenz and Raquel Urtasun · 2012
Earlier work this paper cites.
“Cats and dogs”
Omkar Parkhi, Andrea Vedaldi, Andrew Zisserman and CV Jawahar · 2012
Earlier work this paper cites.
“3d object representations for fine-grained categorization”
Jonathan Krause, Michael Stark, Jia Deng and Li Fei-Fei · 2013
Earlier work this paper cites.
“Fine-grained visual classification of aircraft”
Subhransu Maji, Esa Rahtu, Juho Kannala, Matthew Blaschko and Andrea Vedaldi · 2013
Earlier work this paper cites.
“Food-101–mining discriminative components with random forests”
Lukas Bossard, Matthieu Guillaumin and Luc Van · 2014
Earlier work this paper cites.
“Describing textures in the wild”
Mircea Cimpoi, Subhransu Maji, Iasonas Kokkinos, Sammy Mohamed and Andrea Vedaldi · 2014
Earlier work this paper cites.
“Microsoft coco: Common objects in context”
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár and C Zitnick · 2014
Earlier work this paper cites.
“From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions”
Peter Young, Alice Lai, Micah Hodosh and Julia Hockenmaier · 2014
Earlier work this paper cites.
“Adam: A Method for Stochastic Optimization”
Diederik. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
“Gradient-based hyperparameter optimization through reversible learning”
Dougal Maclaurin, David Duvenaud and Ryan Adams · 2015
Earlier work this paper cites.
“Training Deep Nets with Sublinear Memory Cost”
Tianqi Chen, Bing Xu, Chiyuan Zhang and Carlos Guestrin · 2016
Earlier work this paper cites.
“YFCC100M: The New Data in Multimedia Research”
Bart Thomee, David. Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth and Li-Jia Li · 2016
Earlier work this paper cites.
“Remote sensing image scene classification: Benchmark and state of the art”
Gong Cheng, Junwei Han and Xiaoqiang Lu · 2017
Cited alongside, same era.
“Model-agnostic meta-learning for fast adaptation of deep networks”
Chelsea Finn, Pieter Abbeel and Sergey Levine · 2017
Cited alongside, same era.
“Forward and reverse gradient-based hyperparameter optimization”
Luca Franceschi, Michele Donini, Paolo Frasconi and Massimiliano Pontil · 2017
Cited alongside, same era.
“Badnets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain”
Tianyu Gu, Brendan Dolan-Gavitt and Siddharth Garg · 2017
Cited alongside, same era.
“CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning”
Justin Johnson, Bharath Hariharan, Laurens Van, Li Fei-Fei, C Lawrence and Ross Girshick · 2017
Cited alongside, same era.
“Efficient and modular implicit differentiation”
Mathieu Blondel, Quentin Berthet, Marco Cuturi, Roy Frostig, Stephan Hoyer, Felipe Llinares-López, Fabian Pedregosa and Jean-Philippe Vert · 2022
Later among the works it cites.
“WinoGAViL: Gamified association benchmark to challenge vision-and-language models”
Yonatan Bitton, Nitzan Bitton, Ron Yosef, Yuval Elovici, Mohit Bansal, Gabriel Stanovsky and Roy Schwartz · 2022
Later among the works it cites.
“Implicit differentiation for fast hyperparameter selection in non-smooth convex learning”
Quentin Bertrand, Quentin Klopfenstein, Mathurin Massias, Mathieu Blondel, Samuel Vaiter, Alexandre Gramfort and Joseph Salmon · 2022
Later among the works it cites.
“If Influence Functions are the Answer, Then What is the Question?”
Juhan Bae, Nathan Ng, Alston Lo, Marzyeh Ghassemi and Roger Grosse · 2022
Later among the works it cites.
“Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability”, 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Understanding Black-box Predictions via Influence Functions”
Pang Koh and Percy Liang · 2017
Cited alongside, same era.
“Super-Convergence: Very Fast Training of Neural Networks Using Large Learning Rates”
Leslie. Smith and Topin Nicholay · 2017
Cited alongside, same era.
“From detection of individual metastases to classification of lymph node status at the patient level: the CAMELYON17 challenge”
Peter Bandi, Oscar Geessink, Quirine Manson, Marcory Van, Maschenka Balkenhol, Meyke Hermsen, Babak Bejnordi, Byungjae Lee, Kyunghyun Paeng and Aoxiao Zhong · 2018
Cited alongside, same era.
“Functional Map of the World”
Gordon Christie, Neil Fendley, James Wilson and Ryan Mukherjee · 2018
Cited alongside, same era.
“Darts: Differentiable architecture search”
Hanxiao Liu, Karen Simonyan and Yiming Yang · 2018
Cited alongside, same era.
“CIFAR-10 Fast”, GitHub Repository, 2018
David Page · 2018
Cited alongside, same era.
“Rotation equivariant CNNs for digital pathology”
Bastiaan Veeling, Jasper Linmans, Jim Winkens, Taco Cohen and Max Welling · 2018
Cited alongside, same era.
Jeremy. Cohen, Simran Kaur, Yuanzhi Li, J. Kolter and Ameet Talwalkar · 2022
Later among the works it cites.
“Gradient descent: The ultimate optimizer”
Kartik Chandra, Audrey Xie, Jonathan Ragan-Kelley and Erik Meijer · 2022
Later among the works it cites.
“FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness”, 2022
Tri Dao, Daniel. Fu, Stefano Ermon, Atri Rudra and Christopher Ré · 2022
Later among the works it cites.
“Data Determines Distributional Robustness in Contrastive Language Image Pre-training (CLIP)”
Alex Fang, Gabriel Ilharco, Mitchell Wortsman, Yuhao Wan, Vaishaal Shankar, Achal Dave and Ludwig Schmidt · 2022
Later among the works it cites.
“Training compute-optimal large language models”
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego Casas, Lisa Hendricks, Johannes Welbl and Aidan Clark · 2022
Later among the works it cites.
“ffcv”, https://github.com/libffcv/ffcv/ , 2022
Guillaume Leclerc, Andrew Ilyas, Logan Engstrom, Sung Park, Hadi Salman and Aleksander Madry · 2022
Later among the works it cites.
“Indiscriminate Data Poisoning Attacks on Neural Networks”
Yiwei Lu, Gautam Kamath and Yaoliang Yu · 2022
Later among the works it cites.
“The dollar street dataset: Images representing the geographic and socioeconomic diversity of the world”
William Rojas, Sudnya Diamos, Keertan Kini, David Kanter, Vijay Reddi and Cody Coleman · 2022
Later among the works it cites.
“The curse of unrolling: Rate of differentiating through optimization”
Damien Scieur, Gauthier Gidel, Quentin Bertrand and Fabian Pedregosa · 2022
Later among the works it cites.
“Beyond the imitation game: Quantifying and extrapolating the capabilities of language models”
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Shoeb, Abubakar Abid, Adam Fisch, Adam Brown, Adam Santoro, Aditya Gupta and Adrià Garriga-Alonso · 2022
Later among the works it cites.
“Challenging big-bench tasks and whether chain-of-thought can solve them”
Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Chung, Aakanksha Chowdhery, Quoc Le, Ed Chi and Denny Zhou · 2022
Later among the works it cites.
“Chain-of-thought prompting elicits reasoning in large language models”
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc Le and Denny Zhou · 2022
Later among the works it cites.
“SemDeDup: Data-efficient learning at web-scale through semantic deduplication”
Amro Abbas, Kushal Tirumala, Dániel Simig, Surya Ganguli and Ari Morcos · 2023
Later among the works it cites.
“Free Dolly: Introducing the World’s First Truly Open Instruction-Tuned LLM”, 2023
Mike Conover, Matt Hayes, Ankit Mathur, Jianwei Xie, Jun Wan, Sam Shah, Ali Ghodsi, Patrick Wendell, Matei Zaharia and Reynold Xin · 2023
Later among the works it cites.
“Studying large language model generalization with influence functions”
Roger Grosse, Juhan Bae, Cem Anil, Nelson Elhage, Alex Tamkin, Amirhossein Tajdini, Benoit Steiner, Dustin Li, Esin Durmus and Ethan Perez · 2023
Later among the works it cites.
“The flan collection: Designing data and methods for effective instruction tuning”
Shayne Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Chung, Yi Tay, Denny Zhou, Quoc Le, Barret Zoph and Jason Wei · 2023
Later among the works it cites.
“Exploring the limits of model-targeted indiscriminate data poisoning attacks”
Yiwei Lu, Gautam Kamath and Yaoliang Yu · 2023
Later among the works it cites.
“TRAK: Attributing Model Behavior at Scale”
Sung Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc and Aleksander Madry · 2023
Later among the works it cites.
“Stanford Alpaca: An Instruction-following LLaMA model”
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang and Tatsunori. Hashimoto · 2023
Later among the works it cites.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang and Angela Fan · 2024
Later among the works it cites.
“EcoDatum DataComp-small submission”, https://www.datacomp.ai/dcclip/leaderboard.html , 2024
Team EcoDatum · 2024
Later among the works it cites.
“DsDm: Model-Aware Dataset Selection with Datamodels”, 2024
Logan Engstrom, Axel Feldmann and Aleksander Madry · 2024
Later among the works it cites.
“DataComp: In search of the next generation of multimodal datasets”
Samir Gadre, Gabriel Ilharco, Alex Fang, Jonathan Hayase, Georgios Smyrnis, Thao Nguyen, Ryan Marten, Mitchell Wortsman, Dhruba Ghosh and Jieyu Zhang · 2024
Later among the works it cites.
“94 percent on CIFAR-10 in 3.29 Seconds on a Single GPU”
Keller Jordan · 2024
Later among the works it cites.
“Openassistant conversations-democratizing large language model alignment”
Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi Tam, Keith Stevens, Abdullah Barhoum, Duc Nguyen, Oliver Stanley and Richárd Nagyfi · 2024
Later among the works it cites.
“Deepseek-v3 technical report”
Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang and Chong Ruan · 2024
Later among the works it cites.
“GeoDE: a geographically diverse evaluation dataset for object recognition”
Vikram Ramaswamy, Sing Lin, Dora Zhao, Aaron Adcock, Laurens van Maaten, Deepti Ghadiyaram and Olga Russakovsky · 2024
Later among the works it cites.
“Gemma: Open models based on gemini research and technology”
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Kale and Juliette Love · 2024
Later among the works it cites.
“webdataset”, 2024
Team Webdataset · 2024
Later among the works it cites.
“Less: Selecting influential data for targeted instruction tuning”
Mengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora and Danqi Chen · 2024
Later among the works it cites.