Fetching the paper…
Reading the bibliography…
Training deep neural networks (DNNs) is becoming increasingly more resource- and energy-intensive every year.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R Thompson · 1933
Earlier work this paper cites.
Pareto optimality in multiobjective problems
Yair Censor · 1977
Earlier work this paper cites.
Algorithm as 136: A k-means clustering algorithm
John A Hartigan and Manchek A Wong · 1979
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Tze Leung Lai, Herbert Robbins, et al · 1985
Earlier work this paper cites.
A compendium of conjugate priors
Daniel Fink · 1997
Earlier work this paper cites.
Splice-2 comparative evaluation: Electricity pricing
Michael Harries and New South Wales · 1999
Earlier work this paper cites.
Mining time-changing data streams
Geoff Hulten, Laurie Spencer, and Pedro Domingos · 2001
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
Learning with drift detection
Joao Gama, Pedro Medas, Gladys Castillo, and Pedro Rodrigues · 2004
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Twitter sentiment classification using distant supervision
Alec Go, Richa Bhayani, and Lei Huang · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Moa: Massive online analysis, a framework for stream classification and clustering
Albert Bifet, Geoff Holmes, Bernhard Pfahringer, Philipp Kranen, Hardy Kremer, Timm Jansen, and Thomas Seidl · 2010
Earlier work this paper cites.
An integrated GPU power and performance model
Sunpyo Hong and Hyesoon Kim · 2010
Earlier work this paper cites.
Algorithms for hyper-parameter optimization
James Bergstra, Rémi Bardenet, Yoshua Bengio, and Balázs Kégl · 2011
Earlier work this paper cites.
An empirical evaluation of thompson sampling
Olivier Chapelle and Lihong Li · 2011
Earlier work this paper cites.
A performance and energy consumption analytical model for GPU
Cheng Luo and Reiji Suda · 2011
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
Matthew D Zeiler · 2012
Earlier work this paper cites.
Power capping of CPU-GPU heterogeneous systems through coordinating DVFS and task mapping
Toshiya Komoda, Shingo Hayashi, Takashi Nakada, Shinobu Miwa, and Hiroshi Nakamura · 2013
Earlier work this paper cites.
Stochastic multi-armed-bandit problem with non-stationary rewards
Omar Besbes, Yonatan Gur, and Assaf Zeevi · 2014
Earlier work this paper cites.
Deep speech: Scaling up end-to-end speech recognition
Awni Hannun, Carl Case, Jared Casper, Bryan Catanzaro, Greg Diamos, Erich Elsen, Ryan Prenger, Sanjeev Satheesh, Shubho Sengupta, Adam Coates, et al · 2014
Earlier work this paper cites.
The movielens datasets: History and context
F Maxwell Harper and Joseph A Konstan · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Librispeech: an ASR corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
Deep learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
An overview of gradient descent optimization algorithms
Sebastian Ruder · 2016
Earlier work this paper cites.
Accurate, large minibatch SGD: Training ImageNet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Earlier work this paper cites.
Neural collaborative filtering
Xiangnan He, Lizi Liao, Hanwang Zhang, Liqiang Nie, Xia Hu, and Tat-Seng Chua · 2017
Earlier work this paper cites.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2017
Earlier work this paper cites.
Hyperband: A novel bandit-based approach to hyperparameter optimization
Lisha Li, Kevin Jamieson, Giulia DeSalvo, Afshin Rostamizadeh, and Ameet Talwalkar · 2017
Earlier work this paper cites.
A survey and measurement study of GPU DVFS on energy conservation
Xinxin Mei, Qiang Wang, and Xiaowen Chu · 2017
Cited alongside, same era.
Randomness in neural networks: an overview
Simone Scardapane and Dianhui Wang · 2017
Cited alongside, same era.
TVM: An automated end-to-end optimizing compiler for deep learning
Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, et al · 2018
Cited alongside, same era.
On the computational inefficiency of large batch sizes for stochastic gradient descent
Noah Golmant, Nikita Vemuri, Zhewei Yao, Vladimir Feinberg, Amir Gholami, Kai Rothauge, Michael W Mahoney, and Joseph Gonzalez · 2018
Cited alongside, same era.
Applied machine learning at Facebook: A datacenter infrastructure perspective
Kim Hazelwood, Sarah Bird, David Brooks, Soumith Chintala, Utku Diril, Dmytro Dzhulgakov, Mohamed Fawzy, Bill Jia, Yangqing Jia, Aditya Kalro, et al · 2018
Cited alongside, same era.
Predictive and adaptive failure mitigation to avert production cloud VM interruptions
Sebastien Levy, Randolph Yao, Youjiang Wu, Yingnong Dang, Peng Huang, Zheng Mu, Pu Zhao, Tarun Ramani, Naga Govindaraju, Xukun Li, et al · 2020
Later among the works it cites.
A system for massively parallel hyperparameter tuning
Liam Li, Kevin Jamieson, Afshin Rostamizadeh, Ekaterina Gonina, Jonathan Ben-Tzur, Moritz Hardt, Benjamin Recht, and Ameet Talwalkar · 2020
Later among the works it cites.
ZeRO: Memory optimizations toward training trillion parameter models
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He · 2020
Later among the works it cites.
Green AI
Roy Schwartz, Jesse Dodge, Noah A. Smith, and Oren Etzioni · 2020
Later among the works it cites.
ALERT: Accurate learning for energy and timeliness
Chengcheng Wan, Muhammad Santriaji, Eri Rogers, Henry Hoffmann, Michael Maire, and Shan Lu · 2020
Later among the works it cites.
Blink: Fast and generic collectives for distributed ML
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning under concept drift: A review
Jie Lu, Anjin Liu, Fan Dong, Feng Gu, Joao Gama, and Guangquan Zhang · 2018
Cited alongside, same era.
Shufflenet v2: Practical guidelines for efficient CNN architecture design
Ningning Ma, Xiangyu Zhang, Hai-Tao Zheng, and Jian Sun · 2018
Cited alongside, same era.
An empirical model of large-batch training
Sam McCandlish, Jared Kaplan, Dario Amodei, and OpenAI Dota Team · 2018
Cited alongside, same era.
A tutorial on thompson sampling
Daniel J Russo, Benjamin Van Roy, Abbas Kazerouni, Ian Osband, Zheng Wen, et al · 2018
Cited alongside, same era.
Don’t decay the learning rate, increase the batch size
Samuel L Smith, Pieter-Jan Kindermans, Chris Ying, and Quoc V Le · 2018
Cited alongside, same era.
Analysis of dawnbench, a time-to-accuracy machine learning performance benchmark
Cody Coleman, Daniel Kang, Deepak Narayanan, Luigi Nardi, Tian Zhao, Jian Zhang, Peter Bailis, Kunle Olukotun, Chris Ré, and Matei Zaharia · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Guanhua Wang, Shivaram Venkataraman, Amar Phanishayee, Nikhil Devanur, Jorgen Thelin, and Ion Stoica · 2020
Later among the works it cites.
Benchmarking the performance and energy efficiency of AI accelerators for AI training
Yuxin Wang, Qiang Wang, Shaohuai Shi, Xin He, Zhenheng Tang, Kaiyong Zhao, and Xiaowen Chu · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush · 2020
Later among the works it cites.
Ansor: Generating high-performance tensor programs for deep learning
Lianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu, Cody Hao Yu, Ameer Haj-Ali, Yida Wang, Jun Yang, Danyang Zhuo, Koushik Sen, Joseph E. Gonzalez, and Ion Stoica · 2020
Later among the works it cites.
Indicator-directed dynamic power management for iterative workloads on GPU-accelerated systems
Pengfei Zou, Ang Li, Kevin Barker, and Rong Ge · 2020
Later among the works it cites.
Dub: Dynamic underclocking and bypassing in NoCs for heterogeneous GPU workloads
Srikant Bharadwaj, Shomit Das, Yasuko Eckert, Mark Oskin, and Tushar Krishna · 2021
Later among the works it cites.
AccelWattch: A power modeling framework for modern GPUs
Vijay Kandiah, Scott Peverelle, Mahmoud Khairy, Junrui Pan, Amogh Manjunath, Timothy G Rogers, Tor M Aamodt, and Nikos Hardavellas · 2021
Later among the works it cites.
Grad-match: Gradient matching based data subset selection for efficient deep model training
Krishnateja Killamsetty, S Durga, Ganesh Ramakrishnan, Abir De, and Rishabh Iyer · 2021
Later among the works it cites.
Oort: Efficient federated learning via guided participant selection
Fan Lai, Xiangfeng Zhu, Harsha V Madhyastha, and Mosharaf Chowdhury · 2021
Later among the works it cites.
Bao: Making learned query optimization practical
Ryan Marcus, Parimarjan Negi, Hongzi Mao, Nesime Tatbul, Mohammad Alizadeh, and Tim Kraska · 2021
Later among the works it cites.
A survey on hardware accelerators and optimization techniques for RNNs
Sparsh Mittal and Sumanth Umesh · 2021
Later among the works it cites.
Batchsizer: Power-performance tradeoff for DNN inference
Seyed Morteza Nabavinejad, Sherief Reda, and Masoumeh Ebrahimi · 2021
Later among the works it cites.
Carbon emissions and large neural network training
David Patterson, Joseph Gonzalez, Quoc Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David So, Maud Texier, and Jeff Dean · 2021
Later among the works it cites.
Pollux: Co-adaptive cluster scheduling for goodput-optimized deep learning
Aurick Qiao, Sang Keun Choe, Suhas Jayaram Subramanya, Willie Neiswanger, Qirong Ho, Hao Zhang, Gregory R Ganger, and Eric P Xing · 2021
Later among the works it cites.
EdgeBERT: Sentence-level energy optimizations for latency-aware multi-task NLP inference
Thierry Tambe, Coleman Hooper, Lillian Pentecost, Tianyu Jia, En-Yu Yang, Marco Donato, Victor Sanh, Paul Whatmough, Alexander M Rush, David Brooks, et al · 2021
Later among the works it cites.
Dynamic GPU energy optimization for machine learning training workloads
Farui Wang, Weizhe Zhang, Shichao Lai, Meng Hao, and Zheng Wang · 2021
Later among the works it cites.
PET: Optimizing tensor programs with partially equivalent transformations and automated corrections
Haojie Wang, Jidong Zhai, Mingyu Gao, Zixuan Ma, Shizhi Tang, Liyan Zheng, Yuanzhi Li, Kaiyuan Rong, Yuanyong Chen, and Zhihao Jia · 2021
Later among the works it cites.
GSPMD: general and scalable parallelization for ML computation graphs
Yuanzhong Xu, HyoukJoong Lee, Dehao Chen, Blake Hechtman, Yanping Huang, Rahul Joshi, Maxim Krikun, Dmitry Lepikhin, Andy Ly, Marcello Maggioni, et al · 2021
Later among the works it cites.
Fluid: Resource-aware hyperparameter tuning engine
Peifeng Yu, Jiachen Liu, and Mosharaf Chowdhury · 2021
Later among the works it cites.
Treehouse: A case for carbon-aware datacenter software
Thomas Anderson, Adam Belay, Mosharaf Chowdhury, Asaf Cidon, and Irene Zhang · 2022
Closest in time.
Measuring the carbon intensity of AI in cloud instances
Jesse Dodge, Taylor Prewitt, Remi Tachet des Combes, Erika Odmark, Roy Schwartz, Emma Strubell, Alexandra Sasha Luccioni, Noah A. Smith, Nicole DeCario, and Will Buchanan · 2022
Closest in time.
Learning rates as a function of batch size: A random matrix theory approach to neural network training
Diego Granziol, Stefan Zohren, and Stephen Roberts · 2022
Closest in time.
Cost efficient GPU cluster management for training and inference of deep learning
Dong-Ki Kang, Ki-Beom Lee, and Young-Chon Kim · 2022
Closest in time.
Weixin Liang and James Zou · 2022
Closest in time.
Matchmaker: Data drift mitigation in machine learning for large-scale systems
Ankur Mallick, Kevin Hsieh, Behnaz Arzani, and Gauri Joshi · 2022
Closest in time.
MLaaS in the wild: Workload analysis and scheduling in large-scale heterogeneous GPU clusters
Qizhen Weng, Wencong Xiao, Yinghao Yu, Wei Wang, Cheng Wang, Jian He, Yong Li, Liping Zhang, Wei Lin, and Yu Ding · 2022
Closest in time.
Sustainable AI: Environmental implications, challenges and opportunities
Carole-Jean Wu, Ramya Raghavendra, Udit Gupta, Bilge Acun, Newsha Ardalani, Kiwan Maeng, Gloria Chang, Fiona Aga, Jinshi Huang, Charles Bai, Michael Gschwind, Anurag Gupta, Myle Ott, Anastasia Melnikov, Salvatore Candido, David Brooks, Geeta Chauhan, Benjamin Lee, Hsien-Hsin Lee, Bugra Akyildiz, Maximilian Balandat, Joe Spisak, Ravi Jain, Mike Rabbat, and Kim Hazelwood · 2022
Closest in time.