Fetching the paper…
Reading the bibliography…
The power that machine learning models consume when making predictions can be affected by a model's architecture.
Weight limiting, weight quantisation and generalisation in multi-layer perceptrons
PC Woodland · 1989
Earlier work this paper cites.
Regularization theory and neural networks architectures
Federico Girosi, Michael Jones, and Tomaso Poggio · 1995
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Automatic early stopping using cross validation: quantifying the criteria
Lutz Prechelt · 1998
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Sequence-level knowledge distillation
Yoon Kim and Alexander M Rush · 2016
Earlier work this paper cites.
Inside Volta: The world’s most advanced data center GPU
Luke Durant, Olivier Giroux, Mark Harris, and Nick Stam · 2017
Earlier work this paper cites.
Self-normalizing neural networks
Günter Klambauer, Thomas Unterthiner, Andreas Mayr, and Sepp Hochreiter · 2017
Cited alongside, same era.
AI and Compute
Dario Amodei and Danny Hernandez · 2018
Cited alongside, same era.
NVIDIA tensor core programmability, performance & precision
Stefano Markidis, Steven Wei Der Chien, Erwin Laure, Ivy Bo Peng, and Jeffrey S Vetter · 2018
Cited alongside, same era.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2018
Cited alongside, same era.
Training deep neural networks with 8-bit floating point numbers
Naigang Wang, Jungwook Choi, Daniel Brand, Chia-Yu Chen, and Kailash Gopalakrishnan · 2018
Cited alongside, same era.
End-to-end streaming keyword spotting
Raziel Alvarez and Hyun-Jin Park · 2019
Instruction tables
Agner Fog · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Later among the works it cites.
Roy Schwartz, Jesse Dodge, Noah A. Smith, and Oren Etzioni · 2019
Later among the works it cites.
Energy and policy considerations for deep learning in NLP
Emma Strubell, Ananya Ganesh, and Andrew McCallum · 2019
Later among the works it cites.
Proceedings of the first workshop on simple and efficient natural language processing
Angela Fan, Goran Glavaš, Shafiq Joty, Nafise Sadat Moosav, Vered Shwartz, Alex Wang, and Thomas Wolf · 2020
Closest in time.
Towards the systematic reporting of the energy and carbon footprints of machine learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Peter Henderson, Jieru Hu, Joshua Romoff, Emma Brunskill, Dan Jurafsky, and Joelle Pineau · 2020
Closest in time.