Fetching the paper…
Reading the bibliography…
Recent progress in hardware and methodology for training neural networks has ushered in a new generation of large networks trained on abundant data.
Algorithms for hyper-parameter optimization
James S Bergstra, Rémi Bardenet, Yoshua Bengio, and Balázs Kégl. 2011 · 2011
Earlier work this paper cites.
Random search for hyper-parameter optimization
James Bergstra and Yoshua Bengio. 2012 · 2012
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
Jasper Snoek, Hugo Larochelle, and Ryan P Adams. 2012 · 2012
Earlier work this paper cites.
Neural Machine Translation by Jointly Learning to Align and Translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Thang Luong, Hieu Pham, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
An analysis of deep neural network models for practical applications
Alfredo Canziani, Adam Paszke, and Eugenio Culurciello. 2016 · 2016
Earlier work this paper cites.
Evaluating the energy efficiency of deep convolutional neural networks on cpus and gpus
Da Li, Xinbo Chen, Michela Becchi, and Ziliang Zong. 2016 · 2016
Earlier work this paper cites.
Clicking Clean: Who is winning the race to build a green internet?
Gary Cook, Jude Lee, Tamina Tsai, Ada Kongn, John Deans, Brian Johnson, Elizabeth Jardim, and Brian Johnson. 2017 · 2017
Cited alongside, same era.
Deep biaffine attention for neural dependency parsing
Timothy Dozat and Christopher D. Manning. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Uptime Institute Global Data Center Survey
Rhonda Ascierto. 2018 · 2018
Cited alongside, same era.
Net Public Electricity Generation in Germany in 2018
Bruno Burger. 2019 · 2018
Cited alongside, same era.
Emissions & Generation Resource Integrated Database (eGRID)
EPA. 2018 · 2018
Cited alongside, same era.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Later among the works it cites.
Linguistically-Informed Self-Attention for Semantic Role Labeling
Emma Strubell, Patrick Verga, Daniel Andor, David Weiss, and Andrew McCallum. 2018 · 2018
Later among the works it cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Closest in time.
BERT Meets GPUs
Christopher Forster, Thor Johnsen, Swetha Mandava, Sharath Turuvekere Sreenivas, Deyu Fu, Julie Bernauer, Allison Gray, Sharan Chetlur, and Raul Puri. 2019 · 2019
Closest in time.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
David R. So, Chen Liang, and Quoc V. Le. 2019 · 2019
Closest in time.