Fetching the paper…
Reading the bibliography…
We demonstrate that, hidden within one-layer randomly weighted neural networks, there exist subnetworks that can achieve impressive performance, without ever modifying the weight initializations, on machine translation tasks.
No training required: Exploring random encoders for sentence classification
John Wieting and Douwe Kiela. 2019 · 1901
Earlier work this paper cites.
Weight agnostic neural networks
Adam Gaier and David Ha. 2019 · 1906
Earlier work this paper cites.
Playing the lottery with rewards and multiple languages: lottery tickets in rl and nlp
Haonan Yu, Sergey Edunov, Yuandong Tian, and Ari S Morcos. 2019 · 1906
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Megatron-LM: Training multi-billion parameter language models using gpu model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019 · 1909
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2019 · 1910
Earlier work this paper cites.
Further experiments with papa
A. Gamba, L. Gamberini, G. Palmieri, and R. Sanna. 1961 · 1965
Earlier work this paper cites.
On the capabilities of multilayer perceptrons
Eric B Baum. 1988 · 1988
Earlier work this paper cites.
Feedforward neural networks with random weights
Wouter F Schmidt, Martin A Kraaijveld, and Robert PW Duin. 1992 · 1992
Earlier work this paper cites.
Learning and generalization characteristics of the random vector functional-link net
Yoh-Han Pao, Gwang-Hoon Park, and Dejan J Sobajic. 1994 · 1994
Earlier work this paper cites.
Real-time computing without stable states: A new framework for neural computation based on perturbations
Wolfgang Maass, Thomas Natschläger, and Henry Markram. 2002 · 2002
Earlier work this paper cites.
On the impressive performance of randomly weighted encoders in summarization tasks
Jonathan Pilault, Jaehong Park, and Christopher Pal. 2020 · 2002
Earlier work this paper cites.
Adaptive nonlinear system identification with echo state networks
Herbert Jaeger. 2003 · 2003
Earlier work this paper cites.
Comparing rewinding and fine-tuning in neural network pruning
Alex Renda, Jonathan Frankle, and Michael Carbin. 2020 · 2003
Earlier work this paper cites.
Information-theoretic probing with minimum description length
Elena Voita and Ivan Titov. 2020 · 2003
Earlier work this paper cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 2005
Earlier work this paper cites.
When bert plays the lottery, all tickets are winning
Sai Prasanna, Anna Rogers, and Anna Rumshisky. 2020 · 2005
Earlier work this paper cites.
Masked language modeling for proteins via linearly scalable long-context transformers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Jared Davis, Tamas Sarlos, David Belanger, Lucy Colwell, and Adrian Weller. 2020 · 2006
Earlier work this paper cites.
The lottery ticket hypothesis for pre-trained bert networks
Tianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu, Yang Zhang, Zhangyang Wang, and Michael Carbin. 2020 · 2007
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht. 2008 · 2008
Cited alongside, same era.
Reservoir computing approaches to recurrent neural network training
Mantas Lukoševičius and Herbert Jaeger. 2009 · 2009
Cited alongside, same era.
Weighted sums of random kitchen sinks: Replacing minimization with randomization in learning
Ali Rahimi and Benjamin Recht. 2009 · 2009
Cited alongside, same era.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Léonard, and Aaron Courville. 2013 · 2013
Cited alongside, same era.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013 · 2013
Cited alongside, same era.
Scaling neural machine translation
Myle Ott, Sergey Edunov, David Grangier, and Michael Auli. 2018 · 2018
Later among the works it cites.
Deep equilibrium models
Shaojie Bai, J Zico Kolter, and Vladlen Koltun. 2019 · 2019
Later among the works it cites.
Reducing transformer depth on demand with structured dropout
Angela Fan, Edouard Grave, and Armand Joulin. 2019 · 2019
Later among the works it cites.
Are sixteen heads really better than one?
Paul Michel, Omer Levy, and Graham Neubig. 2019 · 2019
Later among the works it cites.
Deconstructing lottery tickets: Zeros, signs, and the supermask
Hattie Zhou, Janice Lan, Rosanne Liu, and Jason Yosinski. 2019 · 2019
Later among the works it cites.
Linear mode connectivity and the lottery ticket hypothesis
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel Roy, and Michael Carbin. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Findings of the 2014 workshop on statistical machine translation
Ondřej Bojar, Christian Buck, Christian Federmann, Barry Haddow, Philipp Koehn, Johannes Leveling, Christof Monz, Pavel Pecina, Matt Post, Herve Saint-Amand, Radu Soricut, Lucia Specia, and Aleš Tamchyna. 2014 · 2014
Cited alongside, same era.
Report on the 11 th iwslt evaluation campaign , iwslt 2014
M. Cettolo, J. Niehues, S. Stüker, L. Bentivogli, and Marcello Federico. 2015 · 2014
Cited alongside, same era.
GloVe: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014 · 2014
Cited alongside, same era.
Supervised learning of universal sentence representations from natural language inference data
Alexis Conneau, Douwe Kiela, Holger Schwenk, Loic Barrault, and Antoine Bordes. 2017 · 2017
Cited alongside, same era.
Echo state property of deep reservoir computing networks
Claudio Gallicchio and Alessio Micheli. 2017 · 2017
Cited alongside, same era.
First quora dataset release: Question pairs, 2017
Shankar Iyer, Nikhil Dandekar, and Kornl Csernai. 2017 · 2017
Cited alongside, same era.
Randomness in neural networks: an overview
Simone Scardapane and Dianhui Wang. 2017 · 2017
Cited alongside, same era.
Later among the works it cites.
Tinybert: Distilling bert for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2020 · 2020
Later among the works it cites.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020 · 2020
Later among the works it cites.
Train big, then compress: Rethinking model size for efficient training and inference of transformers
Zhuohan Li, Eric Wallace, Sheng Shen, Kevin Lin, Kurt Keutzer, Dan Klein, and Joey Gonzalez. 2020 · 2020
Later among the works it cites.
What’s hidden in a randomly weighted neural network?
Vivek Ramanujan, Mitchell Wortsman, Aniruddha Kembhavi, Ali Farhadi, and Mohammad Rastegari. 2020 · 2020
Later among the works it cites.
Movement pruning: Adaptive sparsity by fine-tuning
Victor Sanh, Thomas Wolf, and Alexander Rush. 2020 · 2020
Later among the works it cites.
Q-bert: Hessian based ultra low precision quantization of bert
Sheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma, Zhewei Yao, Amir Gholami, Michael W Mahoney, and Kurt Keutzer. 2020 · 2020
Later among the works it cites.
Mobilebert: a compact task-agnostic bert for resource-limited devices
Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou. 2020 · 2020
Later among the works it cites.
Supermasks in superposition for continual learning
Mitchell Wortsman, Vivek Ramanujan, Rosanne Liu, Aniruddha Kembhavi, Mohammad Rastegari, Jason Yosinski, and Ali Farhadi. 2020 · 2020
Later among the works it cites.
Playing lottery tickets with vision and language
Zhe Gan, Yen-Chun Chen, Linjie Li, Tianlong Chen, Yu Cheng, Shuohang Wang, and Jingjing Liu. 2021 · 2021
Closest in time.
Random feature attention
Hao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz, Noah Smith, and Lingpeng Kong. 2021 · 2021
Closest in time.
Reservoir transformers
Sheng Shen, Alexei Baevski, Ari S Morcos, Kurt Keutzer, Michael Auli, and Douwe Kiela. 2021 · 2021
Closest in time.
Mlpruning: A multilevel structured pruning framework for transformer-based models
Zhewei Yao, Linjian Ma, Sheng Shen, Kurt Keutzer, and Michael W Mahoney. 2021 · 2021
Closest in time.