Fetching the paper…
Reading the bibliography…
We demonstrate that transformers obtain impressive performance even when some of the layers are randomly initialized and never updated.
No training required: Exploring random encoders for sentence classification
John Wieting and Douwe Kiela. 2019 · 1901
Earlier work this paper cites.
Pay less attention with lightweight and dynamic convolutions
Felix Wu, Angela Fan, Alexei Baevski, Yann N Dauphin, and Michael Auli. 2019 · 1901
Earlier work this paper cites.
Chiyuan Zhang, Samy Bengio, and Yoram Singer. 2019 · 1902
Earlier work this paper cites.
Bert rediscovers the classical nlp pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019 · 1905
Earlier work this paper cites.
Energy and policy considerations for deep learning in nlp
Emma Strubell, Ananya Ganesh, and Andrew McCallum. 2019 · 1906
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Roy Schwartz, Jesse Dodge, Noah A Smith, and Oren Etzioni. 2019 · 1907
Earlier work this paper cites.
Reducing transformer depth on demand with structured dropout
Angela Fan, Edouard Grave, and Armand Joulin. 2019 · 1909
Earlier work this paper cites.
vq-wav2vec: Self-supervised learning of discrete speech representations
Alexei Baevski, Steffen Schneider, and Michael Auli. 2019 · 1910
Earlier work this paper cites.
Maha Elbayad, Jiatao Gu, Edouard Grave, and Michael Auli. 2019 · 1910
Earlier work this paper cites.
Joseph Enguehard, Dan Busbridge, Vitalii Zhelezniak, and Nils Hammerla. 2019 · 1910
Earlier work this paper cites.
Improving transformer models by reordering their sublayers
Ofir Press, Noah A Smith, and Omer Levy. 2019 · 1911
Earlier work this paper cites.
The perceptron: A model for brain functioning. i
Hans-Dieter Block. 1962 · 1962
Earlier work this paper cites.
An outline of a mathematical theory of papa
A Borsellino and A Gamba. 1961 · 1965
Earlier work this paper cites.
Geometrical and statistical properties of systems of linear inequalities with applications in pattern recognition
Thomas M Cover. 1965 · 1965
Earlier work this paper cites.
Further experiments with papa
A. Gamba, L. Gamberini, G. Palmieri, and R. Sanna. 1961 · 1965
Earlier work this paper cites.
Extensions of lipschitz mappings into a hilbert space
William B Johnson and Joram Lindenstrauss. 1984 · 1984
Earlier work this paper cites.
On the capabilities of multilayer perceptrons
Eric B Baum. 1988 · 1988
Earlier work this paper cites.
Feedforward neural networks with random weights
Wouter F Schmidt, Martin A Kraaijveld, and Robert PW Duin. 1992 · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams. 1992 · 1992
Earlier work this paper cites.
Learning and generalization characteristics of the random vector functional-link net
Yoh-Han Pao, Gwang-Hoon Park, and Dejan J Sobajic. 1994 · 1994
Earlier work this paper cites.
Effects of noise on convergence and generalization in recurrent networks
Kam Jim, Bill G Horne, and C Lee Giles. 1995 · 1995
Earlier work this paper cites.
An analysis of noise in recurrent neural networks: convergence and generalization
Kam-Chuen Jim, C Lee Giles, and Bill G Horne. 1996 · 1996
Earlier work this paper cites.
The use of the area under the roc curve in the evaluation of machine learning algorithms
Andrew P Bradley. 1997 · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner. 1998 · 1998
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2001
Earlier work this paper cites.
Reformer: The efficient transformer
Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya. 2020 · 2001
Earlier work this paper cites.
Echo state neural machine translation
Ankush Garg, Yuan Cao, and Qi Ge. 2020 · 2002
Earlier work this paper cites.
Zhuohan Li, Eric Wallace, Sheng Shen, Kevin Lin, Kurt Keutzer, Dan Klein, and Joseph E Gonzalez. 2020 · 2002
Earlier work this paper cites.
Real-time computing without stable states: A new framework for neural computation based on perturbations
Wolfgang Maass, Thomas Natschläger, and Henry Markram. 2002 · 2002
Earlier work this paper cites.
On the impressive performance of randomly weighted encoders in summarization tasks
Jonathan Pilault, Jaehong Park, and Christopher Pal. 2020 · 2002
Earlier work this paper cites.
A primer in bertology: What we know about how bert works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky. 2020 · 2002
Earlier work this paper cites.
Rezero is all you need: Fast convergence at large depth
Thomas Bachlechner, Bodhisattwa Prasad Majumder, Huanru Henry Mao, Garrison W Cottrell, and Julian McAuley. 2020 · 2003
Earlier work this paper cites.
Training batchnorm and only batchnorm: On the expressive power of random features in cnns
Jonathan Frankle, David J Schwab, and Ari S Morcos. 2020 · 2003
Earlier work this paper cites.
Adaptive nonlinear system identification with echo state networks
Herbert Jaeger. 2003 · 2003
Cited alongside, same era.
Longformer: The long-document transformer
Iz Beltagy, Matthew E Peters, and Arman Cohan. 2020 · 2004
Cited alongside, same era.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 2005
Cited alongside, same era.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. 2020 · 2005
Cited alongside, same era.
An introduction to random indexing
Magnus Sahlgren. 2005 · 2005
Cited alongside, same era.
Unconventional Information Processing Systems, Novel Hardware: A Tour D’Horizon
Fatemeh Hadaeghi, Xu He, and Herbert Jaeger. 2017 · 2017
Later among the works it cites.
Decoupled neural interfaces using synthetic gradients
Max Jaderberg, Wojciech Marian Czarnecki, Simon Osindero, Oriol Vinyals, Alex Graves, David Silver, and Koray Kavukcuoglu. 2017 · 2017
Later among the works it cites.
Perceptrons: An introduction to computational geometry
Marvin Minsky and Seymour A Papert. 2017 · 2017
Later among the works it cites.
Event-driven random back-propagation: Enabling neuromorphic deep learning machines
Emre O Neftci, Charles Augustine, Somnath Paul, and Georgios Detorakis. 2017 · 2017
Later among the works it cites.
Regularizing deep neural networks by noise: Its interpretation and optimization
Hyeonwoo Noh, Tackgeun You, Jonghwan Mun, and Bohyung Han. 2017 · 2017
Later among the works it cites.
Randomness in neural networks: an overview
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Synthesizer: Rethinking self-attention in transformer models
Yi Tay, Dara Bahri, Donald Metzler, Da-Cheng Juan, Zhe Zhao, and Che Zheng. 2020a · 2005
Cited alongside, same era.
Masked language modeling for proteins via linearly scalable long-context transformers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Jared Davis, Tamas Sarlos, David Belanger, Lucy Colwell, and Adrian Weller. 2020 · 2006
Cited alongside, same era.
Extreme learning machine: theory and applications
Guang-Bin Huang, Qin-Yu Zhu, and Chee-Kheong Siew. 2006 · 2006
Cited alongside, same era.
Deep encoder, shallow decoder: Reevaluating the speed-quality tradeoff in machine translation
Jungo Kasai, Nikolaos Pappas, Hao Peng, James Cross, and Noah A Smith. 2020 · 2006
Cited alongside, same era.
Transformers are rnns: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. 2020 · 2006
Cited alongside, same era.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Li, Madian Khabsa, Han Fang, and Hao Ma. 2020 · 2006
Cited alongside, same era.
Randomized automatic differentiation
Deniz Oktay, Nick McGreivy, Joshua Aduol, Alex Beatson, and Ryan P Adams. 2020 · 2007
Cited alongside, same era.
Simone Scardapane and Dianhui Wang. 2017 · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Later among the works it cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel R Bowman. 2017 · 2017
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Later among the works it cites.
Classical structured prediction losses for sequence to sequence learning
Sergey Edunov, Myle Ott, Michael Auli, David Grangier, and Marc’Aurelio Ranzato. 2018 · 2018
Later among the works it cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin. 2018 · 2018
Later among the works it cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang. 2018 · 2018
Later among the works it cites.
Scaling neural machine translation
Myle Ott, Sergey Edunov, David Grangier, and Michael Auli. 2018 · 2018
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2018 · 2018
Later among the works it cites.
Deep image prior
Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky. 2018 · 2018
Later among the works it cites.
Language modeling teaches you more than translation does: Lessons learned through auxiliary syntactic task analysis
Kelly Zhang and Samuel Bowman. 2018 · 2018
Later among the works it cites.
Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov. 2019 · 2019
Later among the works it cites.
Gradient descent finds global minima of deep neural networks
Simon Du, Jason Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai. 2019 · 2019
Later among the works it cites.
What does BERT learn about the structure of language?
Ganesh Jawahar, Benoît Sagot, and Djamé Seddah. 2019 · 2019
Later among the works it cites.
Randomly weighted cnns for (music) audio classification
Jordi Pons and Xavier Serra. 2019 · 2019
Later among the works it cites.
Intriguing properties of randomly weighted networks: Generalizing while learning next to nothing
Amir Rosenfeld and John K Tsotsos. 2019 · 2019
Later among the works it cites.
Recent advances in physical reservoir computing: A review
Gouhei Tanaka, Toshiyuki Yamane, Jean Benoit Héroux, Ryosho Nakane, Naoki Kanazawa, Seiji Takeda, Hidetoshi Numata, Daiju Nakano, and Akira Hirose. 2019 · 2019
Later among the works it cites.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M. Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H. Choi, Richard Powell, Timo Ewalds, Petko Georgiev, Junhyuk Oh, Dan Horgan, Manuel Kroiss, Ivo Danihelka, Aja Huang, Laurent Sifre, Trevor Cai, John P. Agapiou, Max Jaderberg, Alexander S. Vezhnevets, Rémi Leblond, Tobias Pohlen, Valentin Dalibard, David Budden, Yury Sulsky, James Molloy, Tom L. Paine, Caglar Gulcehre, Ziyu Wang, Tobias Pfaff, Yuhuai Wu, Roman Ring, Dani Yogatama, Dario Wünsch, Katrina McKinney, Oliver Smith, Tom Schaul, Timothy Lillicrap, Koray Kavukcuoglu, Demis Hassabis, Chris Apps, and David Silver. 2019 · 2019
Later among the works it cites.
Deconstructing lottery tickets: Zeros, signs, and the supermask
Hattie Zhou, Janice Lan, Rosanne Liu, and Jason Yosinski. 2019 · 2019
Later among the works it cites.
PANLP at MEDIQA 2019: Pre-trained language models, transfer learning and knowledge distillation
Wei Zhu, Xiaofeng Zhou, Keqiang Wang, Xun Luo, Xiepeng Li, Yuan Ni, and Guotong Xie. 2019 · 2019
Later among the works it cites.
ETC: Encoding long and structured inputs in transformers
Joshua Ainslie, Santiago Ontanon, Chris Alberti, Vaclav Cvicek, Zachary Fisher, Philip Pham, Anirudh Ravula, Sumit Sanghai, Qifan Wang, and Li Yang. 2020 · 2020
Closest in time.
Deep randomized neural networks
Claudio Gallicchio and Simone Scardapane. 2020 · 2020
Closest in time.
What’s hidden in a randomly weighted neural network?
Vivek Ramanujan, Mitchell Wortsman, Aniruddha Kembhavi, Ali Farhadi, and Mohammad Rastegari. 2020 · 2020
Closest in time.
Q-bert: Hessian based ultra low precision quantization of bert
Sheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma, Zhewei Yao, Amir Gholami, Michael W Mahoney, and Kurt Keutzer. 2020 · 2020
Closest in time.
Mobilebert: a compact task-agnostic bert for resource-limited devices
Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou. 2020 · 2020
Closest in time.
Information-theoretic probing with minimum description length
Elena Voita and Ivan Titov. 2020 · 2020
Closest in time.
Pipetransformer: Automated elastic pipelining for distributed training of transformers
Chaoyang He, Shen Li, Mahdi Soltanolkotabi, and Salman Avestimehr. 2021 · 2021
Closest in time.
Random feature attention
Hao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz, Noah Smith, and Lingpeng Kong. 2021 · 2021
Closest in time.