Fetching the paper…
Reading the bibliography…
This paper systematically explores neural functional networks (NFN) for transformer architectures.
A NEW MEASURE OF RANK CORRELATION
M. G. Kendall · 1938
Earlier work this paper cites.
Learning internal representations by error propagation
David E. Rumelhart, Geoffrey E. Hinton, and Ronald J. Williams · 1986
Earlier work this paper cites.
On the algebraic structure of feedforward network weight spaces
Robert Hecht-Nielsen · 1990
Earlier work this paper cites.
On the geometry of feedforward neural network error surfaces
An Mei Chen, Haw-minn Lu, and Robert Hecht-Nielsen · 1993
Earlier work this paper cites.
Recovering a feed-forward net from its output
Charles Fefferman and Scott Markel · 1993
Earlier work this paper cites.
Functionally equivalent feedforward neural networks
Vera Kurkova and Paul C Kainen · 1994
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Full rank factorization of matrices
Robert Piziak and Patrick L Odell · 1999
Earlier work this paper cites.
Evolution and design of distributed learning rules
Thomas Philip Runarsson and Magnus Thor Jonsson · 2000
Earlier work this paper cites.
Random forests
Leo Breiman · 2001
Earlier work this paper cites.
The mnist database of handwritten digits
Yann LeCun and Corinna Cortes · 2005
Earlier work this paper cites.
Compositional pattern producing networks: A novel abstraction of development
Kenneth O Stanley · 2007
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton · 2012
Earlier work this paper cites.
On the optimization of a synaptic learning rule
Samy Bengio, Yoshua Bengio, Jocelyn Cloutier, and Jan Gecsei · 2013
Earlier work this paper cites.
Learning phrase representations using RNN encoder–decoder for statistical machine translation
Kyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, X. Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott E. Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun · 2015
Earlier work this paper cites.
Learning to learn by gradient descent by gradient descent
Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W Hoffman, David Pfau, Tom Schaul, Brendan Shillingford, and Nando De Freitas · 2016
Earlier work this paper cites.
Xgboost: A scalable tree boosting system
Tianqi Chen and Carlos Guestrin · 2016
Earlier work this paper cites.
A decomposable attention model for natural language inference
Ankur Parikh, Oscar Täckström, Dipanjan Das, and Jakob Uszkoreit · 2016
Earlier work this paper cites.
Lightgbm: A highly efficient gradient boosting decision tree
Guolin Ke, Qi Meng, Thomas Finley, Taifeng Wang, Wei Chen, Weidong Ma, Qiwei Ye, and Tie-Yan Liu · 2017
Earlier work this paper cites.
Structured attention networks
Yoon Kim, Carl Denton, Luong Hoang, and Alexander M. Rush · 2017
Earlier work this paper cites.
A STRUCTURED SELF-ATTENTIVE SENTENCE EMBEDDING
Zhouhan Lin, Minwei Feng, Cicero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou, and Yoshua Bengio · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Accelerating neural architecture search using performance prediction
Bowen Baker, Otkrist Gupta, Ramesh Raskar, and Nikhil Naik · 2018
Cited alongside, same era.
Sensitivity and generalization in neural networks: an empirical study
Roman Novak, Yasaman Bahri, Daniel A Abolafia, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Cited alongside, same era.
Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations
M. Raissi, P. Perdikaris, and G.E. Karniadakis · 2018
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Cited alongside, same era.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Cited alongside, same era.
A fine-tuning approach to belief state modeling
Samuel Sokota, Hengyuan Hu, David J Wu, J Zico Kolter, Jakob Nicolaus Foerster, and Noam Brown · 2021
Later among the works it cites.
From data to functa: Your data point is a function and you can treat it like one
Emilien Dupont, Hyunjik Kim, S. M. Ali Eslami, Danilo Jimenez Rezende, and Dan Rosenbaum · 2022
Later among the works it cites.
Generative models as distributions of functions
Emilien Dupont, Yee Whye Teh, and Arnaud Doucet · 2022
Later among the works it cites.
On the symmetries of deep learning models and their internal representations
Charles Godfrey, Davis Brown, Tegan Emerson, and Henry Kvinge · 2022
Later among the works it cites.
Velo: Training versatile learned optimizers by scaling up
Luke Metz, James Harrison, C Daniel Freeman, Amil Merchant, Lucas Beyer, James Bradbury, Naman Agrawal, Ben Poole, Igor Mordatch, Adam Roberts, et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
Simon Du, Jason Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2019
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin · 2019
Cited alongside, same era.
Turbulence forecasting via neural ode
Gavin D Portwood, Peetak P Mitra, Mateus Dias Ribeiro, Tan Minh Nguyen, Balasubramanya T Nadiga, Juan A Saenz, Michael Chertkov, Animesh Garg, Anima Anandkumar, Andreas Dengel, et al · 2019
Cited alongside, same era.
Efficient numerical algorithms for constructing orthogonal generalized doubly stochastic matrices
Alicja Smoktunowicz, Ryszard Kozera, and Gianluca Oderda · 2019
Cited alongside, same era.
Functional vs. parametric equivalence of relu networks
Phuong Bui Thi Mai and Christoph Lampert · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Cited alongside, same era.
Fast model editing at scale
Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, and Christopher D. Manning · 2022
Later among the works it cites.
Learning to throw with a handful of samples using decision transformers
Maxim Monastirsky, Osher Azulay, and Avishai Sintov · 2022
Later among the works it cites.
Improving transformers with probabilistic attention keys
Tam Minh Nguyen, Tan Minh Nguyen, Dung DD Le, Duy Khuong Nguyen, Viet-Anh Tran, Richard Baraniuk, Nhat Ho, and Stanley Osher · 2022
Later among the works it cites.
Learning to learn with generative models of neural network checkpoints
William Peebles, Ilija Radosavovic, Tim Brooks, Alexei A Efros, and Jitendra Malik · 2022
Later among the works it cites.
Set-based neural network encoding
Bruno Andreis, Soro Bedionita, and Sung Ju Hwang · 2023
Later among the works it cites.
Nern: Learning neural representations for neural networks
Maor Ashkenazi, Zohar Rimon, Ron Vainshtein, Shir Levi, Elad Richardson, Pinchas Mintz, and Eran Treister · 2023
Later among the works it cites.
Spatial functa: Scaling functa to imagenet classification and generation
Matthias Bauer, Emilien Dupont, Andy Brock, Dan Rosenbaum, Jonathan Richard Schwarz, and Hyunjik Kim · 2023
Later among the works it cites.
Hyperdiffusion: Generating implicit neural fields with weight-space diffusion
Ziya Erkoç, Fangchang Ma, Qi Shan, Matthias Nießner, and Angela Dai · 2023
Later among the works it cites.
Deep learning on implicit neural representations of shapes
Luca De Luigi, Adriano Cardace, Riccardo Spezialetti, Pierluigi Zama Ramirez, Samuele Salti, and Luigi Di Stefano · 2023
Later among the works it cites.
Equivariant architectures for learning in deep weight spaces
Aviv Navon, Aviv Shamsian, Idan Achituve, Ethan Fetaya, Gal Chechik, and Haggai Maron · 2023
Later among the works it cites.
Robust speech recognition via large-scale weak supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever · 2023
Later among the works it cites.
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao · 2024
Closest in time.
Scale equivariant graph metanetworks
Ioannis Kalogeropoulos, Giorgos Bouritsas, and Yannis Panagakis · 2024
Closest in time.
Graph neural networks for learning equivariant representations of neural networks
Miltiadis Kofinas, Boris Knyazev, Yan Zhang, Yunlu Chen, Gertjan J. Burghouts, Efstratios Gavves, Cees G. M. Snoek, and David W. Zhang · 2024
Closest in time.
Graph metanetworks for processing diverse neural architectures
Derek Lim, Haggai Maron, Marc T. Law, Jonathan Lorraine, and James Lucas · 2024
Closest in time.
Monomial matrix group equivariant neural functional networks
Hoang Tran, Thieu Vo, Tho Huu, An Nguyen The, and Tan Nguyen · 2024
Closest in time.
Equivariant polynomial functional networks, 2024
Thieu N. Vo, Viet-Hoang Tran, Tho Tran Huu, An Nguyen The, Thanh Tran, Minh-Khoi Nguyen-Nhat, Duy-Tung Pham, and Tan Minh Nguyen · 2024
Closest in time.
Transformer meets twicing: Harnessing unattended residual information
Laziz Abdullaev and Tan Minh Nguyen · 2025
Closest in time.
Unveiling the hidden structure of self-attention via kernel principal component analysis
Rachel SY Teo and Tan Nguyen · 2025
Closest in time.
Demystifying the token dynamics of deep selective state space models
Thieu Vo, Duy-Tung Pham, Xin T. Tong, and Tan Minh Nguyen · 2025
Closest in time.