Fetching the paper…
Reading the bibliography…
Recent empirical studies have identified fixed point iteration phenomena in deep neural networks, where the hidden state tends to stabilize after several layers, showing minimal change in subsequent layers.
A fast fixed-point algorithm for independent component analysis
Aapo Hyvärinen and Erkki Oja · 1997
Earlier work this paper cites.
Fixed point theory and applications
Ravi P Agarwal, Maria Meehan, and Donal O’regan · 2001
Earlier work this paper cites.
Fixed point theory: an introduction
Vasile I Istratescu · 2001
Earlier work this paper cites.
Fixed point theory
Andrzej Granas and James Dugundji · 2003
Earlier work this paper cites.
Theoretical Numerical Analysis: A Functional Analysis Framework
Kendall Atkinson and Weimin Han · 2009
Earlier work this paper cites.
Anderson acceleration for fixed-point iterations
Homer F Walker and Peng Ni · 2011
Earlier work this paper cites.
A quasi-newton acceleration for high-dimensional optimization algorithms
Hua Zhou, David Alexander, and Kenneth Lange · 2011
Earlier work this paper cites.
On a new faster implicit fixed point iterative scheme in convex metric spaces
Renu Chugh, Preety Malik, and Vivek Kumar · 2015
Earlier work this paper cites.
A new approach to the study of fixed point theory for simulation functions
Farshid Khojasteh, Satish Shukla, and Stojan Radenović · 2015
Earlier work this paper cites.
Loopy neural nets: Imitating feedback loops in the human brain
Isaac Caswell, Chuanqi Shen, and Lisa Wang · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Deep equilibrium models
Shaojie Bai, J Zico Kolter, and Vladlen Koltun · 2019
Earlier work this paper cites.
Training deep nets with progressive batch normalization on multi-gpus
Lianke Qin, Yifan Gong, Tianqi Tang, Yutian Wang, and Jiangming Jin · 2019
Earlier work this paper cites.
Recommendation system based on deep learning methods: a systematic review and new directions
Aminu Da’u and Naomie Salim · 2020
Earlier work this paper cites.
Globally convergent type-i anderson acceleration for nonsmooth fixed-point iterations
Junzi Zhang, Brendan O’Donoghue, and Stephen Boyd · 2020
Earlier work this paper cites.
Iterative approximation of fixed points and applications to two-point second-order boundary value problems and to machine learning
Emirhan Hacioglu, Faik Gürsoy, Samet Maldar, Yunus Atalan, and Gradimir V Milovanović · 2021
Earlier work this paper cites.
Differentiable forward and backward fixed-point iteration layers
Younghan Jeon, Minsik Lee, and Jin Young Choi · 2021
Cited alongside, same era.
Deep face recognition: A survey
Mei Wang and Weihong Deng · 2021
Cited alongside, same era.
Locating and editing factual associations in gpt
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2022
Cited alongside, same era.
Dynamic tensor product regression
Aravind Reddy, Zhao Song, and Lichen Zhang · 2022
Cited alongside, same era.
Convergence of deep convolutional neural networks
Yuesheng Xu and Haizhang Zhang · 2022
Cited alongside, same era.
Fine-tune language models to approximate unbiased in-context learning
Timothy Chu, Zhao Song, and Chiwun Yang · 2023
Cited alongside, same era.
Bypass exponential time preprocessing: Fast neural network training via weight-data correlation preprocessing
Josh Alman, Zhao Song, Ruizhe Zhang, and Danyang Zhuo · 2024
Closest in time.
Bypassing the exponential dependency: Looped transformers efficiently learn in-context by multi-step gradient descent, 2024
Bo Chen, Xiaoyu Li, Yingyu Liang, Zhenmei Shi, and Zhao Song · 2024
Closest in time.
Layer skip: Enabling early exit inference and self-speculative decoding
Mostafa Elhoushi, Akshat Shrivastava, Diana Liskovich, Basil Hosmer, Bram Wasti, Liangzhen Lai, Anas Mahmoud, Bilge Acun, Saurabh Agarwal, Ahmed Roman, et al · 2024
Closest in time.
Looped transformers for length generalization
Ying Fan, Yilun Du, Kannan Ramchandran, and Kangwook Lee · 2024
Closest in time.
Can looped transformers learn to implement multi-step gradient descent for in-context learning?
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Convergence of two-layer regression with nonlinear units
Yichuan Deng, Zhao Song, and Shenghao Xie · 2023
Cited alongside, same era.
Unmasking transformers: A theoretical approach to data recovery via attention weights
Yichuan Deng, Zhao Song, Shenghao Xie, and Chiwun Yang · 2023
Cited alongside, same era.
Looped transformers as programmable computers
Angeliki Giannou, Shashank Rajput, Jy-yong Sohn, Kangwook Lee, Jason D Lee, and Dimitris Papailiopoulos · 2023
Cited alongside, same era.
Yeqi Gao, Zhao Song, and Shenghao Xie · 2023
Cited alongside, same era.
OpenAI · 2023
Cited alongside, same era.
Efficient sgd neural network training via sublinear activated neuron identification
Lianke Qin, Zhao Song, and Yuanyuan Yang · 2023
Cited alongside, same era.
Khashayar Gatmiry, Nikunj Saunshi, Sashank J Reddi, Stefanie Jegelka, and Sanjiv Kumar · 2024
Closest in time.
On the expressive power of a variant of the looped transformer
Yihang Gao, Chuanyang Zheng, Enze Xie, Han Shi, Tianyang Hu, Yu Li, Michael K Ng, Zhenguo Li, and Zhaoqiang Liu · 2024
Closest in time.
On statistical rates and provably efficient criteria of latent diffusion transformers (dits)
Jerry Yao-Chieh Hu, Weimin Wu, Zhuoru Li, Sophia Pi, , Zhao Song, and Han Liu · 2024
Closest in time.
A tighter complexity analysis of sparsegpt
Xiaoyu Li, Yingyu Liang, Zhenmei Shi, and Zhao Song · 2024
Closest in time.
Looped relu mlps may be all you need as practical programmable computers
Yingyu Liang, Zhizhou Sha, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Closest in time.
Multi-layer transformers gradient can be approximated in almost linear time
Yingyu Liang, Zhizhou Sha, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Closest in time.
Toward infinite-long prefix in transformer
Yingyu Liang, Zhenmei Shi, Zhao Song, and Chiwun Yang · 2024
Closest in time.
Zhenmei Shi, Yifei Ming, Xuan-Phi Nguyen, Yingyu Liang, and Shafiq Joty · 2024
Closest in time.
Solving attention kernel regression problem via pre-conditioner
Zhao Song, Junze Yin, and Lichen Zhang · 2024
Closest in time.
Training multi-layer over-parametrized neural network in subquadratic time
Zhao Song, Lichen Zhang, and Ruizhe Zhang · 2024
Closest in time.
STanhop: Sparse tandem hopfield model for memory-enhanced time series prediction
Dennis Wu, Jerry Yao-Chieh Hu, Weijian Li, Bo-Yu Chen, and Han Liu · 2024
Closest in time.
Kevin Xu and Issei Sato · 2024
Closest in time.
Uniform convergence of deep neural networks with lipschitz continuous activation functions and variable widths
Yuesheng Xu and Haizhang Zhang · 2024
Closest in time.