Fetching the paper…
Reading the bibliography…
Deep learning models such as the Transformer are often constructed by heuristics and experience.
Convex Analysis
R. Tyrrell Rockafellar · 1970
Earlier work this paper cites.
Projective Geometry
Pierre Samuel · 1988
Earlier work this paper cites.
The concave-convex procedure (CCCP)
Alan L. Yuille and Anand Rangarajan · 2001
Earlier work this paper cites.
A tutorial on energy-based learning
Yann LeCun, Sumit Chopra, Raia Hadsell, M Ranzato, and F Huang · 2006
Earlier work this paper cites.
An overview of bilevel optimization
Benoît Colson, Patrice Marcotte, and Gilles Savard · 2007
Earlier work this paper cites.
Learning fast approximations of sparse coding
Karol Gregor and Yann LeCun · 2010
Earlier work this paper cites.
Proximal splitting methods in signal processing
Patrick L Combettes and Jean-Christophe Pesquet · 2011
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Convergence rates of inexact proximal-gradient methods for convex optimization
Mark Schmidt, Nicolas Le Roux, and Francis R. Bach · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Convex analysis
Patrick Cheridito · 2013
Earlier work this paper cites.
Parsing with compositional vector grammars
Richard Socher, John Bauer, Christopher D. Manning, and Andrew Y. Ng · 2013
Earlier work this paper cites.
Optimization models
Giuseppe C Calafiore and Laurent El Ghaoui · 2014
Earlier work this paper cites.
First-order methods of smooth convex optimization with inexact oracle
Olivier Devolder, François Glineur, and Yurii E. Nesterov · 2014
Earlier work this paper cites.
Understanding alternating minimization for matrix completion
Moritz Hardt · 2014
Earlier work this paper cites.
Deep unfolding: Model-based inspiration of novel deep architectures
John Hershey, Jonathan Le Roux, and Felix Weninger · 2014
Earlier work this paper cites.
Proximal algorithms
Neal Parikh and Stephen Boyd · 2014
Earlier work this paper cites.
GloVe: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning · 2014
Earlier work this paper cites.
A proximal stochastic gradient method with progressive variance reduction
Lin Xiao and Tong Zhang · 2014
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Earlier work this paper cites.
Learning efficient sparse and low rank models
Pablo Sprechmann, Alex Bronstein, and Guillermo Sapiro · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Majorization-minimization algorithms in signal processing, communications, and machine learning
Ying Sun, Prabhu Babu, and Daniel P Palomar · 2016
Cited alongside, same era.
Learning deep ℓ 0 \ell_{0} encoders
Zhangyang Wang, Qing Ling, and Thomas Huang · 2016
Cited alongside, same era.
OptNet: Differentiable optimization as a layer in neural networks
Brandon Amos and J. Zico Kolter · 2017
Cited alongside, same era.
Input convex neural networks
Brandon Amos, Lei Xu, and J. Zico Kolter · 2017
Cited alongside, same era.
A unified view on graph neural networks as graph signal denoising
Yao Ma, Xiaorui Liu, Tong Zhao, Yozen Liu, Jiliang Tang, and Neil Shah · 2020
Later among the works it cites.
Revisiting "over-smoothing" in deep GCNs
Chaoqi Yang, Ruijie Wang, Shuochao Yao, Shengzhong Liu, and Tarek F. Abdelzaher · 2020
Later among the works it cites.
Hongwei Zhang, Tijin Yan, Zenjun Xie, Yuanqing Xia, and Yuan Zhang · 2020
Later among the works it cites.
Geometric deep learning: Grids, groups, graphs, geodesics, and gauges
Michael M. Bronstein, Joan Bruna, Taco Cohen, and Petar Velickovic · 2021
Later among the works it cites.
Graph signal denoising via unrolling networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
From bayesian sparsity to gated recurrent nets
Hao He, Bo Xin, Satoshi Ikehata, and David Wipf · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E. Curtis, and Jorge Nocedal · 2018
Cited alongside, same era.
Inexact proximal gradient methods for non-convex and non-smooth optimization
Bin Gu, De Wang, Zhouyuan Huo, and Heng Huang · 2018
Cited alongside, same era.
Deep bilevel learning
Simon Jenni and Paolo Favaro · 2018
Cited alongside, same era.
Graph attention networks
Petar Velickovic, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio · 2018
Cited alongside, same era.
Siheng Chen and Yonina C Eldar · 2021
Later among the works it cites.
Is attention better than matrix decomposition?
Zhengyang Geng, Meng-Hao Guo, Hongxu Chen, Xia Li, Ke Wei, and Zhouchen Lin · 2021
Later among the works it cites.
Hanxiao Liu, Zihang Dai, David R. So, and Quoc V. Le · 2021
Later among the works it cites.
Elastic graph neural networks
Xiaorui Liu, Wei Jin, Yao Ma, Yaxin Li, Hua Liu, Yiqi Wang, Ming Yan, and Jiliang Tang · 2021
Later among the works it cites.
SOFT: softmax-free transformer with linear complexity
Jiachen Lu, Jinghan Yao, Junge Zhang, Xiatian Zhu, Hang Xu, Weiguo Gao, Chunjing Xu, Tao Xiang, and Li Zhang · 2021
Later among the works it cites.
IGLU: efficient GCN training via lazy updates
S. Deepak Narayanan, Aditya Sinha, Prateek Jain, Purushottam Kar, and Sundararajan Sellamanickam · 2021
Later among the works it cites.
A unified framework for convolution-based graph neural networks, 2021
Xuran Pan, Shiji Song, and Gao Huang · 2021
Later among the works it cites.
Hopfield networks is all you need
Hubert Ramsauer, Bernhard Schäfl, Johannes Lehner, Philipp Seidl, Michael Widrich, Lukas Gruber, Markus Holzleitner, Thomas Adler, David P. Kreil, Michael K. Kopp, Günter Klambauer, Johannes Brandstetter, and Sepp Hochreiter · 2021
Later among the works it cites.
Optimization induced equilibrium networks
Xingyu Xie, Qiuhao Wang, Zenan Ling, Xia Li, Yisen Wang, Guangcan Liu, and Zhouchen Lin · 2021
Later among the works it cites.
Graph neural networks inspired by classical iterative algorithms
Yongyi Yang, Tang Liu, Yangkun Wang, Jinjing Zhou, Quan Gan, Zhewei Wei, Zheng Zhang, Zengfeng Huang, and David Wipf · 2021
Later among the works it cites.
Implicit vs unfolded graph neural networks
Yongyi Yang, Yangkun Wang, Zengfeng Huang, and David Wipf · 2021
Later among the works it cites.
Dirichlet energy constrained learning for deep graph neural networks
Kaixiong Zhou, Xiao Huang, Daochen Zha, Rui Chen, Li Li, Soo-Hyun Choi, and Xia Hu · 2021
Later among the works it cites.
Interpreting and unifying graph neural networks with an optimization framework
Meiqi Zhu, Xiao Wang, Chuan Shi, Houye Ji, and Peng Cui · 2021
Later among the works it cites.
Bregman neural networks
Jordan Frécon, Gilles Gasso, Massimiliano Pontil, and Saverio Salzo · 2022
Closest in time.
cosFormer: Rethinking softmax in attention
Zhen Qin, Weixuan Sun, Hui Deng, Dongxu Li, Yunshen Wei, Baohong Lv, Junjie Yan, Lingpeng Kong, and Yiran Zhong · 2022
Closest in time.
Revisiting over-smoothing in BERT from the perspective of graph
Han Shi, Jiahui Gao, Hang Xu, Xiaodan Liang, Zhenguo Li, Lingpeng Kong, Stephen Lee, and James T Kwok · 2022
Closest in time.