Fetching the paper…
Reading the bibliography…
Transformer architectures have led to remarkable progress in many state-of-art applications.
R. Impagliazzo and R. Paturi, “On the complexity of k-sat,” Journal of Computer and System Sciences , vol. 62, no. 2, pp. 367–375, 2001
2001
Earlier work this paper cites.
R. Impagliazzo, R. Paturi, and F. Zane, “Which problems have strongly exponential complexity?” Journal of Computer and System Sciences , vol. 63, no. 4, p. 512–530, 2001
2001
Earlier work this paper cites.
R. Williams, “A new algorithm for optimal 2-constraint satisfaction and its implications,” Theoretical Computer Science , vol. 348, no. 2, pp. 357–365, 2005
2005
Earlier work this paper cites.
K. Bringmann, “Why walking the dog takes time: Frechet distance has no strongly subquadratic algorithms unless seth fails,” in 2014 IEEE 55th Annual Symposium on Foundations of Computer Science . IEEE, 2014, pp. 661–670
2014
Earlier work this paper cites.
A. Backurs and P. Indyk, “Edit distance cannot be computed in strongly subquadratic time (unless seth is false),” in Proceedings of the forty-seventh annual ACM symposium on Theory of computing , 2015, pp. 51–58
2015
Earlier work this paper cites.
K. Bringmann and M. Künnemann, “Quadratic conditional lower bounds for string problems and dynamic time warping,” in 2015 IEEE 56th Annual Symposium on Foundations of Computer Science . IEEE, 2015, pp. 79–97
2015
Earlier work this paper cites.
A. Abboud, A. Backurs, and V. V. Williams, “Tight hardness results for lcs and other sequence similarity measures,” in 2015 IEEE 56th Annual Symposium on Foundations of Computer Science . IEEE, 2015, pp. 59–78
2015
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
A. Backurs, P. Indyk, and L. Schmidt, “On the fine-grained complexity of empirical risk minimization: Kernel methods and neural networks,” Advances in Neural Information Processing Systems , vol. 30, 2017
2017
Earlier work this paper cites.
P. Indyk, “Beyond p vs. np: quadratic-time hardness for big data problems,” in Proceedings of the 29th ACM Symposium on Parallelism in Algorithms and Architectures , 2017, pp. 1–1
2017
Earlier work this paper cites.
A. Rubinstein, “Hardness of approximate nearest neighbor search,” in Proceedings of the 50th annual ACM SIGACT symposium on theory of computing , 2018, pp. 1260–1268
2018
Earlier work this paper cites.
A. Abboud, A. Backurs, and V. V. Williams, “If the current clique algorithms are optimal, so is valiant’s parser,” SIAM Journal on Computing , vol. 47, no. 6, pp. 2527–2555, 2018
2018
Earlier work this paper cites.
J. D. M.-W. C. Kenton and L. K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of NAACL-HLT , 2019, pp. 4171–4186
2019
Earlier work this paper cites.
C. Yun, S. Bhojanapalli, A. S. Rawat, S. Reddi, and S. Kumar, “Are transformers universal approximators of sequence-to-sequence functions?” in International Conference on Learning Representations , 2019
2019
Earlier work this paper cites.
A. Rubinstein and V. V. Williams, “Seth vs approximation,” ACM SIGACT News , vol. 50, no. 4, pp. 57–76, 2019
2019
Cited alongside, same era.
A. Abboud, V. Cohen-Addad, and H. Houdrougé, “Subquadratic high-dimensional hierarchical clustering,” Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Cited alongside, same era.
N. Kitaev, L. Kaiser, and A. Levskaya, “Reformer: The efficient transformer,” in International Conference on Learning Representations , 2020
2020
Cited alongside, same era.
M. Zaheer, G. Guruganesh, K. A. Dubey, J. Ainslie, C. Alberti, S. Ontanon, P. Pham, A. Ravula, Q. Wang, L. Yang et al. , “Big bird: Transformers for longer sequences,” Advances in Neural Information Processing Systems , vol. 33, pp. 17 283–17 297, 2020
2020
Cited alongside, same era.
J. Lu, J. Yao, J. Zhang, X. Zhu, H. Xu, W. Gao, C. XU, T. Xiang, and L. Zhang, “Soft: Softmax-free transformer with linear complexity,” in Advances in Neural Information Processing Systems , M. Ranzato, A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan, Eds., vol. 34. Curran Associates, Inc., 2021, pp. 21 297–21 309
2021
Later among the works it cites.
Y. Chen, Q. Zeng, H. Ji, and Y. Yang, “Skyformer: Remodel self-attention with gaussian kernel and nystr \ \backslash " om method,” Advances in Neural Information Processing Systems , vol. 34, 2021
2021
Later among the works it cites.
K. Bringmann, “Fine-grained complexity theory: Conditional lower bounds for computational geometry,” in Conference on Computability in Europe . Springer, 2021, pp. 60–70
2021
Later among the works it cites.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 10 012–10 022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al. , “An image is worth 16x16 words: Transformers for image recognition at scale,” in International Conference on Learning Representations , 2021
2021
Cited alongside, same era.
S. Khan, M. Naseer, M. Hayat, S. W. Zamir, F. S. Khan, and M. Shah, “Transformers in vision: A survey,” ACM Computing Surveys (CSUR) , 2021
2021
Cited alongside, same era.
J. Jumper, R. Evans, A. Pritzel, T. Green, M. Figurnov, O. Ronneberger, K. Tunyasuvunakool, R. Bates, A. Žídek, A. Potapenko et al. , “Highly accurate protein structure prediction with alphafold,” Nature , vol. 596, no. 7873, pp. 583–589, 2021
2021
Cited alongside, same era.
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. de Oliveira Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman et al. , “Evaluating large language models trained on code,” CoRR , 2021
2021
Cited alongside, same era.
2021
Later among the works it cites.
Y. Dong, J.-B. Cordonnier, and A. Loukas, “Attention is not all you need: Pure attention loses rank doubly exponentially with depth,” in International Conference on Machine Learning . PMLR, 2021, pp. 2793–2803
2021
Later among the works it cites.
H. Kim, G. Papamakarios, and A. Mnih, “The lipschitz constant of self-attention,” in International Conference on Machine Learning . PMLR, 2021, pp. 5562–5571
2021
Later among the works it cites.
A. Roy, M. Saffar, A. Vaswani, and D. Grangier, “Efficient content-based sparse attention with routing transformers,” Transactions of the Association for Computational Linguistics , vol. 9, pp. 53–68, 2021
2021
Later among the works it cites.
Y. Xiong, Z. Zeng, R. Chakraborty, M. Tan, G. Fung, Y. Li, and V. Singh, “Nyströmformer: A nyström-based algorithm for approximating self-attention,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 35, no. 16, 2021, pp. 14 138–14 148
2021
Later among the works it cites.
V. Likhosherstov, K. M. Choromanski, J. Q. Davis, X. Song, and A. Weller, “Sub-linear memory: How to make performers slim,” Advances in Neural Information Processing Systems , vol. 34, 2021
2021
Later among the works it cites.
H. Peng, N. Pappas, D. Yogatama, R. Schwartz, N. Smith, and L. Kong, “Random feature attention,” in International Conference on Learning Representations , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
K. Bringmann and A. Nusser, “Translating hausdorff is hard: Fine-grained lower bounds for hausdorff distance under translation,” in 37th International Symposium on Computational Geometry , 2021
2021
Later among the works it cites.