Fetching the paper…
Reading the bibliography…
Attention-based models have been a key element of many recent breakthroughs in deep learning.
Quadratic programming as an extension of classical quadratic maximization
Theil, H. and Van de Panne, C · 1960
Earlier work this paper cites.
Least-squares estimation of transformation parameters between two point patterns
Umeyama, S · 1991
Earlier work this paper cites.
Understanding consumption
Deaton, A. et al · 1992
Earlier work this paper cites.
Continuous univariate distributions, volume 1
Johnson, N. L., Kotz, S., and Balakrishnan, N · 1994
Earlier work this paper cites.
Learning and generalization characteristics of the random vector functional-link net
Pao, Y.-H., Park, G.-H., and Sobajic, D. J · 1994
Earlier work this paper cites.
A comparison of methods for trend estimation
Bianchi, M., Boyle, M., and Hollingsworth, D · 1999
Earlier work this paper cites.
Chaos control using least-squares support vector machines
Suykens, J. A. and Vandewalle, J · 1999
Earlier work this paper cites.
Using the nyström method to speed up kernel machines
Williams, C. and Seeger, M · 2000
Earlier work this paper cites.
15.093 Optimization Methods: Lecture 18: Optimality Conditions and Gradient Methods for Unconstrained Optimization
Bertsimas, D · 2009
Earlier work this paper cites.
International economics: Theory and policy
Krugman, P. R. and Obstfeld, M · 2009
Earlier work this paper cites.
The mnist database of handwritten digit images for machine learning research
Deng, L · 2012
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Luong, M.-T., Pham, H., and Manning, C. D · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Matching networks for one shot learning
Vinyals, O., Blundell, C., Lillicrap, T., Wierstra, D., et al · 2016
Earlier work this paper cites.
Hierarchical attention networks for document classification
Yang, Z., Yang, D., Dyer, C., He, X., Smola, A., and Hovy, E · 2016
Earlier work this paper cites.
Optnet: Differentiable optimization as a layer in neural networks
Amos, B. and Kolter, J. Z · 2017
Cited alongside, same era.
Input convex neural networks
Amos, B., Xu, L., and Kolter, J. Z · 2017
Cited alongside, same era.
Extreme entropy machines: robust information theoretic classification
Czarnecki, W. M. and Tabor, J · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Cited alongside, same era.
Prototypical networks for few-shot learning
Snell, J., Swersky, K., and Zemel, R · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al · 2019
Later among the works it cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Later among the works it cites.
Multiplicative interactions and where to find them
Jayakumar, S. M., Czarnecki, W. M., Menick, J., Schwarz, J., Rae, J., Osindero, S., Teh, Y. W., Harley, T., and Pascanu, R · 2020
Later among the works it cites.
A survey of deep meta-learning
Huisman, M., Van Rijn, J. N., and Plaat, A · 2021
Later among the works it cites.
Highly accurate protein structure prediction with alphafold
Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., Tunyasuvunakool, K., Bates, R., Žídek, A., Potapenko, A., et al · 2021
Later among the works it cites.
The lipschitz constant of self-attention
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
JAX: composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q · 2018
Cited alongside, same era.
Neural scene representation and rendering
Eslami, S. A., Jimenez Rezende, D., Besse, F., Viola, F., Morcos, A. S., Garnelo, M., Ruderman, A., Rusu, A. A., Danihelka, I., Gregor, K., et al · 2018
Cited alongside, same era.
A simple neural attentive meta-learner
Mishra, N., Rohaninejad, M., Chen, X., and Abbeel, P · 2018
Cited alongside, same era.
Sentiment analysis by capsules
Wang, Y., Sun, A., Han, J., Liu, Y., and Zhu, X · 2018
Cited alongside, same era.
Hierarchical attention transfer network for cross-domain sentiment classification
Zheng, L., Ying, W., Yu, Z., Qiang, Y., et al · 2018
Cited alongside, same era.
Dota 2 with large scale deep reinforcement learning
Berner, C., Brockman, G., Chan, B., Cheung, V., Debiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., et al · 2019
Cited alongside, same era.
Kim, H., Papamakarios, G., and Mnih, A · 2021
Later among the works it cites.
Pure transformers are powerful graph learners
Kim, J., Nguyen, D. T., Min, S., Cho, S., Lee, M., Lee, H., and Hong, S · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with clip latents
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M · 2022
Later among the works it cites.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E., Ghasemipour, S. K. S., Gontijo-Lopes, R., Ayan, B. K., Salimans, T., Ho, J., Fleet, D. J., and Norouzi, M · 2022
Later among the works it cites.
Chatgpt: Optimizing language models for dialogue, 2022
Schulman, J., Zoph, B., Kim, C., Hilton, J., Menick, J., Weng, J., Uribe, J., Fedus, L., Metz, L., Pokorny, M., et al · 2022
Later among the works it cites.
Language models generalize beyond natural proteins
Verkuil, R., Kabeli, O., Du, Y., Wicky, B. I., Milles, L. F., Dauparas, J., Baker, D., Ovchinnikov, S., Sercu, T., and Rives, A · 2022
Later among the works it cites.
Transformers learn in-context by gradient descent
von Oswald, J., Niklasson, E., Randazzo, E., Sacramento, J., Mordvintsev, A., Zhmoginov, A., and Vladymyrov, M · 2022
Later among the works it cites.
Coin and its various meanings: a rant about the english language
Tschernakki, K. W · 2023
Closest in time.
Neural codec language models are zero-shot text to speech synthesizers
Wang, C., Chen, S., Wu, Y., Zhang, Z., Zhou, L., Liu, S., Chen, Z., Liu, Y., Wang, H., Li, J., et al · 2023
Closest in time.