Fetching the paper…
Reading the bibliography…
Transformers have become the de-facto standard model in artificial intelligence since 2017 despite numerous shortcomings ranging from energy inefficiency to hallucinations.
Kiefer, J., Wolfowitz, J.: Stochastic estimation of the maximum of a regression function. The Annals of Mathematical Statistics (1952)
1952
Earlier work this paper cites.
Skinner, B.F.: Reinforcement today. American Psychologist 13
1958
Earlier work this paper cites.
Kalman, R.E.: A new approach to linear filtering and prediction problems (1960)
1960
Earlier work this paper cites.
Williams, R.J., Peng, J.: Function optimization using connectionist reinforcement learning algorithms. Connection Science (1991)
1991
Earlier work this paper cites.
Maass, W.: Networks of spiking neurons: the third generation of neural network models. Neural networks 10
1997
Earlier work this paper cites.
Schultz, M., Joachims, T.: Learning a distance metric from relative comparisons. Advances in neural information processing systems 16
2003
Earlier work this paper cites.
Bengio, Y., Louradour, J., Collobert, R., Weston, J.: Curriculum learning. In: Proceedings of the 26th annual international conference on machine learning. pp. 41–48 (2009)
2009
Earlier work this paper cites.
Nair, V., Hinton, G.E.: Rectified linear units improve restricted boltzmann machines. In: Int. Conf. on machine learning (ICML-) (2010)
2010
Earlier work this paper cites.
Hinton, G.E., Krizhevsky, A., Wang, S.D.: Transforming auto-encoders. In: Artificial Neural Networks and Machine Learning–ICANN 2011: 21st International Conference on Artificial Neural Networks, Espoo, Finland, June 14-17, 2011, Proceedings, Part I 21. pp. 44–51 (2011)
2011
Earlier work this paper cites.
Schneider, J., Wattenhofer, R.: Trading bit, message, and time complexity of distributed algorithms. In: International Symposium on Distributed Computing. pp. 51–65. Springer (2011)
2011
Earlier work this paper cites.
Yuksel, S.E., Wilson, J.N., Gader, P.D.: Twenty years of mixture of experts. IEEE transactions on neural networks and learning systems 23
2012
Earlier work this paper cites.
Kingma, D.P., Ba, J.: Adam: A method for stochastic optimization. arXiv:1412.6980 (2014)
2014
Earlier work this paper cites.
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., Salakhutdinov, R.: Dropout: a simple way to prevent neural networks from overfitting. The journal of machine learning research (2014)
2014
Earlier work this paper cites.
Ioffe, S., Szegedy, C.: Batch normalization: Accelerating deep network training by reducing internal covariate shift. In: Int. Conf. on machine learning (2015)
2015
Earlier work this paper cites.
Ba, J.L., Kiros, J.R., Hinton, G.E.: Layer normalization. arXiv:1607.06450 (2016)
2016
Earlier work this paper cites.
Goodfellow, I., Bengio, Y., Courville, A.: Deep learning (2016)
2016
Earlier work this paper cites.
Grover, A., Leskovec, J.: node2vec: Scalable feature learning for networks. In: ACM SIGKDD Int. Conf. on Knowledge discovery and data mining (2016)
2016
Earlier work this paper cites.
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: Conf. on computer vision and pattern recognition (2016)
2016
Earlier work this paper cites.
Hendrycks, D., Gimpel, K.: Gaussian error linear units (gelus). arXiv:1606.08415 (2016)
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
Mnih, V., Badia, A.P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., Kavukcuoglu, K.: Asynchronous methods for deep reinforcement learning. In: Int. Conf. on machine learning (2016)
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., Hochreiter, S.: Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems (2017)
2017
Earlier work this paper cites.
Huang, G., Liu, Z., Van Der Maaten, L., Weinberger, K.Q.: Densely connected convolutional networks. In: Conf. on computer vision and pattern recognition (2017)
2017
Earlier work this paper cites.
Lin, T.Y., Goyal, P., Girshick, R., He, K., Dollár, P.: Focal loss for dense object detection. In: Int. Conf. on computer vision (2017)
2017
Earlier work this paper cites.
Loshchilov, I., Hutter, F.: Decoupled weight decay regularization. arXiv:1711.05101 (2017)
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Sabour, S., Frosst, N., Hinton, G.E.: Dynamic routing between capsules. Advances in neural information processing systems 30
2017
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I.: Attention is all you need. Advances in neural information processing systems (2017)
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Wang, F., Jiang, M., Qian, C., Yang, S., Li, C., Zhang, H., Wang, X., Tang, X.: Residual attention network for image classification. In: Conf. on computer vision and pattern recognition (2017)
2017
Earlier work this paper cites.
Xie, S., Girshick, R., Dollár, P., Tu, Z., He, K.: Aggregated residual transformations for deep neural networks. In: Conf. on computer vision and pattern recognition (2017)
2017
Earlier work this paper cites.
Zhu, J.Y., Park, T., Isola, P., Efros, A.A.: Unpaired image-to-image translation using cycle-consistent adversarial networks. In: Int. Conf. on computer vision (2017)
2017
Earlier work this paper cites.
Chen, R.T., Rubanova, Y., Bettencourt, J., Duvenaud, D.K.: Neural ordinary differential equations. Advances in neural information processing systems 31
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Dong, X., Shen, J.: Triplet loss in siamese network for object tracking. In: European Conf. on computer vision (ECCV) (2018)
2018
Earlier work this paper cites.
Ghiasi, G., Lin, T.Y., Le, Q.V.: Dropblock: A regularization method for convolutional networks. Advances in neural information processing systems (2018)
2018
Earlier work this paper cites.
Hinton, G.E., Sabour, S., Frosst, N.: Matrix capsules with em routing. In: International conference on learning representations (2018)
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Mescheder, L., Geiger, A., Nowozin, S.: Which training methods for GANs do actually converge? In: Int. Conf. on machine learning (2018)
2018
Earlier work this paper cites.
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al.: Improving language understanding by generative pre-training (2018)
2018
Earlier work this paper cites.
Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., Chen, L.C.: Mobilenetv2: Inverted residuals and linear bottlenecks. In: Conf. on computer vision and pattern recognition (2018)
2018
Earlier work this paper cites.
Shazeer, N., Stern, M.: Adafactor: Adaptive learning rates with sublinear memory cost. In: Int. Conf. on Machine Learning (2018)
2018
Earlier work this paper cites.
Alom, M.Z., Taha, T.M., Yakopcic, C., Westberg, S., Sidike, P., Nasrin, M.S., Hasan, M., Van Essen, B.C., Awwal, A.A., Asari, V.K.: A state-of-the-art survey on deep learning theory and architectures. electronics (2019)
2019
Earlier work this paper cites.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Kar, K., Kubilius, J., Schmidt, K., Issa, E.B., DiCarlo, J.J.: Evidence that recurrent circuits are critical to the ventral stream’s execution of core object recognition behavior. Nature neuroscience 22
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2021
Later among the works it cites.
Tolstikhin, I.O., Houlsby, N., Kolesnikov, A., Beyer, L., Zhai, X., Unterthiner, T., Yung, J., Steiner, A., Keysers, D., Uszkoreit, J., et al.: Mlp-mixer: An all-mlp architecture for vision. Advances in neural information processing systems 34
2021
Later among the works it cites.
Touvron, H., Cord, M., Sablayrolles, A., Synnaeve, G., Jégou, H.: Going deeper with image transformers. In: Int. Conf. on Computer Vision (2021)
2021
Later among the works it cites.
Zbontar, J., Jing, L., Misra, I., LeCun, Y., Deny, S.: Barlow twins: Self-supervised learning via redundancy reduction. In: Int. Conf. on Machine Learning (2021)
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Misra, D.: Mish: A self regularized non-monotonic activation function. arXiv:1908.08681 (2019)
2019
Cited alongside, same era.
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al.: Language models are unsupervised multitask learners. OpenAI blog (2019)
2019
Cited alongside, same era.
Reddi, S.J., Kale, S., Kumar, S.: On the convergence of adam and beyond. arXiv:1904.09237 (2019)
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Shrestha, A., Mahmood, A.: Review of deep learning algorithms and architectures. IEEE access (2019)
2019
Cited alongside, same era.
2022
Later among the works it cites.
Dubey, S.R., Singh, S.K., Chaudhuri, B.B.: Activation functions in deep learning: A comprehensive survey and benchmark. Neurocomputing 503
2022
Later among the works it cites.
Ericsson, L., Gouk, H., Loy, C.C., Hospedales, T.M.: Self-supervised representation learning: Introduction, advances, and challenges. Signal Processing Magazine (2022)
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
Guo, M.H., Xu, T.X., Liu, J.J., Liu, Z.N., Jiang, P.T., Mu, T.J., Zhang, S.H., Martin, R.R., Cheng, M.M., Hu, S.M.: Attention mechanisms in computer vision: A survey. Computational Visual Media (2022)
2022
Later among the works it cites.
Hamilton, K., Nayak, A., Božić, B., Longo, L.: Is neuro-symbolic ai meeting its promises in natural language processing? a structured review. Semantic Web (Preprint), 1–42 (2022)
2022
Later among the works it cites.
Han, K., Wang, Y., Chen, H., Chen, X., Guo, J., Liu, Z., Tang, Y., Xiao, A., Xu, C., Xu, Y., et al.: A survey on vision transformer. transactions on pattern analysis and machine intelligence (2022)
2022
Later among the works it cites.
Khan, S., Naseer, M., Hayat, M., Zamir, S.W., Khan, F.S., Shah, M.: Transformers in vision: A survey. ACM computing surveys (CSUR) (2022)
2022
Later among the works it cites.
OpenAI: Chatgpt: Optimizing language models for dialogue. https://openai.com/blog/chatgpt/
2022
Later among the works it cites.
2022
Later among the works it cites.
de Santana Correia, A., Colombini, E.L.: Attention, please! a survey of neural attention models in deep learning. Artificial Intelligence Review 55
2022
Later among the works it cites.
Yang, X., Song, Z., King, I., Xu, Z.: A survey on deep semi-supervised learning. Transactions on Knowledge and Data Engineering (2022)
2022
Later among the works it cites.
2023
Later among the works it cites.
Bishop, C.M., Bishop, H.: Deep learning: Foundations and concepts. Springer Nature (2023)
2023
Later among the works it cites.
Google: Palm 2 technical report. https://ai.google/static/documents/palm2techreport.pdf
2023
Later among the works it cites.
2023
Later among the works it cites.
Liu, P., Yuan, W., Fu, J., Jiang, Z., Hayashi, H., Neubig, G.: Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ACM Computing Surveys (2023)
2023
Later among the works it cites.
OpenAI: Gpt-4 technical report (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
Poli, M., Massaroli, S., Nguyen, E., Fu, D.Y., Dao, T., Baccus, S., Bengio, Y., Ermon, S., Ré, C.: Hyena hierarchy: Towards larger convolutional language models. In: International Conference on Machine Learning. pp. 28043–28078 (2023)
2023
Later among the works it cites.
Schneider, J., Vlachos, M.: Reflective-net: Learning from explanations. Data Mining and Knowledge Discovery pp. 1–22 (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Anthropic: The claude 3 model family: Opus, sonnet, haiku. Online (2023), https://www-cdn.anthropic.com/de8ba9b01c9ab7cbabf5c33b80b7bbc618857627/Model_Card_Claude_3.pdf
2024
Closest in time.
2024
Closest in time.
Fu, D., Arora, S., Grogan, J., Johnson, I., Eyuboglu, E.S., Thomas, A., Spector, B., Poli, M., Rudra, A., Ré, C.: Monarch mixer: A simple sub-quadratic gemm-based architecture. Advances in Neural Information Processing Systems 36
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Khoshraftar, S., An, A.: A survey on graph representation learning methods. ACM Transactions on Intelligent Systems and Technology 15
2024
Closest in time.
Mandar Kahade: Gpt-4: 8 models in one. https://www.kdnuggets.com/2023/08/gpt4-8-models-one-secret.html
2024
Closest in time.
Meta: The llama 3 herd of models. Online (2024), https://ai.meta.com/research/publications/the-llama-3-herd-of-models/
2024
Closest in time.
OpenAI: Hello gpt-4o! https://openai.com/index/hello-gpt-4o/
2024
Closest in time.
Rafailov, R., Sharma, A., Mitchell, E., Manning, C.D., Ermon, S., Finn, C.: Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems 36
2024
Closest in time.
2024
Closest in time.
Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Hambro, E., Zettlemoyer, L., Cancedda, N., Scialom, T.: Toolformer: Language models can teach themselves to use tools. Advances in Neural Information Processing Systems 36
2024
Closest in time.
Schneider, J., Meske, C., Kuss, P.: Foundation models: a new paradigm for artificial intelligence. Business & Information Systems Engineering pp. 1–11 (2024)
2024
Closest in time.
Schneider, J., Vlachos, M.: A survey of deep learning: From activations to transformers. In: Proceedings of the International Conference on Agents and Artificial Intelligence (ICAART) (2024)
2024
Closest in time.