Fetching the paper…
Reading the bibliography…
Transformer models have demonstrated impressive capabilities when trained at scale, excelling at difficult cognitive tasks requiring complex reasoning and rational decision-making.
1902
Earlier work this paper cites.
2002
Earlier work this paper cites.
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton et al. , “Mastering the game of go without human knowledge,” Nature , vol. 550, pp. 254–259, 2017. [Online]. Available: https://doi.org/10.1038/nature24270
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in NIPS , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel et al. , “A general reinforcement learning algorithm that masters chess, shogi, and go through self-play,” Science , vol. 362, no. 6419, pp. 1140–1144, 2018
2018
Earlier work this paper cites.
P. Izmailov, D. Podoprikhin, T. Garipov, D. Vetrov, and A. Wilson, “Averaging weights leads to wider optima and better generalization,” in 34th Conference on Uncertainty in Artificial Intelligence 2018, UAI 2018 , ser. 34th Conference on Uncertainty in Artificial Intelligence 2018, UAI 2018, R. Silva, A. Globerson, and A. Globerson, Eds. Association For Uncertainty in Artificial Intelligence (AUAI), 2018, pp. 876–885
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
J. Jumper, R. Evans, A. Pritzel, T. Green, M. Figurnov, O. Ronneberger, K. Tunyasuvunakool, R. Bates, A. Žídek, A. Potapenko et al. , “Highly accurate protein structure prediction with alphafold,” Nature , vol. 596, no. 7873, pp. 583–589, 2021
2021
Cited alongside, same era.
T. McGrath, A. Kapishnikov, N. Tomašev, A. Pearce, M. Wattenberg, D. Hassabis, B. Kim, U. Paquet, and V. Kramnik, “Acquisition of chess knowledge in alphazero,” Proceedings of the National Academy of Sciences , vol. 119, no. 47, p. e2206625119, 2022. [Online]. Available: https://www.pnas.org/doi/abs/10.1073/pnas.2206625119
2022
Cited alongside, same era.
J. Wei, X. Wang, D. Schuurmans, M. Bosma, b. ichter, F. Xia, E. Chi, Q. V. Le, and D. Zhou, “Chain-of-thought prompting elicits reasoning in large language models,” in Advances in Neural Information Processing Systems , S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, Eds., vol. 35. Curran Associates, Inc., 2022, pp. 24 824–24 837. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2022/file/9d5609613524ecf4f15af0f7b31abca4-Paper-Conference.pdf
2022
——, “Other methods implemented in katago,” 2024. [Online]. Available: https://github.com/lightvector/KataGo/blob/master/docs/KataGoMethods.md
2024
Closest in time.
2024
Closest in time.
OpenAI, “Introducing openai o1-preview,” 12 September 2024. [Online]. Available: https://openai.com/index/introducing-openai-o1-preview/
2024
Closest in time.
DeepMind, “Ai achieves silver-medal standard solving international mathematical olympiad problems,” 25 July 2024. [Online]. Available: https://deepmind.google/discover/blog/ai-solves-imo-problems-at-silver-medal-level/
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
OpenAI, “GPT-4 technical report,” arXiv:2303.08774 , 2023
2023
Cited alongside, same era.
J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu, “Roformer: Enhanced transformer with rotary position embedding,” Neurocomput. , vol. 568, no. C, mar 2024. [Online]. Available: https://doi.org/10.1016/j.neucom.2023.127063
2023
Cited alongside, same era.
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, and G. M. et al., “Palm: Scaling language modeling with pathways,” Journal of Machine Learning Research , vol. 24, no. 240, pp. 1–113, 2023. [Online]. Available: http://jmlr.org/papers/v24/22-1144.html
2023
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Closest in time.
J. Puigcerver, C. R. Ruiz, B. Mustafa, and N. Houlsby, “From sparse to soft mixtures of experts,” in The Twelfth International Conference on Learning Representations , 2024. [Online]. Available: https://openreview.net/forum?id=jxpsAj7ltE
2024
Closest in time.
P. Shaw, J. Uszkoreit, and A. Vaswani, “Self-attention with relative position representations,” in Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers) , M. Walker, H. Ji, and A. Stent, Eds. New Orleans, Louisiana: Association for Computational Linguistics, Jun. 2018, pp. 464–468. [Online]. Available: https://aclanthology.org/N18-2074
2074
Closest in time.