Fetching the paper…
Reading the bibliography…
Scaling up the model size and computation has brought consistent performance improvements in supervised learning.
A comprehensive introduction to differential geometry
Spivak, M. D · 1970
Earlier work this paper cites.
An introduction to probability theory and its applications, Volume 2 , volume 81
Feller, W · 1991
Earlier work this paper cites.
A simple weight decay can improve generalization
Krogh, A. and Hertz, J · 1991
Earlier work this paper cites.
Riemannian geometry , volume 2
Do Carmo, M. P. and Flaherty Francis, J · 1992
Earlier work this paper cites.
Hypersphere
Weisstein, E. W · 2002
Earlier work this paper cites.
Solid angle
Weisstein, E. W · 2005
Earlier work this paper cites.
Riemannian manifolds: an introduction to curvature , volume 176
Lee, J. M · 2006
Earlier work this paper cites.
Optimization algorithms on matrix manifolds
Absil, P.-A., Mahony, R., and Sepulchre, R · 2008
Earlier work this paper cites.
Concise formulas for the area and volume of a hyperspherical cap
Li, S · 2010
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
Stochastic gradient descent on riemannian manifolds
Bonnabel, S · 2013
Earlier work this paper cites.
Distributions of angles in random packing on spheres
Cai, T. T., Fan, J., and Jiang, T · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T · 2015
Earlier work this paper cites.
Brockman, G · 2016
Earlier work this paper cites.
Layer normalization
Lei Ba, J., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Salimans, T. and Kingma, D. P · 2016
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A · 2017
Earlier work this paper cites.
Normface: L2 hypersphere embedding for face verification
Wang, F., Xiang, X., Cheng, J., and Yuille, A. L · 2017
Earlier work this paper cites.
Generalization and regularization in dqn
Farebrother, J., Machado, M. C., and Bowling, M · 2018
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
Fujimoto, S., Hoof, H., and Meger, D · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Earlier work this paper cites.
Improving regression performance with distributional losses
Imani, E. and White, M · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Earlier work this paper cites.
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., Casas, D. d. L., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., et al · 2018
Earlier work this paper cites.
Spherical latent spaces for stable variational autoencoders
Xu, J. and Durrett, G · 2018
Earlier work this paper cites.
Learning agile and dynamic motor skills for legged robots
Hwangbo, J., Lee, J., Dosovitskiy, A., Bellicoso, D., Tsounis, V., Koltun, V., and Hutter, M · 2019
Earlier work this paper cites.
Observational overfitting in reinforcement learning
Song, X., Jiang, Y., Tu, S., Du, Y., and Neyshabur, B · 2019
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations
Chen, T., Kornblith, S., Norouzi, M., and Hinton, G · 2020
Earlier work this paper cites.
D4rl: Datasets for deep data-driven reinforcement learning
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Cited alongside, same era.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Kostrikov, I., Yarats, D., and Fergus, R · 2020
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
Controlling overestimation bias with truncated mixture of continuous distributional quantile critics
Kuznetsov, A., Shvechikov, P., Grishin, A., and Vetrov, D · 2020
Cited alongside, same era.
Can increasing input dimensionality improve deep reinforcement learning?
Ota, K., Oiki, T., Jha, D., Mariyama, T., and Nikovski, D · 2020
Cited alongside, same era.
Mastering diverse domains through world models
Hafner, D., Pasukonis, J., Ba, J., and Lillicrap, T · 2023
Later among the works it cites.
Td-mpc2: Scalable, robust world models for continuous control
Hansen, N., Su, H., and Wang, X · 2023
Later among the works it cites.
Idql: Implicit q-learning as an actor-critic method with diffusion policies
Hansen-Estruch, P., Kostrikov, I., Janner, M., Kuba, J. G., and Levine, S · 2023
Later among the works it cites.
Sample-efficient and safe deep reinforcement learning via reset deep ensemble agents
Kim, W., Shin, Y., Park, J., and Sung, Y · 2023
Later among the works it cites.
Efficient deep reinforcement learning requires regulating overfitting
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Understanding contrastive representation learning through alignment and uniformity on the hypersphere
Wang, T. and Isola, P · 2020
Cited alongside, same era.
On layer normalization in the transformer architecture
Xiong, R., Yang, Y., He, D., Zheng, K., Zheng, S., Xing, C., Zhang, H., Lan, Y., Wang, L., and Liu, T · 2020
Cited alongside, same era.
Towards deeper deep reinforcement learning with spectral normalization
Bjorck, N., Gomes, C. P., and Weinberger, K. Q · 2021
Cited alongside, same era.
A minimalist approach to offline reinforcement learning
Fujimoto, S. and Gu, S. S · 2021
Cited alongside, same era.
Spectral normalisation for deep reinforcement learning: an optimisation perspective
Gogianu, F., Berariu, T., Rosca, M. C., Clopath, C., Busoniu, L., and Pascanu, R · 2021
Cited alongside, same era.
Dropout q-functions for doubly efficient reinforcement learning
Hiraoka, T., Imagawa, T., Hashimoto, T., Onishi, T., and Tsuruoka, Y · 2021
Cited alongside, same era.
Offline reinforcement learning with implicit q-learning
Kostrikov, I., Nair, A., and Levine, S · 2021
Cited alongside, same era.
Li, Q., Kumar, A., Kostrikov, I., and Levine, S · 2023
Later among the works it cites.
Understanding plasticity in neural networks
Lyle, C., Zheng, Z., Nikishin, E., Pires, B. A., Pascanu, R., and Dabney, W · 2023
Later among the works it cites.
Revisiting plasticity in visual reinforcement learning: Data, modules and training stages
Ma, G., Li, L., Zhang, S., Liu, Z., Wang, Z., Chen, Y., Shen, L., Wang, X., and Tao, D · 2023
Later among the works it cites.
Bigger, better, faster: Human-level atari with human-level efficiency
Schwarzer, M., Ceron, J. S. O., Courville, A., Bellemare, M. G., Agarwal, R., and Castro, P. S · 2023
Later among the works it cites.
The dormant neuron phenomenon in deep reinforcement learning
Sokar, G., Agarwal, R., Castro, P. S., and Evci, U · 2023
Later among the works it cites.
Drm: Mastering visual reinforcement learning through dormant ratio minimization
Xu, G., Zheng, R., Liang, Y., Wang, X., Yuan, Z., Ji, T., Luo, Y., Liu, X., Yuan, J., Hua, P., et al · 2023
Later among the works it cites.
Crossq: Batch normalization in deep reinforcement learning for greater sample efficiency and simplicity
Bhatt, A., Palenicek, D., Belousov, B., Argus, M., Amiranashvili, A., Brox, T., and Peters, J · 2024
Later among the works it cites.
Streaming deep reinforcement learning finally works
Elsayed, M., Vasan, G., and Mahmood, A. R · 2024
Later among the works it cites.
Stop regressing: Training value functions via classification for scalable deep rl
Farebrother, J., Orbay, J., Vuong, Q., Taïga, A. A., Chebotar, Y., Xiao, T., Irpan, A., Levine, S., Castro, P. S., Faust, A., et al · 2024
Later among the works it cites.
Simplifying deep temporal difference learning
Gallici, M., Fellows, M., Ellis, B., Pou, B., Masmitja, I., Foerster, J. N., and Martin, M · 2024
Later among the works it cites.
Dissecting deep rl with high update ratios: Combatting value divergence
Hussing, M., Voelcker, C. A., Gilitschenski, I., Farahmand, A.-m., and Eaton, E · 2024
Later among the works it cites.
Analyzing and improving the training dynamics of diffusion models
Karras, T., Aittala, M., Lehtinen, J., Hellsten, J., Aila, T., and Laine, S · 2024
Later among the works it cites.
Plasticity loss in deep reinforcement learning: A survey
Klein, T., Miklautz, L., Sidak, K., Plant, C., and Tschiatschek, S · 2024
Later among the works it cites.
ngpt: Normalized transformer with representation learning on the hypersphere
Loshchilov, I., Hsieh, C.-P., Sun, S., and Ginsburg, B · 2024
Later among the works it cites.
Normalization and effective learning rates in reinforcement learning
Lyle, C., Zheng, Z., Khetarpal, K., Martens, J., van Hasselt, H., Pascanu, R., and Dabney, W · 2024
Later among the works it cites.
Naik, A., Wan, Y., Tomar, M., and Sutton, R. S · 2024
Later among the works it cites.
Mixtures of experts unlock parameter scaling for deep rl
Obando-Ceron, J., Sokar, G., Willi, T., Lyle, C., Farebrother, J., Foerster, J., Dziugaite, G. K., Precup, D., and Castro, P. S · 2024
Later among the works it cites.
iqrl–implicitly quantized representations for sample-efficient reinforcement learning
Scannell, A., Kujanpää, K., Zhao, Y., Nakhaei, M., Solin, A., and Pajarinen, J · 2024
Later among the works it cites.
Humanoidbench: Simulated humanoid benchmark for whole-body locomotion and manipulation
Sferrazza, C., Huang, D.-M., Lin, X., Lee, Y., and Abbeel, P · 2024
Later among the works it cites.
Riemannian gradient descent for spherical area-preserving mappings
Sutti, M. and Yueh, M.-H · 2024
Later among the works it cites.
Gymnasium: A standard interface for reinforcement learning environments
Towers, M., Kwiatkowski, A., Terry, J., Balis, J. U., De Cola, G., Deleu, T., Goulão, M., Kallinteris, A., Krimmel, M., KG, A., et al · 2024
Later among the works it cites.
Mad-td: Model-augmented data stabilizes high update ratio rl
Voelcker, C. A., Hussing, M., Eaton, E., Farahmand, A.-m., and Gilitschenski, I · 2024
Later among the works it cites.
Mixture of experts in a mixture of rl settings
Willi, T., Obando-Ceron, J., Foerster, J., Dziugaite, K., and Castro, P. S · 2024
Later among the works it cites.
Efficient online reinforcement learning fine-tuning need not retain offline data
Zhou, Z., Peng, A., Li, Q., Levine, S., and Kumar, A · 2024
Later among the works it cites.
Towards general-purpose model-free reinforcement learning
Fujimoto, S., D’Oro, P., Zhang, A., Tian, Y., and Rabbat, M · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al · 2025
Closest in time.
Scaling off-policy reinforcement learning with batch and weight normalization
Palenicek, D., Vogt, F., and Peters, J · 2025
Closest in time.