Fetching the paper…
Reading the bibliography…
Visual deep reinforcement learning (RL) enables robots to acquire skills from visual input for unstructured tasks.
Adaptive mixtures of local experts
Jacobs, R. A., Jordan, M. I., Nowlan, S. J., and Hinton, G. E · 1991
Earlier work this paper cites.
Hierarchical mixtures of experts and the em algorithm
Jordan, M. I. and Jacobs, R. A · 1994
Earlier work this paper cites.
Mixture of experts: a literature survey
Masoudnia, S. and Ebrahimpour, R · 2014
Earlier work this paper cites.
Adding gradient noise improves learning for very deep networks
Neelakantan, A., Vilnis, L., Le, Q. V., Sutskever, I., Kaiser, L., Kurach, K., and Martens, J · 2015
Earlier work this paper cites.
How to escape saddle points efficiently
Jin, C., Ge, R., Netrapalli, P., Kakade, S. M., and Jordan, M. I · 2017
Earlier work this paper cites.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
Rajeswaran, A., Kumar, V., Gupta, A., Vezzani, G., Schulman, J., Todorov, E., and Levine, S · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J · 2017
Earlier work this paper cites.
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., Casas, D. d. L., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., et al · 2018
Earlier work this paper cites.
A study on overfitting in deep reinforcement learning
Zhang, C., Vinyals, O., Munos, R., and Bengio, S · 2018
Earlier work this paper cites.
Challenges of real-world reinforcement learning
Dulac-Arnold, G., Mankowitz, D., and Hester, T · 2019
Earlier work this paper cites.
Dream to control: Learning behaviors by latent imagination
Hafner, D., Lillicrap, T., Ba, J., and Norouzi, M · 2019
Earlier work this paper cites.
Observational overfitting in reinforcement learning
Song, X., Jiang, Y., Tu, S., Du, Y., and Neyshabur, B · 2019
Earlier work this paper cites.
On warm-starting neural network training
Ash, J. and Adams, R. P · 2020
Earlier work this paper cites.
Mastering atari with discrete world models
Hafner, D., Lillicrap, T., Norouzi, M., and Ba, J · 2020
Earlier work this paper cites.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Kostrikov, I., Yarats, D., and Fergus, R · 2020
Earlier work this paper cites.
Gshard: Scaling giant models with conditional computation and automatic sharding
Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z · 2020
Cited alongside, same era.
Data-efficient reinforcement learning with self-predictive representations
Schwarzer, M., Anand, A., Goel, R., Hjelm, R. D., Courville, A., and Bachman, P · 2020
Cited alongside, same era.
Sim-to-real transfer in deep reinforcement learning for robotics: a survey
Zhao, W., Queralta, J. P., and Westerlund, T · 2020
Cited alongside, same era.
The ingredients of real-world robotic reinforcement learning
Zhu, H., Yu, J., Gupta, A., Shah, D., Hartikainen, K., Singh, A., Kumar, V., and Levine, S · 2020
Cited alongside, same era.
Conflict-averse gradient descent for multi-task learning
Liu, B., Liu, X., Jin, X., Stone, P., and Liu, Q · 2021
On the convergence of stochastic multi-objective gradient manipulation and beyond
Zhou, S., Zhang, W., Jiang, J., Zhong, W., Gu, J., and Zhu, W · 2022
Later among the works it cites.
Alternating gradient descent and mixture-of-experts for integrated multimodal perception
Akbari, H., Kondratyuk, D., Cui, Y., Hornung, R., Wang, H., and Adam, H · 2023
Later among the works it cites.
Mod-squad: Designing mixtures of experts as modular multi-task learners
Chen, Z., Shen, Y., Ding, M., Chen, Z., Zhao, H., Learned-Miller, E. G., and Gan, C · 2023
Later among the works it cites.
Environment agnostic representation for visual reinforcement learning
Choi, H., Lee, H., Jeong, S., and Min, D · 2023
Later among the works it cites.
Mastering diverse domains through world models
Hafner, D., Pasukonis, J., Ba, J., and Lillicrap, T · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Crossing the reality gap: A survey on sim-to-real transferability of robot controllers in reinforcement learning
Salvato, E., Fenu, G., Medvet, E., and Pellegrino, F. A · 2021
Cited alongside, same era.
Avoid overfitting in deep reinforcement learning: Increasing robustness through decentralized control
Schilling, M · 2021
Cited alongside, same era.
Decoupling representation learning from reinforcement learning
Stooke, A., Lee, K., Abbeel, P., and Laskin, M · 2021
Cited alongside, same era.
Mastering visual continuous control: Improved data-augmented reinforcement learning
Yarats, D., Fergus, R., Lazaric, A., and Pinto, L · 2021
Cited alongside, same era.
Stabilizing off-policy deep reinforcement learning from pixels
Cetin, E., Ball, P. J., Roberts, S., and Celiktutan, O · 2022
Cited alongside, same era.
Sample-efficient reinforcement learning by breaking the replay ratio barrier
D’Oro, P., Schwarzer, M., Nikishin, E., Bacon, P.-L., Bellemare, M. G., and Courville, A · 2022
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W., Zoph, B., and Shazeer, N · 2022
Cited alongside, same era.
Improving generalization in visual reinforcement learning via conflict-aware gradient agreement augmentation
Liu, S., Chen, Z., Liu, Y., Wang, Y., Yang, D., Zhao, Z., Zhou, Z., Yi, X., Li, W., Zhang, W., et al · 2023
Later among the works it cites.
Moduleformer: Learning modular large language models from uncurated data
Shen, Y., Zhang, Z., Cao, T., Tan, S., Chen, Z., and Gan, C · 2023
Later among the works it cites.
The dormant neuron phenomenon in deep reinforcement learning
Sokar, G., Agarwal, R., Castro, P. S., and Evci, U · 2023
Later among the works it cites.
Drm: Mastering visual reinforcement learning through dormant ratio minimization
Xu, G., Zheng, R., Liang, Y., Wang, X., Yuan, Z., Ji, T., Luo, Y., Liu, X., Yuan, J., Hua, P., et al · 2023
Later among the works it cites.
Taco: Temporal latent action-driven contrastive loss for visual reinforcement learning
Zheng, R., Wang, X., Sun, Y., Ma, S., Zhao, J., Xu, H., Daumé III, H., and Huang, F · 2023
Later among the works it cites.
Ace: Off-policy actor-critic with causality-aware entropy regularization
Ji, T., Liang, Y., Zeng, Y., Luo, Y., Xu, G., Guo, J., Zheng, R., Huang, F., Sun, F., and Xu, H · 2024
Closest in time.
Spawnnet: Learning generalizable visuomotor skills from pre-trained network
Lin, X., So, J., Mahalingam, S., Liu, F., and Abbeel, P · 2024
Closest in time.
Serl: A software suite for sample-efficient robotic reinforcement learning
Luo, J., Hu, Z., Xu, C., Tan, Y. L., Berg, J., Sharma, A., Schaal, S., Finn, C., Gupta, A., and Levine, S · 2024
Closest in time.
Solving token gradient conflict in mixture-of-experts for large vision-language model
Yang, L., Sheng, D., Cai, C., Yang, F., Li, S., Zhang, D., and Li, X · 2024
Closest in time.