Fetching the paper…
Reading the bibliography…
In this paper, we investigate the fundamental question: To what extent are gradient-based neural architecture search (NAS) techniques applicable to RL? Using the original DARTS as a convenient baseline, we discover that the discrete architectures found can achieve up to 250% performance compared to manual architecture designs on both discrete and continuous action space environments across off-policy and on-policy RL algorithms, at only 3x more computation time.
sharpdarts: Faster and more accurate differentiable architecture search
Hundt, A., Jain, V., and Hager, G. D. (2019) · 1903
Earlier work this paper cites.
DARTS+: improved differentiable architecture search with early stopping
Liang, H., Zhang, S., Sun, J., He, X., Huang, W., Zhuang, K., and Li, Z. (2019) · 1909
Earlier work this paper cites.
Efficient reinforcement learning through evolving neural network topologies
Stanley, K. O. and Miikkulainen, R. (2002) · 2002
Earlier work this paper cites.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Kostrikov, I., Yarats, D., and Fergus, R. (2020) · 2004
Earlier work this paper cites.
Noisy differentiable architecture search
Chu, X., Zhang, B., and Li, X. (2020b) · 2005
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J. (2020) · 2005
Earlier work this paper cites.
Acme: A research framework for distributed reinforcement learning
Hoffman, M., Shahriari, B., Aslanides, J., Barth-Maron, G., Behbahani, F., Norman, T., Abdolmaleki, A., Cassirer, A., Yang, F., Baumli, K., Henderson, S., Novikov, A., Colmenarejo, S. G., Cabi, S., Gulcehre, C., Paine, T. L., Cowie, A., Wang, Z., Piot, B., and de Freitas, N. (2020) · 2006
Earlier work this paper cites.
Neural architecture search without training
Mellor, J., Turner, J., Storkey, A. J., and Crowley, E. J. (2020) · 2006
Earlier work this paper cites.
Automatic data augmentation for generalization in deep reinforcement learning
Raileanu, R., Goldstein, M., Yarats, D., Kostrikov, I., and Fergus, R. (2020) · 2006
Earlier work this paper cites.
DARTS-: robustly stepping out of performance collapse without indicators
Chu, X., Wang, X., Zhang, B., Lu, S., Wei, X., and Yan, J. (2020a) · 2009
Earlier work this paper cites.
A hypercube-based encoding for evolving large-scale neural networks
Stanley, K. O., D’Ambrosio, D. B., and Gauci, J. (2009) · 2009
Earlier work this paper cites.
Optimal algorithms for online convex optimization with multi-point bandit feedback
Agarwal, A., Dekel, O., and Xiao, L. (2010) · 2010
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. A. (2013) · 2013
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M. S., Berg, A. C., and Li, F. (2015) · 2015
Earlier work this paper cites.
Rl$ˆ2$: Fast reinforcement learning via slow reinforcement learning
Duan, Y., Schulman, J., Chen, X., Bartlett, P. L., Sutskever, I., and Abbeel, P. (2016) · 2016
Earlier work this paper cites.
Prioritized experience replay
Schaul, T., Quan, J., Antonoglou, I., and Silver, D. (2016) · 2016
Earlier work this paper cites.
Neural architecture search with reinforcement learning
Zoph, B. and Le, Q. V. (2016) · 2016
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R. (2017) · 2017
Earlier work this paper cites.
Neural episodic control
Pritzel, A., Uria, B., Srinivasan, S., Badia, A. P., Vinyals, O., Hassabis, D., Wierstra, D., and Blundell, C. (2017) · 2017
Earlier work this paper cites.
Evolution strategies as a scalable alternative to reinforcement learning
Salimans, T., Ho, J., Chen, X., and Sutskever, I. (2017) · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017) · 2017
Cited alongside, same era.
Learning to reinforcement learn
Wang, J., Kurth-Nelson, Z., Soyer, H., Leibo, J. Z., Tirumala, D., Munos, R., Blundell, C., Kumaran, D., and Botvinick, M. M. (2017) · 2017
Cited alongside, same era.
IMPALA: scalable distributed deep-rl with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., Legg, S., and Kavukcuoglu, K. (2018) · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S. (2018) · 2018
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M., Modayil, J., van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M. G., and Silver, D. (2018) · 2018
Cited alongside, same era.
Stabilizing differentiable architecture search via perturbation-based regularization
Chen, X. and Hsieh, C. (2020) · 2020
Later among the works it cites.
Fair DARTS: eliminating unfair advantages in differentiable architecture search
Chu, X., Zhou, T., Zhang, B., and Li, J. (2020c) · 2020
Later among the works it cites.
Nas-bench-201: Extending the scope of reproducible neural architecture search
Dong, X. and Yang, Y. (2020) · 2020
Later among the works it cites.
SGAS: sequential greedy architecture search
Li, G., Qian, G., Delgadillo, I. C., Müller, M., Thabet, A. K., and Ghanem, B. (2020) · 2020
Later among the works it cites.
Provably efficient online hyperparameter optimization with population-based bandits
Parker-Holder, J., Nguyen, V., and Roberts, S. J. (2020) · 2020
Later among the works it cites.
Rl-cyclegan: Reinforcement learning aware simulation-to-real
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scalable deep reinforcement learning for vision-based robotic manipulation
Kalashnikov, D., Irpan, A., Pastor, P., Ibarz, J., Herzog, A., Jang, E., Quillen, D., Holly, E., Kalakrishnan, M., Vanhoucke, V., and Levine, S. (2018) · 2018
Cited alongside, same era.
Progressive neural architecture search
Liu, C., Zoph, B., Neumann, M., Shlens, J., Hua, W., Li, L., Fei-Fei, L., Yuille, A. L., Huang, J., and Murphy, K. (2018) · 2018
Cited alongside, same era.
Neural architecture optimization
Luo, R., Tian, F., Qin, T., Chen, E., and Liu, T. (2018) · 2018
Cited alongside, same era.
Efficient neural architecture search via parameter sharing
Pham, H., Guan, M. Y., Zoph, B., Le, Q. V., and Dean, J. (2018) · 2018
Cited alongside, same era.
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., de Las Casas, D., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., Lillicrap, T. P., and Riedmiller, M. A. (2018) · 2018
Cited alongside, same era.
Learning transferable architectures for scalable image recognition
Zoph, B., Vasudevan, V., Shlens, J., and Le, Q. V. (2018) · 2018
Cited alongside, same era.
Proxylessnas: Direct neural architecture search on target task and hardware
Cai, H., Zhu, L., and Han, S. (2019) · 2019
Cited alongside, same era.
Rao, K., Harris, C., Irpan, A., Levine, S., Ibarz, J., and Khansari, M. (2020) · 2020
Later among the works it cites.
Observational overfitting in reinforcement learning
Song, X., Jiang, Y., Tu, S., Du, Y., and Neyshabur, B. (2020) · 2020
Later among the works it cites.
Understanding and robustifying differentiable architecture search
Zela, A., Elsken, T., Saikia, T., Marrakchi, Y., Brox, T., and Hutter, F. (2020) · 2020
Later among the works it cites.
Visionary: Vision architecture discovery for robot learning
Akinola, I., Angelova, A., Lu, Y., Chebotar, Y., Kalashnikov, D., Varley, J., Ibarz, J., and Ryoo, M. S. (2021) · 2021
Closest in time.
Geometry-aware gradient algorithms for neural architecture search
Li, L., Khodak, M., Balcan, N., and Talwalkar, A. (2021) · 2021
Closest in time.
Attention-based partial decoupling of policy and value for generalization in reinforcement learning
Nafi, N. M., Glasscock, C., and Hsu, W. (2021) · 2021
Closest in time.
Decoupling value and policy for generalization in reinforcement learning
Raileanu, R. and Fergus, R. (2021) · 2021
Closest in time.
Song, X., Choromanski, K., Parker-Holder, J., Tang, Y., Peng, D., Jain, D., Gao, W., Pacchiano, A., Sarlós, T., and Yang, Y. (2021) · 2021
Closest in time.
Tang, Y. and Ha, D. (2021) · 2021
Closest in time.
Rethinking architecture selection in differentiable nas
Wang, R., Cheng, M., Chen, X., Tang, X., and Hsieh, C.-J. (2021b) · 2021
Closest in time.
How powerful are performance predictors in neural architecture search?
White, C., Zela, A., Ru, B., Liu, Y., and Hutter, F. (2021) · 2021
Closest in time.
On the importance of hyperparameter optimization for model-based reinforcement learning
Zhang, B., Rajan, R., Pineda, L., Lambert, N. O., Biedenkapp, A., Chua, K., Hutter, F., and Calandra, R. (2021) · 2021
Closest in time.
TF-Agents: A library for reinforcement learning in tensorflow
Guadarrama, S., Korattikara, A., Ramirez, O., Castro, P., Holly, E., Fishman, S., Wang, K., Gonina, E., Wu, N., Kokiopoulou, E., Sbaiz, L., Smith, J., Bartók, G., Berent, J., Harris, C., Vanhoucke, V., and Brevdo, E. (2018) · 2022
Closest in time.
Automated reinforcement learning (autorl): A survey and open problems
Parker-Holder, J., Rajan, R., Song, X., Biedenkapp, A., Miao, Y., Eimer, T., Zhang, B., Nguyen, V., Calandra, R., Faust, A., Hutter, F., and Lindauer, M. (2022) · 2022
Closest in time.
Leveraging procedural generation to benchmark reinforcement learning
Cobbe, K., Hesse, C., Hilton, J., and Schulman, J. (2020) · 2056
Closest in time.