Fetching the paper…
Reading the bibliography…
Recently, there are many efforts attempting to learn useful policies for continuous control in visual reinforcement learning (RL).
Robust Domain Randomization for Reinforcement Learning
Slaoui, R. B., Clements, W. R., Foerster, J. N., and Toth, S. (2019) · 1910
Earlier work this paper cites.
Invertible Gaussian Reparameterization: Revisiting the Gumbel-Softmax
Potapczynski, A., Loaiza-Ganem, G., and Cunningham, J. P. (2019) · 1912
Earlier work this paper cites.
Statistical theory of extreme values and some practical applications : A series of lectures
Gumbel, E. J. (1954) · 1954
Earlier work this paper cites.
Reinforcement Learning Applied to Linear Quadratic Regulation
Bradtke, S. J. (1992) · 1992
Earlier work this paper cites.
Robust Reinforcement Learning
Morimoto, J., and Doya, K. (2001) · 2001
Earlier work this paper cites.
Lyapunov-Constrained Action Sets for Reinforcement Learning
Perkins, T. J., and Barto, A. G. (2001) · 2001
Earlier work this paper cites.
Generalized Gumbel-Softmax Gradient Estimator for Various Discrete Random Variables
Joo, W., Kim, D., Shin, S.-J., and Moon, I.-C. (2020) · 2003
Earlier work this paper cites.
Near-optimal Regret Bounds for Reinforcement Learning
Jaksch, T., Ortner, R., and Auer, P. (2008) · 2008
Earlier work this paper cites.
Policy Gradients with Parameter-Based Exploration for Control
Sehnke, F., Osendorfer, C., Rückstiess, T., Graves, A., Peters, J., and Schmidhuber, J. (2008) · 2008
Earlier work this paper cites.
Extracting and Composing Robust Features with Denoising Autoencoders
Vincent, P., Larochelle, H., Bengio, Y., and Manzagol, P.-A. (2008) · 2008
Earlier work this paper cites.
Optimal Linear-Consensus Algorithms: An LQR Perspective
Cao, Y., and Ren, W. (2010) · 2010
Earlier work this paper cites.
Contextual bandits with similarity information
Slivkins, A. (2011) · 2011
Earlier work this paper cites.
Thompson Sampling for Contextual Bandits with Linear Payoffs
Agrawal, S., and Goyal, N. (2012) · 2012
Earlier work this paper cites.
Auto-encoding Variational Bayes
Kingma, D. P., and Welling, M. (2013) · 2013
Earlier work this paper cites.
Taming the Monster: A Fast and Simple Algorithm for Contextual Bandits
Agarwal, A., Hsu, D. J., Kale, S., Langford, J., Li, L., and Schapire, R. E. (2014) · 2014
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shalev-Shwartz, S., and Ben-David, S. (2014) · 2014
Earlier work this paper cites.
Unsupervised Visual Representation Learning by Context Prediction
Doersch, C., Gupta, A. K., and Efros, A. A. (2015) · 2015
Earlier work this paper cites.
Learning Visual Feature Spaces for Robotic Manipulation with Deep Spatial Autoencoders
Finn, C., Tan, X. Y., Duan, Y., Darrell, T., Levine, S., and Abbeel, P. (2015) · 2015
Earlier work this paper cites.
Learning Continuous Control Policies by Stochastic Value Gradients
Heess, N. M. O., Wayne, G., Silver, D., Lillicrap, T. P., Erez, T., and Tassa, Y. (2015) · 2015
Earlier work this paper cites.
Variational Dropout and the Local Reparameterization Trick
Kingma, D. P., Salimans, T., and Welling, M. (2015) · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2015) · 2016
Earlier work this paper cites.
The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables
Maddison, C. J., Mnih, A., and Teh, Y. W. (2016) · 2016
Earlier work this paper cites.
The Generalized Reparameterization Gradient
Ruiz, F. J. R., Titsias, M. K., and Blei, D. M. (2016) · 2016
Earlier work this paper cites.
Weight Normalization: A Simple Reparameterization to Accelerate Training of Deep Neural Networks
Salimans, T., and Kingma, D. P. (2016) · 2016
Earlier work this paper cites.
Reinforcement Learning with Unsupervised Auxiliary Tasks
Jaderberg, M., Mnih, V., Czarnecki, W. M., Schaul, T., Leibo, J. Z., Silver, D., and Kavukcuoglu, K. (2017) · 2017
Earlier work this paper cites.
Categorical Reparameterization with Gumbel-Softmax
Jang, E., Gu, S., and Poole, B. (2017) · 2017
Earlier work this paper cites.
Accelerating Linear Model Predictive Control by Constraint Removal
Jost, M., Pannocchia, G., and Mönnigmann, M. (2017) · 2017
Earlier work this paper cites.
Unified Deep Supervised Domain Adaptation and Generalization
Motiian, S., Piccirilli, M., Adjeroh, D. A., and Doretto, G. (2017) · 2017
Earlier work this paper cites.
Asymmetric Actor Critic for Image-Based Robot Learning
Pinto, L., Andrychowicz, M., Welinder, P., Zaremba, W., and Abbeel, P. (2017) · 2017
Earlier work this paper cites.
Quantifying Generalization in Reinforcement Learning
Cobbe, K., Klimov, O., Hesse, C., Kim, T., and Schulman, J. (2018) · 2018
Earlier work this paper cites.
Implicit Quantile Networks for Distributional Reinforcement Learning
Dabney, W., Ostrovski, G., Silver, D., and Munos, R. (2018) · 2018
Earlier work this paper cites.
Implicit Reparameterization Gradients
Figurnov, M., Mohamed, S., and Mnih, A. (2018) · 2018
Earlier work this paper cites.
Learning Latent Dynamics for Planning from Pixels
Hafner, D., Lillicrap, T. P., Fischer, I. S., Villegas, R., Ha, D. R., Lee, H., and Davidson, J. (2018) · 2018
Earlier work this paper cites.
Domain Generalization with Adversarial Feature Learning
Li, H., Pan, S. J., Wang, S., and Kot, A. C. (2018) · 2018
Earlier work this paper cites.
Spectral Normalization for Generative Adversarial Networks
Miyato, T., Kataoka, T., Koyama, M., and Yoshida, Y. (2018) · 2018
Earlier work this paper cites.
Foundations of Machine Learning
Mohri, M., Rostamizadeh, A., and Talwalkar, A. (2018) · 2018
Earlier work this paper cites.
Visual reinforcement learning with imagined goals
Nair, A., Pong, V. H., Dalal, M., Bahl, S., Lin, S., and Levine, S. (2018) · 2018
Earlier work this paper cites.
Assessing Generalization in Deep Reinforcement Learning
Packer, C., Gao, K., Kos, J., Krähenbühl, P., Koltun, V., and Song, D. X. (2018) · 2018
Earlier work this paper cites.
Sim-to-Real Transfer of Robotic Control with Dynamics Randomization
Peng, X. B., Andrychowicz, M., Zaremba, W., and Abbeel, P. (2017) · 2018
Cited alongside, same era.
Lipschitz Regularity of Deep Neural Networks: Analysis and Efficient Estimation
Scaman, K., and Virmaux, A. (2018) · 2018
Cited alongside, same era.
Recent Advances in Autoencoder-Based Representation Learning
Tschannen, M., Bachem, O., and Lucic, M. (2018) · 2018
Cited alongside, same era.
Hierarchical Approaches for Reinforcement Learning in Parameterized Action Space
Wei, E., Wicke, D., and Luke, S. (2018) · 2018
Cited alongside, same era.
Variance Reduction Properties of the Reparameterization Trick
Xu, M., Quiroz, M., Kohn, R., and Sisson, S. A. (2018) · 2018
Cited alongside, same era.
SECANT: Self-Expert Cloning for Zero-Shot Generalization of Visual Policies
Fan, L. J., Wang, G., Huang, D.-A., Yu, Z., Fei-Fei, L., Zhu, Y., and Anandkumar, A. (2021) · 2021
Later among the works it cites.
Generalization in Reinforcement Learning by Soft Data Augmentation
Hansen, N., and Wang, X. (2020) · 2021
Later among the works it cites.
A Review of the Gumbel-max Trick and its Extensions for Discrete Stochasticity in Machine Learning
Huijben, I. A. M., Kool, W., Paulus, M. B., and van Sloun, R. J. G. (2021) · 2021
Later among the works it cites.
Learning Generalized Gumbel-max Causal Mechanisms
Lorberbom, G., Johnson, D. D., Maddison, C. J., Tarlow, D., and Hazan, T. (2021) · 2021
Later among the works it cites.
On The Effect of Auxiliary Tasks on Representation Dynamics
Lyle, C., Rowland, M., Ostrovski, G., and Dabney, W. (2021) · 2021
Later among the works it cites.
THDA: Treasure Hunt Data Augmentation for Semantic Navigation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zambaldi, V. F., Raposo, D., Santoro, A., Bapst, V., Li, Y., Babuschkin, I., Tuyls, K., Reichert, D. P., Lillicrap, T. P., Lockhart, E., Shanahan, M., Langston, V., Pascanu, R., Botvinick, M. M., Vinyals, O., and Battaglia, P. W. (2018) · 2018
Cited alongside, same era.
A Dissection of Overfitting and Generalization in Continuous Reinforcement Learning
Zhang, A., Ballas, N., and Pineau, J. (2018) · 2018
Cited alongside, same era.
A Study on Overfitting in Deep Reinforcement Learning
Zhang, C., Vinyals, O., Munos, R., and Bengio, S. (2018) · 2018
Cited alongside, same era.
Closing the Sim-to-Real Loop: Adapting Simulation Randomization with Real World Experience
Chebotar, Y., Handa, A., Makoviychuk, V., Macklin, M., Issac, J., Ratliff, N. D., and Fox, D. (2018) · 2019
Cited alongside, same era.
Analysing Deep Reinforcement Learning Agents Trained with Domain Randomisation
Dai, T., Arulkumaran, K., Tukra, S., Behbahani, F. M. P., and Bharath, A. A. (2019) · 2019
Cited alongside, same era.
Efficient and Accurate Estimation of Lipschitz Constants for Deep Neural Networks
Fazlyab, M., Robey, A., Hassani, H., Morari, M., and Pappas, G. J. (2019) · 2019
Cited alongside, same era.
Neural logic reinforcement learning
Jiang, Z., and Luo, S. (2019) · 2019
Cited alongside, same era.
Maksymets, O., Cartillier, V., Gokaslan, A., Wijmans, E., Galuba, W., Lee, S., and Batra, D. (2021) · 2021
Later among the works it cites.
Automatic Data Augmentation for Generalization in Reinforcement Learning
Raileanu, R., Goldstein, M., Yarats, D., Kostrikov, I., and Fergus, R. (2021) · 2021
Later among the works it cites.
The Distracting Control Suite - A Challenging Benchmark for Reinforcement Learning from Pixels
Stone, A., Ramirez, O., Konolige, K., and Jonschkowski, R. (2021) · 2021
Later among the works it cites.
Unsupervised Visual Attention and Invariance for Reinforcement Learning
Wang, X., Lian, L., and Yu, S. X. (2021) · 2021
Later among the works it cites.
Reinforcement Learning with Prototypical Representations
Yarats, D., Fergus, R., Lazaric, A., and Pinto, L. (2021) · 2021
Later among the works it cites.
Mastering Atari Games with Limited Data
Ye, W., Liu, S.-W., Kurutach, T., Abbeel, P., and Gao, Y. (2021) · 2021
Later among the works it cites.
Learning Invariant Representations for Reinforcement Learning without Reconstruction
Zhang, A., McAllister, R. T., Calandra, R., Gal, Y., and Levine, S. (2021) · 2021
Later among the works it cites.
Robust Reinforcement Learning on State Observations with Learned Optimal Adversary
Zhang, H., Chen, H., Boning, D. S., and Hsieh, C.-J. (2021) · 2021
Later among the works it cites.
Domain Generalization: A Survey
Zhou, K., Liu, Z., Qiao, Y., Xiang, T., and Loy, C. C. (2021) · 2021
Later among the works it cites.
Look where you look! Saliency-guided Q-networks for generalization in visual Reinforcement Learning
Bertoin, D., Zouitine, A., Zouitine, M., and Rachelson, E. (2022) · 2022
Later among the works it cites.
Galois: Boosting deep reinforcement learning via generalizable logic synthesis
Cao, Y., Li, Z., Yang, T., Zhang, H., Zheng, Y., Li, Y., Hao, J., and Liu, Y. (2022) · 2022
Later among the works it cites.
When Does Contrastive Visual Representation Learning Work?
Cole, E., Yang, X. S., Wilber, K., Aodha, O. M., and Belongie, S. J. (2021) · 2022
Later among the works it cites.
Temporal Difference Learning for Model Predictive Control
Hansen, N., Wang, X., and Su, H. (2022) · 2022
Later among the works it cites.
Spectrum Random Masking for Generalization in Image-based Reinforcement Learning
Huang, Y., Peng, P., Zhao, Y., Chen, G., and Tian, Y. (2022) · 2022
Later among the works it cites.
On the Generalization of Representations in Reinforcement Learning
Le Lan, C., Tu, S., Oberman, A., Agarwal, R., and Bellemare, M. G. (2022) · 2022
Later among the works it cites.
Instance-optimal PAC Algorithms for Contextual Bandits
Li, Z., Ratliff, L. J., Nassif, H., Jamieson, K. G., and Jain, L. P. (2022) · 2022
Later among the works it cites.
Learning Smooth Neural Functions via Lipschitz Regularization
Liu, H.-T. D., Williams, F., Jacobson, A., Fidler, S., and Litany, O. (2022) · 2022
Later among the works it cites.
Data Augmentation for Manipulation
Mitrano, P., and Berenson, D. (2022) · 2022
Later among the works it cites.
TRAIL: Near-Optimal Imitation Learning with Suboptimal Data
Yang, M., Levine, S., and Nachum, O. (2022) · 2022
Later among the works it cites.
Mastering visual continuous control: Improved data-augmented reinforcement learning
Yarats, D., Fergus, R., Lazaric, A., and Pinto, L. (2022) · 2022
Later among the works it cites.
Pre-Trained Image Encoder for Generalizable Visual Reinforcement Learning
Yuan, Z., Xue, Z., Yuan, B., Wang, X., Wu, Y., Gao, Y., and Xu, H. (2022) · 2022
Later among the works it cites.
Visual Reinforcement Learning With Self-Supervised 3D Representations
Ze, Y., Hansen, N., Chen, Y., Jain, M., and Wang, X. (2022) · 2022
Later among the works it cites.
Zheng, W., Sharan, S. P., Fan, Z., Wang, K., Xi, Y., and Wang, Z. (2022) · 2022
Later among the works it cites.
Interpretable and explainable logical policies via neurally guided symbolic abstraction
Delfosse, Q., Shindo, H., Dhami, D. S., and Kersting, K. (2023) · 2023
Later among the works it cites.
Automatic noise filtering with dynamic sparse training in deep reinforcement learning
Grooten, B., Sokar, G., Dohare, S., Mocanu, E., Taylor, M. E., Pechenizkiy, M., and Mocanu, D. C. (2023) · 2023
Later among the works it cites.
Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., Dollár, P., and Girshick, R. B. (2023) · 2023
Later among the works it cites.
Normalization Enhances Generalization in Visual Reinforcement Learning
Li, L., Lyu, J., Ma, G., Wang, Z., Yang, Z., Li, X., and Li, Z. (2023) · 2023
Later among the works it cites.
Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence?
Majumdar, A., Yadav, K., Arnaud, S., Ma, Y. J., Chen, C., Silwal, S., Jain, A., Berges, V.-P., Abbeel, P., Malik, J., Batra, D., Lin, Y., Maksymets, O., Rajeswaran, A., and Meier, F. (2023) · 2023
Later among the works it cites.
Read and reap the rewards: Learning to play atari with the help of instruction manuals
Wu, Y., Fan, Y., Liang, P. P., Azaria, A., Li, Y.-F., and Mitchell, T. M. (2023) · 2023
Later among the works it cites.
Interpretable concept bottlenecks to align reinforcement learning agents
Delfosse, Q., Sztwiertnia, S., Stammer, W., Rothermel, M., and Kersting, K. (2024) · 2024
Closest in time.
Insight: End-to-end neuro-symbolic visual reinforcement learning with language explanations
Luo, L., Zhang, G., Xu, H., Yang, Y., Fang, C., and Li, Q. (2024) · 2024
Closest in time.
Off-policy rl algorithms can be sample-efficient for continuous control via sample multiple reuse
Lyu, J., Wan, L., Li, X., and Lu, Z. (2024) · 2024
Closest in time.