Fetching the paper…
Reading the bibliography…
This review addresses the problem of learning abstract representations of the measurement data in the context of Deep Reinforcement Learning (DRL).
Wiley-interscience New York, 1972
H. Kwakernaak and R. Sivan, Linear optimal control systems · 1972
Earlier work this paper cites.
Wiley London, 1974
P. Eykhoff, System identification · 1974
Earlier work this paper cites.
CRC Press, 1975
A. E. Bryson, Applied optimal control: optimization, estimation and control · 1975
Earlier work this paper cites.
Addison-Wesley Longman Publishing Co., Inc., 1984
P. H. Winston, Artificial intelligence · 1984
Earlier work this paper cites.
F. J. Pineda, “Generalization of back-propagation to recurrent neural networks,” Physical review letters
1987
Earlier work this paper cites.
S. Wold, K. Esbensen, and P. Geladi, “Principal component analysis,” Chemometrics and intelligent laboratory systems
1987
Earlier work this paper cites.
Y. Le Cun and F. Fogelman-Soulié, “Modèles connexionnistes de l’apprentissage,” Intellectica
1987
Earlier work this paper cites.
K. Hornik, M. Stinchcombe, and H. White, “Multilayer feedforward networks are universal approximators,” Neural networks
1989
Earlier work this paper cites.
M. L. Puterman, “Markov decision processes,” Handbooks in operations research and management science
1990
Earlier work this paper cites.
J. Schmidhuber, “Curious model-building control systems,” in Proc. international joint conference on neural networks
1991
Earlier work this paper cites.
R. S. Sutton, A. G. Barto, and R. J. Williams, “Reinforcement learning is direct adaptive optimal control,” IEEE control systems magazine
1992
Earlier work this paper cites.
CRC Press, 1993
F. L. Chernousko, State estimation for dynamic systems · 1993
Earlier work this paper cites.
Courier Corporation, 1994
R. F. Stengel, Optimal control and estimation · 1994
Earlier work this paper cites.
Y. LeCun, Y. Bengio, et al
1995
Earlier work this paper cites.
McGraw-hill New York, 1997
T. M. Mitchell and T. M. Mitchell, Machine learning · 1997
Earlier work this paper cites.
B. Schölkopf, A. Smola, and K.-R. Müller, “Kernel principal component analysis,” in International conference on artificial neural networks
1997
Earlier work this paper cites.
R. Sutton and A. Barto, “Reinforcement Learning: An Introduction,” IEEE Transactions on Neural Networks
1998
Earlier work this paper cites.
L. Ljung, “System identification,” in Signal analysis and prediction
1998
Earlier work this paper cites.
W. H. Fleming and W. M. McEneaney, “A max-plus-based algorithm for a hamilton–jacobi–bellman equation of nonlinear filtering,” SIAM Journal on Control and Optimization
2000
Earlier work this paper cites.
J. B. Rawlings, “Tutorial overview of model predictive control,” IEEE control systems magazine
2000
Earlier work this paper cites.
S. Thrun, “Probabilistic robotics,” Communications of the ACM
2002
Earlier work this paper cites.
B. Ravindran and A. G. Barto, “Model minimization in hierarchical reinforcement learning,” in International Symposium on Abstraction, Reformulation, and Approximation
2002
Earlier work this paper cites.
R. Vilalta and Y. Drissi, “A perspective view and survey of meta-learning,” Artificial intelligence review
2002
Earlier work this paper cites.
R. Givan, T. Dean, and M. Greig, “Equivalence notions and model minimization in markov decision processes,” Artificial Intelligence
2003
Earlier work this paper cites.
Courier Corporation, 2004
D. E. Kirk, Optimal control theory: an introduction · 2004
Earlier work this paper cites.
D. J. Lucia, P. S. Beran, and W. A. Silva, “Reduced-order modeling: new approaches for computational physics,” Progress in aerospace sciences
2004
Earlier work this paper cites.
B. Ravindran and A. G. Barto, “Approximate homomorphisms: A framework for non-exact minimization in markov decision processes,” 2004
2004
Earlier work this paper cites.
N. Ferns, P. Panangaden, and D. Precup, “Metrics for finite markov decision processes.,” in UAI
2004
Earlier work this paper cites.
S. Chopra, R. Hadsell, and Y. LeCun, “Learning a similarity metric discriminatively, with application to face verification,” in 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’05)
2005
Earlier work this paper cites.
Springer, 2006
W. M. McEneaney, Max-plus methods for nonlinear control and estimation · 2006
Earlier work this paper cites.
Springer, 2006
C. E. Rasmussen, C. K. Williams, et al · 2006
Earlier work this paper cites.
Springer, 2008
W. H. Schilders, H. A. Van der Vorst, and J. Rommes, Model order reduction: theory, research aspects and applications · 2008
Earlier work this paper cites.
J. Taylor, D. Precup, and P. Panagaden, “Bounding performance loss in approximate mdp homomorphisms,” Advances in Neural Information Processing Systems
2008
Earlier work this paper cites.
P. Vincent, H. Larochelle, Y. Bengio, and P.-A. Manzagol, “Extracting and composing robust features with denoising autoencoders,” in Proceedings of the 25th international conference on Machine learning
2008
Earlier work this paper cites.
L. Van der Maaten and G. Hinton, “Visualizing data using t-sne.,” Journal of machine learning research
2008
Earlier work this paper cites.
F. L. Lewis and D. Vrabie, “Reinforcement learning and adaptive dynamic programming for feedback control,” IEEE circuits and systems magazine
2009
Earlier work this paper cites.
S. Lange and M. Riedmiller, “Deep auto-encoder neural networks in reinforcement learning,” in The 2010 International Joint Conference on Neural Networks (IJCNN)
2010
Earlier work this paper cites.
Springer Science & Business Media, 2011
B. R. Noack, M. Morzynski, and G. Tadmor, Reduced-order modelling for flow control · 2011
Earlier work this paper cites.
John Wiley & Sons, 2012
F. L. Lewis, D. Vrabie, and V. L. Syrmos, Optimal control · 2012
Earlier work this paper cites.
F. L. Lewis, D. Vrabie, and K. G. Vamvoudakis, “Reinforcement learning and feedback control: Using natural decision methods to design optimal adaptive controllers,” IEEE Control Systems Magazine
2012
Earlier work this paper cites.
Athena scientific, 2012
D. Bertsekas, Dynamic programming and optimal control: Volume I · 2012
Earlier work this paper cites.
A. Graves, “Supervised sequence labelling,” in Supervised sequence labelling with recurrent neural networks
2012
Earlier work this paper cites.
J. Mattner, S. Lange, and M. Riedmiller, “Learn to swing up and balance a real pole based on raw visual input data,” in International Conference on Neural Information Processing
2012
Earlier work this paper cites.
P. Zhang, Y. Ren, and B. Zhang, “A new embedding quality assessment method for manifold learning,” Neurocomputing
2012
Earlier work this paper cites.
E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model-based control,” in 2012 IEEE/RSJ international conference on intelligent robots and systems
2012
Earlier work this paper cites.
Courier Corporation, 2013
M. Athans and P. L. Falb, Optimal control: an introduction to the theory and its applications · 2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
Y. Bengio, A. Courville, and P. Vincent, “Representation learning: A review and new perspectives,” IEEE Transactions on Pattern Analysis and Machine Intelligence
2013
Earlier work this paper cites.
John Wiley & Sons, 2013
F. L. Lewis and D. Liu, Reinforcement learning and approximate dynamic programming for feedback control · 2013
Earlier work this paper cites.
J. Kober, J. A. Bagnell, and J. Peters, “Reinforcement learning in robotics: A survey,” The International Journal of Robotics Research
2013
Earlier work this paper cites.
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling, “The arcade learning environment: An evaluation platform for general agents,” Journal of Artificial Intelligence Research
2013
Earlier work this paper cites.
T. Lassila, A. Manzoni, A. Quarteroni, and G. Rozza, “Model order reduction in fluid dynamics: challenges and perspectives,” Reduced Order Methods for modeling and computational reduction
2014
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Stochastic gradient vb and the variational auto-encoder,” in Second International Conference on Learning Representations, ICLR
2014
Earlier work this paper cites.
R. Jonschkowski and O. Brock, “State representation learning in robotics: Using prior knowledge about physical interaction.,” in Robotics: Science and Systems
2014
Earlier work this paper cites.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” nature
2015
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al
2015
Earlier work this paper cites.
W. Böhmer, J. T. Springenberg, J. Boedecker, M. Riedmiller, and K. Obermayer, “Autonomous Learning of State Representations for Control: An Emerging Field Aims to Autonomously Learn State Representations for Reinforcement Learning Agents from Their Real-World Sensor Observations,” KI - Kunstliche Intelligenz
2015
Earlier work this paper cites.
R. Goroshin, M. Mathieu, and Y. Lecun, “Learning to linearize under uncertainty,” in Advances in Neural Information Processing Systems
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
R. Jonschkowski and O. Brock, “Learning state representations with robotic priors,” Autonomous Robots
2015
Earlier work this paper cites.
M. Watter, J. T. Springenberg, J. Boedecker, and M. Riedmiller, “Embed to control: a locally linear latent dynamics model for control from raw images,” in Proceedings of the 28th International Conference on Neural Information Processing Systems-Volume 2
2015
Earlier work this paper cites.
N. Wahlström, T. B. Schön, and M. P. Desienroth, “From pixels to torques: Policy learning with deep dynamical models,” in Deep Learning Workshop at the 32nd International Conference on Machine Learning (ICML 2015), July 10-11, Lille, France
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
R. G. Krishnan, U. Shalit, and D. Sontag, “Deep kalman filters,” arXiv preprint arXiv:1511.05121
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
S. L. Brunton, J. L. Proctor, and J. N. Kutz, “Discovering governing equations from data by sparse identification of nonlinear dynamical systems,” Proceedings of the national academy of sciences
2016
Earlier work this paper cites.
MIT press, 2016
I. Goodfellow, Y. Bengio, and A. Courville, Deep learning · 2016
Earlier work this paper cites.
I. Higgins, L. Matthey, A. Pal, C. Burgess, X. Glorot, M. Botvinick, S. Mohamed, and A. Lerchner, “beta-vae: Learning basic visual concepts with a constrained variational framework,” in International conference on learning representations
2016
Earlier work this paper cites.
H. Van Hoof, N. Chen, M. Karl, P. van der Smagt, and J. Peters, “Stable reinforcement learning with autoencoders for tactile and visual data,” in 2016 IEEE/RSJ international conference on intelligent robots and systems (IROS)
2016
Cited alongside, same era.
P. Agrawal, A. Nair, P. Abbeel, J. Malik, and S. Levine, “Learning to poke by poking: Experiential learning of intuitive physics,” Advances in Neural Information Processing Systems
2016
Cited alongside, same era.
M. Kempka, M. Wydmuch, G. Runc, J. Toczek, and W. Jaśkowski, “Vizdoom: A doom-based ai research platform for visual reinforcement learning,” in 2016 IEEE conference on computational intelligence and games (CIG)
2016
Cited alongside, same era.
C. Finn, X. Y. Tan, Y. Duan, T. Darrell, S. Levine, and P. Abbeel, “Deep spatial autoencoders for visuomotor learning,” in 2016 IEEE International Conference on Robotics and Automation (ICRA)
2016
Cited alongside, same era.
D. Hafner, T. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson, “Learning latent dynamics for planning from pixels,” in International Conference on Machine Learning
2019
Later among the works it cites.
D. Hafner, T. Lillicrap, J. Ba, and M. Norouzi, “Dream to control: Learning behaviors by latent imagination,” in International Conference on Learning Representations
2019
Later among the works it cites.
2019
Later among the works it cites.
Y. Chandak, G. Theocharous, J. Kostas, S. Jordan, and P. Thomas, “Learning action representations for reinforcement learning,” in International Conference on Machine Learning
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Höfer, A. Raffin, R. Jonschkowski, O. Brock, and F. Stulp, “Unsupervised learning of state representations for multiple tasks,” in Deep Learning Workshop at the Conference on Neural Information Processing Systems (NIPS)
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2016
Cited alongside, same era.
J. Munk, J. Kober, and R. Babuška, “Learning state representation for deep actor-critic control,” in 2016 IEEE 55th Conference on Decision and Control (CDC)
2016
Cited alongside, same era.
2016
Cited alongside, same era.
R. Houthooft, X. Chen, Y. Duan, J. Schulman, F. De Turck, and P. Abbeel, “Vime: Variational information maximizing exploration,” Advances in neural information processing systems
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2019
Later among the works it cites.
D. Pathak, D. Gandhi, and A. Gupta, “Self-supervised exploration via disagreement,” in International conference on machine learning
2019
Later among the works it cites.
F. Tao and Q. Qi, “Make more digital twins,” Nature
2019
Later among the works it cites.
Z. Xiong, Y. Zhang, D. Niyato, R. Deng, P. Wang, and L.-C. Wang, “Deep reinforcement learning for mobile 5g and beyond: Fundamentals, applications, and challenges,” IEEE Vehicular Technology Magazine
2019
Later among the works it cites.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, et al
2019
Later among the works it cites.
J. Sanders, A. Proutière, and S.-Y. Yun, “Clustering in block markov chains,” The Annals of Statistics
2020
Later among the works it cites.
E. van der Pol, T. Kipf, F. A. Oliehoek, and M. Welling, “Plannable approximations to mdp homomorphisms: Equivariance under actions,” in Proceedings of the 19th International Conference on Autonomous Agents and MultiAgent Systems
2020
Later among the works it cites.
2020
Later among the works it cites.
M. Laskin, A. Srinivas, and P. Abbeel, “Curl: Contrastive unsupervised representations for reinforcement learning,” in International Conference on Machine Learning
2020
Later among the works it cites.
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of visual representations,” in International conference on machine learning
2020
Later among the works it cites.
O. Henaff, “Data-efficient image recognition with contrastive predictive coding,” in International Conference on Machine Learning
2020
Later among the works it cites.
A. X. Lee, A. Nagabandi, P. Abbeel, and S. Levine, “Stochastic latent actor-critic: Deep reinforcement learning with a latent variable model,” Advances in Neural Information Processing Systems
2020
Later among the works it cites.
P. S. Castro, “Scalable methods for computing state similarity in deterministic markov decision processes,” in Proceedings of the AAAI Conference on Artificial Intelligence
2020
Later among the works it cites.
K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick, “Momentum contrast for unsupervised visual representation learning,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
Z. D. Guo, B. A. Pires, B. Piot, J.-B. Grill, F. Altché, R. Munos, and M. G. Azar, “Bootstrap latent-predictive representations for multitask reinforcement learning,” in International Conference on Machine Learning
2020
Later among the works it cites.
A. Merckling, A. Coninx, L. Cressot, S. Doncieux, and N. Perrin, “State representation learning from demonstration,” in International Conference on Machine Learning, Optimization, and Data Science
2020
Later among the works it cites.
2020
Later among the works it cites.
G. Yang, A. Zhang, A. Morcos, J. Pineau, P. Abbeel, and R. Calandra, “Plan2vec: Unsupervised representation learning by latent plans,” in Learning for Dynamics and Control
2020
Later among the works it cites.
P. J. Pritz, L. Ma, and K. K. Leung, “Joint state-action embedding for efficient reinforcement learning,” arXiv e-prints
2020
Later among the works it cites.
R. Y. Tao, V. François-Lavet, and J. Pineau, “Novelty search in representational space for sample efficient exploration,” Advances in Neural Information Processing Systems
2020
Later among the works it cites.
2020
Later among the works it cites.
M. C. Machado, M. G. Bellemare, and M. Bowling, “Count-based exploration with the successor representation,” in Proceedings of the AAAI Conference on Artificial Intelligence
2020
Later among the works it cites.
A. Mosavi, Y. Faghan, P. Ghamisi, P. Duan, S. F. Ardabili, E. Salwana, and S. S. Band, “Comprehensive review of deep reinforcement learning methods and applications in economics,” Mathematics
2020
Later among the works it cites.
E. van der Pol, D. Worrall, H. van Hoof, F. Oliehoek, and M. Welling, “Mdp homomorphic networks: Group symmetries in reinforcement learning,” Advances in Neural Information Processing Systems
2020
Later among the works it cites.
2020
Later among the works it cites.
S. Fresca, L. Dede’, and A. Manzoni, “A comprehensive deep learning-based approach to reduced order modeling of nonlinear time-dependent parametrized pdes,” Journal of Scientific Computing
2021
Later among the works it cites.
University of Twente, 2021
N. Botteghi, Robotics deep reinforcement learning with loose prior knowledge · 2021
Later among the works it cites.
A. Stooke, K. Lee, P. Abbeel, and M. Laskin, “Decoupling representation learning from reinforcement learning,” in International Conference on Machine Learning
2021
Later among the works it cites.
2021
Later among the works it cites.
N. Botteghi, K. Alaa, M. Poel, B. Sirmacek, C. Brune, A. Mersha, and S. Stramigioli, “Low dimensional state representation learning with robotics priors in continuous action spaces,” in IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
2021
Later among the works it cites.
D. Yarats, A. Zhang, I. Kostrikov, B. Amos, J. Pineau, and R. Fergus, “Improving sample efficiency in model-free reinforcement learning from images,” in Proceedings of the AAAI Conference on Artificial Intelligence
2021
Later among the works it cites.
T. Kim, Y. Park, Y. Park, S. H. Lee, and I. H. Suh, “Acceleration of actor-critic deep reinforcement learning for visual grasping by state representation learning based on a preprocessed input image,” in 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
2021
Later among the works it cites.
N. Botteghi, R. Obbink, D. Geijs, M. Poel, B. Sirmacek, C. Brune, A. Mersha, and S. Stramigioli, “Low dimensional state representation learning with reward-shaped priors,” in 2020 25th International Conference on Pattern Recognition (ICPR)
2021
Later among the works it cites.
J. Shang and M. S. Ryoo, “Self-supervised disentangled representation learning for third-person imitation learning,” in 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
2021
Later among the works it cites.
R. Rafailov, T. Yu, A. Rajeswaran, and C. Finn, “Offline reinforcement learning from images with latent space models,” in Learning for Dynamics and Control
2021
Later among the works it cites.
2021
Later among the works it cites.
M. Kemertas and T. Aumentado-Armstrong, “Towards robust bisimulation metric learning,” Advances in Neural Information Processing Systems
2021
Later among the works it cites.
2021
Later among the works it cites.
K. Kim, M. W. Lee, Y. Kim, J. Ryu, M. Lee, and B.-T. Zhang, “Goal-aware cross-entropy for multi-target reinforcement learning,” Advances in Neural Information Processing Systems
2021
Later among the works it cites.
C. Allen, N. Parikh, O. Gottesman, and G. Konidaris, “Learning markov state abstractions for deep reinforcement learning,” Advances in Neural Information Processing Systems
2021
Later among the works it cites.
G. Papoudakis, F. Christianos, and S. Albrecht, “Agent modelling under partial observability for deep reinforcement learning,” Advances in Neural Information Processing Systems
2021
Later among the works it cites.
2021
Later among the works it cites.
B. van der Heijden, L. Ferranti, J. Kober, and R. Babuška, “Deepkoco: Efficient latent planning with a task-relevant koopman representation,” in 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
Y. Seo, L. Chen, J. Shin, H. Lee, P. Abbeel, and K. Lee, “State entropy maximization with random encoders for efficient exploration,” in International Conference on Machine Learning
2021
Later among the works it cites.
H. Liu and P. Abbeel, “Behavior from the void: Unsupervised active pre-training,” Advances in Neural Information Processing Systems
2021
Later among the works it cites.
D. Yarats, R. Fergus, A. Lazaric, and L. Pinto, “Reinforcement learning with prototypical representations,” in International Conference on Machine Learning
2021
Later among the works it cites.
P. Garnier, J. Viquerat, J. Rabault, A. Larcher, A. Kuhnle, and E. Hachem, “A review on deep reinforcement learning for fluid mechanics,” Computers & Fluids
2021
Later among the works it cites.
C. Yu, J. Liu, S. Nemati, and G. Yin, “Reinforcement learning in healthcare: A survey,” ACM Computing Surveys (CSUR)
2021
Later among the works it cites.
L. Wells and T. Bednarz, “Explainable ai and reinforcement learning—a systematic review of current approaches and trends,” Frontiers in artificial intelligence
2021
Later among the works it cites.
R. Agarwal, M. Schwarzer, P. S. Castro, A. C. Courville, and M. Bellemare, “Deep reinforcement learning at the edge of the statistical precipice,” Advances in Neural Information Processing Systems
2021
Later among the works it cites.
M. Abdar, F. Pourpanah, S. Hussain, D. Rezazadegan, L. Liu, M. Ghavamzadeh, P. Fieguth, X. Cao, A. Khosravi, U. R. Acharya, et al
2021
Later among the works it cites.
Cambridge University Press, 2022
S. L. Brunton and J. N. Kutz, Data-driven science and engineering: Machine learning, dynamical systems, and control · 2022
Closest in time.
Q. Liu, A. Chung, C. Szepesvári, and C. Jin, “When is partially observable reinforcement learning not scary?,” in Conference on Learning Theory
2022
Closest in time.
B. You, O. Arenz, Y. Chen, and J. Peters, “Integrating contrastive learning with dynamic models for reinforcement learning from images,” Neurocomputing
2022
Closest in time.
S. Parisi, A. Rajeswaran, S. Purushwalkam, and A. Gupta, “The unsurprising effectiveness of pre-trained vision models for control,” in International Conference on Machine Learning
2022
Closest in time.
V. Uc-Cetina, N. Navarro-Guerrero, A. Martin-Gonzalez, C. Weber, and S. Wermter, “Survey on reinforcement learning for language processing,” Artificial Intelligence Review
2022
Closest in time.
2022
Closest in time.