Fetching the paper…
Reading the bibliography…
The ability to separate signal from noise, and reason with clean abstractions, is critical to intelligence.
Dream to control: Learning behaviors by latent imagination
Hafner, D., Lillicrap, T., Ba, J., and Norouzi, M · 1912
Earlier work this paper cites.
The nature of explanation , volume 445
Craik, K. J. W · 1952
Earlier work this paper cites.
Why the law of effect will not go away
Dennett, D. C · 1975
Earlier work this paper cites.
An adaptive network that constructs and uses and internal model of its world
Sutton, R. S · 1981
Earlier work this paper cites.
Finding structure in time
Elman, J. L · 1990
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
Sutton, R. S · 1991
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Puterman, M. L · 1994
Earlier work this paper cites.
Object recognition from local scale-invariant features
Lowe, D. G · 1999
Earlier work this paper cites.
Equivalence notions and model minimization in markov decision processes
Givan, R., Dean, T., and Greig, M · 2003
Earlier work this paper cites.
Metrics for finite markov decision processes
Ferns, N., Panangaden, P., and Precup, D · 2004
Earlier work this paper cites.
Proto-value functions: A laplacian framework for learning representation and control in markov decision processes
Mahadevan, S. and Maggioni, M · 2007
Earlier work this paper cites.
Core knowledge
Spelke, E. S. and Kinzler, K. D · 2007
Earlier work this paper cites.
An analysis of single-layer networks in unsupervised feature learning
Coates, A., Ng, A., and Lee, H · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Mikolov, T., Chen, K., Corrado, G., and Dean, J · 2013
Earlier work this paper cites.
Decaf: A deep convolutional activation feature for generic visual recognition
Donahue, J., Jia, Y., Vinyals, O., Hoffman, J., Zhang, N., Tzeng, E., and Darrell, T · 2014
Earlier work this paper cites.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D · 2015
Cited alongside, same era.
What makes imagenet good for transfer learning?
Huh, M., Agrawal, P., and Efros, A. A · 2016
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T · 2017
Cited alongside, same era.
Ha, D. and Schmidhuber, J · 2018
Cited alongside, same era.
Soft actor-critic algorithms and applications
Haarnoja, T., Zhou, A., Hartikainen, K., Tucker, G., Ha, S., Tan, J., Kumar, V., Zhu, H., Gupta, A., Abbeel, P., et al · 2018
Cited alongside, same era.
FLAMBE: Structural complexity and representation learning of low rank mdps
Agarwal, A., Kakade, S., Krishnamurthy, A., and Sun, W · 2020
Later among the works it cites.
Scalable methods for computing state similarity in deterministic markov decision processes
Castro, P. S · 2020
Later among the works it cites.
Learning to Simulate Dynamic Environments with GameGAN
Kim, S. W., Zhou, Y., Philion, J., Torralba, A., and Fidler, S · 2020
Later among the works it cites.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Kostrikov, I., Yarats, D., and Fergus, R · 2020
Later among the works it cites.
Predictive information accelerates learning in rl
Lee, K.-H., Fischer, I., Liu, A., Guo, Y., Lee, H., Canny, J., and Guadarrama, S · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Oord, A. v. d., Li, Y., and Vinyals, O · 2018
Cited alongside, same era.
Densepose: Dense human pose estimation in the wild
Rıza Alp Güler, Natalia Neverova, I. K · 2018
Cited alongside, same era.
A geometric perspective on optimal representations for reinforcement learning
Bellemare, M., Dabney, W., Dadashi, R., Ali Taiga, A., Castro, P. S., Le Roux, N., Schuurmans, D., Lattimore, T., and Lyle, C · 2019
Cited alongside, same era.
Provably efficient rl with rich observations via latent state decoding
Du, S., Krishnamurthy, A., Jiang, N., Agarwal, A., Dudik, M., and Langford, J · 2019
Cited alongside, same era.
Deepmdp: Learning continuous latent space models for representation learning
Gelada, C., Kumar, S., Buckman, J., Nachum, O., and Bellemare, M. G · 2019
Cited alongside, same era.
Mask-guided portrait editing with conditional gans
Gu, S., Bao, J., Yang, H., Chen, D., Wen, F., and Yuan, L · 2019
Cited alongside, same era.
Real2sim: Visco-elastic parameter estimation from dynamic motion
Hahn, D., Banzet, P., Bern, J. M., and Coros, S · 2019
Cited alongside, same era.
Sample complexity of reinforcement learning using linearly combined model ensembles
Modi, A., Jiang, N., Tewari, A., and Singh, S · 2020
Later among the works it cites.
A short note on the kinetics-700-2020 human action dataset
Smaira, L., Carreira, J., Noland, E., Clancy, E., Wu, A., and Zisserman, A · 2020
Later among the works it cites.
dm_control: Software and tasks for continuous control
Tunyasuvunakool, S., Muldal, A., Doron, Y., Liu, S., Bohez, S., Merel, J., Erez, T., Lillicrap, T., Heess, N., and Tassa, Y · 2020
Later among the works it cites.
Understanding contrastive representation learning through alignment and uniformity on the hypersphere
Wang, T. and Isola, P · 2020
Later among the works it cites.
Learning invariant representations for reinforcement learning without reconstruction
Zhang, A., McAllister, R., Calandra, R., Gal, Y., and Levine, S · 2020
Later among the works it cites.
Provable rl with exogenous distractors via multistep inverse dynamics
Efroni, Y., Misra, D., Krishnamurthy, A., Agarwal, A., and Langford, J · 2021
Later among the works it cites.
Eysenbach, B., Salakhutdinov, R., and Levine, S · 2021
Later among the works it cites.
Learning task informed abstractions
Fu, X., Yang, G., Agrawal, P., and Jaakkola, T · 2021
Later among the works it cites.
RoboDesk: A multi-task reinforcement learning benchmark
Kannan, H., Hafner, D., Finn, C., and Erhan, D · 2021
Later among the works it cites.
Mastering visual continuous control: Improved data-augmented reinforcement learning
Yarats, D., Fergus, R., Lazaric, A., and Pinto, L · 2021
Later among the works it cites.
iNeRF: Inverting neural radiance fields for pose estimation
Yen-Chen, L., Florence, P., Barron, J. T., Rodriguez, A., Isola, P., and Lin, T.-Y · 2021
Later among the works it cites.