Fetching the paper…
Reading the bibliography…
This paper contributes a new approach for distributional reinforcement learning which elucidates a clean separation of transition structure and reward in the learning process.
Temporal credit assignment in reinforcement learning
Sutton, R. S · 1984
Earlier work this paper cites.
Learning from delayed rewards
Watkins, C · 1989
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
Dayan, P · 1993
Earlier work this paper cites.
Robot learning from demonstration
Atkeson, C. G. and Schaal, S · 1997
Earlier work this paper cites.
Conditional value-at-risk for general loss distributions
Rockafellar, R. T. and Uryasev, S · 2002
Earlier work this paper cites.
Nonparametric quantile estimation
Takeuchi, I., Le, Q. V., Sears, T. D., and Smola, A. J · 2006
Earlier work this paper cites.
Optimal transport: old and new , volume 338
Villani, C · 2008
Earlier work this paper cites.
Universal kernels on non-standard input spaces
Christmann, A. and Steinwart, I · 2010
Earlier work this paper cites.
Nonparametric return distribution approximation for reinforcement learning
Morimura, T., Sugiyama, M., Kashima, H., Hachiya, H., and Tanaka, T · 2010
Earlier work this paper cites.
A kernel two-sample test
Gretton, A., Borgwardt, K. M., Rasch, M. J., Schölkopf, B., and Smola, A · 2012
Earlier work this paper cites.
Energy statistics: A class of statistics based on distances
Székely, G. J. and Rizzo, M. L · 2013
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Earlier work this paper cites.
Markov decision processes: Discrete stochastic dynamic programming
Puterman, M. L · 2014
Earlier work this paper cites.
Design principles of the hippocampal cognitive map
Stachenfeld, K. L., Botvinick, M. M., and Gershman, S. J · 2014
Earlier work this paper cites.
Two-stage sampled learning theory on distributions
Szabo, Z., Gretton, A., Poczos, B., and Sriperumbudur, B · 2015
Earlier work this paper cites.
Brownian motion, martingales, and stochastic calculus
Le Gall, J.-F · 2016
Earlier work this paper cites.
Successor features for transfer in reinforcement learning
Barreto, A., Dabney, W., Munos, R., Hunt, J. J., Schaul, T., van Hasselt, H. P., and Silver, D · 2017
Earlier work this paper cites.
MMD GAN: towards deeper understanding of moment matching network
Li, C., Chang, W., Cheng, Y., Yang, Y., and Póczos, B · 2017
Earlier work this paper cites.
The successor representation in human reinforcement learning
Momennejad, I., Russek, E. M., Cheong, J. H., Botvinick, M. M., Daw, N. D., and Gershman, S. J · 2017
Earlier work this paper cites.
The hippocampus as a predictive map
Stachenfeld, K. L., Botvinick, M. M., and Gershman, S. J · 2017
Earlier work this paper cites.
Demystifying MMD GANs
Binkowski, M., Sutherland, D. J., Arbel, M., and Gretton, A · 2018
Earlier work this paper cites.
Universal successor features approximators
Borsa, D., Barreto, A., Quan, J., Mankowitz, D. J., van Hasselt, H., Munos, R., Silver, D., and Schaul, T · 2018
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q · 2018
Cited alongside, same era.
Eigenoption discovery through the deep successor representation
Machado, M. C., Rosenbaum, C., Guo, X., Liu, M., Tesauro, G., and Campbell, M · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Cited alongside, same era.
Invertible residual networks
Behrmann, J., Grathwohl, W., Chen, R. T. Q., Duvenaud, D., and Jacobsen, J · 2019
Cited alongside, same era.
Distributional multivariate policy evaluation and exploration with the Bellman GAN
Freirich, D., Shimkin, T., Meir, R., and Tamar, A · 2019
Cited alongside, same era.
Discovering options for exploration by minimizing cover time
Jinnai, Y., Park, J. W., Abel, D., and Konidaris, G · 2019
Distributional reinforcement learning via moment matching
Nguyen-Tang, T., Gupta, S., and Venkatesh, S · 2021
Later among the works it cites.
Learning One Representation to Optimize All Rewards
Touati, A. and Ollivier, Y · 2021
Later among the works it cites.
Seaborn: statistical data visualization
Waskom, M. L · 2021
Later among the works it cites.
Safe distributional reinforcement learning
Zhang, J. and Weng, P · 2021
Later among the works it cites.
Some Principled Methods for Deep Reinforcement Learning
Blier, L · 2022
Later among the works it cites.
Contrastive learning as goal-conditioned reinforcement learning
Eysenbach, B., Zhang, T., Levine, S., and Salakhutdinov, R · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A neurally plausible model learns successor representations in partially observable environments
Vértes, E. and Sahani, M · 2019
Cited alongside, same era.
Fully parameterized quantile function for distributional reinforcement learning
Yang, D., Zhao, L., Lin, Z., Qin, T., Bian, J., and Liu, T.-Y · 2019
Cited alongside, same era.
Selective dyna-style planning under limited model capacity
Abbas, Z., Sokota, S., Talvitie, E., and White, M · 2020
Cited alongside, same era.
A distributional analysis of sampling-based reinforcement learning algorithms
Amortila, P., Precup, D., Panangaden, P., and Bellemare, M. G · 2020
Cited alongside, same era.
The DeepMind JAX Ecosystem, 2020
Babuschkin, I., Baumli, K., Bell, A., Bhupatiraju, S., Bruce, J., Buchlovsky, P., Budden, D., Cai, T., Clark, A., Danihelka, I., Fantacci, C., Godwin, J., Jones, C., Hemsley, R., Hennigan, T., Hessel, M., Hou, S., Kapturowski, S., Keck, T., Kemaev, I., King, M., Kunesch, M., Martens, L., Merzic, H., Mikulik, V., Norman, T., Quan, J., Papamakarios, G., Ring, R., Ruiz, F., Sanchez, A., Schneider, R., Sezener, E., Spencer, S., Srinivasan, S., Wang, L., Stokowiec, W., and Viola, F · 2020
Cited alongside, same era.
Fast reinforcement learning with generalized policy updates
Barreto, A., Hou, S., Borsa, D., Silver, D., and Precup, D · 2020
Cited alongside, same era.
Discovering faster matrix multiplication algorithms with reinforcement learning
Fawzi, A., Balog, M., Huang, A., Hubert, T., Romera-Paredes, B., Barekatain, M., Novikov, A., Ruiz, F. J. R., Schrittwieser, J., Swirszcz, G., Silver, D., Hassabis, D., and Kohli, P · 2022
Later among the works it cites.
Investigating compounding prediction errors in learned dynamics models
Lambert, N., Pister, K., and Calandra, R · 2022
Later among the works it cites.
On the generalization of representations in reinforcement learning
Le Lan, C., Tu, S., Oberman, A., Agarwal, R., and Bellemare, M. G · 2022
Later among the works it cites.
Einops: Clear and reliable tensor manipulations with Einstein-like notation
Rogozhnikov, A · 2022
Later among the works it cites.
Generalised policy improvement with geometric policy composition
Thakoor, S., Rowland, M., Borsa, D., Dabney, W., Munos, R., and Barreto, A · 2022
Later among the works it cites.
One-step distributional reinforcement learning
Achab, M., Alamo, R., Djilali, Y. A. D., Fedyanin, K., and Moulines, E · 2023
Later among the works it cites.
Distributional Reinforcement Learning
Bellemare, M. G., Dabney, W., and Rowland, M · 2023
Later among the works it cites.
Combining behaviors with the successor features keyboard
Carvalho, W., Saraiva, A., Filos, A., Lampinen, A. K., Matthey, L., Lewis, R., Lee, H., Singh, S., Rezende, D. J., and Zoran, D · 2023
Later among the works it cites.
Proto-value networks: Scaling representation learning with auxiliary tasks
Farebrother, J., Greaves, J., Agarwal, R., Le Lan, C., Goroshin, R., Castro, P. S., and Bellemare, M. G · 2023
Later among the works it cites.
Reinforcement learning from passive data via latent intentions
Ghosh, D., Bhateja, C. A., and Levine, S · 2023
Later among the works it cites.
Flax: A neural network library and ecosystem for JAX, 2023
Heek, J., Levskaya, A., Oliver, A., Ritter, M., Rondepierre, B., Steiner, A., and van Zee, M · 2023
Later among the works it cites.
Maximum state entropy exploration using predecessor and successor representations
Jain, A. K., Lehnert, L., Rish, I., and Berseth, G · 2023
Later among the works it cites.
Temporal abstraction in reinforcement learning with the successor representation
Machado, M. C., Barreto, A., Precup, D., and Bowling, M · 2023
Later among the works it cites.
Dual RL: Unification and new methods for reinforcement and imitation learning
Sikchi, H., Zheng, Q., Zhang, A., and Niekum, S · 2023
Later among the works it cites.
Does zero-shot reinforcement learning exist?
Touati, A., Rapin, J., and Ollivier, Y · 2023
Later among the works it cites.
Fast imitation via behavior foundation models
Pirotta, M., Tirinzoni, A., Touati, A., Lazaric, A., and Ollivier, Y · 2024
Closest in time.