Fetching the paper…
Reading the bibliography…
Credit assignment in reinforcement learning is the problem of measuring an action's influence on future rewards.
Steps toward artificial intelligence
Minsky, M · 1961
Earlier work this paper cites.
Some guidelines and guarantees for common random numbers
Glasserman, P. and Yao, D. D · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y · 2000
Earlier work this paper cites.
Neural networks for machine learning lecture 6a overview of mini-batch gradient descent
Hinton, G., Srivastava, N., and Swersky, K · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Stochastic backpropagation and approximate inference in deep generative models
Rezende, D. J., Mohamed, S., and Wierstra, D · 2014
Earlier work this paper cites.
Gated feedback recurrent neural networks
Chung, J., Gulcehre, C., Cho, K., and Bengio, Y · 2015
Earlier work this paper cites.
Learning continuous control policies by stochastic value gradients
Heess, N., Wayne, G., Silver, D., Lillicrap, T., Erez, T., and Tassa, Y · 2015
Earlier work this paper cites.
Stochastic gradient estimation with finite differences
Buesing, L., Weber, T., and Mohamed, S · 2016
Earlier work this paper cites.
Domain-adversarial training of neural networks
Ganin, Y., Ustinova, E., Ajakan, H., Germain, P., Larochelle, H., Laviolette, F., Marchand, M., and Lempitsky, V · 2016
Earlier work this paper cites.
Hindsight experience replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, O. P., and Zaremba, W · 2017
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R · 2017
Earlier work this paper cites.
Rauber, P., Ummadisingu, A., Mutz, F., and Schmidhuber, J · 2017
Earlier work this paper cites.
Adversarial discriminative domain adaptation
Tzeng, E., Hoffman, J., Saenko, K., and Darrell, T · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Counterfactual multi-agent policy gradients
Foerster, J. N., Farquhar, G., Afouras, T., Nardelli, N., and Whiteson, S · 2018
Cited alongside, same era.
Ha, D. and Schmidhuber, J · 2018
Cited alongside, same era.
Variance reduction for reinforcement learning in input-driven environments
Mao, H., Venkatakrishnan, S. B., Schwarzkopf, M., and Alizadeh, M · 2018
Cited alongside, same era.
Model-based reinforcement learning for atari
Kaiser, L., Babaeizadeh, M., Milos, P., Osinski, B., Campbell, R. H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., et al · 2019
Later among the works it cites.
Counterfactual off-policy evaluation with gumbel-max structural causal models
Oberst, M. and Sontag, D · 2019
Later among the works it cites.
Normalizing flows for probabilistic modeling and inference
Papamakarios, G., Nalisnick, E., Rezende, D. J., Mohamed, S., and Lakshminarayanan, B · 2019
Later among the works it cites.
Stabilizing transformers for reinforcement learning
Parisotto, E., Song, H. F., Rae, J. W., Pascanu, R., Gulcehre, C., Jayakumar, S. M., Jaderberg, M., Kaufman, R. L., Clark, A., Noury, S., et al · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rezende, D. J. and Viola, F · 2018
Cited alongside, same era.
Variance reduction for policy gradient with action-dependent factorized baselines
Wu, C., Rajeswaran, A., Duan, Y., Kumar, V., Bayen, A. M., Kakade, S., Mordatch, I., and Abbeel, P · 2018
Cited alongside, same era.
Rudder: Return decomposition for delayed rewards
Arjona-Medina, J. A., Gillhofer, M., Widrich, M., Unterthiner, T., Brandstetter, J., and Hochreiter, S · 2019
Cited alongside, same era.
Woulda, coulda, shoulda: Counterfactually-guided policy search
Buesing, L., Weber, T., Zwols, Y., Racaniere, S., Guez, A., Lespiau, J.-B., and Heess, N · 2019
Cited alongside, same era.
Credit assignment as a proxy for transfer in reinforcement learning
Ferret, J., Marinier, R., Geist, M., and Pietquin, O · 2019
Cited alongside, same era.
Recurrent independent mechanisms
Goyal, A., Lamb, A., Hoffmann, J., Sodhani, S., Levine, S., Bengio, Y., and Schölkopf, B · 2019
Cited alongside, same era.
Value-driven hindsight modelling
Guez, A., Viola, F., Weber, T., Buesing, L., Kapturowski, S., Precup, D., Silver, D., and Heess, N · 2019
Cited alongside, same era.
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., et al · 2019
Later among the works it cites.
Grandmaster level in starcraft II using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al · 2019
Later among the works it cites.
Credit assignment techniques in stochastic computation graphs
Weber, T., Heess, N., Buesing, L., and Silver, D · 2019
Later among the works it cites.
Variance reduced advantage estimation with δ \delta -hindsight credit assignment
Young, K · 2019
Later among the works it cites.
Independence-aware advantage estimation
Zhang, P., Zhao, L., Liu, G., Bian, J., Huang, M., Qin, T., and Tie-Yan, L · 2019
Later among the works it cites.
Learning dexterous in-hand manipulation
Andrychowicz, O. M., Baker, B., Chociej, M., Jozefowicz, R., McGrew, B., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., et al · 2020
Closest in time.
Bica, I., Alaa, A. M., Jordon, J., and van der Schaar, M · 2020
Closest in time.
Designing optimal dynamic treatment regimes: A causal reinforcement learning approach
Zhang, J · 2020
Closest in time.
Towards practical credit assignment for deep reinforcement learning
Alipov, V., Simmons-Edler, R., Putintsev, N., Kalinin, P., and Vetrov, D · 2021
Closest in time.
Posterior value functions: Hindsight baselines for policy gradient methods
Nota, C., Thomas, P., and Da Silva, B. C · 2021
Closest in time.
Policy gradients incorporating the future
Venuto, D., Lau, E., Precup, D., and Nachum, O · 2021
Closest in time.