Fetching the paper…
Reading the bibliography…
We study the stochastic shortest path (SSP) problem in reinforcement learning with linear function approximation, where the transition kernel is represented as a linear mixture of unknown models.
An analysis of stochastic shortest path problems
Bertsekas, D. P. and Tsitsiklis, J. N · 1991
Earlier work this paper cites.
Convex optimization
Boyd, S., Boyd, S. P., and Vandenberghe, L · 2004
Earlier work this paper cites.
On the speed of convergence of value iteration on stochastic shortest-path problems
Bonet, B · 2007
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, T., Ortner, R., and Auer, P · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y., Pál, D., and Szepesvári, C · 2011
Earlier work this paper cites.
An empirical evaluation of thompson sampling
Chapelle, O. and Li, L · 2011
Earlier work this paper cites.
Heuristic search for generalized stochastic shortest path mdps
Kolobov, A., Mausam, M., Weld, D. S., and Geffner, H · 2011
Earlier work this paper cites.
Dynamic programming and optimal control: Volume I , volume 1
Bertsekas, D · 2012
Earlier work this paper cites.
Pac bounds for discounted mdps
Lattimore, T. and Hutter, M · 2012
Earlier work this paper cites.
Autonomous exploration for navigating in mdps
Lim, S. H. and Auer, P · 2012
Earlier work this paper cites.
Stochastic shortest path problems under weak conditions
Bertsekas, D. P. and Yu, H · 2013
Earlier work this paper cites.
(more) efficient reinforcement learning via posterior sampling
Osband, I., Russo, D., and Van Roy, B · 2013
Earlier work this paper cites.
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, P., and Zaremba, W · 2017
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., and Munos, R · 2017
Earlier work this paper cites.
Is q-learning provably efficient?
Jin, C., Allen-Zhu, Z., Bubeck, S., and Jordan, M. I · 2018
Earlier work this paper cites.
Is a good representation sufficient for sample efficient reinforcement learning?
Du, S. S., Kakade, S. M., Wang, R., and Yang, L. F · 2019
Earlier work this paper cites.
Planning with goal-conditioned policies
Nasiriany, S., Pong, V. H., Lin, S., and Levine, S · 2019
Earlier work this paper cites.
Sample-optimal parametric q-learning using linearly additive features
Yang, L. and Wang, M · 2019
Cited alongside, same era.
Regret minimization for reinforcement learning by evaluating the optimal bias function
Zhang, Z. and Ji, X · 2019
Cited alongside, same era.
Model-based reinforcement learning with value-targeted regression
Ayoub, A., Jia, Z., Szepesvari, C., Wang, M., and Yang, L · 2020
Cited alongside, same era.
Provably efficient exploration in policy optimization
Cai, Q., Yang, Z., Jin, C., and Wang, Z · 2020
Cited alongside, same era.
Near-optimal regret bounds for stochastic shortest path
Cohen, A., Kaplan, H., Mansour, Y., and Rosenberg, A · 2020
Cited alongside, same era.
Risk-sensitive reinforcement learning: Near-optimal risk-sample tradeoff in regret
Chen, L. and Luo, H · 2021
Closest in time.
Minimax regret for stochastic shortest path
Cohen, A., Efroni, Y., Mansour, Y., and Rosenberg, A · 2021
Closest in time.
Online learning for stochastic shortest path model via posterior sampling
Jafarnia-Jahromi, M., Chen, L., Jain, R., and Luo, H · 2021
Closest in time.
Corruption-robust exploration in episodic reinforcement learning
Lykouris, T., Simchowitz, M., Slivkins, A., and Sun, W · 2021
Closest in time.
Variance-aware off-policy evaluation with linear function approximation
Min, Y., Wang, T., Zhou, D., and Gu, Q · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fei, Y., Yang, Z., Chen, Y., Wang, Z., and Xie, Q · 2020
Cited alongside, same era.
The stochastic shortest path problem: a polyhedral combinatorics perspective
Guillot, M. and Stauffer, G · 2020
Cited alongside, same era.
Model-based reinforcement learning with value-targeted regression
Jia, Z., Yang, L., Szepesvari, C., and Wang, M · 2020
Cited alongside, same era.
Provably efficient reinforcement learning with linear function approximation
Jin, C., Yang, Z., Wang, Z., and Jordan, M. I · 2020
Cited alongside, same era.
Sample complexity of reinforcement learning using linearly combined model ensembles
Modi, A., Jiang, N., Tewari, A., and Singh, S · 2020
Cited alongside, same era.
Stochastic shortest path with adversarially changing costs
Rosenberg, A. and Mansour, Y · 2020
Cited alongside, same era.
Near-optimal regret bounds for stochastic shortest path
Rosenberg, A., Cohen, A., Mansour, Y., and Kaplan, H · 2020
Cited alongside, same era.
Regret bounds for stochastic shortest path problems with linear function approximation
Vial, D., Parulekar, A., Shakkottai, S., and Srikant, R · 2021
Closest in time.
Wang, T., Zhou, D., and Gu, Q · 2021
Closest in time.
Improved variance-aware confidence sets for linear bandits and linear mixture mdp
Zhang, Z., Yang, J., Ji, X., and Du, S. S · 2021
Closest in time.
Policy optimization for stochastic shortest path
Chen, L., Luo, H., and Rosenberg, A · 2022
Closest in time.
Cascaded gaps: Towards gap-dependent regret for risk-sensitive reinforcement learning
Fei, Y. and Xu, R · 2022
Closest in time.
Efficient risk-averse reinforcement learning
Greenberg, I., Chow, Y., Ghavamzadeh, M., and Mannor, S · 2022
Closest in time.
Nearly optimal algorithms for linear contextual bandits with adversarial corruptions
He, J., Zhou, D., Zhang, T., and Gu, Q · 2022
Closest in time.
Robust risk-aware reinforcement learning
Jaimungal, S., Pesenti, S. M., Wang, Y. S., and Tatsat, H · 2022
Closest in time.
Nearly minimax optimal regret for learning infinite-horizon average-reward mdps with linear function approximation
Wu, Y., Zhou, D., and Gu, Q · 2022
Closest in time.
Offline stochastic shortest path: Learning, evaluation and towards optimality
Yin, M., Chen, W., Wang, M., and Wang, Y.-X · 2022
Closest in time.
Corruption-robust offline reinforcement learning
Zhang, X., Chen, Y., Zhu, X., and Sun, W · 2022
Closest in time.