Fetching the paper…

Exploration-exploitation trade-off for continuous-time episodic reinforcement learning with linear-convex models · Around