Fetching the paper…
Reading the bibliography…
Recently, generative AI and reinforcement learning (RL) have been redefining what is possible for AI agents that take information flows as input and produce intelligent behavior.
Fine-tuning language models from human preferences
Ziegler, D.M., et al., 2019 · 1909
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R.S., 1988 · 1988
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R., Barto, A., 1998 · 1998
Earlier work this paper cites.
D4rl: Datasets for deep data-driven reinforcement learning
Fu, J., et al., 2020 · 2004
Earlier work this paper cites.
Efficient reductions for imitation learning, in: Teh, Y.W., Titterington, M. (Eds.), ICAIS, PMLR, Chia Laguna Resort, Sardinia, Italy. pp. 661–668
Ross, S., Bagnell, D., 2010 · 2010
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., et al., 2013 · 2013
Earlier work this paper cites.
Path planning with modified a star algorithm for a mobile robot
Duchoň, F., et al., 2014 · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., et al., 2015 · 2015
Earlier work this paper cites.
Rusu, A.A., et al., 2015 · 2015
Earlier work this paper cites.
Generalized grounding graphs: A probabilistic framework for understanding grounded commands
Kollar, T., et al., 2017 · 2017
Earlier work this paper cites.
Recurrent world models facilitate policy evolution
Ha, D., Schmidhuber, J., 2018 · 2018
Earlier work this paper cites.
Deterministic sampling-based motion planning: Optimality, complexity, and performance
Janson, L., et al., 2018 · 2018
Earlier work this paper cites.
Language as a cognitive tool to imagine goals in curiosity driven exploration
Colas, C., et al., 2020 · 2020
Earlier work this paper cites.
On the opportunities and risks of foundation models
Bommasani, R., et al., 2021 · 2021
Earlier work this paper cites.
Classifier-free diffusion guidance, in: NeurIPS Workshop on Deep Generative Models and Downstream Applications
Ho, J., Salimans, T., 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision, in: ICML, Pmlr. pp. 8748–8763
Radford, A., et al., 2021 · 2021
Earlier work this paper cites.
Zero-shot text-to-image generation, in: ICML, Pmlr. pp. 8821–8831
Ramesh, A., et al., 2021 · 2021
Earlier work this paper cites.
A predictive safety filter for learning-based control of constrained nonlinear dynamical systems
Wabersich, K.P., Zeilinger, M.N., 2021 · 2021
Earlier work this paper cites.
Advances in preference-based reinforcement learning: A review, in: 2022 IEEE SMC, pp. 2527–2532
Abdelkareem, Y., et al., 2022 · 2022
Earlier work this paper cites.
Do as i can, not as i say: Grounding language in robotic affordances
Ahn, M., et al., 2022 · 2022
Earlier work this paper cites.
Is conditional generative modeling all you need for decision-making?
Ajay, A., et al., 2022 · 2022
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J.B., et al., 2022 · 2022
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
Brohan, A., et al., 2022 · 2022
Earlier work this paper cites.
Offline reinforcement learning via high-fidelity generative behavior modeling
Chen, H., et al., 2022 · 2022
Earlier work this paper cites.
Can foundation models perform zero-shot task specification for robot manipulation?, in: LDCC, PMLR. pp. 893–905
Cui, Y., et al., 2022 · 2022
Earlier work this paper cites.
A review of safe reinforcement learning: Methods, theory and applications
Gu, S., et al., 2022 · 2022
Earlier work this paper cites.
Planning with diffusion for flexible behavior synthesis, in: ICML
Janner, M., et al., 2022 · 2022
Earlier work this paper cites.
Pre-training for robots: Offline rl enables learning new tasks in a handful of trials
Kumar, A., et al., 2022 · 2022
Earlier work this paper cites.
Vip: Towards universal visual reward and representation via value-implicit pre-training
Ma, Y.J., et al., 2022 · 2022
Earlier work this paper cites.
Zero-shot reward specification via grounded natural language, in: ICML, PMLR. pp. 14743–14752
Mahmoudieh, P., et al., 2022 · 2022
Earlier work this paper cites.
Learning language-conditioned robot behavior from offline data and crowd-sourced annotation, in: CoRL, PMLR. pp. 1303–1315
Nair, S., et al., 2022 · 2022
Earlier work this paper cites.
Reed, S., et al., 2022 · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models, in: Proceedings of the IEEE/CVF, pp. 10684–10695
Rombach, R., et al., 2022 · 2022
Earlier work this paper cites.
Cliport: What and where pathways for robotic manipulation, in: CoRL, PMLR. pp. 894–906
Shridhar, M., et al., 2022 · 2022
Earlier work this paper cites.
Anymorph: Learning transferable polices by inferring agent morphology, in: ICML, PMLR. pp. 21677–21691
Trabucco, B., et al., 2022 · 2022
Earlier work this paper cites.
Diffusion policies as an expressive policy class for offline reinforcement learning
Wang, Z., et al., 2022 · 2022
Earlier work this paper cites.
Multi-agent reinforcement learning is a sequence modeling problem
Wen, M., et al., 2022 · 2022
Earlier work this paper cites.
Socratic models: Composing zero-shot multimodal reasoning with language
Zeng, A., et al., 2022 · 2022
Earlier work this paper cites.
A review of motion planning algorithms for intelligent robots
Zhou, C., et al., 2022 · 2022
Earlier work this paper cites.
Achiam, J., et al., 2023 · 2023
Earlier work this paper cites.
Language reward modulation for pretraining reinforcement learning
Adeniji, A., et al., 2023 · 2023
Earlier work this paper cites.
Learning reward functions for robotic manipulation by observing humans, in: 2023 IEEE ICRA, IEEE. pp. 5006–5012
Alakuijala, M., et al., 2023 · 2023
Earlier work this paper cites.
Vision-language models as a source of rewards
Baumli, K., et al., 2023 · 2023
Earlier work this paper cites.
A survey of meta-reinforcement learning
Beck, J., et al., 2023 · 2023
Earlier work this paper cites.
Robotic offline rl from internet videos via value-function pre-training
Bhateja, C., et al., 2023 · 2023
Earlier work this paper cites.
Pact: Perception-action causal transformer for autoregressive robotics pre-training, in: 2023 IEEE/RSJ IROS, IEEE. pp. 3621–3627
Bonatti, R., et al., 2023 · 2023
Cited alongside, same era.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Brohan, A., et al., 2023 · 2023
Cited alongside, same era.
Latte: Language trajectory transformer, in: 2023 IEEE ICRA, IEEE. pp. 7287–7294
Bucker, A., et al., 2023 · 2023
Cited alongside, same era.
Grounding large language models in interactive environments with online reinforcement learning, in: ICML, PMLR. pp. 3676–3713
Carta, T., et al., 2023 · 2023
Cited alongside, same era.
Q-transformer: Scalable offline reinforcement learning via autoregressive q-functions, in: CoRL, PMLR. pp. 3909–3928
Chebotar, Y., et al., 2023 · 2023
Cited alongside, same era.
Survey on large language model-enhanced reinforcement learning: Concept, taxonomy, and methods
Cao, Y., et al., 2024 · 2024
Closest in time.
Offline-to-online reinforcement learning for image-based grasping with scarce demonstrations, in: CoRL Workshop on Mastering Robot Manipulation in a World of Abundant Data
Chan, B., et al., 2024 · 2024
Closest in time.
A survey of robotic language grounding
Cohen, V., et al., 2024 · 2024
Closest in time.
Plan-seq-learn: Language model guided rl for solving long horizon robotics tasks
Dalal, M., et al., 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Open-vocabulary queryable scene representations for real world planning, in: 2023 IEEE ICRA, pp. 11509–11522
Chen, B., et al., 2023a · 2023
Cited alongside, same era.
Diffusion policy: Visuomotor policy learning via action diffusion
Chi, C., et al., 2023 · 2023
Cited alongside, same era.
Accelerating reinforcement learning of robotic manipulations via feedback from large language models
Chu, K., et al., 2023 · 2023
Cited alongside, same era.
Augmenting autotelic agents with large language models, in: CLLA, PMLR. pp. 205–226
Colas, C., et al., 2023 · 2023
Cited alongside, same era.
Towards a unified agent with foundation models
Di Palo, N., et al., 2023 · 2023
Cited alongside, same era.
Consistency models as a rich and efficient policy class for reinforcement learning
Ding, Z., Jin, C., 2023 · 2023
Cited alongside, same era.
Foundation models in robotics: Applications, challenges, and the future
Firoozi, R., et al., 2023 · 2023
Cited alongside, same era.
Dubey, A., et al., 2024 · 2024
Closest in time.
Video prediction models as rewards for reinforcement learning
Escontrela, A., et al., 2024 · 2024
Closest in time.
Scaling rectified flow transformers for high-resolution image synthesis, in: ICML
Esser, P., et al., 2024 · 2024
Closest in time.
Conformal alignment: Knowing when to trust foundation models with guarantees
Gui, Y., et al., 2024 · 2024
Closest in time.
Td-mpc2: Scalable, robust world models for continuous control
Hansen, N., et al., 2024 · 2024
Closest in time.
Hassan, M., et al., 2024 · 2024
Closest in time.
Diffusion model is an effective planner and data synthesizer for multi-task reinforcement learning
He, H., et al., 2024 · 2024
Closest in time.
Efficient diffusion policies for offline reinforcement learning
Kang, B., et al., 2024 · 2024
Closest in time.
Rl-gpt: Integrating reinforcement learning and code-as-policy
Liu, S., et al., 2024 · 2024
Closest in time.
Precise and dexterous robotic manipulation via human-in-the-loop reinforcement learning
Luo, J., et al., 2024 · 2024
Closest in time.
Where are we in the search for an artificial visual cortex for embodied intelligence?
Majumdar, A., et al., 2024 · 2024
Closest in time.
Zero-shot safety prediction for autonomous robots with foundation world models
Mao, Z., et al., 2024 · 2024
Closest in time.
Policy agnostic rl: Offline rl and online rl fine-tuning of any class and backbone
Mark, M.S., et al., 2024 · 2024
Closest in time.
Multimodal foundation world models for generalist embodied agents
Mazzaglia, P., et al., 2024 · 2024
Closest in time.
Extracting reward functions from diffusion models
Nuti, F., et al., 2024 · 2024
Closest in time.
Octo: An open-source generalist robot policy, in: Proceedings of RSS, Delft, Netherlands
Oier, M., et al., 2024 · 2024
Closest in time.
Puma: deep metric imitation learning for stable motion primitives
Pérez-Dattari, R., et al., 2024 · 2024
Closest in time.
Interpreting and improving diffusion models from an optimization perspective, in: ICML, JMLR.org
Permenter, F., Yuan, C., 2024 · 2024
Closest in time.
Diffusion policy policy optimization, in: arXiv preprint arXiv:2409.00588
Ren, A.Z., et al., 2024 · 2024
Closest in time.
Roboclip: One demonstration is enough to learn robot policies
Sontakke, S., et al., 2024 · 2024
Closest in time.
Prompt, plan, perform: Llm-based humanoid control via quantized imitation learning, in: 2024 IEEE ICRA, IEEE. pp. 16236–16242
Sun, J., et al., 2024 · 2024
Closest in time.
Tian, R., et al., 2024 · 2024
Closest in time.
Intrinsic language-guided exploration for complex long-horizon robotic manipulation tasks, in: 2024 IEEE ICRA, IEEE. pp. 7493–7500
Triantafyllidis, E., et al., 2024 · 2024
Closest in time.
Continual learning and catastrophic forgetting
Van de Ven, G.M., et al., 2024 · 2024
Closest in time.
Code as reward: Empowering reinforcement learning with vlms
Venuto, D., et al., 2024 · 2024
Closest in time.
Fine-tuning reinforcement learning models is secretly a forgetting mitigation problem
Wołczyk, M., et al., 2024 · 2024
Closest in time.
ivideogpt: Interactive videogpts are scalable world models
Wu, J., et al., 2024 · 2024
Closest in time.
Rldg: Robotic generalist policy distillation via reinforcement learning
Xu, C., et al., 2024 · 2024
Closest in time.
Full-order sampling-based mpc for torque-level locomotion control via diffusion-style annealing
Xue, H., et al., 2024 · 2024
Closest in time.
Robot fine-tuning made easy: Pre-training rewards and policies for autonomous real-world reinforcement learning, in: 2024 IEEE ICRA, IEEE. pp. 4804–4811
Yang, J., et al., 2024 · 2024
Closest in time.
Envgen: Generating and adapting environments via llms for training embodied agents, in: COLM
Zala, A., et al., 2024 · 2024
Closest in time.
Vlmpc: Vision-language model predictive control for robotic manipulation, in: RSS
Zhao, W., et al., 2024 · 2024
Closest in time.
Unpacking failure modes of generative policies: Runtime monitoring of consistency and progress, in: CoRL, PMLR. pp. 689–723
Agia, C., et al., 2025 · 2025
Closest in time.
Fdpp: Fine-tune diffusion policy with human preference
Chen, Y., et al., 2025 · 2025
Closest in time.
Improving vision-language-action model with online reinforcement learning
Guo, Y., et al., 2025 · 2025
Closest in time.
Refined policy distillation: From vla generalists to rl experts
Jülg, T., et al., 2025 · 2025
Closest in time.
What can rl bring to vla generalization? an empirical study
Liu, J., et al., 2025 · 2025
Closest in time.
Generalizing safety beyond collision-avoidance via latent-space reachability analysis
Nakamura, K., et al., 2025 · 2025
Closest in time.
Strengthening generative robot policies through predictive world modeling
Qi, H., et al., 2025 · 2025
Closest in time.
Safediffuser: Safe planning with diffusion probabilistic models, in: ICLR
Xiao, W., et al., 2025 · 2025
Closest in time.