Fetching the paper…
Reading the bibliography…
Visual reinforcement learning (RL), which makes decisions directly from high-dimensional visual inputs, has demonstrated significant potential in various domains.
Overfitting and undercomputing in machine learning
Tom Dietterich · 1995
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Peter Auer · 2002
Earlier work this paper cites.
Dueling network architectures for deep reinforcement learning
Ziyu Wang, Tom Schaul, Matteo Hessel, Hado Hasselt, Marc Lanctot, and Nando Freitas · 2003
Earlier work this paper cites.
Learning invariant representations for reinforcement learning without reconstruction
Amy Zhang, Rowan McAllister, Roberto Calandra, Yarin Gal, and Sergey Levine · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
How does mixup help with robustness and generalization?
Linjun Zhang, Zhun Deng, Kenji Kawaguchi, Amirata Ghorbani, and James Zou · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Earlier work this paper cites.
Contextual markov decision processes
Assaf Hallak, Dotan Di Castro, and Shie Mannor · 2015
Earlier work this paper cites.
Deep recurrent q-learning for partially observable mdps
Matthew Hausknecht and Peter Stone · 2015
Earlier work this paper cites.
Hidden parameter markov decision processes: A semiparametric regression approach for discovering latent task parametrizations
Finale Doshi-Velez and George Konidaris · 2016
Earlier work this paper cites.
Pac reinforcement learning with rich observations
Akshay Krishnamurthy, Alekh Agarwal, and John Langford · 2016
Earlier work this paper cites.
Understanding data augmentation for classification: when to warp?
Sebastien C Wong, Adam Gatt, Victor Stamatescu, and Mark D McDonnell · 2016
Earlier work this paper cites.
Charles Beattie, Joel Z Leibo, Denis Teplyashin, Tom Ward, Marcus Wainwright, Heinrich Küttler, Andrew Lefrancq, Simon Green, Víctor Valdés, Amir Sadik, et al · 2016
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Domain randomization for transferring deep neural networks from simulation to the real world
Josh Tobin, Rachel Fong, Alex Ray, Jonas Schneider, Wojciech Zaremba, and Pieter Abbeel · 2017
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz · 2017
Earlier work this paper cites.
Random erasing data augmentation
Zhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li, and Yi Yang · 2017
Earlier work this paper cites.
Arbitrary style transfer in real-time with adaptive instance normalization
Xun Huang and Serge Belongie · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Earlier work this paper cites.
Data augmentation generative adversarial networks
A Antoniou · 2017
Earlier work this paper cites.
An introduction to deep reinforcement learning
Vincent François-Lavet, Peter Henderson, Riashat Islam, Marc G Bellemare, Joelle Pineau, et al · 2018
Earlier work this paper cites.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2018
Earlier work this paper cites.
Scalable deep reinforcement learning for vision-based robotic manipulation
Dmitry Kalashnikov, Alex Irpan, Peter Pastor, Julian Ibarz, Alexander Herzog, Eric Jang, Deirdre Quillen, Ethan Holly, Mrinal Kalakrishnan, Vincent Vanhoucke, et al · 2018
Earlier work this paper cites.
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, et al · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Sim-to-real transfer of robotic control with dynamics randomization
Xue Bin Peng, Marcin Andrychowicz, Wojciech Zaremba, and Pieter Abbeel · 2018
Earlier work this paper cites.
Towards sample efficient reinforcement learning
Yang Yu · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Earlier work this paper cites.
Generalization and regularization in dqn
Jesse Farebrother, Marlos C Machado, and Michael Bowling · 2018
Earlier work this paper cites.
Data augmentation instead of explicit regularization
Alex Hernández-García and Peter König · 2018
Earlier work this paper cites.
Forward noise adjustment scheme for data augmentation
Francisco J Moreno-Barea, Fiammetta Strazzera, José M Jerez, Daniel Urda, and Leonardo Franco · 2018
Earlier work this paper cites.
Data augmentation via latent space interpolation for image classification
Xiaofeng Liu, Yang Zou, Lingsheng Kong, Zhihui Diao, Junliang Yan, Jun Wang, Site Li, Ping Jia, and Jane You · 2018
Earlier work this paper cites.
Visual reinforcement learning with imagined goals
Ashvin V Nair, Vitchyr Pong, Murtaza Dalal, Shikhar Bahl, Steven Lin, and Sergey Levine · 2018
Earlier work this paper cites.
Visualizing and understanding atari agents
Samuel Greydanus, Anurag Koul, Jonathan Dodge, and Alan Fern · 2018
Earlier work this paper cites.
Quantifying generalization in reinforcement learning
Karl Cobbe, Oleg Klimov, Chris Hesse, Taehoon Kim, and John Schulman · 2019
Earlier work this paper cites.
Unsupervised state representation learning in atari
Ankesh Anand, Evan Racah, Sherjil Ozair, Yoshua Bengio, Marc-Alexandre Côté, and R Devon Hjelm · 2019
Earlier work this paper cites.
A survey on image data augmentation for deep learning
Connor Shorten and Taghi M Khoshgoftaar · 2019
Earlier work this paper cites.
Observational overfitting in reinforcement learning
Xingyou Song, Yiding Jiang, Stephen Tu, Yilun Du, and Behnam Neyshabur · 2019
Earlier work this paper cites.
Provably efficient rl with rich observations via latent state decoding
Simon Du, Akshay Krishnamurthy, Nan Jiang, Alekh Agarwal, Miroslav Dudik, and John Langford · 2019
Earlier work this paper cites.
Generalization in reinforcement learning with selective noise injection and information bottleneck
Maximilian Igl, Kamil Ciosek, Yingzhen Li, Sebastian Tschiatschek, Cheng Zhang, Sam Devlin, and Katja Hofmann · 2019
Earlier work this paper cites.
Autoaugment: Learning augmentation strategies from data
Ekin D Cubuk, Barret Zoph, Dandelion Mane, Vijay Vasudevan, and Quoc V Le · 2019
Earlier work this paper cites.
Network randomization: A simple technique for generalization in deep reinforcement learning
Kimin Lee, Kibok Lee, Jinwoo Shin, and Honglak Lee · 2019
Earlier work this paper cites.
Cutmix: Regularization strategy to train strong classifiers with localizable features
Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo · 2019
Earlier work this paper cites.
Towards interpretable reinforcement learning using attention augmented agents
Alexander Mott, Daniel Zoran, Mike Chrzanowski, Daan Wierstra, and Danilo Jimenez Rezende · 2019
Earlier work this paper cites.
Leveraging human guidance for deep reinforcement learning tasks
Ruohan Zhang, Faraz Torabi, Lin Guan, Dana H Ballard, and Peter Stone · 2019
Earlier work this paper cites.
Self-ensembling with gan-based data augmentation for domain adaptation in semantic segmentation
Jaehoon Choi, Taekyung Kim, and Changick Kim · 2019
Earlier work this paper cites.
Virtual big data for gan based data augmentation
Hadi Mansourifar, Lin Chen, and Weidong Shi · 2019
Earlier work this paper cites.
A geometric perspective on optimal representations for reinforcement learning
Marc Bellemare, Will Dabney, Robert Dadashi, Adrien Ali Taiga, Pablo Samuel Castro, Nicolas Le Roux, Dale Schuurmans, Tor Lattimore, and Clare Lyle · 2019
Earlier work this paper cites.
On mutual information maximization for representation learning
Michael Tschannen, Josip Djolonga, Paul K Rubenstein, Sylvain Gelly, and Mario Lucic · 2019
Earlier work this paper cites.
Augmix: A simple data processing method to improve robustness and uncertainty
Dan Hendrycks, Norman Mu, Ekin D Cubuk, Barret Zoph, Justin Gilmer, and Balaji Lakshminarayanan · 2019
Earlier work this paper cites.
Fast task inference with variational intrinsic successor features
Steven Hansen, Will Dabney, Andre Barreto, Tom Van de Wiele, David Warde-Farley, and Volodymyr Mnih · 2019
Earlier work this paper cites.
Model-based reinforcement learning for atari
Lukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski, Roy H Campbell, Konrad Czechowski, Dumitru Erhan, Chelsea Finn, Piotr Kozakowski, Sergey Levine, et al · 2019
Earlier work this paper cites.
Time matters in regularizing deep networks: Weight decay and data augmentation affect early learning dynamics, matter little near convergence
Aditya Sharad Golatkar, Alessandro Achille, and Stefano Soatto · 2019
Earlier work this paper cites.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Denis Yarats, Ilya Kostrikov, and Rob Fergus · 2020
Earlier work this paper cites.
Regularization matters in policy optimization-an empirical study on continuous control
Zhuang Liu, Xuanlin Li, Bingyi Kang, and Trevor Darrell · 2020
Earlier work this paper cites.
Data-efficient reinforcement learning with self-predictive representations
Max Schwarzer, Ankesh Anand, Rishab Goel, R Devon Hjelm, Aaron Courville, and Philip Bachman · 2020
Earlier work this paper cites.
Deep reinforcement and infomax learning
Bogdan Mazoure, Remi Tachet des Combes, Thang Long Doan, Philip Bachman, and R Devon Hjelm · 2020
Cited alongside, same era.
Mastering atari with discrete world models
Danijar Hafner, Timothy P Lillicrap, Mohammad Norouzi, and Jimmy Ba · 2020
Cited alongside, same era.
Sample-efficient reinforcement learning via counterfactual-based data augmentation
Chaochao Lu, Biwei Huang, Ke Wang, José Miguel Hernández-Lobato, Kun Zhang, and Bernhard Schölkopf · 2020
Cited alongside, same era.
Invariant transform experience replay: Data augmentation for deep reinforcement learning
Yijiong Lin, Jiancong Huang, Matthieu Zimmer, Yisheng Guan, Juan Rojas, and Paul Weng · 2020
Cited alongside, same era.
robosuite: A modular simulation framework and benchmark for robot learning
Yuke Zhu, Josiah Wong, Ajay Mandlekar, and Roberto Martín-Martín · 2020
Cited alongside, same era.
Masked contrastive representation learning for reinforcement learning
Jinhua Zhu, Yingce Xia, Lijun Wu, Jiajun Deng, Wengang Zhou, Tao Qin, Tie-Yan Liu, and Houqiang Li · 2022
Closest in time.
Integrating contrastive learning with dynamic models for reinforcement learning from images
Bang You, Oleg Arenz, Youping Chen, and Jan Peters · 2022
Closest in time.
Cic: Contrastive intrinsic control for unsupervised skill discovery
Michael Laskin, Hao Liu, Xue Bin Peng, Denis Yarats, Aravind Rajeswaran, and Pieter Abbeel · 2022
Closest in time.
Pre-trained image encoder for generalizable visual reinforcement learning
Zhecheng Yuan, Zhengrong Xue, Bo Yuan, Xueqian Wang, Yi Wu, Yang Gao, and Huazhe Xu · 2022
Closest in time.
The unsurprising effectiveness of pre-trained vision models for control
Simone Parisi, Aravind Rajeswaran, Senthil Purushwalkam, and Abhinav Gupta · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Understanding artificial intelligence based radiology studies: What is overfitting?
Simukayi Mutasa, Shawn Sun, and Richard Ha · 2020
Cited alongside, same era.
Modals: Modality-agnostic automated data augmentation in the latent space
Tsz-Him Cheung and Dit-Yan Yeung · 2020
Cited alongside, same era.
Auto-encoder-based generative models for data augmentation on regression problems
Hiroshi Ohno · 2020
Cited alongside, same era.
Domain generalization with mixstyle
Kaiyang Zhou, Yongxin Yang, Yu Qiao, and Tao Xiang · 2020
Cited alongside, same era.
Arda: automatic relational data augmentation for machine learning
Nadiia Chepurko, Ryan Marcus, Emanuel Zgraggen, Raul Castro Fernandez, Tim Kraska, and David Karger · 2020
Cited alongside, same era.
Differentiable automatic data augmentation
Yonggang Li, Guosheng Hu, Yongtao Wang, Timothy Hospedales, Neil M Robertson, and Yongxin Yang · 2020
Cited alongside, same era.
A survey of regularization strategies for deep models
Reza Moradi, Reza Berangi, and Behrouz Minaei · 2020
Cited alongside, same era.
Tete Xiao, Ilija Radosavovic, Trevor Darrell, and Jitendra Malik · 2022
Closest in time.
Neural posterior domain randomization
Fabio Muratore, Theo Gruner, Florian Wiese, Boris Belousov, Michael Gienger, and Jan Peters · 2022
Closest in time.
Object detection using sim2real domain randomization for robotic applications
Dániel Horváth, Gábor Erdős, Zoltán Istenes, Tomáš Horváth, and Sándor Földi · 2022
Closest in time.
Block contextual mdps for continual learning
Shagun Sodhani, Franziska Meier, Joelle Pineau, and Amy Zhang · 2022
Closest in time.
Learn continuously, act discretely: Hybrid action-space reinforcement learning for optimal execution
Feiyang Pan, Tongzhe Zhang, Ling Luo, Jia He, and Shuoling Liu · 2022
Closest in time.
Invariance learning in deep neural networks with differentiable laplace approximations
Alexander Immer, Tycho FA van der Ouderaa, Vincent Fortuin, Gunnar Rätsch, and Mark van der Wilk · 2022
Closest in time.
Regularising for invariance to data augmentation improves supervised learning
Aleksander Botev, Matthias Bauer, and Soham De · 2022
Closest in time.
Local feature swapping for generalization in reinforcement learning
David Bertoin and Emmanuel Rachelson · 2022
Closest in time.
Compound domain generalization via meta-knowledge encoding
Chaoqi Chen, Jiongcheng Li, Xiaoguang Han, Xiaoqing Liu, and Yizhou Yu · 2022
Closest in time.
Rethinking the augmentation module in contrastive learning: Learning hierarchical augmentation invariance with expanded views
Junbo Zhang and Kaisheng Ma · 2022
Closest in time.
Banafsheh Rafiee, Jun Jin, Jun Luo, and Adam White · 2022
Closest in time.
Does self-supervised learning really improve reinforcement learning from pixels?
Xiang Li, Jinghuan Shang, Srijan Das, and Michael S Ryoo · 2022
Closest in time.
Reinforcement learning with automated auxiliary loss search
Tairan He, Yuge Zhang, Kan Ren, Minghuan Liu, Che Wang, Weinan Zhang, Yuqing Yang, and Dongsheng Li · 2022
Closest in time.
Generalizing reinforcement learning through fusing self-supervised learning into intrinsic motivation
Keyu Wu, Min Wu, Zhenghua Chen, Yuecong Xu, and Xiaoli Li · 2022
Closest in time.
Sin: Semantic inference network for few-shot streaming label learning
Zhen Wang, Liu Liu, Yiqun Duan, and Dacheng Tao · 2022
Closest in time.
On pre-training for visuo-motor control: Revisiting a learning-from-scratch baseline
Nicklas Hansen, Zhecheng Yuan, Yanjie Ze, Tongzhou Mu, Aravind Rajeswaran, Hao Su, Huazhe Xu, and Xiaolong Wang · 2022
Closest in time.
R3m: A universal visual representation for robot manipulation
Suraj Nair, Aravind Rajeswaran, Vikash Kumar, Chelsea Finn, and Abhinav Gupta · 2022
Closest in time.
Empirical evaluation and theoretical analysis for representation learning: A survey
Kento Nozawa and Issei Sato · 2022
Closest in time.
The primacy bias in deep reinforcement learning
Evgenii Nikishin, Max Schwarzer, Pierluca D’Oro, Pierre-Luc Bacon, and Aaron Courville · 2022
Closest in time.
Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble
Seunghyun Lee, Younggyo Seo, Kimin Lee, Pieter Abbeel, and Jinwoo Shin · 2022
Closest in time.
Boosting offline reinforcement learning via data rebalancing
Yang Yue, Bingyi Kang, Xiao Ma, Zhongwen Xu, Gao Huang, and Shuicheng Yan · 2022
Closest in time.
The challenges of exploration for offline reinforcement learning
Nathan Lambert, Markus Wulfmeier, William Whitney, Arunkumar Byravan, Michael Bloesch, Vibhavari Dasagi, Tim Hertweck, and Martin Riedmiller · 2022
Closest in time.
Bigger, better, faster: Human-level atari with human-level efficiency
Max Schwarzer, Johan Samir Obando Ceron, Aaron Courville, Marc G Bellemare, Rishabh Agarwal, and Pablo Samuel Castro · 2023
Closest in time.
Diffusion model as representation learner
Xingyi Yang and Xinchao Wang · 2023
Closest in time.
Diffusion models for reinforcement learning: A survey
Zhengbang Zhu, Hanye Zhao, Haoran He, Yichao Zhong, Shenyu Zhang, Yong Yu, and Weinan Zhang · 2023
Closest in time.
Scaling robot learning with semantically imagined experience
Tianhe Yu, Ted Xiao, Austin Stone, Jonathan Tompson, Anthony Brohan, Su Wang, Jaspiar Singh, Clayton Tan, Jodilyn Peralta, Brian Ichter, et al · 2023
Closest in time.
Genaug: Retargeting behaviors to unseen situations via generative augmentation
Zoey Chen, Sho Kiami, Abhishek Gupta, and Vikash Kumar · 2023
Closest in time.
Towards understanding how data augmentation works with imbalanced data
Damien A Dablain and Nitesh V Chawla · 2023
Closest in time.
The benefits of mixup for feature learning
Difan Zou, Yuan Cao, Yuanzhi Li, and Quanquan Gu · 2023
Closest in time.
The dormant neuron phenomenon in deep reinforcement learning
Ghada Sokar, Rishabh Agarwal, Pablo Samuel Castro, and Utku Evci · 2023
Closest in time.
Maintaining plasticity via regenerative regularization
Saurabh Kumar, Henrik Marklund, and Benjamin Van Roy · 2023
Closest in time.
Understanding plasticity in neural networks
Clare Lyle, Zeyu Zheng, Evgenii Nikishin, Bernardo Avila Pires, Razvan Pascanu, and Will Dabney · 2023
Closest in time.
Loss of plasticity in continual deep reinforcement learning
Zaheer Abbas, Rosie Zhao, Joseph Modayil, Adam White, and Marlos C Machado · 2023
Closest in time.
Foundation models in robotics: Applications, challenges, and the future
Roya Firoozi, Johnathan Tucker, Stephen Tian, Anirudha Majumdar, Jiankai Sun, Weiyu Liu, Yuke Zhu, Shuran Song, Ashish Kapoor, Karol Hausman, et al · 2023
Closest in time.
Embodiedgpt: Vision-language pre-training via embodied chain of thought
Yao Mu, Qinglong Zhang, Mengkang Hu, Wenhai Wang, Mingyu Ding, Jun Jin, Bin Wang, Jifeng Dai, Yu Qiao, and Ping Luo · 2023
Closest in time.
A survey on offline reinforcement learning: Taxonomy, review, and open problems
Rafael Figueiredo Prudencio, Marcos ROA Maximo, and Esther Luna Colombini · 2023
Closest in time.
Towards robust offline-to-online reinforcement learning via uncertainty and smoothness
Xiaoyu Wen, Xudong Yu, Rui Yang, Chenjia Bai, and Zhen Wang · 2023
Closest in time.
Domain randomization for sim2real transfer of automatically generated grasping datasets
Johann Huber, François Hélénon, Hippolyte Watrelot, Faïz Ben Amar, and Stéphane Doncieux · 2024
Closest in time.
Learning to manipulate anywhere: A visual generalizable framework for reinforcement learning
Zhecheng Yuan, Tianming Wei, Shuiqi Cheng, Gu Zhang, Yuanpei Chen, and Huazhe Xu · 2024
Closest in time.
Toward understanding generative data augmentation
Chenyu Zheng, Guoqiang Wu, and Chongxuan Li · 2024
Closest in time.
Synthetic experience replay
Cong Lu, Philip Ball, Yee Whye Teh, and Jack Parker-Holder · 2024
Closest in time.
Diffusion model is an effective planner and data synthesizer for multi-task reinforcement learning
Haoran He, Chenjia Bai, Kang Xu, Zhuoran Yang, Weinan Zhang, Dong Wang, Bin Zhao, and Xuelong Li · 2024
Closest in time.
A recipe for unbounded data augmentation in visual reinforcement learning
Abdulaziz Almuzairee, Nicklas Hansen, and Henrik I Christensen · 2024
Closest in time.
Bridging state and history representations: Understanding self-predictive rl
Tianwei Ni, Benjamin Eysenbach, Erfan Seyedsalehi, Michel Ma, Clement Gehring, Aditya Mahajan, and Pierre-Luc Bacon · 2024
Closest in time.
Enhancing reinforcement learning via transformer-based state predictive representations
Minsong Liu, Yuanheng Zhu, Yaran Chen, and Dongbin Zhao · 2024
Closest in time.
Neural metamorphosis
inchao Wang Xingyi Yang · 2024
Closest in time.
Revisiting data augmentation in deep reinforcement learning
Jianshu Hu, Yunpeng Jiang, and Paul Weng · 2024
Closest in time.
Deep reinforcement learning with plasticity injection
Evgenii Nikishin, Junhyuk Oh, Georg Ostrovski, Clare Lyle, Razvan Pascanu, Will Dabney, and André Barreto · 2024
Closest in time.
Normalization and effective learning rates in reinforcement learning
Clare Lyle, Zeyu Zheng, Khimya Khetarpal, James Martens, Hado van Hasselt, Razvan Pascanu, and Will Dabney · 2024
Closest in time.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Closest in time.
Kalie: Fine-tuning vision-language models for open-world manipulation without robot data
Grace Tang, Swetha Rajkumar, Yifei Zhou, Homer Rich Walke, Sergey Levine, and Kuan Fang · 2024
Closest in time.
Revisiting the minimalist approach to offline reinforcement learning
Denis Tarasov, Vladislav Kurenkov, Alexander Nikulin, and Sergey Kolesnikov · 2024
Closest in time.
Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning
Mitsuhiko Nakamoto, Simon Zhai, Anikait Singh, Max Sobol Mark, Yi Ma, Chelsea Finn, Aviral Kumar, and Sergey Levine · 2024
Closest in time.
Reset & distill: A recipe for overcoming negative transfer in continual reinforcement learning
Hongjoon Ahn, Jinu Hyeon, Youngmin Oh, Bosun Hwang, and Taesup Moon · 2024
Closest in time.