Fetching the paper…
Reading the bibliography…
While the exploration for embodied AI has spanned multiple decades, it remains a persistent challenge to endow agents with human-level intelligence, including perception, learning, reasoning, decision-making, control, and generalization capabilities, so that they can perform general-purpose tasks in open, unstructured, and dynamic environments.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y Ng, Daishi Harada, and Stuart Russell · 1999
Earlier work this paper cites.
A systematic review on supervised and unsupervised machine learning algorithms for data science
Mohamed Alloghani, Dhiya Al-Jumeily, Jamila Mustafina, et al · 2020
Earlier work this paper cites.
A comprehensive survey on transfer learning
Fuzhen Zhuang, Zhiyuan Qi, Keyu Duan, et al · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, et al · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, et al · 2021
Earlier work this paper cites.
Model-free safe control for zero-violation reinforcement learning
Weiye Zhao, Tairan He, and Changliu Liu · 2021
Earlier work this paper cites.
Safe learning in robotics: From learning-based control to safe reinforcement learning
Lukas Brunke, Melissa Greeff, Adam W Hall, et al · 2022
Earlier work this paper cites.
Task-driven out-of-distribution detection with statistical guarantees for robot learning
Alec Farid, Sushant Veer, and Anirudha Majumdar · 2022
Earlier work this paper cites.
Ego4d: Around the world in 3,000 hours of egocentric video
Kristen Grauman, Andrew Westbury, Eugene Byrne, et al · 2022
Earlier work this paper cites.
Language models as zero-shot planners: Extracting actionable knowledge for embodied agents
Wenlong Huang, Pieter Abbeel, et al · 2022
Earlier work this paper cites.
Inner monologue: Embodied reasoning through planning with language models
Wenlong Huang, Fei Xia, Ted Xiao, et al · 2022
Earlier work this paper cites.
Neuro-symbolic procedural planning with commonsense prompting
Yujie Lu, Weixi Feng, Wanrong Zhu, et al · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, et al · 2022
Earlier work this paper cites.
Red teaming language models with language models
Ethan Perez, Saffron Huang, Francis Song, et al · 2022
Earlier work this paper cites.
Laion-5b: An open large-scale dataset for training next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, et al · 2022
Earlier work this paper cites.
Emergent abilities of large language models
Jason Wei, Yi Tay, Rishi Bommasani, et al · 2022
Earlier work this paper cites.
A cookbook of self-supervised learning
Randall Balestriero, Mark Ibrahim, Vlad Sobal, et al · 2023
Earlier work this paper cites.
Inverse dynamics pretraining learns good representations for multitask imitation
David Brandfonbrener, Ofir Nachum, and Joan Bruna · 2023
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
Anthony Brohan, Noah Brown, Justice Carbajal, et al · 2023
Earlier work this paper cites.
Do as i can, not as i say: Grounding language in robotic affordances
Anthony Brohan, Yevgen Chebotar, Chelsea Finn, et al · 2023
Earlier work this paper cites.
Q-transformer: Scalable offline reinforcement learning via autoregressive q-functions
Yevgen Chebotar, Quan Vuong, et al · 2023
Earlier work this paper cites.
Cerebras-gpt: Open compute-optimal language models trained on the cerebras wafer-scale cluster
Nolan Dey, Gurpreet Gosal, Hemant Khachane, et al · 2023
Earlier work this paper cites.
Palm-e: an embodied multimodal language model
Danny Driess, Fei Xia, Mehdi S. M. Sajjadi, et al · 2023
Cited alongside, same era.
Physically grounded vision-language models for robotic manipulation
Jensen Gao, Bidipta Sarkar, Fei Xia, et al · 2023
Cited alongside, same era.
Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives
Kristen Grauman, Andrew Westbury, Lorenzo Torresani, et al · 2023
Cited alongside, same era.
Maniskill2: A unified benchmark for generalizable manipulation skills
Jiayuan Gu, Fanbo Xiang, Xuanlin Li, et al · 2023
Cited alongside, same era.
Kazuki Hori, Kanata Suzuki, and Tetsuya Ogata · 2023
Cited alongside, same era.
OpenAI · 2023
Later among the works it cites.
Open X-Embodiment: Robotic learning datasets and RT-X models
Abhishek Padalkar, Acorn Pooley, Ajinkya Jain, et al · 2023
Later among the works it cites.
Open-ended instructable embodied agents with memory-augmented large language models
Gabriel Sarch, Yue Wu, Michael J Tarr, and Katerina Fragkiadaki · 2023
Later among the works it cites.
Semantic mechanical search with large vision and language models
Satvik Sharma, Huang Huang, Kaushik Shivakumar, et al · 2023
Later among the works it cites.
Progprompt: Generating situated robot task plans using large language models
Ishika Singh, Valts Blukis, Arsalan Mousavian, et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Look before you leap: Unveiling the power of gpt-4v in robotic vision-language planning
Yingdong Hu, Fanqi Lin, Tong Zhang, et al · 2023
Cited alongside, same era.
What went wrong? closing the sim-to-real gap via differentiable causal discovery
Peide Huang, Xilun Zhang, Ziang Cao, et al · 2023
Cited alongside, same era.
Instruct2act: Mapping multi-modality instructions to robotic actions with large language model
Siyuan Huang, Zhengkai Jiang, Hao Dong, Yu Qiao, Peng Gao, and Hongsheng Li · 2023
Cited alongside, same era.
Grounded decoding: Guiding text generation with grounded models for robot control
Wenlong Huang, Fei Xia, Dhruv Shah, et al · 2023
Cited alongside, same era.
Skill transformer: A monolithic policy for mobile manipulation
Xiaoyu Huang, Dhruv Batra, Akshara Rai, et al · 2023
Cited alongside, same era.
Human-oriented representation learning for robotic manipulation
Mingxiao Huo, Mingyu Ding, et al · 2023
Cited alongside, same era.
Vima: General robot manipulation with multimodal prompts
Yunfan Jiang, Agrim Gupta, Zichen Zhang, et al · 2023
Cited alongside, same era.
Hugo Touvron, Thibaut Lavril, Gautier Izacard, et al · 2023
Later among the works it cites.
Is imitation all you need? generalized decision-making with dual-phase training
Yao Wei, Yanchao Sun, Ruijie Zheng, et al · 2023
Later among the works it cites.
Omniobject3d: Large-vocabulary 3d object dataset for realistic perception, reconstruction and generation
Tong Wu, Jiarui Zhang, et al · 2023
Later among the works it cites.
A comprehensive survey of image augmentation techniques for deep learning
Mingle Xu, Sook Yoon, Alvaro Fuentes, et al · 2023
Later among the works it cites.
Track anything: Segment anything meets videos
Jinyu Yang, Mingqi Gao, Zhe Li, et al · 2023
Later among the works it cites.
Yunhao Yang, Cyrus Neary, and Ufuk Topcu · 2023
Later among the works it cites.
Weirui Ye, Yunsheng Zhang, Mengchen Wang, et al · 2023
Later among the works it cites.
Scaling robot learning with semantically imagined experience
Tianhe Yu, Ted Xiao, Austin Stone, et al · 2023
Later among the works it cites.
Language to rewards for robotic skill synthesis
Wenhao Yu, Nimrod Gileadi, et al · 2023
Later among the works it cites.
Isr-llm: Iterative self-refined large language model for long-horizon sequential task planning
Zhehua Zhou, Jiayang Song, Kunpeng Yao, Zhan Shu, and Lei Ma · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny · 2023
Later among the works it cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Brianna Zitkovich, Tianhe Yu, et al · 2023
Later among the works it cites.
Fmb: a functional manipulation benchmark for generalizable robotic learning
Jianlan Luo, Charles Xu, Fangchen Liu, Liam Tan, Zipeng Lin, Jeffrey Wu, Pieter Abbeel, and Sergey Levine · 2024
Closest in time.
Object-centric instruction augmentation for robotic manipulation
Junjie Wen, Yichen Zhu, Minjie Zhu, et al · 2024
Closest in time.
Language-conditioned robotic manipulation with fast and slow thinking
Minjie Zhu, Yichen Zhu, et al · 2024
Closest in time.
Llava- ϕ \phi : Efficient multi-modal assistant with small language model
Yichen Zhu, Minjie Zhu, Ning Liu, et al · 2024
Closest in time.