Fetching the paper…
Reading the bibliography…
Scaling deep learning to massive and diverse internet data has driven remarkable breakthroughs in domains such as video generation and natural language processing.
“Learning to identify critical states for reinforcement learning from videos”
Haozhe Liu et al · 1965
Earlier work this paper cites.
“Mind Children: The Future of Robot and Human Intelligence”
Hans Moravec · 1988
Earlier work this paper cites.
“Mobile robots for planetary exploration”
Klaus Schilling and Christoph. Jungius · 1995
Earlier work this paper cites.
“Cramming more components onto integrated circuits”
Gordon Moore · 1998
Earlier work this paper cites.
“Imagenet: A large-scale hierarchical image database”
Jia Deng et al · 2009
Earlier work this paper cites.
“Mujoco: A physics engine for model-based control”
Emanuel Todorov, Tom Erez and Yuval Tassa · 2012
Earlier work this paper cites.
“Explaining and Harnessing Adversarial Examples”
Ian. Goodfellow, Jonathon Shlens and Christian Szegedy · 2014
Earlier work this paper cites.
“Generative adversarial imitation learning”
Jonathan Ho and Stefano Ermon · 2016
Earlier work this paper cites.
“Robot learning”
Jan Peters et al · 2016
Earlier work this paper cites.
“Structure-from-motion revisited”
Johannes Schonberger and Jan-Michael Frahm · 2016
Earlier work this paper cites.
“Stochastic variational video prediction”
Mohammad Babaeizadeh et al · 2017
Earlier work this paper cites.
“The “Something Something” Video Database for Learning and Evaluating Visual Common Sense”
Raghav Goyal et al · 2017
Earlier work this paper cites.
“Ai2-thor: An interactive 3d environment for visual ai”
Eric Kolve et al · 2017
Earlier work this paper cites.
“Proximal policy optimization algorithms”
John Schulman et al · 2017
Earlier work this paper cites.
“Third-person imitation learning”
Bradly Stadie, Pieter Abbeel and Ilya Sutskever · 2017
Earlier work this paper cites.
“Domain randomization for transferring deep neural networks from simulation to the real world”
Joshua Tobin et al · 2017
Earlier work this paper cites.
“Attention is all you need”
Ashish Vaswani et al · 2017
Earlier work this paper cites.
“Playing hard exploration games by watching YouTube”
Yusuf Aytar et al · 2018
Earlier work this paper cites.
“Scaling egocentric vision: The epic-kitchens dataset”
Dima Damen et al · 2018
Earlier work this paper cites.
“Imitating Latent Policies from Observation”
Ashley. Edwards, Himanshu Sahni, Yannick Schroecker and Charles Isbell · 2018
Earlier work this paper cites.
“Soft actor-critic algorithms and applications”
Tuomas Haarnoja et al · 2018
Earlier work this paper cites.
“The DARPA robotics challenge finals: Results and perspectives”
Eric Krotkov et al · 2018
Earlier work this paper cites.
“Representation learning with contrastive predictive coding”
Aaron Oord, Yazhe Li and Oriol Vinyals · 2018
Earlier work this paper cites.
“SFV: Reinforcement Learning of Physical Skills from Videos”
Xue Peng et al · 2018
Earlier work this paper cites.
“Learning what you can do before doing anything”
Oleh Rybkin et al · 2018
Earlier work this paper cites.
“Time-contrastive networks: Self-supervised learning from video”
Pierre Sermanet et al · 2018
Earlier work this paper cites.
“Reinforcement learning: An introduction”
Richard Sutton and Andrew Barto · 2018
Earlier work this paper cites.
“Behavioral Cloning from Observation”
Faraz Torabi, Garrett Warnell and Peter Stone · 2018
Earlier work this paper cites.
“Generative Adversarial Imitation from Observation”
Faraz Torabi, Garrett Warnell and Peter Stone · 2018
Earlier work this paper cites.
“NerveNet: Learning Structured Policy with Graph Neural Networks”
Tingwu Wang, Renjie Liao, Jimmy Ba and Sanja Fidler · 2018
Earlier work this paper cites.
“Solving rubik’s cube with a robot hand”
Ilge Akkaya et al · 2019
Earlier work this paper cites.
“Perceptual values from observation”
Ashley Edwards and Charles Isbell · 2019
Earlier work this paper cites.
“Benchmarking Neural Network Robustness to Common Corruptions and Perturbations”
Dan Hendrycks and Thomas. Dietterich · 2019
Earlier work this paper cites.
“Deep lagrangian networks: Using physics as model prior for deep learning”
Michael Lutter, Christian Ritter and Jan Peters · 2019
Earlier work this paper cites.
“HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips”
Antoine Miech et al · 2019
Earlier work this paper cites.
“Keyframing the Future: Keyframe Discovery for Visual Prediction and Planning”
Karl Pertsch et al · 2019
Earlier work this paper cites.
“Language models are unsupervised multitask learners”
Alec Radford et al · 2019
Earlier work this paper cites.
“Learning Predictive Models From Observation and Interaction”
Karl Schmeckpeper et al · 2019
Earlier work this paper cites.
“AVID: Learning Multi-Stage Tasks via Pixel-Level Translation of Human Videos”
Laura Smith et al · 2019
Earlier work this paper cites.
“The Bitter Lesson”
Richard Sutton · 2019
Earlier work this paper cites.
“Recent Advances in Imitation Learning from Observation”
Faraz Torabi, Garrett Warnell and Peter Stone · 2019
Earlier work this paper cites.
“Robel: Robotics benchmarks for learning with low-cost robots”
Michael Ahn et al · 2020
Earlier work this paper cites.
“Development of quadruped walking robots: A review”
Priyaranjan Biswal and Prases Mohanty · 2020
Earlier work this paper cites.
“Language Models are Few-Shot Learners”
Tom. Brown et al · 2020
Earlier work this paper cites.
“Semantic Visual Navigation by Watching YouTube Videos”
Matthew Chang, Arjun Gupta and Saurabh Gupta · 2020
Earlier work this paper cites.
“Neural Topological SLAM for Visual Navigation”
Devendra Chaplot, Ruslan Salakhutdinov, Abhinav Gupta and Saurabh Gupta · 2020
Earlier work this paper cites.
“Estimating Q(s, s’) with Deep Deterministic Dynamics Gradients”
Ashley. Edwards et al · 2020
Earlier work this paper cites.
“Denoising Diffusion Probabilistic Models”
Jonathan Ho, Ajay Jain and P. Abbeel · 2020
Earlier work this paper cites.
“Scaling laws for neural language models”
Jared Kaplan et al · 2020
Earlier work this paper cites.
“Conservative q-learning for offline reinforcement learning”
Aviral Kumar, Aurick Zhou, George Tucker and Sergey Levine · 2020
Earlier work this paper cites.
“The nethack learning environment”
Heinrich Küttler et al · 2020
Earlier work this paper cites.
“Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems”
Sergey Levine, Aviral Kumar, G. Tucker and Justin Fu · 2020
Earlier work this paper cites.
“Learning latent plans from play”
Corey Lynch et al · 2020
Earlier work this paper cites.
“Learning agile robotic locomotion skills by imitating animals”
Xue Peng et al · 2020
Earlier work this paper cites.
“Recent Advances in Robot Learning from Demonstration”
Harish Ravichandar, Athanasios. Polydoros, Sonia Chernova and Aude Billard · 2020
Earlier work this paper cites.
“Frankmocap: Fast monocular 3d hand and body motion capture by regression and integration”
Yu Rong, Takaaki Shiratori and Hanbyul Joo · 2020
Earlier work this paper cites.
“Reinforcement Learning with Videos: Combining Offline Observations with Interaction”
Karl Schmeckpeper et al · 2020
Earlier work this paper cites.
“Mastering atari, go, chess and shogi by planning with a learned model”
Julian Schrittwieser et al · 2020
Earlier work this paper cites.
“Understanding Human Hands in Contact at Internet Scale”
Dandan Shan, Jiaqi Geng, Michelle Shu and David. Fouhey · 2020
Earlier work this paper cites.
“Concept2Robot: Learning manipulation concepts from instructions and human demonstrations”
Lin Shao et al · 2020
Earlier work this paper cites.
“Graph-structured visual imitation”
Maximilian Sieb et al · 2020
Earlier work this paper cites.
“Learning video representations from textual web supervision”
Jonathan Stroud et al · 2020
Earlier work this paper cites.
“Multi-expert learning of adaptive legged locomotion”
Chuanyu Yang et al · 2020
Earlier work this paper cites.
“Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning”
Tianhe Yu et al · 2020
Earlier work this paper cites.
“Mopo: Model-based offline policy optimization”
Tianhe Yu et al · 2020
Earlier work this paper cites.
“Sim-to-real transfer in deep reinforcement learning for robotics: a survey”
Wenshuai Zhao, Jorgeña Queralta and Tomi Westerlund · 2020
Earlier work this paper cites.
“Frozen in time: A joint video and image encoder for end-to-end retrieval”
Max Bain, Arsha Nagrani, Gül Varol and Andrew Zisserman · 2021
Earlier work this paper cites.
“On the opportunities and risks of foundation models”
Rishi Bommasani et al · 2021
Earlier work this paper cites.
“Learning Generalizable Robotic Reward Functions from "In-The-Wild" Human Videos”
Annie. Chen, Suraj Nair and Chelsea Finn · 2021
Earlier work this paper cites.
“An Empirical Study of Training Self-Supervised Vision Transformers”
Xinlei Chen, Saining Xie and Kaiming He · 2021
Earlier work this paper cites.
“Ego4D: Around the World in 3,000 Hours of Egocentric Video”
Kristen Grauman et al · 2021
Earlier work this paper cites.
“When is Unsupervised Disentanglement Possible?”, 2021
Daniel. Horan, Eitan Richardson and Yair Weiss · 2021
Earlier work this paper cites.
“Lora: Low-rank adaptation of large language models”
Edward Hu et al · 2021
Earlier work this paper cites.
“How to train your robot with deep reinforcement learning: lessons we have learned”
Julian Ibarz et al · 2021
Earlier work this paper cites.
“VOILA: Visual-Observation-Only Imitation Learning for Autonomous Navigation”
Haresh Karnan, Garrett Warnell, Xuesu Xiao and Peter Stone · 2021
Earlier work this paper cites.
“Adversarial Skill Chaining for Long-Horizon Robot Manipulation via Terminal State Regularization”
Youngwoon Lee, Joseph. Lim, Anima Anandkumar and Yuke Zhu · 2021
Earlier work this paper cites.
“Isaac gym: High performance gpu-based physics simulation for robot learning”
Viktor Makoviychuk et al · 2021
Earlier work this paper cites.
“Shaping embodied agent behavior with activity-context priors from egocentric video”
Tushar Nagarajan and Kristen Grauman · 2021
Earlier work this paper cites.
“DexMV: Imitation Learning for Dexterous Manipulation from Human Videos”
Yuzhe Qin et al · 2021
Earlier work this paper cites.
“High-Resolution Image Synthesis with Latent Diffusion Models”
Robin Rombach et al · 2021
Earlier work this paper cites.
“Self-Supervised Disentangled Representation Learning for Third-Person Imitation Learning”
Jinghuan Shang and Michael. Ryoo · 2021
Earlier work this paper cites.
“Learning by Watching: Physical Imitation of Manipulation Skills from Human Videos”
Haoyu Xiong et al · 2021
Earlier work this paper cites.
“VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding”
Hu Xu et al · 2021
Earlier work this paper cites.
“VideoGPT: Video Generation using VQ-VAE and Transformers”
Wilson Yan, Yunzhi Zhang, P. Abbeel and A. Srinivas · 2021
Earlier work this paper cites.
“DMotion: Robotic Visuomotor Control with Unsupervised Forward Model Learned from Videos”
Haoqi Yuan et al · 2021
Earlier work this paper cites.
“XIRL: Cross-embodiment Inverse Reinforcement Learning”
Kevin Zakka et al · 2021
Earlier work this paper cites.
“Merlot: Multimodal neural script knowledge models”
Rowan Zellers et al · 2021
Earlier work this paper cites.
“Plas: Latent action space for offline reinforcement learning”
Wenxuan Zhou, Sujay Bajracharya and David Held · 2021
Earlier work this paper cites.
“Do As I Can, Not As I Say: Grounding Language in Robotic Affordances”
Michael Ahn et al · 2022
Earlier work this paper cites.
“Flamingo: a visual language model for few-shot learning”
Jean-Baptiste Alayrac et al · 2022
Earlier work this paper cites.
“Unmanned Aerial Vehicles: A Literature Review”
Shaaban Ali et al · 2022
Earlier work this paper cites.
“Human-to-Robot Imitation in the Wild”
Shikhar Bahl, Abhi Gupta and Deepak Pathak · 2022
Earlier work this paper cites.
“Constitutional ai: Harmlessness from ai feedback”
Yuntao Bai et al · 2022
Earlier work this paper cites.
“Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos”
Bowen Baker et al · 2022
Earlier work this paper cites.
“RT-1: Robotics Transformer for Real-World Control at Scale”
Anthony Brohan et al · 2022
Earlier work this paper cites.
“MaskGIT: Masked Generative Image Transformer”
Huiwen Chang et al · 2022
Earlier work this paper cites.
“Learning Value Functions from Undirected State-only Experience”
Matthew Chang, Arjun Gupta and Saurabh Gupta · 2022
Earlier work this paper cites.
“From Play to Policy: Conditional Behavior Generation from Uncurated Robot Data”
Zichen Cui, Yibin Wang, Nur(Mahi) Shafiullah and Lerrel Pinto · 2022
Earlier work this paper cites.
“Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100”
Dima Damen et al · 2022
Earlier work this paper cites.
“Epic-kitchens visor benchmark: Video segmentations and object relations”
Ahmad Darkhalil et al · 2022
Earlier work this paper cites.
“ProcTHOR: Large-Scale Embodied AI Using Procedural Generation”
Matt Deitke et al · 2022
Earlier work this paper cites.
“Minedojo: Building open-ended embodied agents with internet-scale knowledge”
Linxi Fan et al · 2022
Earlier work this paper cites.
“Simvp: Simpler yet better video prediction”
Zhangyang Gao, Cheng Tan, Lirong Wu and Stan Li · 2022
Earlier work this paper cites.
“Long video generation with time-agnostic vqgan and time-sensitive transformer”
Songwei Ge et al · 2022
Earlier work this paper cites.
“OmniMAE: Single Model Masked Pretraining on Images and Videos”
Rohit Girdhar et al · 2022
Earlier work this paper cites.
“Maskvit: Masked visual pre-training for video prediction”
Agrim Gupta et al · 2022
Earlier work this paper cites.
“Video diffusion models”
Jonathan Ho et al · 2022
Earlier work this paper cites.
“Imagen video: High definition video generation with diffusion models”
Jonathan Ho et al · 2022
Earlier work this paper cites.
“Diffusion models for video prediction and infilling”
Tobias Höppe et al · 2022
Earlier work this paper cites.
“Inner Monologue: Embodied Reasoning through Planning with Language Models”
Wenlong Huang et al · 2022
Earlier work this paper cites.
“Hierarchical Reinforcement Learning for Precise Soccer Shooting Skills using a Quadrupedal Robot”
Yandong Ji et al · 2022
Earlier work this paper cites.
“Graph Inverse Reinforcement Learning from Diverse Videos”
Sateesh Kumar et al · 2022
Earlier work this paper cites.
“A path towards autonomous machine intelligence version 0.9. 2, 2022-06-27”
Yann LeCun · 2022
Cited alongside, same era.
“Composing ensembles of pre-trained models via iterative consensus”
Shuang Li et al · 2022
Cited alongside, same era.
“VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training”
Yecheng Ma et al · 2022
Cited alongside, same era.
“DexVIP: Learning Dexterous Grasping with Human Hand Pose Priors from Video”
Priyanka Mandikal and Kristen Grauman · 2022
Cited alongside, same era.
“Multi-Object Navigation with dynamically learned neural implicit representations”
Pierre Marza, Laëtitia Matignon, Olivier Simonin and Christian Wolf · 2022
Cited alongside, same era.
“Self-supervised learning for videos: A survey”
Madeline Schiappa, Yogesh Rawat and Mubarak Shah · 2023
Later among the works it cites.
“Learning to Act without Actions”
Dominik Schmidt and Minqi Jiang · 2023
Later among the works it cites.
“RoboVQA: Multimodal Long-Horizon Reasoning for Robotics”
Pierre Sermanet et al · 2023
Later among the works it cites.
Nur Shafiullah et al · 2023
Later among the works it cites.
“Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action”
Dhruv Shah, Błażej Osiński and Sergey Levine · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks”
Oier Mees, Lukas Hermann, Erick Rosete-Beas and Wolfram Burgard · 2022
Cited alongside, same era.
“Learning audio-video modalities from image captions”
Arsha Nagrani et al · 2022
Cited alongside, same era.
“R3M: A Universal Visual Representation for Robot Manipulation”
Suraj Nair et al · 2022
Cited alongside, same era.
“Pathways language model (palm): Scaling to 540 billion parameters for breakthrough performance”
Sharan Narang and Aakanksha Chowdhery · 2022
Cited alongside, same era.
“Training language models to follow instructions with human feedback”
Long Ouyang et al · 2022
Cited alongside, same era.
“The Unsurprising Effectiveness of Pre-Trained Vision Models for Control”
Simone Parisi, Aravind Rajeswaran, Senthil Purushwalkam and Abhinav Gupta · 2022
Cited alongside, same era.
“Evolving Curricula with Regret-Based Environment Design”
Jack Parker-Holder et al · 2022
Cited alongside, same era.
Sneha Silwal et al · 2023
Later among the works it cites.
“How many videos are there on YouTube?”
Alice Sjöberg · 2023
Later among the works it cites.
“RoboCLIP: One Demonstration is Enough to Learn Robot Policies”
Sumedh Sontakke et al · 2023
Later among the works it cites.
“Video understanding with large language models: A survey”
Yunlong Tang et al · 2023
Later among the works it cites.
“Human-timescale adaptation in an open-ended task space”
Adaptive Team et al · 2023
Later among the works it cites.
“Gemini: a family of highly capable multimodal models”
Gemini Team et al · 2023
Later among the works it cites.
“Octo: An open-source generalist robot policy”, 2023
Octo Team et al · 2023
Later among the works it cites.
“PLEX: Making the Most of the Available Data for Robotic Manipulation Pretraining”
Garrett Thomas et al · 2023
Later among the works it cites.
“Video-Guided Skill Discovery”
Manan Tomar et al · 2023
Later among the works it cites.
“Unidexgrasp++: Improving dexterous grasping policy learning via geometry-aware curriculum and iterative generalist-specialist learning”
Weikang Wan et al · 2023
Later among the works it cites.
“MimicPlay: Long-Horizon Imitation Learning by Watching Human Play”
Chen Wang et al · 2023
Later among the works it cites.
“Manipulate by Seeing: Creating Manipulation Controllers from Pre-Trained Representations”
Jianren Wang et al · 2023
Later among the works it cites.
“VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking”
Limin Wang et al · 2023
Later among the works it cites.
“Optimal goal-reaching reinforcement learning via quasimetric learning”
Tongzhou Wang, Antonio Torralba, Phillip Isola and Amy Zhang · 2023
Later among the works it cites.
“InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation”
Yi Wang et al · 2023
Later among the works it cites.
“Robogen: Towards unleashing infinite data for automated robot learning via generative simulation”
Yufei Wang et al · 2023
Later among the works it cites.
“Any-point Trajectory Modeling for Policy Learning”
Chuan Wen et al · 2023
Later among the works it cites.
“Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation”
Hongtao Wu et al · 2023
Later among the works it cites.
“Pre-training Contextualized World Models with In-the-wild Videos for Reinforcement Learning”
Jialong Wu, Haoyu Ma, Chao Deng and Mingsheng Long · 2023
Later among the works it cites.
“Foundations for Transfer in Reinforcement Learning: A Taxonomy of Knowledge Modalities”
Markus Wulfmeier et al · 2023
Later among the works it cites.
“Towards Generalist Robots: A Promising Paradigm via Generative Simulation”, 2023
Zhou Xian et al · 2023
Later among the works it cites.
“Decomposing the generalization gap in imitation learning for visual robotic manipulation”
Annie Xie, Lisa Lee, Ted Xiao and Chelsea Finn · 2023
Later among the works it cites.
“XSkill: Cross Embodiment Skill Discovery”
Mengda Xu et al · 2023
Later among the works it cites.
Jingyun Yang et al · 2023
Later among the works it cites.
“Probabilistic Adaptation of Text-to-Video Models”
Mengjiao Yang et al · 2023
Later among the works it cites.
“Learning Interactive Real-World Simulators”
Mengjiao Yang et al · 2023
Later among the works it cites.
“Foundation Models for Decision Making: Problems, Methods, and Opportunities”
Sherry Yang et al · 2023
Later among the works it cites.
Weirui Ye et al · 2023
Later among the works it cites.
“Language Model Beats Diffusion – Tokenizer is Key to Visual Generation”, 2023
Lijun Yu et al · 2023
Later among the works it cites.
“Scaling robot learning with semantically imagined experience”
Tianhe Yu et al · 2023
Later among the works it cites.
“Language to rewards for robotic skill synthesis”
Wenhao Yu et al · 2023
Later among the works it cites.
“Hierarchical generative modelling for autonomous robots”
Kai Yuan, Noor Sajid, Karl Friston and Zhibin Li · 2023
Later among the works it cites.
“Video-llama: An instruction-tuned audio-visual language model for video understanding”
Hang Zhang, Xin Li and Lidong Bing · 2023
Later among the works it cites.
“Large language models as commonsense knowledge for large-scale task planning”
Zirui Zhao, Wee Lee and David Hsu · 2023
Later among the works it cites.
“Guiding Online Reinforcement Learning with Action-Free Offline Pretraining”
Deyao Zhu, Yuhui Wang, Jürgen Schmidhuber and Mohamed Elhoseiny · 2023
Later among the works it cites.
Ziwen Zhuang et al · 2023
Later among the works it cites.
“A definition of continual reinforcement learning”
David Abel et al · 2024
Later among the works it cites.
“Autort: Embodied foundation models for large scale orchestration of robotic agents”
Michael Ahn et al · 2024
Later among the works it cites.
“Foundational challenges in assuring alignment and safety of large language models”
Usman Anwar et al · 2024
Later among the works it cites.
“Memory Consolidation Enables Long-Context Video Understanding”
Ivana Balažević et al · 2024
Later among the works it cites.
“Lumiere: A space-time diffusion model for video generation”
Omer Bar-Tal et al · 2024
Later among the works it cites.
“Rt-h: Action hierarchies using language”
Suneel Belkhale et al · 2024
Later among the works it cites.
“Position: Scaling Simulation is Neither Necessary Nor Sufficient for In-the-Wild Robot Manipulation”
Homanga Bharadhwaj · 2024
Later among the works it cites.
“Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models”
Nils Blank et al · 2024
Later among the works it cites.
“Video generation models as world simulators”, 2024
Tim Brooks et al · 2024
Later among the works it cites.
“Genie: Generative Interactive Environments”, 2024
Jake Bruce et al · 2024
Later among the works it cites.
“Rh20t-p: A primitive-level robotic dataset towards composable generalization agents”
Zeren Chen et al · 2024
Later among the works it cites.
“Learning Robotic Manipulation Policies from Point Clouds with Conditional Flow Matching”
Eugenio Chisari et al · 2024
Later among the works it cites.
“Compositional Generative Modeling: A Single Model is Not All You Need”
Yilun Du and Leslie Kaelbling · 2024
Later among the works it cites.
Abhimanyu Dubey et al · 2024
Later among the works it cites.
“Learning by Watching: A Review of Video-based Learning Approaches for Robot Manipulation”
Chrisantus Eze and Christopher Crick · 2024
Later among the works it cites.
Maxence Faldor, Jenny Zhang, Antoine Cully and Jeff Clune · 2024
Later among the works it cites.
“Imitation Learning: A Survey of Learning Methods, Environments and Metrics”
Nathan Gavenski, Odinaldo Rodrigues and Michael Luck · 2024
Later among the works it cites.
Lin Guan et al · 2024
Later among the works it cites.
“Large-Scale Actionless Video Pre-Training via Discrete Diffusion for Efficient Policy Learning”
Haoran He et al · 2024
Later among the works it cites.
“Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers”
Vidhi Jain et al · 2024
Later among the works it cites.
“All Neural Networks, All Autonomous, All 1X Speed” Accessed: 2024-04-10, https://www.1x.tech/discover/all-neural-networks-all-autonomous-all-1x-speed , 2024
Eric Jang · 2024
Later among the works it cites.
“Visual Representation Learning with Stochastic Frame Prediction”
Huiwon Jang et al · 2024
Later among the works it cites.
Albert Jiang et al · 2024
Later among the works it cites.
“Miradata: A large-scale video dataset with long durations and structured captions”
Xuan Ju et al · 2024
Later among the works it cites.
“DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset”
Alexander Khazatsky et al · 2024
Later among the works it cites.
“OpenVLA: An Open-Source Vision-Language-Action Model”
Moo Kim et al · 2024
Later among the works it cites.
“Will Scaling Solve Robotics?” Accessed: 2024-10-08, 2024
Nishanth Kumar · 2024
Later among the works it cites.
“RoboHive: A Unified Framework for Robot Learning”
Vikash Kumar et al · 2024
Later among the works it cites.
“Learning to Walk from Three Minutes of Real-World Data with Semi-structured Dynamics Models”
Jacob Levy, Tyler Westenbroek and David Fridovich-Keil · 2024
Later among the works it cites.
Chengshu Li et al · 2024
Later among the works it cites.
“SDS–See it, Do it, Sorted: Quadruped Skill Synthesis from Single Video Demonstration”
Jeffrey Li, Maria Stamatopoulou and Dimitrios Kanoulas · 2024
Later among the works it cites.
“NuminaMath TIR”
Jia Li et al · 2024
Later among the works it cites.
“Okami: Teaching humanoid robots manipulation skills through single video imitation”
Jinhan Li et al · 2024
Later among the works it cites.
“Visual Robotic Manipulation with Depth-Aware Pretraining”
Jinming Li et al · 2024
Later among the works it cites.
“Ag2Manip: Learning Novel Manipulation Skills with Agent-Agnostic Visual and Action Representations”
Puhao Li et al · 2024
Later among the works it cites.
“Evaluating Real-World Robot Manipulation Policies in Simulation”
Xuanlin Li et al · 2024
Later among the works it cites.
“Environment Curriculum Generation via Large Language Models”
William Liang et al · 2024
Later among the works it cites.
“SpawnNet: Learning Generalizable Visuomotor Skills from Pre-trained Network”
Xingyu Lin et al · 2024
Later among the works it cites.
“Libero: Benchmarking knowledge transfer for lifelong robot learning”
Bo Liu et al · 2024
Later among the works it cites.
“World Model on Million-Length Video And Language With Blockwise RingAttention”, 2024
Hao Liu, Wilson Yan, Matei Zaharia and Pieter Abbeel · 2024
Later among the works it cites.
“Enhancing Robotic Manipulation with AI Feedback from Multimodal Large Language Models”
Jinyi Liu et al · 2024
Later among the works it cites.
“Foundation Models for Video Understanding: A Survey”
Neelu Madan et al · 2024
Later among the works it cites.
“RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots”
Soroush Nasiriany et al · 2024
Later among the works it cites.
“PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs”
Soroush Nasiriany et al · 2024
Later among the works it cites.
“Hiql: Offline goal-conditioned rl with latent states as actions”
Seohong Park, Dibya Ghosh, Benjamin Eysenbach and Sergey Levine · 2024
Later among the works it cites.
“Q-slam: Quadric representations for monocular slam”
Chensheng Peng et al · 2024
Later among the works it cites.
“THE COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation”
Wilbert Pumacay et al · 2024
Later among the works it cites.
“Humanoid Locomotion as Next Token Prediction”
Ilija Radosavovic et al · 2024
Later among the works it cites.
“Understanding Long Videos in One Multimodal Language Model Pass”
Kanchana Ranasinghe, Xiang Li, Kumara Kahatapitiya and Michael Ryoo · 2024
Later among the works it cites.
“Neural Scaling Laws for Embodied AI”
Sebastian Sartor and Neil Thompson · 2024
Later among the works it cites.
“Body Transformer: Leveraging Robot Embodiment for Policy Learning”
Carmelo Sferrazza et al · 2024
Later among the works it cites.
“TraveLER: A Multi-LMM Agent Framework for Video Question-Answering”
Chuyi Shang et al · 2024
Later among the works it cites.
“Yell At Your Robot: Improving On-the-Fly from Language Corrections”
Lucy Shi et al · 2024
Later among the works it cites.
“Introducing RFM-1: Giving robots human-like reasoning capabilities” Accessed: 2024-03-29, https://covariant.ai/insights/introducing-rfm-1-giving-robots-human-like-reasoning-capabilities/ , 2024
Andrew Sohn et al · 2024
Later among the works it cites.
“GeRM: A Generalist Robotic Model with Mixture-of-experts for Quadruped Robot”
Wenxuan Song et al · 2024
Later among the works it cites.
“AutoMate: Specialist and Generalist Assembly Policies over Diverse Geometries”
Bingjie Tang et al · 2024
Later among the works it cites.
“Deep reinforcement learning for robotics: A survey of real-world successes”
Chen Tang et al · 2024
Later among the works it cites.
“Internvideo2: Scaling video foundation models for multimodal video understanding”
Yi Wang et al · 2024
Later among the works it cites.
“HelpSteer2: Open-source dataset for training top-performing reward models”
Zhilin Wang et al · 2024
Later among the works it cites.
“Pandora: Towards General World Model with Natural Language Actions and Video States”
Jiannan Xiang et al · 2024
Later among the works it cites.
Zhixuan Xu et al · 2024
Later among the works it cites.
“Video as the New Language for Real-World Decision Making”
Sherry Yang et al · 2024
Later among the works it cites.
“SimEndoGS: Efficient Data-driven Scene Simulation using Robotic Surgery Videos via Physics-embedded 3D Gaussians”, 2024
Zhenya Yang, Kai Chen, Yonghao Long and Qi Dou · 2024
Later among the works it cites.
“Latent Action Pretraining from Videos”, 2024
Seonghyeon Ye et al · 2024
Later among the works it cites.
“General Flow as Foundation Affordance for Scalable Robot Learning”
Chengbo Yuan, Chuan Wen, Tong Zhang and Yang Gao · 2024
Later among the works it cites.
“Mm-llms: Recent advances in multimodal large language models”
Duzhen Zhang et al · 2024
Later among the works it cites.
“VideoPrism: A Foundational Visual Encoder for Video Understanding”, 2024
Long Zhao et al · 2024
Later among the works it cites.
“3D-VLA: A 3D Vision-Language-Action Generative World Model”
Haoyu Zhen et al · 2024
Later among the works it cites.
“RoboDreamer: Learning Compositional World Models for Robot Imagination”
Siyuan Zhou et al · 2024
Later among the works it cites.
“EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video”
Ryan Hoque et al · 2025
Later among the works it cites.
“Recent Advances and Challenges in Industrial Robotics: A Systematic Review of Technological Trends and Emerging Applications”
Claudio Urrea and John Kern · 2025
Later among the works it cites.