Fetching the paper…
Reading the bibliography…
This document serves as a position paper that outlines the authors' vision for a potential pathway towards generalist robots.
An introduction to ray tracing
A. S. Glassner · 1989
Earlier work this paper cites.
Real-time ray tracing using nvidia optix
H. Ludvigsen and A. C. Elster · 2010
Earlier work this paper cites.
Folding and crumpling adaptive sheets
R. Narain, T. Pfaff, and J. F. O’Brien · 2013
Earlier work this paper cites.
Adaptive tearing and cracking of thin sheets
T. Pfaff, R. Narain, J. M. De Joya, and J. F. O’Brien · 2014
Earlier work this paper cites.
Augmented mpm for phase-change and varied materials
A. Stomakhin, C. Schroeder, C. Jiang, L. Chai, J. Teran, and A. Selle · 2014
Earlier work this paper cites.
Domain randomization for transferring deep neural networks from simulation to the real world
J. Tobin, R. Fong, A. Ray, J. Schneider, W. Zaremba, and P. Abbeel · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Closed-chain manipulation of large objects by multi-arm robotic systems
Z. Xian, P. Lertkultanon, and Q.-C. Pham · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Driving policy transfer via modularity and abstraction
M. Müller, A. Dosovitskiy, B. Ghanem, and V. Koltun · 2018
Earlier work this paper cites.
Sim-to-real transfer of robotic control with dynamics randomization
X. B. Peng, M. Andrychowicz, W. Zaremba, and P. Abbeel · 2018
Earlier work this paper cites.
Can robots assemble an ikea chair?
F. Suárez-Ruiz, X. Zhou, and Q.-C. Pham · 2018
Earlier work this paper cites.
Fast fluid simulations with sparse volumes on the gpu
K. Wu, N. Truong, C. Yuksel, and R. Hoetzlein · 2018
Earlier work this paper cites.
Domain randomization for macromolecule structure classification and segmentation in electron cyro-tomograms
C. Che, Z. Xian, X. Zeng, X. Gao, and M. Xu · 2019
Earlier work this paper cites.
A thermomechanical material point method for baking and cooking
M. Ding, X. Han, S. Wang, T. F. Gast, and J. M. Teran · 2019
Earlier work this paper cites.
Taichi: a language for high-performance computation on spatially sparse data structures
Y. Hu, T.-M. Li, L. Anderson, J. Ragan-Kelley, and F. Durand · 2019
Earlier work this paper cites.
The bitter lesson
R. Sutton · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
E. Kaufmann, A. Loquercio, R. Ranftl, M. Müller, V. Koltun, and D. Scaramuzza · 2020
Earlier work this paper cites.
Pifuhd: Multi-level pixel-aligned implicit function for high-resolution 3d human digitization
S. Saito, T. Simon, J. Saragih, and H. Joo · 2020
Earlier work this paper cites.
Graph-structured visual imitation
M. Sieb, Z. Xian, A. Huang, O. Kroemer, and K. Fragkiadaki · 2020
Earlier work this paper cites.
3d-oes: Viewpoint-invariant object-factorized environment simulators
H.-Y. F. Tung, Z. Xian, M. Prabhudesai, S. Lal, and K. Fragkiadaki · 2020
Earlier work this paper cites.
Anisompm: Animating anisotropic damage mechanics: Supplemental document
J. Wolper, Y. Chen, M. Li, Y. Fang, Z. Qu, J. Lu, M. Cheng, and C. Jiang · 2020
Earlier work this paper cites.
Sapien: A simulated part-based interactive environment
F. Xiang, Y. Qin, K. Mo, Y. Xia, H. Zhu, F. Liu, M. Liu, H. Jiang, Y. Yuan, H. Wang, et al · 2020
Earlier work this paper cites.
Accelerating robotic reinforcement learning via parameterized action primitives
M. Dalal, D. Pathak, and R. R. Salakhutdinov · 2021
Earlier work this paper cites.
Plasticinelab: A soft-body manipulation benchmark with differentiable physics
Z. Huang, Y. Hu, T. Du, S. Zhou, H. Su, J. B. Tenenbaum, and C. Gan · 2021
Earlier work this paper cites.
Rma: Rapid motor adaptation for legged robots
A. Kumar, Z. Fu, D. Pathak, and J. Malik · 2021
Earlier work this paper cites.
Learning high-speed flight in the wild
A. Loquercio, E. Kaufmann, R. Ranftl, M. Müller, V. Koltun, and D. Scaramuzza · 2021
Earlier work this paper cites.
What matters in learning from offline human demonstrations for robot manipulation
A. Mandlekar, D. Xu, J. Wong, S. Nasiriany, C. Wang, R. Kulkarni, L. Fei-Fei, S. Savarese, Y. Zhu, and R. Martín-Martín · 2021
Earlier work this paper cites.
Nerf: Representing scenes as neural radiance fields for view synthesis
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng · 2021
Earlier work this paper cites.
The surprising effectiveness of representation learning for visual imitation
J. Pari, N. M. Shafiullah, S. P. Arunachalam, and L. Pinto · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Mlp-mixer: An all-mlp architecture for vision
I. O. Tolstikhin, N. Houlsby, A. Kolesnikov, L. Beyer, X. Zhai, T. Unterthiner, J. Yung, A. Steiner, D. Keysers, J. Uszkoreit, et al · 2021
Earlier work this paper cites.
Y. Wang, W. Wang, S. Joty, and S. C. Hoi · 2021
Earlier work this paper cites.
Hyperdynamics: Meta-learning object and agent dynamics with hypernetworks
Z. Xian, S. Lal, H.-Y. Tung, E. A. Platanios, and K. Fragkiadaki · 2021
Earlier work this paper cites.
Visual imitation made easy
S. Young, D. Gandhi, S. Tulsiani, A. Gupta, P. Abbeel, and L. Pinto · 2021
Earlier work this paper cites.
Transporter networks: Rearranging the visual world for robotic manipulation
A. Zeng, P. Florence, J. Tompson, S. Welker, J. Chien, M. Attarian, T. Armstrong, I. Krasin, D. Duong, V. Sindhwani, et al · 2021
Earlier work this paper cites.
World model as a graph: Learning latent landmarks for planning
L. Zhang, G. Yang, and B. C. Stadie · 2021
Earlier work this paper cites.
Do as i can, not as i say: Grounding language in robotic affordances
M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, et al · 2022
Earlier work this paper cites.
Human-to-robot imitation in the wild
S. Bahl, A. Gupta, and D. Pathak · 2022
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al · 2022
Cited alongside, same era.
Open-vocabulary queryable scene representations for real world planning
B. Chen, F. Xia, B. Ichter, K. Rao, K. Gopalakrishnan, M. S. Ryoo, A. Stone, and D. Kappler · 2022
Cited alongside, same era.
A system for general in-hand object re-orientation
T. Chen, J. Xu, and P. Agrawal · 2022
Cited alongside, same era.
Towards human-level bimanual dexterous manipulation with reinforcement learning
Y. Chen, T. Wu, S. Wang, X. Feng, J. Jiang, Z. Lu, S. McAleer, H. Dong, S.-C. Zhu, and Y. Yang · 2022
Cited alongside, same era.
From play to policy: Conditional behavior generation from uncurated robot data
Z. J. Cui, Y. Wang, N. Muhammad, L. Pinto, et al · 2022
3dgen: Triplane latent diffusion for textured mesh generation
A. Gupta, W. Xiong, Y. Nie, I. Jones, and B. Oğuz · 2023
Closest in time.
Learning agile soccer skills for a bipedal robot with deep reinforcement learning
T. Haarnoja, B. Moran, G. Lever, S. H. Huang, D. Tirumala, M. Wulfmeier, J. Humplik, S. Tunyasuvunakool, N. Y. Siegel, R. Hafner, et al · 2023
Closest in time.
Text2room: Extracting textured 3d meshes from 2d text-to-image models
L. Höllein, A. Cao, A. Owens, J. Johnson, and M. Nießner · 2023
Closest in time.
Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models
R. Huang, J. Huang, D. Yang, Y. Ren, L. Liu, M. Li, Z. Ye, J. Liu, X. Yin, and Z. Zhao · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Objaverse: A universe of annotated 3d objects
M. Deitke, D. Schwenk, J. Salvador, L. Weihs, O. Michel, E. VanderBilt, L. Schmidt, K. Ehsani, A. Kembhavi, and A. Farhadi · 2022
Cited alongside, same era.
Procthor: Large-scale embodied ai using procedural generation
M. Deitke, E. VanderBilt, A. Herrasti, L. Weihs, J. Salvador, K. Ehsani, W. Han, E. Kolve, A. Farhadi, A. Kembhavi, et al · 2022
Cited alongside, same era.
Coupling vision and proprioception for navigation of legged robots
Z. Fu, A. Kumar, A. Agarwal, H. Qi, J. Malik, and D. Pathak · 2022
Cited alongside, same era.
Navigating to objects in the real world
T. Gervet, S. Chintala, D. Batra, J. Malik, and D. S. Chaplot · 2022
Cited alongside, same era.
Inner monologue: Embodied reasoning through planning with language models
W. Huang, F. Xia, T. Xiao, H. Chan, J. Liang, P. Florence, A. Zeng, J. Tompson, I. Mordatch, Y. Chebotar, et al · 2022
Cited alongside, same era.
Q-attention: Enabling efficient learning for vision-based robotic manipulation
S. James and A. J. Davison · 2022
Cited alongside, same era.
Competition-level code generation with alphacode
Y. Li, D. Choi, J. Chung, N. Kushman, J. Schrittwieser, R. Leblond, T. Eccles, J. Keeling, F. Gimeno, A. Dal Lago, et al · 2022
Cited alongside, same era.
H. Jun and A. Nichol · 2023
Closest in time.
Scaling up gans for text-to-image synthesis
M. Kang, J.-Y. Zhu, R. Zhang, J. Park, E. Shechtman, S. Paris, and T. Park · 2023
Closest in time.
Dall-e-bot: Introducing web-scale diffusion models to robotics
I. Kapelyukh, V. Vosylius, and E. Johns · 2023
Closest in time.
Text2video-zero: Text-to-image diffusion models are zero-shot video generators
L. Khachatryan, A. Movsisyan, V. Tadevosyan, R. Henschel, Z. Wang, S. Navasardyan, and H. Shi · 2023
Closest in time.
Otter: A multi-modal model with in-context instruction tuning
B. Li, Y. Zhang, L. Chen, J. Wang, J. Yang, and Z. Liu · 2023
Closest in time.
Behavior-1k: A benchmark for embodied ai with 1,000 everyday activities and realistic simulation
C. Li, R. Zhang, J. Wong, C. Gokmen, S. Srivastava, R. Martín-Martín, C. Wang, G. Levine, M. Lingelbach, J. Sun, et al · 2023
Closest in time.
Text2motion: From natural language instructions to feasible plans
K. Lin, C. Agia, T. Migimatsu, M. Pavone, and J. Bohg · 2023
Closest in time.
Audioldm: Text-to-audio generation with latent diffusion models
H. Liu, Z. Chen, Y. Yuan, X. Mei, X. Liu, D. Mandic, W. Wang, and M. D. Plumbley · 2023
Closest in time.
Zero-1-to-3: Zero-shot one image to 3d object
R. Liu, R. Wu, B. Van Hoorick, P. Tokmakov, S. Zakharov, and C. Vondrick · 2023
Closest in time.
Meshdiffusion: Score-based generative 3d mesh modeling
Z. Liu, Y. Feng, M. J. Black, D. Nowrouzezahrai, L. Paull, and W. Liu · 2023
Closest in time.
Videofusion: Decomposed diffusion models for high-quality video generation
Z. Luo, D. Chen, Y. Zhang, Y. Huang, L. Wang, Y. Shen, D. Zhao, J. Zhou, and T. Tan · 2023
Closest in time.
Realfusion: 360 { \{ \ \backslash deg } \} reconstruction of any object from a single image
L. Melas-Kyriazi, C. Rupprecht, I. Laina, and A. Vedaldi · 2023
Closest in time.
Alan: Autonomously exploring robotic agents in the real world
R. Mendonca, S. Bahl, and D. Pathak · 2023
Closest in time.
Chatgpt plugins
OpenAI · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Learning humanoid locomotion with transformers
I. Radosavovic, T. Xiao, B. Zhang, T. Darrell, J. Malik, and K. Sreenath · 2023
Closest in time.
Toolformer: Language models can teach themselves to use tools
T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, L. Zettlemoyer, N. Cancedda, and T. Scialom · 2023
Closest in time.
Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface
Y. Shen, K. Song, X. Tan, D. Li, W. Lu, and Y. Zhuang · 2023
Closest in time.
Perceiver-actor: A multi-task transformer for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2023
Closest in time.
Open-world object manipulation using pre-trained vision-language models
A. Stone, T. Xiao, Y. Lu, K. Gopalakrishnan, K.-H. Lee, Q. Vuong, P. Wohlhart, B. Zitkovich, F. Xia, C. Finn, et al · 2023
Closest in time.
Vipergpt: Visual inference via python execution for reasoning
D. Surís, S. Menon, and C. Vondrick · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto · 2023
Closest in time.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Closest in time.
Mimicplay: Long-horizon imitation learning by watching human play
C. Wang, L. Fan, J. Sun, R. Zhang, L. Fei-Fei, D. Xu, Y. Zhu, and A. Anandkumar · 2023
Closest in time.
Softzoo: A soft robot co-design benchmark for locomotion in diverse environments
T.-H. Wang, P. Ma, A. E. Spielberg, Z. Xian, H. Zhang, J. B. Tenenbaum, D. Rus, and C. Gan · 2023
Closest in time.
Tidybot: Personalized robot assistance with large language models
J. Wu, R. Antonova, A. Kan, M. Lepert, A. Zeng, S. Song, J. Bohg, S. Rusinkiewicz, and T. Funkhouser · 2023
Closest in time.
Read and reap the rewards: Learning to play atari with the help of instruction manuals
Y. Wu, Y. Fan, P. P. Liang, A. Azaria, Y. Li, and T. M. Mitchell · 2023
Closest in time.
Fluidlab: A differentiable environment for benchmarking complex fluid manipulation
Z. Xian, B. Zhu, Z. Xu, H.-Y. Tung, A. Torralba, K. Fragkiadaki, and C. Gan · 2023
Closest in time.
Efficient tactile simulation with differentiability for robotic manipulation
J. Xu, S. Kim, T. Chen, A. R. Garcia, P. Agrawal, W. Matusik, and S. Sueda · 2023
Closest in time.
Roboninja: Learning an adaptive cutting policy for multi-material objects
Z. Xu, Z. Xian, X. Lin, C. Chi, Z. Huang, C. Gan, and S. Song · 2023
Closest in time.
Mm-react: Prompting chatgpt for multimodal reasoning and action
Z. Yang, L. Li, J. Wang, K. Lin, E. Azarnasab, F. Ahmed, Z. Liu, C. Liu, M. Zeng, and L. Wang · 2023
Closest in time.
Affordance diffusion: Synthesizing hand-object interactions
Y. Ye, X. Li, A. Gupta, S. De Mello, S. Birchfield, J. Song, S. Tulsiani, and S. Liu · 2023
Closest in time.
Scaling robot learning with semantically imagined experience
T. Yu, T. Xiao, A. Stone, J. Tompson, A. Brohan, S. Wang, J. Singh, C. Tan, J. Peralta, B. Ichter, et al · 2023
Closest in time.
Close the optical sensing domain gap by physics-grounded active stereo sensor simulation
X. Zhang, R. Chen, A. Li, F. Xiang, Y. Qin, J. Gu, Z. Ling, M. Liu, P. Zeng, S. Han, et al · 2023
Closest in time.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny · 2023
Closest in time.