Fetching the paper…
Reading the bibliography…
Robotic manipulation systems operating in diverse, dynamic environments must exhibit three critical abilities: multitask interaction, generalization to unseen scenarios, and spatial memory.
The development of embodied cognition: Six lessons from babies
L. Smith and M. Gasser · 2005
Earlier work this paper cites.
Towards an episodic memory for cognitive robots
S. Jockel, M. Weser, D. Westhoff, and J. Zhang · 2008
Earlier work this paper cites.
Rgb-d mapping: Using kinect-style depth cameras for dense 3d modeling of indoor environments
P. Henry, M. Krainin, E. Herbst, X. Ren, and D. Fox · 2012
Earlier work this paper cites.
Probabilistic data association for semantic slam
S. L. Bowman, N. Atanasov, K. Daniilidis, and G. J. Pappas · 2017
Earlier work this paper cites.
Reinforcement learning with augmented data
M. Laskin, K. Lee, A. Stooke, L. Pinto, P. Abbeel, and A. Srinivas · 2020
Earlier work this paper cites.
Curl: Contrastive unsupervised representations for reinforcement learning
M. Laskin, A. Srinivas, and P. Abbeel · 2020
Earlier work this paper cites.
Object goal navigation using goal-oriented semantic exploration
D. S. Chaplot, D. P. Gandhi, A. Gupta, and R. R. Salakhutdinov · 2020
Earlier work this paper cites.
Rlbench: The robot learning benchmark & learning environment
S. James, Z. Ma, D. R. Arrojo, and A. J. Davison · 2020
Earlier work this paper cites.
Transporter networks: Rearranging the visual world for robotic manipulation
A. Zeng, P. Florence, J. Tompson, S. Welker, J. Chien, M. Attarian, T. Armstrong, I. Krasin, D. Duong, V. Sindhwani, et al · 2021
Earlier work this paper cites.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
D. Yarats, I. Kostrikov, and R. Fergus · 2021
Earlier work this paper cites.
Rrl: Resnet as representation for reinforcement learning
R. Shah and V. Kumar · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Earlier work this paper cites.
Roformer: Enhanced transformer with rotary position embedding
J. Su, Y. Lu, S. Pan, B. Wen, and Y. Liu · 2021
Earlier work this paper cites.
Coarse-to-fine q-attention: Efficient learning for visual robotic manipulation via discretisation
S. James, K. Wada, T. Laidlow, and A. J. Davison · 2021
Earlier work this paper cites.
Cliport: What and where pathways for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2022
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al · 2022
Earlier work this paper cites.
Coarse-to-fine q-attention with learned path ranking
S. James and P. Abbeel · 2022
Earlier work this paper cites.
Vip: Towards universal visual reward and representation via value-implicit pre-training
Y. J. Ma, S. Sodhani, D. Jayaraman, O. Bastani, V. Kumar, and A. Zhang · 2022
Earlier work this paper cites.
R3m: A universal visual representation for robot manipulation
S. Nair, A. Rajeswaran, V. Kumar, C. Finn, and A. Gupta · 2022
Earlier work this paper cites.
Vrl3: A data-driven framework for visual deep reinforcement learning
C. Wang, X. Luo, K. Ross, and D. Li · 2022
Cited alongside, same era.
Equivariant q q learning in spatial action spaces
D. Wang, R. Walters, X. Zhu, and R. Platt · 2022
Cited alongside, same era.
Partially observable markov decision processes in robotics: A survey
M. Lauri, D. Hsu, and J. Pajarinen · 2022
Cited alongside, same era.
Bc-z: Zero-shot task generalization with robotic imitation learning
E. Jang, A. Irpan, M. Khansari, D. Kappler, F. Ebert, C. Lynch, S. Levine, and C. Finn · 2022
Cited alongside, same era.
Instruction-driven history-aware policies for robotic manipulations
P.-L. Guhur, S. Chen, R. G. Pinel, M. Tapaswi, I. Laptev, and C. Schmid · 2022
Cited alongside, same era.
Dinov2: Learning robust visual features without supervision
M. Oquab, T. Darcet, T. Moutakanni, H. Q. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, M. Assran, N. Ballas, W. Galuba, R. Howes, P.-Y. B. Huang, S.-W. Li, I. Misra, M. G. Rabbat, V. Sharma, G. Synnaeve, H. Xu, H. Jégou, J. Mairal, P. Labatut, A. Joulin, and P. Bojanowski · 2023
Later among the works it cites.
The colosseum: A benchmark for evaluating generalization for robotic manipulation
W. Pumacay, I. Singh, J. Duan, R. Krishna, J. Thomason, and D. Fox · 2024
Later among the works it cites.
Manipulate-anything: Automating real-world robots using vision-language models
J. Duan, W. Yuan, W. Pumacay, Y. R. Wang, K. Ehsani, D. Fox, and R. Krishna · 2024
Later among the works it cites.
Rvt-2: Learning precise manipulation from few demonstrations
A. Goyal, V. Blukis, J. Xu, Y. Guo, Y.-W. Chao, and D. Fox · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Perceiver-actor: A multi-task transformer for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2023
Cited alongside, same era.
Rvt: Robotic view transformer for 3d object manipulation
A. Goyal, J. Xu, Y. Guo, V. Blukis, Y.-W. Chao, and D. Fox · 2023
Cited alongside, same era.
Learning fine-grained bimanual manipulation with low-cost hardware
T. Z. Zhao, V. Kumar, S. Levine, and C. Finn · 2023
Cited alongside, same era.
Diffusion policy: Visuomotor policy learning via action diffusion
C. Chi, Z. Xu, S. Feng, E. Cousineau, Y. Du, B. Burchfiel, R. Tedrake, and S. Song · 2023
Cited alongside, same era.
Polarnet: 3d point clouds for language-guided robotic manipulation
S. Chen, R. Garcia, C. Schmid, and I. Laptev · 2023
Cited alongside, same era.
M2t2: Multi-task masked transformer for object-centric pick and place
W. Yuan, A. Murali, A. Mousavian, and D. Fox · 2023
Cited alongside, same era.
Act3d: Infinite resolution action detection transformer for robotic manipulation
T. Gervet, Z. Xian, N. Gkanatsios, and K. Fragkiadaki · 2023
Cited alongside, same era.
Theia: Distilling diverse vision foundation models for robot learning
J. Shang, K. Schmeckpeper, B. B. May, M. V. Minniti, T. Kelestemur, D. Watkins, and L. Herlant · 2024
Later among the works it cites.
Sam-e: Leveraging visual foundation model with sequence imitation for embodied manipulation
J. Zhang, C. Bai, H. He, W. Xia, Z. Wang, B. Zhao, X. Li, and X. Li · 2024
Later among the works it cites.
Composing pre-trained object-centric representations for robotics from "what" and "where" foundation models
J. Shi, J. Qian, Y. J. Ma, and D. Jayaraman · 2024
Later among the works it cites.
Task-oriented hierarchical object decomposition for visuomotor control
J. Qian, Y. Li, B. Bucher, and D. Jayaraman · 2024
Later among the works it cites.
Copa: General robotic manipulation through spatial constraints of parts with foundation models
H. Huang, F. Lin, Y. Hu, S. Wang, and Y. Gao · 2024
Later among the works it cites.
Dynamem: Online dynamic spatio-semantic memory for open world mobile manipulation
P. Liu, Z. Guo, M. Warke, S. Chintala, C. Paxton, N. M. M. Shafiullah, and L. Pinto · 2024
Later among the works it cites.
Splat-mover: Multi-stage, open-vocabulary robotic manipulation via editable gaussian splatting
O. Shorinwa, J. Tucker, A. Smith, A. Swann, T. Chen, R. Firoozi, M. D. Kennedy, and M. Schwager · 2024
Later among the works it cites.
Out of sight, still in mind: Reasoning and planning about unobserved objects with video tracking enabled memory models
Y. Huang, J. Yuan, C. Kim, P. Pradhan, B. Chen, L. Fuxin, and T. Hermans · 2024
Later among the works it cites.
Sam 2: Segment anything in images and videos
N. Ravi, V. Gabeur, Y.-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. Rädle, C. Rolland, L. Gustafson, et al · 2024
Later among the works it cites.
Peract2: Benchmarking and learning for robotic bimanual manipulation tasks
M. Grotz, M. Shridhar, Y.-W. Chao, T. Asfour, and D. Fox · 2024
Later among the works it cites.
Rotary position embedding for vision transformer
B. Heo, S. Park, D. Han, and S. Yun · 2024
Later among the works it cites.
Autoregressive action sequence learning for robotic manipulation
X. Zhang, Y. Liu, H. Chang, L. Schramm, and A. Boularias · 2024
Later among the works it cites.
3d diffuser actor: Policy diffusion with 3d scene representations
T.-W. Ke, N. Gkanatsios, and K. Fragkiadaki · 2024
Later among the works it cites.
Towards generalizable vision-language robotic manipulation: A benchmark and llm-guided 3d policy
R. Garcia, S. Chen, and C. Schmid · 2024
Later among the works it cites.
L. Yang, B. Kang, Z. Huang, Z. Zhao, X. Xu, J. Feng, and H. Zhao · 2024
Later among the works it cites.