Fetching the paper…
Reading the bibliography…
Robots rely heavily on sensors, especially RGB and depth cameras, to perceive and interact with the world.
Robotics
K. S. Fu, R. C. Gonzalez, and C. G. Lee · 1983
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
D. A. Pomerleau · 1988
Earlier work this paper cites.
Depth-based tracking with physical constraints for robot manipulation
T. Schmidt, K. Hertkorn, R. Newcombe, Z. Marton, M. Suppa, and D. Fox · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al · 2015
Earlier work this paper cites.
Shapenet: An information-rich 3d model repository
A. X. Chang, T. Funkhouser, L. Guibas, P. Hanrahan, Q. Huang, Z. Li, S. Savarese, M. Savva, S. Song, H. Su, et al · 2015
Earlier work this paper cites.
Learning state representations with robotic priors
R. Jonschkowski and O. Brock · 2015
Earlier work this paper cites.
Voxnet: A 3d convolutional neural network for real-time object recognition
D. Maturana and S. Scherer · 2015
Earlier work this paper cites.
Actor-mimic: Deep multitask and transfer reinforcement learning
E. Parisotto, J. L. Ba, and R. Salakhutdinov · 2015
Earlier work this paper cites.
Pointnet: Deep learning on point sets for 3d classification and segmentation
C. R. Qi, H. Su, K. Mo, and L. J. Guibas · 2017
Earlier work this paper cites.
Pointnet++: Deep hierarchical feature learning on point sets in a metric space
C. R. Qi, L. Yi, H. Su, and L. J. Guibas · 2017
Earlier work this paper cites.
Semantic scene completion from a single depth image
S. Song, F. Yu, A. Zeng, A. X. Chang, M. Savva, and T. Funkhouser · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Reinforcement learning of active vision for manipulating objects under occlusions
R. Cheng, A. Agarwal, and K. Fragkiadaki · 2018
Earlier work this paper cites.
Pointfusion: Deep sensor fusion for 3d bounding box estimation
D. Xu, D. Anguelov, and A. Jain · 2018
Earlier work this paper cites.
Fusing bird’s eye view lidar point cloud and front view camera image for 3d object detection
Z. Wang, W. Zhan, and M. Tomizuka · 2018
Earlier work this paper cites.
Deep continuous fusion for multi-sensor 3d object detection
M. Liang, B. Yang, S. Wang, and R. Urtasun · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
A. v. d. Oord, Y. Li, and O. Vinyals · 2018
Earlier work this paper cites.
Multi-task learning using uncertainty to weigh losses for scene geometry and semantics
A. Kendall, Y. Gal, and R. Cipolla · 2018
Earlier work this paper cites.
Robonet: Large-scale multi-robot learning
S. Dasari, F. Ebert, S. Tian, S. Nair, B. Bucher, K. Schmeckpeper, S. Singh, S. Levine, and C. Finn · 2019
Earlier work this paper cites.
Revisiting point cloud classification: A new benchmark dataset and classification model on real-world data
M. A. Uy, Q.-H. Pham, B.-S. Hua, T. Nguyen, and S.-K. Yeung · 2019
Earlier work this paper cites.
Deepmdp: Learning continuous latent space models for representation learning
C. Gelada, S. Kumar, J. Buckman, O. Nachum, and M. G. Bellemare · 2019
Earlier work this paper cites.
Dream to control: Learning behaviors by latent imagination
D. Hafner, T. Lillicrap, J. Ba, and M. Norouzi · 2019
Earlier work this paper cites.
Multi-task multi-sensor fusion for 3d object detection
M. Liang, B. Yang, Y. Chen, R. Hu, and R. Urtasun · 2019
Earlier work this paper cites.
Densefusion: 6d object pose estimation by iterative dense fusion
C. Wang, D. Xu, Y. Zhu, R. Martín-Martín, C. Lu, L. Fei-Fei, and S. Savarese · 2019
Earlier work this paper cites.
Multi-task deep reinforcement learning with popart
M. Hessel, H. Soyer, L. Espeholt, W. Czarnecki, S. Schmitt, and H. van Hasselt · 2019
Earlier work this paper cites.
kpam: Keypoint affordances for category-level robotic manipulation
L. Manuelli, W. Gao, P. Florence, and R. Tedrake · 2019
Earlier work this paper cites.
Learning robust, real-time, reactive robotic grasping
D. Morrison, P. Corke, and J. Leitner · 2020
Earlier work this paper cites.
Pointcontrast: Unsupervised pre-training for 3d point cloud understanding
S. Xie, J. Gu, D. Guo, C. R. Qi, L. Guibas, and O. Litany · 2020
Earlier work this paper cites.
Rlbench: The robot learning benchmark & learning environment
S. James, Z. Ma, D. R. Arrojo, and A. J. Davison · 2020
Cited alongside, same era.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
I. Kostrikov, D. Yarats, and R. Fergus · 2020
Cited alongside, same era.
Learning invariant representations for reinforcement learning without reconstruction
A. Zhang, R. McAllister, R. Calandra, Y. Gal, and S. Levine · 2020
Cited alongside, same era.
Graspnet-1billion: A large-scale benchmark for general object grasping
H.-S. Fang, C. Wang, M. Gou, and C. Lu · 2020
Cited alongside, same era.
Pointpainting: Sequential fusion for 3d object detection
S. Vora, A. H. Lang, B. Helou, and O. Beijbom · 2020
Cited alongside, same era.
Point-m2ae: multi-scale masked autoencoders for hierarchical point cloud pre-training
R. Zhang, Z. Guo, P. Gao, R. Fang, B. Zhao, D. Wang, Y. Qiao, and H. Li · 2022
Later among the works it cites.
Learning 3d representations from 2d pre-trained models via image-to-point masked autoencoders
R. Zhang, L. Wang, Y. Qiao, P. Gao, and H. Li · 2022
Later among the works it cites.
Pointnext: Revisiting pointnet++ with improved training and scaling strategies
G. Qian, Y. Li, H. Peng, J. Mai, H. A. A. K. Hammoud, M. Elhoseiny, and B. Ghanem · 2022
Later among the works it cites.
Perceiver-actor: A multi-task transformer for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2022
Later among the works it cites.
Pre-trained image encoder for generalizable visual reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A novel depth and color feature fusion framework for 6d object pose estimation
G. Zhou, Y. Yan, D. Wang, and Q. Chen · 2020
Cited alongside, same era.
6-pack: Category-level 6d pose tracker with anchor-based keypoints
C. Wang, R. Martín-Martín, D. Xu, J. Lv, C. Lu, L. Fei-Fei, S. Savarese, and Y. Zhu · 2020
Cited alongside, same era.
Multi-task learning with deep neural networks: A survey
M. Crawshaw · 2020
Cited alongside, same era.
Gradient surgery for multi-task learning
T. Yu, S. Kumar, A. Gupta, S. Levine, K. Hausman, and C. Finn · 2020
Cited alongside, same era.
Understanding and improving information transfer in multi-task learning
S. Wu, H. R. Zhang, and C. Ré · 2020
Cited alongside, same era.
Bridge data: Boosting generalization of robotic skills with cross-domain datasets
F. Ebert, Y. Yang, K. Schmeckpeper, B. Bucher, G. Georgakis, K. Daniilidis, C. Finn, and S. Levine · 2021
Cited alongside, same era.
A comprehensive study of 3-d vision-based robot manipulation
Y. Cong, R. Chen, B. Ma, H. Liu, D. Hou, and C. Yang · 2021
Cited alongside, same era.
Z. Yuan, Z. Xue, B. Yuan, X. Wang, Y. Wu, Y. Gao, and H. Xu · 2022
Later among the works it cites.
Vip: Towards universal visual reward and representation via value-implicit pre-training
Y. J. Ma, S. Sodhani, D. Jayaraman, O. Bastani, V. Kumar, and A. Zhang · 2022
Later among the works it cites.
Instruction-following agents with jointly pre-trained vision-language models
H. Liu, L. Lee, K. Lee, and P. Abbeel · 2022
Later among the works it cites.
Instruction-driven history-aware policies for robotic manipulations
P.-L. Guhur, S. Chen, R. Garcia, M. Tapaswi, I. Laptev, and C. Schmid · 2022
Later among the works it cites.
Coarse-to-fine q-attention: Efficient learning for visual robotic manipulation via discretisation
S. James, K. Wada, T. Laidlow, and A. J. Davison · 2022
Later among the works it cites.
Learning generalizable dexterous manipulation from human grasp affordance
Y.-H. Wu, J. Wang, and X. Wang · 2022
Later among the works it cites.
Dexpoint: Generalizable point cloud reinforcement learning for sim-to-real dexterous manipulation
Y. Qin, B. Huang, Z.-H. Yin, H. Su, and X. Wang · 2022
Later among the works it cites.
Frame mining: a free lunch for learning robotic manipulation from 3d point clouds
M. Liu, X. Li, Z. Ling, Y. Li, and H. Su · 2022
Later among the works it cites.
Cliport: What and where pathways for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2022
Later among the works it cites.
Simple but effective: Clip embeddings for embodied ai
A. Khandelwal, L. Weihs, R. Mottaghi, and A. Kembhavi · 2022
Later among the works it cites.
Extract free dense labels from clip
C. Zhou, C. C. Loy, and B. Dai · 2022
Later among the works it cites.
Q-attention: Enabling efficient learning for vision-based robotic manipulation
S. James and A. J. Davison · 2022
Later among the works it cites.
Semantic-aware fine-grained correspondence
Y. Hu, R. Wang, K. Zhang, and Y. Gao · 2022
Later among the works it cites.
Ulip: Learning unified representation of language, image and point cloud for 3d understanding
L. Xue, M. Gao, C. Xing, R. Martín-Martín, J. Wu, C. Xiong, R. Xu, J. C. Niebles, and S. Savarese · 2022
Later among the works it cites.
Auto-lambda: Disentangling dynamic task relationships
S. Liu, S. James, A. J. Davison, and E. Johns · 2022
Later among the works it cites.
Behavior transformers: Cloning k k modes with one stone
N. M. Shafiullah, Z. Cui, A. A. Altanzaya, and L. Pinto · 2022
Later among the works it cites.
You only demonstrate once: Category-level manipulation from single visual demonstration
B. Wen, W. Lian, K. Bekris, and S. Schaal · 2022
Later among the works it cites.
https://www.intel.com/content/www/us/en/architecture-and-technology/realsense-overview.html , 2023
Realsense · 2023
Closest in time.
For pre-trained vision models in motor control, not all policy learning methods are created equal
Y. Hu, R. Wang, L. E. Li, and Y. Gao · 2023
Closest in time.
Language-driven representation learning for robotics
S. Karamcheti, S. Nair, A. S. Chen, T. Kollar, C. Finn, D. Sadigh, and P. Liang · 2023
Closest in time.
Where are we in the search for an artificial visual cortex for embodied intelligence?
A. Majumdar, K. Yadav, S. Arnaud, Y. J. Ma, C. Chen, S. Silwal, A. Jain, V.-P. Berges, P. Abbeel, J. Malik, et al · 2023
Closest in time.
Multi-view masked world models for visual robotic manipulation
Y. Seo, J. Kim, S. James, K. Lee, J. Shin, and P. Abbeel · 2023
Closest in time.
Programmatically grounded, compositionally generalizable robotic manipulation
R. Wang, J. Mao, J. Hsu, H. Zhao, J. Wu, and Y. Gao · 2023
Closest in time.
Imitating human behaviour with diffusion models
T. Pearce, T. Rashid, A. Kanervisto, D. Bignell, M. Sun, R. Georgescu, S. V. Macua, S. Z. Tan, I. Momennejad, K. Hofmann, et al · 2023
Closest in time.