Fetching the paper…
Reading the bibliography…
Humans use all of their senses to accomplish different tasks in everyday activities.
Research on advanced assembly automation
J. L. Nevins and D. E. Whitney · 1977
Earlier work this paper cites.
Hybrid position/force control of manipulators
M. Raibert and J. Craig · 1981
Earlier work this paper cites.
Quasi-static assembly of compliantly supported rigid parts
D. E. Whitney et al · 1982
Earlier work this paper cites.
A survey of robot learning from demonstration
B. D. Argall, S. Chernova, M. Veloso, and B. Browning · 2009
Earlier work this paper cites.
Interactive learning of the acoustic properties of household objects
J. Sinapov, M. Wiemer, and A. Stoytchev · 2009
Earlier work this paper cites.
Multimodal deep learning
J. Ngiam, A. Khosla, M. Kim, J. Nam, H. Lee, and A. Y. Ng · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2015
Earlier work this paper cites.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
The curious robot: Learning visual representations via physical interactions
L. Pinto, D. Gandhi, Y. Han, Y.-L. Park, and A. Gupta · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Acoustics based terrain classification for legged robots
J. Christie and N. Kottege · 2016
Earlier work this paper cites.
The feeling of success: Does touch sensing help predict grasp outcomes?
R. Calandra, A. Owens, M. Upadhyaya, W. Yuan, J. Lin, E. H. Adelson, and S. Levine · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Gelsight: High-resolution robot tactile sensors for estimating geometry and force
W. Yuan, S. Dong, and E. H. Adelson · 2017
Earlier work this paper cites.
An algorithmic perspective on imitation learning
T. Osa, J. Pajarinen, G. Neumann, J. A. Bagnell, P. Abbeel, J. Peters, et al · 2018
Earlier work this paper cites.
More than a feeling: Learning to grasp and regrasp using vision and touch
R. Calandra, A. Owens, D. Jayaraman, J. Lin, W. Yuan, J. Malik, E. H. Adelson, and S. Levine · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Audio-visual scene analysis with self-supervised multisensory features
A. Owens and A. A. Efros · 2018
Earlier work this paper cites.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
P. Anderson, Q. Wu, D. Teney, J. Bruce, M. Johnson, N. Sünderhauf, I. Reid, S. Gould, and A. Van Den Hengel · 2018
Cited alongside, same era.
Toward robotic manipulation
M. T. Mason · 2018
Cited alongside, same era.
Nonprehensile dynamic manipulation: A survey
F. Ruggiero, V. Lippiello, and B. Siciliano · 2018
Cited alongside, same era.
Time-contrastive networks: Self-supervised learning from video
P. Sermanet, C. Lynch, Y. Chebotar, J. Hsu, E. Jang, S. Schaal, S. Levine, and G. Brain · 2018
Cited alongside, same era.
Learning audio feedback for estimating amount and flow of granular material
S. Clarke, T. Rhodes, C. G. Atkeson, and O. Kroemer · 2018
Cited alongside, same era.
Making sense of vision and touch: Self-supervised learning of multimodal representations for contact-rich tasks
A review of robot learning for manipulation: Challenges, representations, and algorithms
O. Kroemer, S. Niekum, and G. D. Konidaris · 2021
Later among the works it cites.
Multibench: Multiscale benchmarks for multimodal representation learning
P. P. Liang, Y. Lyu, X. Fan, Z. Wu, Y. Cheng, J. Wu, L. Y. Chen, P. Wu, M. A. Lee, Y. Zhu, et al · 2021
Later among the works it cites.
Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text
H. Akbari, L. Yuan, R. Qian, W.-H. Chuang, S.-F. Chang, Y. Cui, and B. Gong · 2021
Later among the works it cites.
Learning to set waypoints for audio-visual navigation
C. Chen, S. Majumder, A.-H. Ziad, R. Gao, S. Kumar Ramakrishnan, and K. Grauman · 2021
Later among the works it cites.
Threedworld: A platform for interactive multi-modal physical simulation
C. Gan, J. Schwartz, S. Alter, M. Schrimpf, J. Traer, J. De Freitas, J. Kubilius, A. Bhandwaldar, N. Haber, M. Sano, et al · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. A. Lee, Y. Zhu, K. Srinivasan, P. Shah, S. Savarese, L. Fei-Fei, A. Garg, and J. Bohg · 2019
Cited alongside, same era.
Connecting touch and vision via cross-modal prediction
Y. Li, J.-Y. Zhu, R. Tedrake, and A. Torralba · 2019
Cited alongside, same era.
Making sense of audio vibration for liquid height estimation in robotic pouring
H. Liang, S. Li, X. Ma, N. Hendrich, T. Gerkmann, F. Sun, and J. Zhang · 2019
Cited alongside, same era.
Tactile-based insertion for dense box-packing
S. Dong and A. Rodriguez · 2019
Cited alongside, same era.
Deep dynamics models for learning dexterous manipulation
A. Nagabandi, K. Konolige, S. Levine, and V. Kumar · 2020
Cited alongside, same era.
Learning dexterous in-hand manipulation
O. M. Andrychowicz, B. Baker, M. Chociej, R. Jozefowicz, B. McGrew, J. Pachocki, A. Petron, M. Plappert, G. Powell, A. Ray, et al · 2020
Cited alongside, same era.
Stressd: Sim-to-real from sound for stochastic dynamics
C. Matl, Y. Narang, D. Fox, R. Bajcsy, and F. Ramos · 2020
Cited alongside, same era.
Attention bottlenecks for multimodal fusion
A. Nagrani, S. Yang, A. Arnab, A. Jansen, C. Schmid, and C. Sun · 2021
Later among the works it cites.
Geometry-aware multi-task learning for binaural audio generation from video
R. Garg, R. Gao, and K. Grauman · 2021
Later among the works it cites.
Zero-shot text-to-image generation
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever · 2021
Later among the works it cites.
Diffimpact: Differentiable rendering and identification of impact sounds
S. Clarke, N. Heravi, M. Rau, R. Gao, J. Wu, D. James, and J. Bohg · 2021
Later among the works it cites.
Active 3d shape reconstruction from vision and touch
E. Smith, D. Meger, L. Pineda, R. Calandra, J. Malik, A. Romero Soriano, and M. Drozdzal · 2021
Later among the works it cites.
Tactile-rl for insertion: Generalization to objects of unknown geometry
S. Dong, D. K. Jha, D. Romeres, S. Kim, D. Nikovski, and A. Rodriguez · 2021
Later among the works it cites.
Objectfolder: A dataset of objects with implicit visual, auditory, and tactile representations
R. Gao, Y.-Y. Chang, S. Mall, L. Fei-Fei, and J. Wu · 2021
Later among the works it cites.
Gelsight wedge: Measuring high-resolution 3d contact geometry with a compact robot finger
S. Wang, Y. She, B. Romero, and E. Adelson · 2021
Later among the works it cites.
Robocraft: Learning to see, simulate, and shape elasto-plastic objects with graph networks
H. Shi, H. Xu, Z. Huang, Y. Li, and J. Wu · 2022
Closest in time.
Play it by ear: Learning skills amidst occlusion through audio-visual imitation learning
M. Du, O. Lee, S. Nair, and C. Finn · 2022
Closest in time.
Merlot reserve: Multimodal neural script knowledge through vision and language and sound
R. Zellers, J. Lu, X. Lu, Y. Yu, Y. Zhao, M. Salehi, A. Kusupati, J. Hessel, A. Farhadi, and Y. Choi · 2022
Closest in time.
Audio-adaptive activity recognition across video domains
Y. Zhang, H. Doughty, L. Shao, and C. G. Snoek · 2022
Closest in time.
Hierarchical text-conditional image generation with clip latents
A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen · 2022
Closest in time.
Shapemap 3-d: Efficient shape mapping through dense touch and vision
S. Suresh, Z. Si, J. G. Mangelson, W. Yuan, and M. Kaess · 2022
Closest in time.
Objectfolder 2.0: A multisensory object dataset for sim2real transfer
R. Gao, Z. Si, Y.-Y. Chang, S. Clarke, J. Bohg, L. Fei-Fei, W. Yuan, and J. Wu · 2022
Closest in time.