Fetching the paper…
Reading the bibliography…
Although pre-training on a large amount of data is beneficial for robot learning, current paradigms only perform large-scale pretraining for visual representations, whereas representations for other modalities are trained from scratch.
E. Donlon, S. Dong, M. Liu, J. Li, E. Adelson, and A. Rodriguez, “Gelslim: A high-resolution, compact, robust, and calibrated tactile-sensing finger,” in 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2018, pp. 1927–1934
1934
Earlier work this paper cites.
T. Bhattacharjee, A. Jain, S. Vaish, M. D. Killpack, and C. C. Kemp, “Tactile sensing over articulated joints with stretchable sensors,” in 2013 World Haptics Conference (WHC) . IEEE, 2013, pp. 103–108
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
Y. Aytar, C. Vondrick, and A. Torralba, “Soundnet: Learning sound representations from unlabeled video,” Advances in neural information processing systems , vol. 29, 2016
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in 2017 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2017, pp. 776–780
2017
Earlier work this paper cites.
W. Yuan, S. Dong, and E. H. Adelson, “Gelsight: High-resolution robot tactile sensors for estimating geometry and force,” Sensors , vol. 17, no. 12, p. 2762, 2017
2017
Earlier work this paper cites.
R. Arandjelovic and A. Zisserman, “Look, listen and learn,” in Proceedings of the IEEE International Conference on Computer Vision , 2017, pp. 609–617
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
R. Calandra, A. Owens, D. Jayaraman, J. Lin, W. Yuan, J. Malik, E. H. Adelson, and S. Levine, “More than a feeling: Learning to grasp and regrasp using vision and touch,” IEEE Robotics and Automation Letters , vol. 3, no. 4, pp. 3300–3307, 2018
2018
Earlier work this paper cites.
A. Murali, Y. Li, D. Gandhi, and A. Gupta, “Learning to grasp without seeing,” in International Symposium on Experimental Robotics . Springer, 2018, pp. 375–386
2018
Earlier work this paper cites.
S. Clarke, T. Rhodes, C. G. Atkeson, and O. Kroemer, “Learning audio feedback for estimating amount and flow of granular material,” Proceedings of Machine Learning Research , vol. 87, 2018
2018
Earlier work this paper cites.
R. Arandjelovic and A. Zisserman, “Objects that sound,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 435–451
2018
Earlier work this paper cites.
K. Zhang, M. Sharma, M. Veloso, and O. Kroemer, “Leveraging multimodal haptic sensory data for robust cutting,” in 2019 IEEE-RAS 19th International Conference on Humanoid Robots (Humanoids) . IEEE, 2019, pp. 409–416
2019
Earlier work this paper cites.
M. A. Lee, Y. Zhu, K. Srinivasan, P. Shah, S. Savarese, L. Fei-Fei, A. Garg, and J. Bohg, “Making sense of vision and touch: Self-supervised learning of multimodal representations for contact-rich tasks,” in 2019 International Conference on Robotics and Automation (ICRA) . IEEE, 2019, pp. 8943–8950
2019
Earlier work this paper cites.
W. Li, J. Konstantinova, Y. Noh, Z. Ma, A. Alomainy, and K. Althoefer, “An elastomer-based flexible optical force and tactile sensor,” in 2019 2nd IEEE International Conference on Soft Robotics (RoboSoft) . IEEE, 2019, pp. 361–366
2019
Cited alongside, same era.
B. Sundaralingam, A. S. Lambert, A. Handa, B. Boots, T. Hermans, S. Birchfield, N. Ratliff, and D. Fox, “Robust learning of tactile force estimation through robot interaction,” in 2019 International Conference on Robotics and Automation (ICRA) . IEEE, 2019, pp. 9035–9042
2019
Cited alongside, same era.
H. Zhao, C. Gan, W.-C. Ma, and A. Torralba, “The sound of motions,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 1735–1744
2019
Cited alongside, same era.
2020
Cited alongside, same era.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
V. Dean, S. Tulsiani, and A. Gupta, “See, hear, explore: Curiosity via audio-visual association,” Advances in Neural Information Processing Systems , vol. 33, pp. 14 961–14 972, 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
M. Lambeta, P.-W. Chou, S. Tian, B. Yang, B. Maloon, V. R. Most, D. Stroud, R. Santos, A. Byagowi, G. Kammerer et al. , “Digit: A novel design for a low-cost compact high-resolution tactile sensor with application to in-hand manipulation,” IEEE Robotics and Automation Letters , vol. 5, no. 3, pp. 3838–3845, 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Cited alongside, same era.
2021
Cited alongside, same era.
P. Morgado, N. Vasconcelos, and I. Misra, “Audio-visual instance discrimination with cross-modal agreement,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 12 475–12 486
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2022
Later among the works it cites.
S. Clarke, N. Heravi, M. Rau, R. Gao, J. Wu, D. James, and J. Bohg, “Diffimpact: Differentiable rendering and identification of impact sounds,” in Conference on Robot Learning . PMLR, 2022, pp. 662–673
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liu et al. , “Ego4d: Around the world in 3,000 hours of egocentric video,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 18 995–19 012
2022
Later among the works it cites.
V. Dean, D. K. Toyama, and D. Precup, “Don’t freeze your embedding: Lessons from policy finetuning in environment transfer,” in ICLR Workshop on Agent Learning in Open-Endedness , 2022. [Online]. Available: https://openreview.net/forum?id=HBHMrQD-LZc
2022
Later among the works it cites.
2022
Later among the works it cites.
2023
Later among the works it cites.
I. Radosavovic, T. Xiao, S. James, P. Abbeel, J. Malik, and T. Darrell, “Real-world robot learning with masked visual pre-training,” in Conference on Robot Learning . PMLR, 2023, pp. 416–426
2023
Later among the works it cites.
2023
Later among the works it cites.