Fetching the paper…
Reading the bibliography…
Cross-modal distillation has been widely used to transfer knowledge across different modalities, enriching the representation of the target unimodal one.
“Imagenet: A large-scale hierarchical image database,”
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei, · 2009
Earlier work this paper cites.
“Ucf101: A dataset of 101 human actions classes from videos in the wild,”
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah, · 2012
Earlier work this paper cites.
“Activitynet: A large-scale video benchmark for human activity understanding,”
Bernard Ghanem Fabian Caba Heilbron, Victor Escorcia and Juan Carlos Niebles, · 2015
Earlier work this paper cites.
“Distilling the knowledge in a neural network,”
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean, · 2015
Earlier work this paper cites.
“Cross modal distillation for supervision transfer,”
Saurabh Gupta, Judy Hoffman, and Jitendra Malik, · 2016
Earlier work this paper cites.
“Soundnet: Learning sound representations from unlabeled video,”
Yusuf Aytar, Carl Vondrick, and Antonio Torralba, · 2016
Earlier work this paper cites.
“Audio set: An ontology and human-labeled dataset for audio events,”
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter, · 2017
Earlier work this paper cites.
“Look, listen and learn,”
Relja Arandjelovic and Andrew Zisserman, · 2017
Earlier work this paper cites.
“A closer look at spatiotemporal convolutions for action recognition,”
Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri, · 2018
Cited alongside, same era.
“Squeeze-and-excitation networks,”
Jie Hu, Li Shen, and Gang Sun, · 2018
Cited alongside, same era.
“Contrastive representation distillation,”
Yonglong Tian, Dilip Krishnan, and Phillip Isola, · 2019
Cited alongside, same era.
“Panns: Large-scale pretrained audio neural networks for audio pattern recognition,”
Qiuqiang Kong, Yin Cao, Turab Iqbal, Yuxuan Wang, Wenwu Wang, and Mark D. Plumbley, · 2019
Cited alongside, same era.
“Relational knowledge distillation,”
Wonpyo Park, Dongju Kim, Yan Lu, and Minsu Cho, · 2019
Cited alongside, same era.
“Learning from the master: Distilling cross-modal advanced knowledge for lip reading,”
Sucheng Ren, Yong Du, Jianming Lv, Guoqiang Han, and Shengfeng He, · 2021
Later among the works it cites.
“Distilling audio-visual knowledge by compositional contrastive learning,”
Yanbei Chen, Yongqin Xian, A Koepke, Ying Shan, and Zeynep Akata, · 2021
Later among the works it cites.
“Class-aware sounding objects localization via audiovisual correspondence,”
Di Hu, Yake Wei, Rui Qian, Weiyao Lin, Ruihua Song, and Ji-Rong Wen, · 2021
Later among the works it cites.
“Learning cross-modal retrieval with noisy labels,”
Peng Hu, Xi Peng, Hongyuan Zhu, Liangli Zhen, and Jie Lin, · 2021
Later among the works it cites.
“Learning in audio-visual context: A review, analysis, and new perspective,”
Yake Wei, Di Hu, Yapeng Tian, and Xuelong Li, · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yonglong Tian, Dilip Krishnan, and Phillip Isola, · 2020
Cited alongside, same era.
“Revisiting knowledge distillation via label smoothing regularization,”
Li Yuan, Francis EH Tay, Guilin Li, Tao Wang, and Jiashi Feng, · 2020
Cited alongside, same era.
“Robust audio-visual instance discrimination,”
Pedro Morgado, Ishan Misra, and Nuno Vasconcelos, · 2021
Cited alongside, same era.
“Learning to answer questions in dynamic audio-visual scenarios,”
Guangyao Li, Yake Wei, Yapeng Tian, Chenliang Xu, Ji-Rong Wen, and Di Hu, · 2022
Later among the works it cites.
“Balanced multimodal learning via on-the-fly gradient modulation,”
Xiaokang Peng, Yake Wei, Andong Deng, Dong Wang, and Di Hu, · 2022
Later among the works it cites.
“Mmcosine: Multi-modal cosine loss towards balanced audio-visual fine-grained learning,”
Ruize Xu, Ruoxuan Feng, Shi-Xiong Zhang, and Di Hu, · 2023
Closest in time.