Fetching the paper…
Reading the bibliography…
This report introduces our novel method named STHG for the Audio-Visual Diarization task of the Ego4D Challenge 2023.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Multimodal speaker diarization for meetings using volume-evaluated srp-phat and video analysis
P Cabañas-Molero, M Lucena, José Manuel Fuertes, Pedro Vera-Candeas, and Nicolás Ruiz-Reyes · 2018
Earlier work this paper cites.
ESPnet: End-to-end speech processing toolkit
Shinji Watanabe, Takaaki Hori, Shigeki Karita, Tomoki Hayashi, Jiro Nishitoba, Yuya Unno, Nelson Yalta, Jahn Heymann, Matthew Wiesner, Nanxin Chen, Adithya Renduchintala, and Tsubasa Ochiai · 2018
Earlier work this paper cites.
Who said that?: Audio-visual speaker diarisation of real-world meetings
Joon Son Chung, Bong-Jin Lee, and Icksang Han · 2019
Earlier work this paper cites.
Multimodal speaker diarization of real-world meetings using d-vectors with spatial features
Wonjune Kang, Brandon C Roy, and Wesley Chow · 2020
Earlier work this paper cites.
Chime-6 challenge: Tackling multispeaker speech recognition for unsegmented recordings
Shinji Watanabe, Michael Mandel, Jon Barker, Emmanuel Vincent, Ashish Arora, Xuankai Chang, Sanjeev Khudanpur, Vimal Manohar, Daniel Povey, Desh Raj, et al · 2020
Cited alongside, same era.
Is someone speaking? exploring long-term temporal features for audio-visual active speaker detection
Ruijie Tao, Zexu Pan, Rohan Kumar Das, Xinyuan Qian, Mike Zheng Shou, and Haizhou Li · 2021
Cited alongside, same era.
Silero vad: pre-trained enterprise-grade voice activity detector (vad), number detector and language classifier
Silero Team · 2021
Cited alongside, same era.
Ava-avd: Audio-visual speaker diarization in the wild
Eric Zhongcong Xu, Zeyang Song, Chao Feng, Mang Ye, and Mike Zheng Shou · 2021
Cited alongside, same era.
Ego4d: Around the world in 3,000 hours of egocentric video
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al · 2022
Later among the works it cites.
Intel labs at ego4d challenge 2022: A better baseline for audio-visual diarization
Kyle Min · 2022
Later among the works it cites.
Learning long-term spatial-temporal graphs for active speaker detection
Kyle Min, Sourya Roy, Subarna Tripathi, Tanaya Guha, and Somdeb Majumdar · 2022
Later among the works it cites.
Using active speaker faces for diarization in tv shows
Rahul Sharma and Shrikanth Narayanan · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…