Fetching the paper…
Reading the bibliography…
This report describes our approach for the Audio-Visual Diarization (AVD) task of the Ego4D Challenge 2022.
Segmentation of tv shows into scenes using speaker diarization and speech recognition
Hervé Bredin · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Look who’s talking: visual identification of the active speaker in multi-party human-robot interaction
Kalin Stefanov, Akihiro Sugimoto, and Jonas Beskow · 2016
Earlier work this paper cites.
Deep lip reading: a comparison of models and an online application
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman · 2018
Earlier work this paper cites.
Multimodal speaker diarization for meetings using volume-evaluated srp-phat and video analysis
P Cabañas-Molero, M Lucena, José Manuel Fuertes, Pedro Vera-Candeas, and Nicolás Ruiz-Reyes · 2018
Earlier work this paper cites.
Audio-Visual Analysis In the Framework of Humans Interacting with Robots
Israel Dejene Gebru · 2018
Cited alongside, same era.
Who said that?: Audio-visual speaker diarisation of real-world meetings
Joon Son Chung, Bong-Jin Lee, and Icksang Han · 2019
Cited alongside, same era.
Advances in online audio-visual meeting transcription
Takuya Yoshioka, Igor Abramovski, Cem Aksoylar, Zhuo Chen, Moshe David, Dimitrios Dimitriadis, Yifan Gong, Ilya Gurvich, Xuedong Huang, Yan Huang, et al · 2019
Cited alongside, same era.
Multimodal speaker diarization of real-world meetings using d-vectors with spatial features
Wonjune Kang, Brandon C Roy, and Wesley Chow · 2020
Cited alongside, same era.
Ava active speaker: An audio-visual dataset for active speaker detection
Joseph Roth, Sourish Chaudhuri, Ondrej Klejch, Radhika Marvin, Andrew Gallagher, Liat Kaver, Sharadh Ramaswamy, Arkadiusz Stopczynski, Cordelia Schmid, Zhonghua Xi, et al · 2020
Cited alongside, same era.
Silero vad: pre-trained enterprise-grade voice activity detector (vad), number detector and language classifier
Silero Team · 2021
Later among the works it cites.
Ava-avd: Audio-visual speaker diarization in the wild
Eric Zhongcong Xu, Zeyang Song, Chao Feng, Mang Ye, and Mike Zheng Shou · 2021
Later among the works it cites.
Ego4d: Around the world in 3,000 hours of egocentric video
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al · 2022
Closest in time.
Intel labs at activitynet challenge 2022: Spell for long-term active speaker detection
Kyle Min, Sourya Roy, Subarna Tripathi, Tanaya Guha, and Somdeb Majumdar · 2022
Closest in time.
Learning long-term spatial-temporal graphs for active speaker detection
Kyle Min, Sourya Roy, Subarna Tripathi, Tanaya Guha, and Somdeb Majumdar · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Is someone speaking? exploring long-term temporal features for audio-visual active speaker detection
Ruijie Tao, Zexu Pan, Rohan Kumar Das, Xinyuan Qian, Mike Zheng Shou, and Haizhou Li · 2021
Cited alongside, same era.
Closest in time.