Fetching the paper…
Reading the bibliography…
Active speaker detection (ASD) in videos with multiple speakers is a challenging task as it requires learning effective audiovisual features and spatial-temporal correlations over long temporal windows.
Cutler, R., Davis, L.: Look who’s talking: Speaker detection using video and audio correlation. In: 2000 IEEE International Conference on Multimedia and Expo. ICME2000. Proceedings. Latest Advances in the Fast Changing World of Multimedia (Cat. No. 00TH8532). vol. 3, pp. 1589–1592. IEEE (2000)
2000
Earlier work this paper cites.
Everingham, M., Sivic, J., Zisserman, A.: Hello! my name is… buffy”–automatic naming of characters in tv video. In: BMVC. vol. 2, p. 6 (2006)
2006
Earlier work this paper cites.
Everingham, M., Sivic, J., Zisserman, A.: Taking the bite out of automated naming of characters in tv video. Image and Vision Computing 27
2009
Earlier work this paper cites.
Nair, V., Hinton, G.E.: Rectified linear units improve restricted boltzmann machines. In: Proceedings of the 27th international conference on machine learning (ICML-10). pp. 807–814 (2010)
2010
Earlier work this paper cites.
Chatfield, K., Simonyan, K., Vedaldi, A., Zisserman, A.: Return of the devil in the details: Delving deep into convolutional nets. In: Proceedings of the British Machine Vision Conference. BMVA Press (2014)
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
Chakravarty, P., Tuytelaars, T.: Cross-modal supervision for learning active speaker detection in video. In: European Conference on Computer Vision. pp. 285–301. Springer (2016)
2016
Earlier work this paper cites.
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 770–778 (2016)
2016
Earlier work this paper cites.
Stefanov, K., Sugimoto, A., Beskow, J.: Look who’s talking: visual identification of the active speaker in multi-party human-robot interaction. In: Proceedings of the 2nd Workshop on Advancements in Social Signal Processing for Multimodal Interaction. pp. 22–27 (2016)
2016
Earlier work this paper cites.
Hamilton, W., Ying, Z., Leskovec, J.: Inductive representation learning on large graphs. Advances in neural information processing systems 30
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Loshchilov, I., Hutter, F.: Sgdr: Stochastic gradient descent with warm restarts. In: International Conference on Learning Representations (ICLR) (2017)
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Stefanov, K., Beskow, J., Salvi, G.: Vision-based active speaker detection in multiparty interaction. In: Grounding Language Understanding GLU2017 August 25, 2017, KTH Royal Institute of Technology, Stockholm, Sweden (2017)
2017
Cited alongside, same era.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I.: Attention is all you need. Advances in neural information processing systems 30
2017
Cited alongside, same era.
2018
Cited alongside, same era.
Ravanelli, M., Bengio, Y.: Speaker recognition from raw waveform with sincnet. In: 2018 IEEE Spoken Language Technology Workshop (SLT). pp. 1021–1028. IEEE (2018)
2018
Cited alongside, same era.
Nagarajan, T., Li, Y., Feichtenhofer, C., Grauman, K.: Ego-topo: Environment affordances from egocentric video. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 163–172 (2020)
2020
Later among the works it cites.
Roth, J., Chaudhuri, S., Klejch, O., Marvin, R., Gallagher, A., Kaver, L., Ramaswamy, S., Stopczynski, A., Schmid, C., Xi, Z., et al.: Ava active speaker: An audio-visual dataset for active speaker detection. In: ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 4492–4496. IEEE (2020)
2020
Later among the works it cites.
2020
Later among the works it cites.
Shirian, A., Tripathi, S., Guha, T.: Learnable graph inception network for emotion recognition. IEEE Transactions on Multimedia (2020)
2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Kopuklu, O., Kose, N., Gunduz, A., Rigoll, G.: Resource efficient 3d convolutional neural networks. In: Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops. pp. 0–0 (2019)
2019
Cited alongside, same era.
Lin, J., Gan, C., Han, S.: Tsm: Temporal shift module for efficient video understanding. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 7083–7093 (2019)
2019
Cited alongside, same era.
Wang, Y., Sun, Y., Liu, Z., Sarma, S.E., Bronstein, M.M., Solomon, J.M.: Dynamic graph cnn for learning on point clouds. Acm Transactions On Graphics (tog) 38
2019
Cited alongside, same era.
Zhang, Y.H., Xiao, J., Yang, S., Shan, S.: Multi-task learning for audio-visual active speaker detection. The ActivityNet Large-Scale Activity Recognition Challenge pp. 1–4 (2019)
2019
Cited alongside, same era.
Zhang, Y., Tokmakov, P., Hebert, M., Schmid, C.: A structured model for action detection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 9975–9984 (2019)
2019
Cited alongside, same era.
Alcazar, J.L., Heilbron, F.C., Mai, L., Perazzi, F., Lee, J.Y., Arbelaez, P., Ghanem, B.: Active Speakers in Context. In: Proc. IEEE Conf. Comput. Vis. Pattern Recognit. p. 12465–12474 (2020)
2020
Cited alongside, same era.
Later among the works it cites.
2021
Later among the works it cites.
Köpüklü, O., Taseska, M., Rigoll, G.: How to Design a Three-Stage Architecture for Audio-Visual Active Speaker Detection in the Wild. In: Proc. Internal Conference on Computer Vision (Jun 2021)
2021
Later among the works it cites.
León-Alcázar, J., Heilbron, F.C., Thabet, A., Ghanem, B.: MAAS: Multi-modal Assignation for Active Speaker Detection. In: Internal Conference on Computer Vision (2021)
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
Tan, R., Xu, H., Saenko, K., Plummer, B.A.: Logan: Latent graph co-attention network for weakly-supervised video moment retrieval. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. pp. 2083–2092 (2021)
2021
Later among the works it cites.
Tao, R., Pan, Z., Das, R.K., Qian, X., Shou, M.Z., Li, H.: Is someone speaking? exploring long-term temporal features for audio-visual active speaker detection. In: Proceedings of the 29th ACM International Conference on Multimedia. pp. 3927–3935 (2021)
2021
Later among the works it cites.
Zhang, Y., Liang, S., Yang, S., Liu, X., Wu, Z., Shan, S., Chen, X.: Unicon: Unified context network for robust active speaker detection. In: Proceedings of the 29th ACM International Conference on Multimedia. pp. 3964–3972 (2021)
2021
Later among the works it cites.
Min, K., Roy, S., Tripathi, S., Guha, T., Majumdar, S.: Intel labs at activitynet challenge 2022: Spell for long-term active speaker detection. The ActivityNet Large-Scale Activity Recognition Challenge (2022), https://research.google.com/ava/2022/S2_SPELL_ActivityNet_Challenge_2022.pdf
2022
Closest in time.