Fetching the paper…
Reading the bibliography…
Successful active speaker detection requires a three-stage pipeline: (i) audio-visual encoding for all speakers in the clip, (ii) inter-speaker relation modeling between a reference speaker and the background speakers within each frame, and (iii) temporal modeling for the reference speaker.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Look who’s talking: Speaker detection using video and audio correlation
Ross Cutler and Larry Davis · 2000
Earlier work this paper cites.
Audio-visual segmentation and “the cocktail party effect”
Trevor Darrell, John W Fisher, and Paul Viola · 2000
Earlier work this paper cites.
Visual categorization with bags of keypoints
Gabriella Csurka, Christopher Dance, Lixin Fan, Jutta Willamowski, and Cédric Bray · 2004
Earlier work this paper cites.
On space-time interest points
Ivan Laptev · 2005
Earlier work this paper cites.
Learning realistic human actions from movies
Ivan Laptev, Marcin Marszalek, Cordelia Schmid, and Benjamin Rozenfeld · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Improving the fisher kernel for large-scale image classification
Florent Perronnin, Jorge Sánchez, and Thomas Mensink · 2010
Earlier work this paper cites.
3d convolutional neural networks for human action recognition
Shuiwang Ji, Wei Xu, Ming Yang, and Kai Yu · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Action recognition with improved trajectories
Heng Wang and Cordelia Schmid · 2013
Earlier work this paper cites.
On the properties of neural machine translation: Encoder-decoder approaches
Kyunghyun Cho, Bart Van Merriënboer, Dzmitry Bahdanau, and Yoshua Bengio · 2014
Earlier work this paper cites.
End-to-end learning for music audio
Sander Dieleman and Benjamin Schrauwen · 2014
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Who’s speaking? audio-supervised classification of active speakers in video
Punarjay Chakravarty, Sayeh Mirzaei, Tinne Tuytelaars, and Hugo Van hamme · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Cited alongside, same era.
Empirical evaluation of rectified activations in convolutional network
Bing Xu, Naiyan Wang, Tianqi Chen, and Mu Li · 2015
Cited alongside, same era.
Deep clustering: Discriminative embeddings for segmentation and separation
John R Hershey, Zhuo Chen, Jonathan Le Roux, and Shinji Watanabe · 2016
Cited alongside, same era.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Cited alongside, same era.
Audio-visual speaker diarization based on spatiotemporal bayesian fusion
Israel D Gebru, Sileye Ba, Xiaofei Li, and Radu Horaud · 2017
Cited alongside, same era.
An end-to-end multimodal voice activity detection using wavenet encoder and residual networks
I. Ariav and I. Cohen · 2019
Later among the works it cites.
Naver at activitynet challenge 2019–task b active speaker detection (ava)
Joon Son Chung · 2019
Later among the works it cites.
Slowfast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2019
Later among the works it cites.
Comparative analysis of cnn-based spatiotemporal reasoning in videos
Okan Köpüklü, Fabian Herzog, and Gerhard Rigoll · 2019
Later among the works it cites.
Resource efficient 3d convolutional neural networks
Okan Kopuklu, Neslihan Kose, Ahmet Gunduz, and Gerhard Rigoll · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jongpil Lee, Taejun Kim, Jiyoung Park, and Juhan Nam · 2017
Cited alongside, same era.
Multimodal gesture recognition based on the resc3d network
Qiguang Miao, Yunan Li, Wanli Ouyang, Zhenxin Ma, Xin Xu, Weikang Shi, and Xiaochun Cao · 2017
Cited alongside, same era.
Learning spatio-temporal representation with pseudo-3d residual networks
Zhaofan Qiu, Ting Yao, and Tao Mei · 2017
Cited alongside, same era.
Voxceleb2: Deep speaker recognition
Joon Son Chung, Arsha Nagrani, and Andrew Zisserman · 2018
Cited alongside, same era.
Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation
Ariel Ephrat, Inbar Mosseri, Oran Lang, Tali Dekel, Kevin Wilson, Avinatan Hassidim, William T. Freeman, and Michael Rubinstein · 2018
Cited alongside, same era.
Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet?
Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh · 2018
Cited alongside, same era.
Motion fused frames: Data level fusion strategy for hand gesture recognition
Okan Kopuklu, Neslihan Kose, and Gerhard Rigoll · 2018
Cited alongside, same era.
The jester dataset: A large-scale video dataset of human gestures
Joanna Materzynska, Guillaume Berger, Ingo Bax, and Roland Memisevic · 2019
Later among the works it cites.
Moments in time dataset: one million videos for event understanding
Mathew Monfort, Alex Andonian, Bolei Zhou, Kandan Ramakrishnan, Sarah Adel Bargal, Tom Yan, Lisa Brown, Quanfu Fan, Dan Gutfreund, Carl Vondrick, et al · 2019
Later among the works it cites.
Voicefilter: Targeted voice separation by speaker-conditioned spectrogram masking
Quan Wang, Hannah Muckenhirn, Kevin Wilson, Prashant Sridhar, Zelin Wu, John R Hershey, Rif A Saurous, Ron J Weiss, Ye Jia, and Ignacio Lopez Moreno · 2019
Later among the works it cites.
Active speakers in context
Juan León Alcázar, Fabian Caba, Long Mai, Federico Perazzi, Joon-Young Lee, Pablo Arbeláez, and Bernard Ghanem · 2020
Later among the works it cites.
Listen to look: Action recognition by previewing audio
Ruohan Gao, Tae-Hyun Oh, Kristen Grauman, and Lorenzo Torresani · 2020
Later among the works it cites.
Lightweight end-to-end speech recognition from raw audio data using sinc-convolutions
Ludwig Kürzinger, Nicolas Lindae, Palle Klewitz, and Gerhard Rigoll · 2020
Later among the works it cites.
Multichannel speech enhancement by raw waveform-mapping using fully convolutional networks
Chang-Le Liu, Sze-Wei Fu, You-Jin Li, Jen-Wei Huang, Hsin-Min Wang, and Yu Tsao · 2020
Later among the works it cites.
Small-footprint keyword spotting on raw audio data with sinc-convolutions
Simon Mittermaier, Ludwig Kürzinger, Bernd Waschneck, and Gerhard Rigoll · 2020
Later among the works it cites.
Speech2action: Cross-modal supervision for action recognition
Arsha Nagrani, Chen Sun, David Ross, Rahul Sukthankar, Cordelia Schmid, and Andrew Zisserman · 2020
Later among the works it cites.
Ava active speaker: An audio-visual dataset for active speaker detection
Joseph Roth, Sourish Chaudhuri, Ondrej Klejch, Radhika Marvin, Andrew Gallagher, Liat Kaver, Sharadh Ramaswamy, Arkadiusz Stopczynski, Cordelia Schmid, Zhonghua Xi, et al · 2020
Later among the works it cites.
Crossmodal learning for audio-visual speech event localization
Rahul Sharma, Krishna Somandepalli, and Shrikanth Narayanan · 2020
Later among the works it cites.
Maas: Multi-modal assignation for active speaker detection
Juan León-Alcázar, Fabian Caba Heilbron, Ali Thabet, and Bernard Ghanem · 2021
Closest in time.