Fetching the paper…
Reading the bibliography…
Active speaker detection requires a solid integration of multi-modal cues.
Phoneme recognition using time-delay neural networks
Alex Waibel, Toshiyuki Hanazawa, Geoffrey Hinton, Kiyohiro Shikano, and Kevin J Lang · 1989
Earlier work this paper cites.
Look who’s talking: Speaker detection using video and audio correlation
Ross Cutler and Larry Davis · 2000
Earlier work this paper cites.
Visual speech recognition with loosely synchronized feature streams
Kate Saenko, Karen Livescu, Michael Siracusa, Kevin Wilson, James Glass, and Trevor Darrell · 2005
Earlier work this paper cites.
Dimensionality reduction by learning an invariant mapping
Raia Hadsell, Sumit Chopra, and Yann LeCun · 2006
Earlier work this paper cites.
An overview of automatic speaker diarization systems
Sue E Tranter and Douglas A Reynolds · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Taking the bite out of automated naming of characters in tv video
Mark Everingham, Josef Sivic, and Andrew Zisserman · 2009
Earlier work this paper cites.
Multimodal deep learning
Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and Andrew Y Ng · 2011
Earlier work this paper cites.
Speaker diarization: A review of recent research
Xavier Anguera, Simon Bozonnet, Nicholas Evans, Corinne Fredouille, Gerald Friedland, and Oriol Vinyals · 2012
Earlier work this paper cites.
Unsupervised methods for speaker diarization: An integrated and iterative approach
Stephen H Shum, Najim Dehak, Réda Dehak, and James R Glass · 2013
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei · 2014
Earlier work this paper cites.
Who’s speaking? audio-supervised classification of active speakers in video
Punarjay Chakravarty, Sayeh Mirzaei, Tinne Tuytelaars, and Hugo Van hamme · 2015
Earlier work this paper cites.
A method for stochastic optimization
D Kinga and J Ba Adam · 2015
Earlier work this paper cites.
Active speaker detection with audio-visual co-training
Punarjay Chakravarty, Jeroen Zegers, Tinne Tuytelaars, et al · 2016
Earlier work this paper cites.
Out of time: automated lip sync in the wild
Joon Son Chung and Andrew Zisserman · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Identity mappings in deep residual networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Structural-rnn: Deep learning on spatio-temporal graphs
Ashesh Jain, Amir R Zamir, Silvio Savarese, and Ashutosh Saxena · 2016
Earlier work this paper cites.
Semi-supervised classification with graph convolutional networks
Thomas N Kipf and Max Welling · 2016
Earlier work this paper cites.
Lip reading sentences in the wild
Joon Son Chung, Andrew Senior, Oriol Vinyals, and Andrew Zisserman · 2017
Earlier work this paper cites.
Speaker diarization using deep neural network embeddings
Daniel Garcia-Romero, David Snyder, Gregory Sell, Daniel Povey, and Alan McCree · 2017
Earlier work this paper cites.
Voxceleb: a large-scale speaker identification dataset
Arsha Nagrani, Joon Son Chung, and Andrew Zisserman · 2017
Cited alongside, same era.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Cited alongside, same era.
3d graph neural networks for rgbd semantic segmentation
Xiaojuan Qi, Renjie Liao, Jiaya Jia, Sanja Fidler, and Raquel Urtasun · 2017
Cited alongside, same era.
Bimodal recurrent neural network for audiovisual voice activity detection
Fei Tao and Carlos Busso · 2017
Cited alongside, same era.
Simple online and realtime tracking with a deep association metric
Nicolai Wojke, Alex Bewley, and Dietrich Paulus · 2017
Cited alongside, same era.
Scene graph generation by iterative message passing
Graph r-cnn for scene graph generation
Jianwei Yang, Jiasen Lu, Stefan Lee, Dhruv Batra, and Devi Parikh · 2018
Later among the works it cites.
Naver at activitynet challenge 2019–task b active speaker detection (ava)
Joon Son Chung · 2019
Later among the works it cites.
Perfect match: Improved cross-modal embeddings for audio-visual synchronisation
Soo-Whan Chung, Joon Son Chung, and Hong-Goo Kang · 2019
Later among the works it cites.
Retinaface: Single-stage dense face localisation in the wild
Jiankang Deng, Jia Guo, Yuxiang Zhou, Jinke Yu, Irene Kotsia, and Stefanos Zafeiriou · 2019
Later among the works it cites.
Fast graph representation learning with PyTorch Geometric
Matthias Fey and Jan E. Lenssen · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Danfei Xu, Yuke Zhu, Christopher B Choy, and Li Fei-Fei · 2017
Cited alongside, same era.
Active learning based constrained clustering for speaker diarization
Chengzhu Yu and John HL Hansen · 2017
Cited alongside, same era.
Voxceleb2: Deep speaker recognition
Joon Son Chung, Arsha Nagrani, and Andrew Zisserman · 2018
Cited alongside, same era.
Image generation from scene graphs
Justin Johnson, Agrim Gupta, and Li Fei-Fei · 2018
Cited alongside, same era.
On learning associations of faces and voices
Changil Kim, Hijung Valentina Shin, Tae-Hyun Oh, Alexandre Kaspar, Mohamed Elgharib, and Wojciech Matusik · 2018
Cited alongside, same era.
Factorizable net: an efficient subgraph-based framework for scene graph generation
Yikang Li, Wanli Ouyang, Bolei Zhou, Jianping Shi, Chao Zhang, and Xiaogang Wang · 2018
Cited alongside, same era.
Learnable pins: Cross-modal embeddings for person identity
Arsha Nagrani, Samuel Albanie, and Andrew Zisserman · 2018
Cited alongside, same era.
Georgia Gkioxari, Jitendra Malik, and Justin Johnson · 2019
Later among the works it cites.
Neural predictive coding using convolutional neural networks toward unsupervised learning of speaker characteristics
Arindam Jati and Panayiotis Georgiou · 2019
Later among the works it cites.
Deepgcns: Can gcns go as deep as cnns?
Guohao Li, Matthias Muller, Ali Thabet, and Bernard Ghanem · 2019
Later among the works it cites.
Deepgcns: Making gcns go as deep as cnns, 2019
Guohao Li, Matthias Müller, Guocheng Qian, Itzel C. Delgadillo, Abdulellah Abualshour, Ali Thabet, and Bernard Ghanem · 2019
Later among the works it cites.
Sgas: Sequential greedy architecture search, 2019
Guohao Li, Guocheng Qian, Itzel C. Delgadillo, Matthias Müller, Ali Thabet, and Bernard Ghanem · 2019
Later among the works it cites.
Ava-activespeaker: An audio-visual dataset for active speaker detection
Joseph Roth, Sourish Chaudhuri, Ondrej Klejch, Radhika Marvin, Andrew Gallagher, Liat Kaver, Sharadh Ramaswamy, Arkadiusz Stopczynski, Cordelia Schmid, Zhonghua Xi, et al · 2019
Later among the works it cites.
Dynamic graph cnn for learning on point clouds
Yue Wang, Yongbin Sun, Ziwei Liu, Sanjay E Sarma, Michael M Bronstein, and Justin M Solomon · 2019
Later among the works it cites.
Point clouds learning with attention-based graph convolution networks
Zhuyang Xie, Junzhou Chen, and Bo Peng · 2019
Later among the works it cites.
Fully supervised speaker diarization
Aonan Zhang, Quan Wang, Zhenyao Zhu, John Paisley, and Chong Wang · 2019
Later among the works it cites.
Self-supervised learning of audio-visual objects from video
Triantafyllos Afouras, Andrew Owens, Joon Son Chung, and Andrew Zisserman · 2020
Later among the works it cites.
Active speakers in context
Juan Leon Alcazar, Fabian Caba, Long Mai, Federico Perazzi, Joon-Young Lee, Pablo Arbelaez, and Bernard Ghanem · 2020
Later among the works it cites.
Spot the conversation: speaker diarisation in the wild
Joon Son Chung, Jaesung Huh, Arsha Nagrani, Triantafyllos Afouras, and Andrew Zisserman · 2020
Later among the works it cites.
Deepergcn: All you need to train deeper gcns
Guohao Li, Chenxin Xiong, Ali Thabet, and Bernard Ghanem · 2020
Later among the works it cites.
Crossmodal learning for audio-visual speech event localization
Rahul Sharma, Krishna Somandepalli, and Shrikanth Narayanan · 2020
Later among the works it cites.
G-tad: Sub-graph localization for temporal action detection
Mengmeng Xu, Chen Zhao, David S Rojas, Ali Thabet, and Bernard Ghanem · 2020
Later among the works it cites.