Fetching the paper…
Reading the bibliography…
The objective of this paper is to learn representations of speaker identity without access to manually annotated data.
“Darpa timit acoustic-phonetic continous speech corpus cd-rom. nist speech disc 1-1.1,”
John S Garofolo, Lori F Lamel, William M Fisher, Jonathan G Fiscus, and David S Pallett, · 1993
Earlier work this paper cites.
“Rapid and accurate spoken term detection,”
David RH Miller et al., · 2007
Earlier work this paper cites.
“Variational recurrent auto-encoders,”
Otto Fabius and Joost R van Amersfoort, · 2014
Earlier work this paper cites.
“Return of the devil in the details: Delving deep into convolutional nets,”
Ken Chatfield, Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman, · 2014
Earlier work this paper cites.
“Deep convolutional inverse graphics network,”
Tejas D Kulkarni, William F Whitney, Pushmeet Kohli, and Josh Tenenbaum, · 2015
Earlier work this paper cites.
“A recurrent latent variable model for sequential data,”
Junyoung Chung, Kyle Kastner, Laurent Dinh, Kratarth Goel, Aaron C Courville, and Yoshua Bengio, · 2015
Earlier work this paper cites.
“Simultaneous deep transfer across domains and tasks,”
Eric Tzeng, Judy Hoffman, Trevor Darrell, and Kate Saenko, · 2015
Earlier work this paper cites.
“Out of time: automated lip sync in the wild,”
Joon Son Chung and Andrew Zisserman, · 2016
Earlier work this paper cites.
“Audio word2vec: Unsupervised learning of audio segment representations using sequence-to-sequence autoencoder,”
Yu-An Chung, Chao-Chung Wu, Chia-Hao Shen, Hung-Yi Lee, and Lin-Shan Lee, · 2016
Earlier work this paper cites.
“Infogan: Interpretable representation learning by information maximizing generative adversarial nets,”
Xi Chen, Yan Duan, Rein Houthooft, John Schulman, Ilya Sutskever, and Pieter Abbeel, · 2016
Earlier work this paper cites.
“Hierarchical multiscale recurrent neural networks,”
Junyoung Chung, Sungjin Ahn, and Yoshua Bengio, · 2016
Cited alongside, same era.
“Sequential neural models with stochastic layers,”
Marco Fraccaro, Søren Kaae Sønderby, Ulrich Paquet, and Ole Winther, · 2016
Cited alongside, same era.
“Soundnet: Learning sound representations from unlabeled video,”
Yusuf Aytar, Carl Vondrick, and Antonio Torralba, · 2016
Cited alongside, same era.
“Ambient sound provides supervision for visual learning,”
Andrew Owens, Jiajun Wu, Josh H McDermott, William T Freeman, and Antonio Torralba, · 2016
Cited alongside, same era.
“Adversarial multi-task learning of deep neural networks for robust speech recognition.,”
Yusuke Shinohara, · 2016
Cited alongside, same era.
“VoxCeleb: a large-scale speaker identification dataset,”
“Text-independent speaker verification using 3d convolutional neural networks,”
Amirsina Torfi, Jeremy Dawson, and Nasser M Nasrabadi, · 2018
Later among the works it cites.
“VoxCeleb2: Deep speaker recognition,”
Joon Son Chung, Arsha Nagrani, and Andrew Zisserman, · 2018
Later among the works it cites.
“Emotion recognition in speech using cross-modal transfer in the wild,”
Samuel Albanie, Arsha Nagrani, Andrea Vedaldi, and Andrew Zisserman, · 2018
Later among the works it cites.
“Cooperative learning of audio and video models from self-supervised synchronization,”
Bruno Korbar, Du Tran, and Lorenzo Torresani, · 2018
Later among the works it cites.
“Seeing voices and hearing faces: Cross-modal biometric matching,”
Arsha Nagrani, Samuel Albanie, and Andrew Zisserman, · 2018
Later among the works it cites.
“Learnable PINs: Cross-modal embeddings for person identity,”
Arsha Nagrani, Samuel Albanie, and Andrew Zisserman, · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Arsha Nagrani, Joon Son Chung, and Andrew Zisserman, · 2017
Cited alongside, same era.
“Neural discrete representation learning,”
Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu, · 2017
Cited alongside, same era.
“Look, listen and learn,”
Relja Arandjelovic and Andrew Zisserman, · 2017
Cited alongside, same era.
“See, hear, and read: Deep aligned representations,”
Yusuf Aytar, Carl Vondrick, and Antonio Torralba, · 2017
Cited alongside, same era.
“State-of-the-art speech recognition with sequence-to-sequence models,”
Chung-Cheng Chiu et al., · 2018
Cited alongside, same era.
Later among the works it cites.
“On learning associations of faces and voices,”
Changil Kim, Hijung Valentina Shin, Tae-Hyun Oh, Alexandre Kaspar, Mohamed Elgharib, and Wojciech Matusik, · 2018
Later among the works it cites.
“Turning a blind eye: Explicit removal of biases and variation from deep neural network embeddings,”
Mohsan Alvi, Andrew Zisserman, and Christoffer Nellåker, · 2018
Later among the works it cites.
“Utterance-level aggregation for speaker recognition in the wild,”
Weidi Xie, Arsha Nagrani, Joon Son Chung, and Andrew Zisserman, · 2019
Later among the works it cites.
“Perfect match: Improved cross-modal embeddings for audio-visual synchronisation,”
Soo-Whan Chung, Joon Son Chung, and Hong-Goo Kang, · 2019
Later among the works it cites.