Fetching the paper…
Reading the bibliography…
The objective of this paper is to separate a target speaker's speech from a mixture of two speakers using a deep audio-visual speech separation network.
“Signal estimation from modified short-time fourier transform,”
Daniel Griffin and Jae Lim, · 1984
Earlier work this paper cites.
“Front-end factor analysis for speaker verification,”
Najim Dehak, Patrick J Kenny, Réda Dehak, Pierre Dumouchel, and Pierre Ouellet, · 2010
Earlier work this paper cites.
“Selective cortical representation of attended speaker in multi-talker speech perception,”
Nima Mesgarani and Edward F Chang, · 2012
Earlier work this paper cites.
“Return of the devil in the details: Delving deep into convolutional nets,”
K. Chatfield, K. Simonyan, A. Vedaldi, and A. Zisserman, · 2014
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Facenet: A unified embedding for face recognition and clustering,”
Florian Schroff, Dmitry Kalenichenko, and James Philbin, · 2015
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals, · 2016
Earlier work this paper cites.
“Deep speech 2: End-to-end speech recognition in English and Mandarin,”
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Qiang Cheng, Guoliang Chen, et al., · 2016
Earlier work this paper cites.
“Deep clustering: Discriminative embeddings for segmentation and separation,”
John R Hershey, Zhuo Chen, Jonathan Le Roux, and Shinji Watanabe, · 2016
Earlier work this paper cites.
“Permutation invariant training of deep models for speaker-independent multi-talker speech separation,”
Dong Yu, Morten Kolbæk, Zheng-Hua Tan, and Jesper Jensen, · 2017
Earlier work this paper cites.
“Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks,”
Morten Kolbæk, Dong Yu, Zheng-Hua Tan, and Jesper Jensen, · 2017
Earlier work this paper cites.
“A structured self-attentive sentence embedding,”
Zhouhan Lin, Minwei Feng, Cicero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou, and Yoshua Bengio, · 2017
Earlier work this paper cites.
“State-of-the-art speech recognition with sequence-to-sequence models,”
Chung-Cheng Chiu, Tara N Sainath, Yonghui Wu, Rohit Prabhavalkar, Patrick Nguyen, Zhifeng Chen, Anjuli Kannan, Ron J Weiss, Kanishka Rao, Ekaterina Gonina, et al., · 2018
Cited alongside, same era.
“Tasnet: time-domain audio separation network for real-time, single-channel speech separation,”
Yi Luo and Nima Mesgarani, · 2018
Cited alongside, same era.
“The conversation: Deep audio-visual speech enhancement,”
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman, · 2018
Cited alongside, same era.
“Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation,”
Ariel Ephrat, Inbar Mosseri, Oran Lang, Tali Dekel, Kevin Wilson, Avinatan Hassidim, William T. Freeman, and Michael Rubinstein, · 2018
Cited alongside, same era.
“Audio-visual scene analysis with self-supervised multisensory features,”
Andrew Owens and Alexei A Efros, · 2018
Cited alongside, same era.
“VoxCeleb2: Deep speaker recognition,”
Joon Son Chung, Arsha Nagrani, and Andrew Zisserman, · 2018
Later among the works it cites.
“VoiceFilter: Targeted Voice Separation by Speaker-Conditioned Spectrogram Masking,”
Quan Wang, Hannah Muckenhirn, Kevin Wilson, Prashant Sridhar, Zelin Wu, John R. Hershey, Rif A. Saurous, Ron J. Weiss, Ye Jia, and Ignacio Lopez Moreno, · 2019
Later among the works it cites.
“Time-domain speaker extraction network,”
Chenglin Xu, Wei Rao, Eng Siong Chng, and Haizhou Li, · 2019
Later among the works it cites.
“Audio–visual deep clustering for speech separation,”
Rui Lu, Zhiyao Duan, and Changshui Zhang, · 2019
Later among the works it cites.
“Multimodal speakerbeam: Single channel target speech extraction with audio-visual speaker clues,”
Tsubasa Ochiai, Marc Delcroix, Keisuke Kinoshita, Atsunori Ogawa, and Tomohiro Nakatani, · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Seeing voices and hearing faces: Cross-modal biometric matching,”
Arsha Nagrani, Samuel Albanie, and Andrew Zisserman, · 2018
Cited alongside, same era.
“Learnable pins: Cross-modal embeddings for person identity,”
Arsha Nagrani, Samuel Albanie, and Andrew Zisserman, · 2018
Cited alongside, same era.
“On learning associations of faces and voices,”
Changil Kim, Hijung Valentina Shin, Tae-Hyun Oh, Alexandre Kaspar, Mohamed Elgharib, and Wojciech Matusik, · 2018
Cited alongside, same era.
“Learning to lip read words by watching videos,”
Joon Son Chung and Andrew Zisserman, · 2018
Cited alongside, same era.
“Speaker-independent speech separation with deep attractor network,”
Yi Luo, Zhuo Chen, and Nima Mesgarani, · 2018
Cited alongside, same era.
“X-vectors: Robust dnn embeddings for speaker recognition,”
David Snyder, Daniel Garcia-Romero, Gregory Sell, Daniel Povey, and Sanjeev Khudanpur, · 2018
Cited alongside, same era.
“Visual speech enhancement,”
Aviv Gabbay, Asaph Shamir, and Shmuel Peleg, · 2018
Cited alongside, same era.
“Speech2face: Learning the face behind a voice,”
Tae-Hyun Oh, Tali Dekel, Changil Kim, Inbar Mosseri, William T Freeman, Michael Rubinstein, and Wojciech Matusik, · 2019
Later among the works it cites.
“Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation,”
Yi Luo and Nima Mesgarani, · 2019
Later among the works it cites.
“Speakerbeam: Speaker aware neural network for target speaker extraction in speech mixtures,”
Kateřina Žmolíková, Marc Delcroix, Keisuke Kinoshita, Tsubasa Ochiai, Tomohiro Nakatani, Lukáš Burget, and Jan Černockỳ, · 2019
Later among the works it cites.
“My lips are concealed: Audio-visual speech enhancement through obstructions,”
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman, · 2019
Later among the works it cites.
“Perfect match: Self-supervised embeddings for cross-modal retrieval,”
Soo-Whan Chung, Joon Son Chung, and Hong-Goo Kang, · 2020
Closest in time.
“The sound of my voice: Speaker representation loss for target voice separation,”
Seongkyu Mun, Soyeon Choe, Jaesung Huh, and Joon Son Chung, · 2020
Closest in time.
“Disentangled speech embeddings using cross-modal self-supervision,”
Arsha Nagrani, Joon Son Chung, Samuel Albanie, and Andrew Zisserman, · 2020
Closest in time.