Fetching the paper…
Reading the bibliography…
The task of isolating a target singing voice in music videos has useful applications.
Performance measurement in blind audio source separation
E. Vincent, R. Gribonval, and C. Fevotte · 2005
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
Complex ratio masking for monaural speech separation
Donald S Williamson, Yuxuan Wang, and DeLiang Wang · 2015
Earlier work this paper cites.
Synthesizing normalized faces from facial identity features
Forrester Cole, David Belanger, Dilip Krishnan, Aaron Sarna, Inbar Mosseri, and William T Freeman · 2017
Earlier work this paper cites.
Towards estimating the upper bound of visual-speech recognition: The visual lip-reading feasibility database
Adriana Fernandez-Lopez, Oriol Martinez, and Federico M Sukno · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Earlier work this paper cites.
Singing voice separation with deep U-Net convolutional networks
A Jansson, E Humphrey, N Montecchio, R Bittner, A Kumar, and T Weyde · 2017
Earlier work this paper cites.
The MUSDB18 corpus for music separation, December 2017
Zafar Rafii, Antoine Liutkus, Fabian-Robert Stöter, Stylianos Ioannis Mimilakis, and Rachel Bittner · 2017
Earlier work this paper cites.
The conversation: Deep audio-visual speech enhancement
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman · 2018
Earlier work this paper cites.
Feature-wise transformations
Vincent Dumoulin, Ethan Perez, Nathan Schucher, Florian Strub, Harm de Vries, Aaron Courville, and Yoshua Bengio · 2018
Earlier work this paper cites.
Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation
Ariel Ephrat, Inbar Mosseri, Oran Lang, Tali Dekel, Kevin Wilson, Avinatan Hassidim, William T Freeman, and Michael Rubinstein · 2018
Earlier work this paper cites.
Visual speech enhancement
Aviv Gabbay, Asaph Shamir, and Shmuel Peleg · 2018
Earlier work this paper cites.
On learning associations of faces and voices
Changil Kim, Hijung Valentina Shin, Tae-Hyun Oh, Alexandre Kaspar, Mohamed Elgharib, and Wojciech Matusik · 2018
Earlier work this paper cites.
Audio-visual scene analysis with self-supervised multisensory features
Andrew Owens and Alexei A Efros · 2018
Earlier work this paper cites.
Wave-u-net: A multi-scale neural network for end-to-end audio source separation
Daniel Stoller, Sebastian Ewert, and Simon Dixon · 2018
Earlier work this paper cites.
MMDenseLSTM: An efficient combination of convolutional and recurrent neural networks for audio source separation
Naoya Takahashi, Nabarun Goswami, and Yuki Mitsufuji · 2018
Cited alongside, same era.
A closer look at spatiotemporal convolutions for action recognition
Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri · 2018
Cited alongside, same era.
Spatial temporal graph convolutional networks for skeleton-based action recognition
Sijie Yan, Yuanjun Xiong, and Dahua Lin · 2018
Cited alongside, same era.
The sound of pixels
Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh McDermott, and Antonio Torralba · 2018
Cited alongside, same era.
Co-separating sounds of visual objects
Ruohan Gao and Kristen Grauman · 2019
Cited alongside, same era.
Conditioned-u-net: Introducing a control mechanism in the u-net for multiple source separations
Multi-Modal Analysis for Music Performances
Bochen Li · 2020
Later among the works it cites.
Deep audio-visual speech separation with attention mechanism
Chenda Li and Yanmin Qian · 2020
Later among the works it cites.
Content based singing voice source separation via strong conditioning using aligned phonemes
Gabriel Meseguer-Brocal and Geoffroy Peeters · 2020
Later among the works it cites.
An overview of deep-learning-based audio-visual speech enhancement and separation
Daniel Michelsanti, Zheng-Hua Tan, Shi-Xiong Zhang, Yong Xu, Meng Yu, Dong Yu, and Jesper Jensen · 2020
Later among the works it cites.
Facial landmark-based emotion recognition via directed graph neural network
Quang Tran Ngoc, Seunghyun Lee, and Byung Cheol Song · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gabriel Meseguer-Brocal and Geoffroy Peeters · 2019
Cited alongside, same era.
Face landmark-based speaker-independent audio-visual speech enhancement in multi-talker environments
Giovanni Morrone, Sonia Bergamaschi, Luca Pasa, Luciano Fadiga, Vadim Tikhanoff, and Leonardo Badino · 2019
Cited alongside, same era.
Speech2face: Learning the face behind a voice
Tae-Hyun Oh, Tali Dekel, Changil Kim, Inbar Mosseri, William T Freeman, Michael Rubinstein, and Wojciech Matusik · 2019
Cited alongside, same era.
End-to-end sound source separation conditioned on instrument labels
Olga Slizovskaia, Leo Kim, Gloria Haro, and Emilia Gomez · 2019
Cited alongside, same era.
Time domain audio visual speech separation
Jian Wu, Yong Xu, Shi-Xiong Zhang, Lian-Wu Chen, Meng Yu, Lei Xie, and Dong Yu · 2019
Cited alongside, same era.
Recursive visual sound separation using minus-plus net
Xudong Xu, Bo Dai, and Dahua Lin · 2019
Cited alongside, same era.
The sound of motions
Hang Zhao, Chuang Gan, Wei-Chiu Ma, and Antonio Torralba · 2019
Cited alongside, same era.
Viet-Nhat Nguyen, Mostafa Sadeghi, Elisa Ricci, and Xavier Alameda-Pineda · 2020
Later among the works it cites.
Deep learning based source separation applied to choir ensembles
Darius Petermann, Pritish Chandna, Helena Cuesta, Jordi Bonada, and Emilia Gomez · 2020
Later among the works it cites.
Meta-learning extractors for music source separation
David Samuel, Aditya Ganeshan, and Jason Naradowsky · 2020
Later among the works it cites.
Visually guided sound source separation using cascaded opponent filter network
Lingyu Zhu and Esa Rahtu · 2020
Later among the works it cites.
Visualvoice: Audio-visual speech separation with cross-modal consistency
Ruohan Gao and Kristen Grauman · 2021
Closest in time.
Looking into your speech: Learning cross-modal affinity for audio-visual speech separation
Jiyoung Lee, Soo-Whan Chung, Sunok Kim, Hong-Goo Kang, and Kwanghoon Sohn · 2021
Closest in time.
Sams-net: A sliced attention-based neural network for music source separation
Tingle Li, Jiawei Chen, Haowen Hou, and Ming Li · 2021
Closest in time.
Face-gcn: A graph convolutional network for 3d dynamic face identification/recognition
Konstantinos Papadopoulos, Anis Kacem, Abdelrahman Shabayek, and Djamila Aouada · 2021
Closest in time.
Conditioned source separation for musical instrument performances
Olga Slizovskaia, Gloria Haro, and Emilia Gómez · 2021
Closest in time.