Fetching the paper…
Reading the bibliography…
Recent progress in deep learning has enabled many advances in sound separation and visual scene understanding.
A framework for multiple-instance learning
Oded Maron and Tomás Lozano-Pérez · 1998
Earlier work this paper cites.
Audio vision: Using audio-visual synchrony to locate sounds
John R Hershey and Javier R Movellan · 2000
Earlier work this paper cites.
Audio-visual sound separation via hidden Markov models
John R Hershey and Michael Casey · 2002
Earlier work this paper cites.
Performance measurement in blind audio source separation
Emmanuel Vincent, Rémi Gribonval, and Cédric Févotte · 2006
Earlier work this paper cites.
Separating sounds from a single image
Lingyu Zhu and Esa Rahtu · 2007
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Yfcc100m: The new data in multimedia research
Bart Thomee, David A Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li · 2016
Earlier work this paper cites.
Look, listen and learn
Relja Arandjelovic and Andrew Zisserman · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Jort F Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Earlier work this paper cites.
Attentional pooling for action recognition
Rohit Girdhar and Deva Ramanan · 2017
Earlier work this paper cites.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam · 2017
Earlier work this paper cites.
Looking to listen at the cocktail party: a speaker-independent audio-visual model for speech separation
Ariel Ephrat, Inbar Mosseri, Oran Lang, Tali Dekel, Kevin Wilson, Avinatan Hassidim, William T Freeman, and Michael Rubinstein · 2018
Earlier work this paper cites.
Learning to separate object sounds by watching unlabeled video
Ruohan Gao, Rogerio Feris, and Kristen Grauman · 2018
Earlier work this paper cites.
Jointly discovering visual objects and spoken words from raw sensory input
David Harwath, Adria Recasens, Dídac Surís, Galen Chuang, Antonio Torralba, and James Glass · 2018
Earlier work this paper cites.
Audio-visual speech enhancement using multimodal deep convolutional neural networks
Jen-Cheng Hou, Syu-Siang Wang, Ying-Hui Lai, Yu Tsao, Hsiu-Wen Chang, and Hsin-Min Wang · 2018
Cited alongside, same era.
Cooperative learning of audio and video models from self-supervised synchronization
Bruno Korbar, Du Tran, and Lorenzo Torresani · 2018
Cited alongside, same era.
Audio-visual scene analysis with self-supervised multisensory features
Andrew Owens and Alexei A Efros · 2018
Cited alongside, same era.
Learning to localize sound source in visual scenes
Arda Senocak, Tae-Hyun Oh, Junsik Kim, Ming-Hsuan Yang, and In So Kweon · 2018
Cited alongside, same era.
Audio-visual event localization in unconstrained videos
Yapeng Tian, Jing Shi, Bochen Li, Zhiyao Duan, and Chenliang Xu · 2018
Cited alongside, same era.
The sound of pixels
Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh McDermott, and Antonio Torralba · 2018
Dual attention matching for audio-visual event localization
Yu Wu, Linchao Zhu, Yan Yan, and Yi Yang · 2019
Later among the works it cites.
Recursive visual sound separation using minus-plus net
Xudong Xu, Bo Dai, and Dahua Lin · 2019
Later among the works it cites.
The sound of motions
Hang Zhao, Chuang Gan, Wei-Chiu Ma, and Antonio Torralba · 2019
Later among the works it cites.
Self-supervised learning of audio-visual objects from video
Triantafyllos Afouras, Andrew Owens, Joon Son Chung, and Andrew Zisserman · 2020
Closest in time.
Music gesture for visual sound separation
Chuang Gan, Deng Huang, Hang Zhao, Joshua B Tenenbaum, and Antonio Torralba · 2020
Closest in time.
Curriculum audiovisual learning
Di Hu, Zheng Wang, Haoyi Xiong, Dong Wang, Feiping Nie, and Dejing Dou · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Unsupervised training of a deep clustering model for multichannel blind source separation
Lukas Drude, Daniel Hasenklever, and Reinhold Haeb-Umbach · 2019
Cited alongside, same era.
Co-separating sounds of visual objects
Ruohan Gao and Kristen Grauman · 2019
Cited alongside, same era.
Universal sound separation
Ilya Kavalerov, Scott Wisdom, Hakan Erdogan, Brian Patton, Kevin Wilson, Jonathan Le Roux, and John R. Hershey · 2019
Cited alongside, same era.
SDR–half-baked or well done?
Jonathan Le Roux, Scott Wisdom, Hakan Erdogan, and John R. Hershey · 2019
Cited alongside, same era.
Self-supervised audio-visual co-segmentation
Andrew Rouditchenko, Hang Zhao, Chuang Gan, Josh McDermott, and Antonio Torralba · 2019
Cited alongside, same era.
Bootstrapping single-channel source separation via unsupervised spatial clustering on stereo mixtures
Prem Seetharaman, Gordon Wichern, Jonathan Le Roux, and Bryan Pardo · 2019
Cited alongside, same era.
Closest in time.
Coincidence, categorization, and consolidation: Learning to recognize sounds with minimal supervision
Aren Jansen, Daniel PW Ellis, Shawn Hershey, R Channing Moore, Manoj Plakal, Ashok C Popat, and Rif A Saurous · 2020
Closest in time.
Source separation with weakly labelled data: An approach to computational auditory scene analysis
Qiuqiang Kong, Yuxuan Wang, Xuchen Song, Yin Cao, Wenwu Wang, and Mark D Plumbley · 2020
Closest in time.
An overview of deep-learning-based audio-visual speech enhancement and separation
Daniel Michelsanti, Zheng-Hua Tan, Shi-Xiong Zhang, Yong Xu, Meng Yu, Dong Yu, and Jesper Jensen · 2020
Closest in time.
Listen to what you want: Neural network-based universal sound selector
Tsubasa Ochiai, Marc Delcroix, Yuma Koizumi, Hiroaki Ito, Keisuke Kinoshita, and Shoko Araki · 2020
Closest in time.
Finding strength in weakness: Learning to separate sounds with weak supervision
F. Pishdadian, G. Wichern, and J. Le Roux · 2020
Closest in time.
Improving universal sound separation using sound classification
Efthymios Tzinis, Scott Wisdom, John R. Hershey, Aren Jansen, and Daniel P. W. Ellis · 2020
Closest in time.
Unsupervised sound separation using mixture invariant training
Scott Wisdom, Efthymios Tzinis, Hakan Erdogan, Ron J. Weiss, Kevin Wilson, and John R. Hershey · 2020
Closest in time.
What’s all the fuss about free universal sound separation data?
Scott Wisdom, Hakan Erdogan, Daniel Ellis, Romain Serizel, Nicolas Turpault, Eduardo Fonseca, Justin Salamon, Prem Seetharaman, and John Hershey · 2021
Closest in time.