Fetching the paper…
Reading the bibliography…
Associating sound and its producer in complex audiovisual scene is a challenging task, especially when we are lack of annotated training data.
A review of the cocktail party effect
Barry Arons · 1992
Earlier work this paper cites.
Learning and development in neural networks: the importance of starting small
Jeffrey L. Elman · 1993
Earlier work this paper cites.
The expectation-maximization algorithm
Todd K Moon · 1996
Earlier work this paper cites.
Multisensory integration: space, time and superadditivity
Nicholas P Holmes and Charles Spence · 2005
Earlier work this paper cites.
Pixels that sound
E Kidron, Y. Y Schechner, and M Elad · 2005
Earlier work this paper cites.
Curriculum learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston · 2009
Earlier work this paper cites.
Source-filter based clustering for monaural blind source separation
Martin Spiertz and Volker Gnann · 2009
Earlier work this paper cites.
Natural language processing (almost) from scratch
Ronan Collobert, Koray Kavukcuoglu, Jason Weston, Leon Bottou, Pavel Kuksa, and Michael Karlen · 2011
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Siamese neural networks for one-shot image recognition
Gregory Koch, Richard Zemel, and Ruslan Salakhutdinov · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
Soundnet: Learning sound representations from unlabeled video
Yusuf Aytar, Carl Vondrick, and Antonio Torralba · 2016
Cited alongside, same era.
Multi-modal curriculum learning for semi-supervised image classification
C. Gong, D. Tao, S. J. Maybank, W. Liu, G. Kang, and J. Yang · 2016
Cited alongside, same era.
Ambient sound provides supervision for visual learning
Andrew Owens, Jiajun Wu, Josh H McDermott, William T Freeman, and Antonio Torralba · 2016
Cited alongside, same era.
Learning deep features for discriminative localization
Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba · 2016
Cited alongside, same era.
Look, listen and learn
Relja Arandjelovic and Andrew Zisserman · 2017
Cited alongside, same era.
Relja Arandjelović and Andrew Zisserman · 2017
Learning to separate object sounds by watching unlabeled video
Ruohan Gao, Rogerio Feris, and Kristen Grauman · 2018
Later among the works it cites.
Co-training of audio and video representations from self-supervised temporal synchronization
Bruno Korbar, Du Tran, and Lorenzo Torresani · 2018
Later among the works it cites.
Cassl: Curriculum accelerated self-supervised learning
Adithyavairavan Murali, Lerrel Pinto, Dhiraj Gandhi, and Abhinav Gupta · 2018
Later among the works it cites.
Audio-visual scene analysis with self-supervised multisensory features
Andrew Owens and Alexei A Efros · 2018
Later among the works it cites.
Learning to localize sound source in visual scenes
Arda Senocak, Tae-Hyun Oh, Junsik Kim, Ming-Hsuan Yang, and In So Kweon · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A curriculum learning method for improved noise robustness in automatic speech recognition
Stefan Braun, Daniel Neil, and Shih Chii Liu · 2017
Cited alongside, same era.
Audio set: An ontology and human-labeled dataset for audio events
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Cited alongside, same era.
Cnn architectures for large-scale audio classification
Shawn Hershey, Sourish Chaudhuri, Daniel PW Ellis, Jort F Gemmeke, Aren Jansen, R Channing Moore, Manoj Plakal, Devin Platt, Rif A Saurous, Bryan Seybold, et al · 2017
Cited alongside, same era.
Ariel Ephrat, Inbar Mosseri, Oran Lang, Tali Dekel, Kevin Wilson, Avinatan Hassidim, William T Freeman, and Michael Rubinstein · 2018
Cited alongside, same era.
Lecture notes on data science: Soft k-means clustering
Christian Bauckhage
Cited in the paper.
Later among the works it cites.
Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh McDermott, and Antonio Torralba · 2018
Later among the works it cites.
Co-separating sounds of visual objects
Ruohan Gao and Kristen Grauman · 2019
Later among the works it cites.
Deep multimodal clustering for unsupervised audiovisual learning
Di Hu, Feiping Nie, and Xuelong Li · 2019
Later among the works it cites.
Recursive visual sound separation using minus-plus net
Xudong Xu, Bo Dai, and Dahua Lin · 2019
Later among the works it cites.
Hang Zhao, Chuang Gan, Wei-Chiu Ma, and Antonio Torralba · 2019
Later among the works it cites.