Fetching the paper…
Reading the bibliography…
The sound of crashing waves, the roar of fast-moving cars -- sound conveys important information about the objects in our surroundings.
Gaver WW (1993) What in the world do we hear?: An ecological approach to auditory event perception. Ecological psychology 5(1):1–29
1993
Earlier work this paper cites.
de Sa VR (1994b) Minimizing disagreement for self-supervised classification. In: Proceedings of the 1993 Connectionist Models Summer School, Psychology Press, p 300
1993
Earlier work this paper cites.
Indyk P, Motwani R (1998) Approximate nearest neighbors: towards removing the curse of dimensionality. In: ACM Symposium on Theory of Computing
1998
Earlier work this paper cites.
Hershey JR, Movellan JR (1999) Audio vision: Using audio-visual synchrony to locate sounds. In: Advances in Neural Information Processing Systems
1999
Earlier work this paper cites.
Fisher III JW, Darrell T, Freeman WT, Viola PA (2000) Learning joint statistical models for audio-visual fusion and segregation. In: Advances in Neural Information Processing Systems
2000
Earlier work this paper cites.
Slaney M, Covell M (2000) Facesync: A linear operator for measuring synchronization of video facial images and audio tracks. In: Advances in Neural Information Processing Systems
2000
Earlier work this paper cites.
Leung T, Malik J (2001) Representing and recognizing the visual appearance of materials using three-dimensional textons. International Journal of Computer Vision 43(1):29–44
2001
Earlier work this paper cites.
Kidron E, Schechner YY, Elad M (2005) Pixels that sound. In: IEEE Conference on Computer Vision and Pattern Recognition
2005
Earlier work this paper cites.
Smith L, Gasser M (2005) The development of embodied cognition: Six lessons from babies. Artificial life 11(1-2):13–29
2005
Earlier work this paper cites.
Eronen AJ, Peltonen VT, Tuomi JT, Klapuri AP, Fagerlund S, Sorsa T, Lorho G, Huopaniemi J (2006) Audio-based context recognition. IEEE/ACM Transactions on Audio Speech and Language Processing 14(1):321–329
2006
Earlier work this paper cites.
Deng J, Dong W, Socher R, Li LJ, Li K, Fei-Fei L (2009) Imagenet: A large-scale hierarchical image database. In: IEEE Conference on Computer Vision and Pattern Recognition
2009
Earlier work this paper cites.
Mobahi H, Collobert R, Weston J (2009) Deep learning from temporal coherence in video. In: International Conference on Machine Learning
2009
Earlier work this paper cites.
Salakhutdinov R, Hinton G (2009) Semantic hashing. International Journal of Approximate Reasoning 50(7):969–978
2009
Earlier work this paper cites.
Weiss Y, Torralba A, Fergus R (2009) Spectral hashing. In: Advances in Neural Information Processing Systems
2009
Earlier work this paper cites.
Everingham M, Van Gool L, Williams CK, Winn J, Zisserman A (2010) The pascal visual object classes (voc) challenge. International Journal of Computer Vision 88(2):303–338
2010
Earlier work this paper cites.
Lee K, Ellis DP, Loui AC (2010) Detecting local semantic concepts in environmental sounds using markov model based clustering. In: IEEE International Conference on Acoustics, Speech, and Signal Processing
2010
Earlier work this paper cites.
Xiao J, Hays J, Ehinger KA, Oliva A, Torralba A (2010) Sun database: Large-scale scene recognition from abbey to zoo. In: IEEE Conference on Computer Vision and Pattern Recognition
2010
Earlier work this paper cites.
Ellis DP, Zeng X, McDermott JH (2011) Classifying soundtracks with audio texture features. In: IEEE International Conference on Acoustics, Speech, and Signal Processing
2011
Earlier work this paper cites.
McDermott JH, Simoncelli EP (2011) Sound texture perception via statistics of the auditory periphery: evidence from sound synthesis. Neuron 71(5):926–940
2011
Cited alongside, same era.
Ngiam J, Khosla A, Kim M, Nam J, Lee H, Ng AY (2011) Multimodal deep learning. In: International Conference on Machine Learning
2011
Cited alongside, same era.
Krizhevsky A, Sutskever I, Hinton GE (2012) Imagenet classification with deep convolutional neural networks. In: Advances in Neural Information Processing Systems
2012
Cited alongside, same era.
Le QV, Ranzato MA, Monga R, Devin M, Chen K, Corrado GS, Dean J, Ng AY (2012) Building high-level features using large scale unsupervised learning. In: International Conference on Machine Learning
2012
Cited alongside, same era.
Srivastava N, Salakhutdinov RR (2012) Multimodal learning with deep boltzmann machines. In: Advances in Neural Information Processing Systems
Mishkin D, Matas J (2015) All you need is a good init. arXiv preprint arXiv:151106422
2015
Later among the works it cites.
Oquab M, Bottou L, Laptev I, Sivic J (2015) Is object localization for free?-weakly-supervised learning with convolutional neural networks. In: IEEE Conference on Computer Vision and Pattern Recognition
2015
Later among the works it cites.
Thomee B, Shamma DA, Friedland G, Elizalde B, Ni K, Poland D, Borth D, Li LJ (2015) The new data and new challenges in multimedia research. arXiv preprint arXiv:150301817
2015
Later among the works it cites.
Wang X, Gupta A (2015) Unsupervised learning of visual representations using videos. In: IEEE International Conference on Computer Vision
2015
Later among the works it cites.
Zhou B, Khosla A, Lapedriza A, Oliva A, Torralba A (2015) Object detectors emerge in deep scene cnns. In: International Conference on Learning Representations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2012
Cited alongside, same era.
Andrew G, Arora R, Bilmes JA, Livescu K (2013) Deep canonical correlation analysis. In: International Conference on Machine Learning
2013
Cited alongside, same era.
Dosovitskiy A, Springenberg JT, Riedmiller M, Brox T (2014) Discriminative unsupervised feature learning with convolutional neural networks. In: Advances in Neural Information Processing Systems
2014
Cited alongside, same era.
Jia Y, Shelhamer E, Donahue J, Karayev S, Long J, Girshick R, Guadarrama S, Darrell T (2014) Caffe: Convolutional architecture for fast feature embedding. In: ACM Multimedia Conference
2014
Cited alongside, same era.
Lin M, Chen Q, Yan S (2014) Network in network. International Conference on Learning Representations
2014
Cited alongside, same era.
Zhou B, Lapedriza A, Xiao J, Torralba A, Oliva A (2014) Learning deep features for scene recognition using places database. In: Advances in Neural Information Processing Systems
2014
Cited alongside, same era.
Agrawal P, Carreira J, Malik J (2015) Learning to see by moving. In: IEEE International Conference on Computer Vision
2015
Cited alongside, same era.
Doersch C, Gupta A, Efros AA (2015) Unsupervised visual representation learning by context prediction. In: IEEE International Conference on Computer Vision
2015
Cited alongside, same era.
2015
Later among the works it cites.
Aytar Y, Vondrick C, Torralba A (2016) Soundnet: Learning sound representations from unlabeled video. In: Advances in Neural Information Processing Systems
2016
Later among the works it cites.
Gupta S, Hoffman J, Malik J (2016) Cross modal distillation for supervision transfer. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
2016
Later among the works it cites.
Isola P, Zoran D, Krishnan D, Adelson EH (2016) Learning visual groups from co-occurrences in space and time. In: International Conference on Learning Representations, Workshop
2016
Later among the works it cites.
Krähenbühl P, Doersch C, Donahue J, Darrell T (2016) Data-dependent initializations of convolutional neural networks. In: International Conference on Learning Representations
2016
Later among the works it cites.
Zhang R, Isola P, Efros AA (2016) Colorful image colorization. In: European Conference on Computer Vision, Springer, pp 649–666
2016
Later among the works it cites.
2016
Later among the works it cites.
Arandjelović R, Zisserman A (2017) Look, listen and learn. ICCV
2017
Closest in time.
Bau D, Zhou B, Khosla A, Oliva A, Torralba A (2017) Network dissection: Quantifying interpretability of deep visual representations. CVPR
2017
Closest in time.
Gemmeke JF, Ellis DP, Freedman D, Jansen A, Lawrence W, Moore RC, Plakal M, Ritter M (2017) Audio set: An ontology and human-labeled dartaset for audio events. In: IEEE International Conference on Acoustics, Speech, and Signal Processing
2017
Closest in time.
Pathak D, Girshick R, Dollár P, Darrell T, Hariharan B (2017) Learning features by watching objects move. In: CVPR
2017
Closest in time.
Zhang R, Isola P, Efros AA (2017) Split-brain autoencoders: Unsupervised learning by cross-channel prediction. In: CVPR
2017
Closest in time.
Doersch C, Zisserman A (2017) Multi-task self-supervised visual learning. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp 2051–2060
2060
Closest in time.