Fetching the paper…
Reading the bibliography…
In this paper our objectives are, first, networks that can embed audio and visual inputs into a common space that is suitable for cross-modal retrieval; and second, a network that can localize the object that sounds in an image, given the audio signal.
Learning classification from unlabelled data
de Sa, V.R.: · 1994
Earlier work this paper cites.
Solving the multiple instance problem with axis-parallel rectangles
Dietterich, T.G., Lathrop, R.H., Lozano-Perez, T.: · 1997
Earlier work this paper cites.
Object recognition as machine translation: Learning a lexicon for a fixed image vocabulary
Duygulu, P., Barnard, K., de Freitas, J.F.G., Forsyth, D.A.: · 2002
Earlier work this paper cites.
Matching words and pictures
Barnard, K., Duygulu, P., de Freitas, N., Forsyth, D., Blei, D., Jordan, M.: · 2003
Earlier work this paper cites.
Pixels that sound
Kidron, E., Schechner, Y.Y., Elad, M.: · 2005
Earlier work this paper cites.
Learning a similarity metric discriminatively, with application to face verification
Chopra, S., Hadsell, R., LeCun, Y.: · 2005
Earlier work this paper cites.
A duality based approach for realtime TV-L1 optical flow
Zach, C., Pock, T., Bischof, H.: · 2007
Earlier work this paper cites.
Audio-visual fusion and tracking with multilevel iterative decoding: Framework and experimental evaluation
Shivappa, S.T., Rao, B.D., Trivedi, M.M.: · 2010
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
Frome, A., Corrado, G.S., Shlens, J., Bengio, S., Dean, J., Ranzato, M.A., Mikolov, T.: · 2013
Earlier work this paper cites.
Discriminative unsupervised feature learning with convolutional neural networks
Dosovitskiy, A., Springenberg, J.T., Riedmiller, M., Brox, T.: · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Simonyan, K., Zisserman, A.: · 2014
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
Xu, K., Ba, J., Kiros, R., Courville, A., Salakhutdinov, R., Zemel, R., Bengio, Y.: · 2015
Earlier work this paper cites.
Unsupervised visual representation learning by context prediction
Doersch, C., Gupta, A., Efros, A.A.: · 2015
Earlier work this paper cites.
Learning to see by moving
Agrawal, P., Carreira, J., Malik, J.: · 2015
Earlier work this paper cites.
Unsupervised learning of visual representations using videos
Wang, X., Gupta, A.: · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S., Szegedy, C.: · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Simonyan, K., Zisserman, A.: · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D.P., Ba, J.: · 2015
Cited alongside, same era.
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., Rabinovich, A.: · 2015
Cited alongside, same era.
ESC: Dataset for environmental sound classification
Piczak, K.J.: · 2015
Cited alongside, same era.
Context encoders: Feature learning by inpainting
Pathak, D., Krähenbühl, P., Donahue, J., Darrell, T., Efros, A.A.: · 2016
Later among the works it cites.
Unsupervised learning of visual representations by solving jigsaw puzzles
Noroozi, M., Favaro, P.: · 2016
Later among the works it cites.
Learning deep structure-preserving image-text embeddings
Wang, L., Li, Y., Lazebnik, S.: · 2016
Later among the works it cites.
Learning deep features for discriminative localization
Zhou, B., Khosla, A., Lapedriza, A., Oliva, A., Torralba, A.: · 2016
Later among the works it cites.
Look, listen and learn
Arandjelović, R., Zisserman, A.: · 2017
Closest in time.
See, hear, and read: Deep aligned representations
Aytar, Y., Vondrick, C., Torralba, A.: · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Is object localization for free? – weakly-supervised learning with convolutional neural networks
Oquab, M., Bottou, L., Laptev, I., Sivic, J.: · 2015
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., Bengio, Y.: · 2015
Cited alongside, same era.
SoundNet: Learning sound representations from unlabeled video
Aytar, Y., Vondrick, C., Torralba, A.: · 2016
Cited alongside, same era.
Unsupervised learning of spoken language with visual context
Harwath, D., Torralba, A., Glass, J.R.: · 2016
Cited alongside, same era.
Ambient sound provides supervision for visual learning
Owens, A., Jiajun, W., McDermott, J., Freeman, W., Torralba, A.: · 2016
Cited alongside, same era.
Visually indicated sounds
Owens, A., Isola, P., McDermott, J.H., Torralba, A., Adelson, E.H., Freeman, W.T.: · 2016
Cited alongside, same era.
Closest in time.
Self-supervised video representation learning with odd-one-out networks
Fernando, B., Bilen, H., Gavves, E., Gould, S.: · 2017
Closest in time.
Multi-task self-supervised visual learning
Doersch, C., Zisserman, A.: · 2017
Closest in time.
Audio Set: An ontology and human-labeled dataset for audio events
Gemmeke, J.F., Ellis, D.P.W., Freedman, D., Jansen, A., Lawrence, W., Moore, R.C., Plakal, M., Ritter, M.: · 2017
Closest in time.
NetVLAD: CNN architecture for weakly supervised place recognition
Arandjelović, R., Gronat, P., Torii, A., Pajdla, T., Sivic, J.: · 2017
Closest in time.
CBVMR: Content-Based Video-Music Retrieval Using Soft Intra-Modal Structure Constraint
Hong, S., Im, W., S. Yang, H.: · 2018
Closest in time.
On learning association of sound source and visual scenes
Senocak, A., Oh, T.H., Kim, J., Yang, M.H., Kweon, I.S.: · 2018
Closest in time.
The sound of pixels
Zhao, H., Gan, C., Rouditchenko, A., Vondrick, C., McDermott, J., Torralba, A.: · 2018
Closest in time.
Audio-visual scene analysis with self-supervised multisensory features
Owens, A., Efros, A.A.: · 2018
Closest in time.