Fetching the paper…
Reading the bibliography…
The seen birds twitter, the running cars accompany with noise, etc.
Polysensory properties of neurons in the anterior bank of the caudal superior temporal sulcus of the macaque monkey
K. Hikosaka, E. Iwai, H. Saito, and K. Tanaka · 1988
Earlier work this paper cites.
Activation of auditory cortex during silent lipreading
G. A. Calvert, E. T. Bullmore, M. J. Brammer, R. Campbell, S. C. Williams, P. K. McGuire, P. W. Woodruff, S. D. Iversen, and A. S. David · 1997
Earlier work this paper cites.
Determination of number of clusters in k-means clustering and application in colour image segmentation
S. Ray and R. H. Turi · 1999
Earlier work this paper cites.
Multisensory integration: space, time and superadditivity
N. P. Holmes and C. Spence · 2005
Earlier work this paper cites.
Blind audiovisual source separation based on sparse redundant representations
A. L. Casanovas, G. Monaci, P. Vandergheynst, and R. Gribonval · 2010
Earlier work this paper cites.
The pascal visual object classes (voc) challenge
M. Everingham, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman · 2010
Earlier work this paper cites.
Data clustering: 50 years beyond k-means
A. K. Jain · 2010
Earlier work this paper cites.
Multisensory perceptual learning and sensory substitution
M. J. Proulx, D. J. Brown, A. Pasqualotto, and P. Meijer · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Learning to see by moving
P. Agrawal, J. Carreira, and J. Malik · 2015
Earlier work this paper cites.
Unsupervised visual representation learning by context prediction
C. Doersch, A. Gupta, and A. A. Efros · 2015
Earlier work this paper cites.
Data-dependent initializations of convolutional neural networks
P. Krähenbühl, C. Doersch, J. Donahue, and T. Darrell · 2015
Earlier work this paper cites.
Environmental sound classification with convolutional neural networks
K. J. Piczak · 2015
Cited alongside, same era.
Esc: Dataset for environmental sound classification
K. J. Piczak · 2015
Cited alongside, same era.
Unsupervised learning of visual representations using videos
X. Wang and A. Gupta · 2015
Cited alongside, same era.
Character-level convolutional networks for text classification
X. Zhang, J. Zhao, and Y. LeCun · 2015
Cited alongside, same era.
Soundnet: Learning sound representations from unlabeled video
Y. Aytar, C. Vondrick, and A. Torralba · 2016
Cited alongside, same era.
Unsupervised learning of spoken language with visual context
D. Harwath, A. Torralba, and J. Glass · 2016
DCASE 2017 challenge setup: Tasks, datasets and baseline system
T. Heittola and A. Mesaros · 2017
Later among the works it cites.
Cnn architectures for large-scale audio classification
S. Hershey, S. Chaudhuri, D. P. Ellis, J. F. Gemmeke, A. Jansen, R. C. Moore, M. Plakal, D. Platt, R. A. Saurous, B. Seybold, et al · 2017
Later among the works it cites.
Audio event detection using multiple-input convolutional neural network
I.-Y. Jeong, S. Lee, Y. Han, and K. Lee · 2017
Later among the works it cites.
Neuroevolution for sound event detection in real life audio: A pilot study
C. Kroos and M. D. Plumbley · 2017
Later among the works it cites.
Deep binary reconstruction for cross-modal hashing
X. Li, D. Hu, and F. Nie · 2017
Later among the works it cites.
Dynamic routing between capsules
S. Sabour, N. Frosst, and G. E. Hinton · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Ambient sound provides supervision for visual learning
A. Owens, J. Wu, J. H. McDermott, W. T. Freeman, and A. Torralba · 2016
Cited alongside, same era.
A report on sound event detection with different binaural features
S. Adavanne and T. Virtanen · 2017
Cited alongside, same era.
Look, listen and learn
R. Arandjelovic and A. Zisserman · 2017
Cited alongside, same era.
R. Arandjelović and A. Zisserman · 2017
Cited alongside, same era.
See, hear, and read: Deep aligned representations
Y. Aytar, C. Vondrick, and A. Torralba · 2017
Cited alongside, same era.
Learning word-like units from joint audio-visual analysis
D. Harwath and J. R. Glass · 2017
Cited alongside, same era.
A. Ephrat, I. Mosseri, O. Lang, T. Dekel, K. Wilson, A. Hassidim, W. T. Freeman, and M. Rubinstein · 2018
Closest in time.
Learning to separate object sounds by watching unlabeled video
R. Gao, R. Feris, and K. Grauman · 2018
Closest in time.
Audio-visual scene analysis with self-supervised multisensory features
A. Owens and A. A. Efros · 2018
Closest in time.
Learning to localize sound source in visual scenes
A. Senocak, T.-H. Oh, J. Kim, M.-H. Yang, and I. S. Kweon · 2018
Closest in time.
An optimization view on dynamic routing between capsules
D. Wang and Q. Liu · 2018
Closest in time.
H. Zhao, C. Gan, A. Rouditchenko, C. Vondrick, J. McDermott, and A. Torralba · 2018
Closest in time.