Fetching the paper…
Reading the bibliography…
Cross-modal representation learning allows to integrate information from different modalities into one representation.
Representation learning: A review and new perspectives
Bengio, Y., Courville, A., and Vincent, P · 1939
Earlier work this paper cites.
A learning algorithm for boltzmann machines
Ackley, D. H., Hinton, G. E., and Sejnowski, T. J · 1985
Earlier work this paper cites.
Distributed Representations , pp. 77–109
Hinton, G. E., McClelland, J. L., and Rumelhart, D. E · 1986
Earlier work this paper cites.
Learning Internal Representations by Error Propagation , pp. 318–362
Rumelhart, D. E., Hinton, G. E., and Williams, R. J · 1986
Earlier work this paper cites.
Auto-association by multilayer perceptrons and singular value decomposition
Bourlard, H. and Kamp, Y · 1988
Earlier work this paper cites.
Nonlinear principal component analysis using autoassociative neural networks
Kramer, M. A · 1991
Earlier work this paper cites.
Learning factorial codes by predictability minimization
Schmidhuber, J · 1992
Earlier work this paper cites.
Autoencoders, minimum description length and helmholtz free energy
Hinton, G. E. and Zemel, R. S · 1993
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Lecun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
Vincent, P., Larochelle, H., Bengio, Y., and Manzagol, P.-A · 2008
Earlier work this paper cites.
Large scale image annotation: learning to rank with joint word-image embeddings
Weston, J., Bengio, S., and Usunier, N · 2010
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Cited alongside, same era.
Wavenet: A generative model for raw audio
Oord, A. v. d., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Unsupervised representation learning with deep convolutional generative adversarial networks
Radford, A., Metz, L., and Chintala, S · 2016
Cited alongside, same era.
Generative adversarial text-to-image synthesis
Reed, S., Akata, Z., Yan, X., Logeswaran, L., Schiele, B., and Lee, H · 2016
Cited alongside, same era.
Look, listen and learn
Arandjelovic, R. and Zisserman, A · 2017
Disentangling by partitioning: A representation learning framework for multimodal sensory data
Hsu, W.-N. and Glass, J · 2018
Later among the works it cites.
Spectral normalization for generative adversarial networks
Miyato, T., Kataoka, T., Koyama, M., and Yoshida, Y · 2018
Later among the works it cites.
Large scale GAN training for high fidelity natural image synthesis
Brock, A., Donahue, J., and Simonyan, K · 2019
Later among the works it cites.
Cross-modal learning with adversarial samples
Li, C., Gao, S., Deng, C., Xie, D., and Liu, W · 2019
Later among the works it cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Lu, J., Batra, D., Parikh, D., and Lee, S · 2019
Later among the works it cites.
Translating visual art into music
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Wasserstein gan
Arjovsky, M., Chintala, S., and Bottou, L · 2017
Cited alongside, same era.
Deep cross-modal audio-visual generation
Chen, L., Srivastava, S., Duan, Z., and Xu, C · 2017
Cited alongside, same era.
Adversarial feature learning
Donahue, J., Krähenbühl, P., and Darrell, T · 2017
Cited alongside, same era.
Neural discrete representation learning
van den Oord, A., Vinyals, O., and Kavukcuoglu, K · 2017
Cited alongside, same era.
Objects that sound
Arandjelovic, R. and Zisserman, A · 2018
Cited alongside, same era.
Cmcgan: A uniform framework for cross-modal visual-audio mutual generation
Hao, W., Zhang, Z., and Guan, H · 2018
Cited alongside, same era.
Müller-Eberstein, M. and van Noord, N · 2019
Later among the works it cites.
Generating diverse high-fidelity images with vq-vae-2
Razavi, A., van den Oord, A., and Vinyals, O · 2019
Later among the works it cites.
Variational mixture-of-experts autoencoders for multi-modal deep generative models
Shi, Y., N, S., Paige, B., and Torr, P · 2019
Later among the works it cites.
Videobert: A joint model for video and language representation learning
Sun, C., Myers, A., Vondrick, C., Murphy, K., and Schmid, C · 2019
Later among the works it cites.
Disjoint mapping network for cross-modal matching of voices and faces
Wen, Y., Ismail, M. A., Liu, W., Raj, B., and Singh, R · 2019
Later among the works it cites.
Adaptive cross-modal few-shot learning
Xing, C., Rostamzadeh, N., Oreshkin, B., and O. Pinheiro, P. O · 2019
Later among the works it cites.