Fetching the paper…
Reading the bibliography…
Learning generative models that span multiple data modalities, such as vision and language, is often motivated by the desire to learn more useful, generalisable representations that faithfully capture common underlying factors between the modalities.
How diagrams can improve reasoning
M. I. Bauer and P. N. Johnson-Laird · 1993
Earlier work this paper cites.
Grounding language in perception
J. M. Siskind · 1994
Earlier work this paper cites.
Training products of experts by minimizing contrastive divergence
G. E. Hinton · 2002
Earlier work this paper cites.
Matching words and pictures
K. Barnard, P. Duygulu, D. Forsyth, N. d. Freitas, D. M. Blei, and M. I. Jordan · 2003
Earlier work this paper cites.
Modeling annotated data
D. M. Blei and M. I. Jordan · 2003
Earlier work this paper cites.
Grounded cognition
L. W. Barsalou · 2008
Earlier work this paper cites.
Explicit encoding of multimodal percepts by single neurons in the human brain
R. Q. Quiroga, A. Kraskov, C. Koch, and I. Fried · 2009
Earlier work this paper cites.
The neural basis of multisensory integration in the midbrain: its organization and maturation
B. E. Stein, T. R. Stanford, and B. A. Rowland · 2009
Earlier work this paper cites.
Multimodal deep learning
J. Ngiam, A. Khosla, M. Kim, J. Nam, H. Lee, and A. Y. Ng · 2011
Earlier work this paper cites.
The Caltech-UCSD Birds-200-2011 Dataset
C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie · 2011
Earlier work this paper cites.
Monte Carlo statistical methods
C. Robert and G. Casella · 2013
Earlier work this paper cites.
A convolutional neural network for modelling sentences
N. Kalchbrenner, E. Grefenstette, and P. Blunsom · 2014
Earlier work this paper cites.
Adam: a method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2014
Earlier work this paper cites.
Learning grounded meaning representations with autoencoders
C. Silberer and M. Lapata · 2014
Earlier work this paper cites.
Deep domain confusion: Maximizing for domain invariance
E. Tzeng, J. Hoffman, N. Zhang, K. Saenko, and T. Darrell · 2014
Earlier work this paper cites.
From perception to conception: learning multisensory representations
I. Yildirim · 2014
Earlier work this paper cites.
Importance weighted autoencoders
Y. Burda, R. Grosse, and R. Salakhutdinov · 2015
Earlier work this paper cites.
Unsupervised domain adaptation by backpropagation
Y. Ganin and V. Lempitsky · 2015
Cited alongside, same era.
Learning transferable features with deep adaptation networks
M. Long, Y. Cao, J. Wang, and M. I. Jordan · 2015
Cited alongside, same era.
Learning structured output representation using deep conditional generative models
K. Sohn, H. Lee, and X. Yan · 2015
Cited alongside, same era.
Domain separation networks
K. Bousmalis, G. Trigeorgis, N. Silberman, D. Krishnan, and D. Erhan · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Precomputed real-time texture synthesis with markovian generative adversarial networks
C. Li and M. Wand · 2016
Cited alongside, same era.
Variational methods for conditional multimodal deep learning
G. Pandey and A. Dukkipati · 2017
Later among the works it cites.
Sticking the landing: Simple, lower-variance gradient estimators for variational inference
G. Roeder, Y. Wu, and D. K. Duvenaud · 2017
Later among the works it cites.
Joint multimodal learning with deep generative models
M. Suzuki, K. Nakayama, and Y. Matsuo · 2017
Later among the works it cites.
Unsupervised cross-domain image generation
Y. Taigman, A. Polyak, and L. Wolf · 2017
Later among the works it cites.
Towards deeper understanding of variational autoencoding models
S. Zhao, J. Song, and S. Ermon · 2017
Later among the works it cites.
Common object representations for visual recognition and production
J. E. Fan, D. Yamins, and N. B. Turk-Browne · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Unsupervised domain adaptation with residual transfer networks
M. Long, H. Zhu, J. Wang, and M. I. Jordan · 2016
Cited alongside, same era.
Convolutional neural network language models
N.-Q. Pham, G. Kruszewski, and G. Boleda · 2016
Cited alongside, same era.
Variational autoencoder for deep learning of images, labels and captions
Y. Pu, Z. Gan, R. Henao, X. Yuan, C. Li, A. Stevens, and L. Carin · 2016
Cited alongside, same era.
Deep variational canonical correlation analysis
W. Wang, X. Yan, H. Lee, and K. Livescu · 2016
Cited alongside, same era.
Generative image modeling using style and structure adversarial networks
X. Wang and A. Gupta · 2016
Cited alongside, same era.
Attribute2image: Conditional image generation from visual attributes
X. Yan, J. Yang, K. Sohn, and H. Lee · 2016
Cited alongside, same era.
Later among the works it cites.
Auto-encoding sequential monte carlo
T. A. Le, M. Igl, T. Rainforth, T. Jin, and F. Wood · 2018
Later among the works it cites.
The mythos of model interpretability
Z. C. Lipton · 2018
Later among the works it cites.
Tighter variational bounds are not necessarily better
T. Rainforth, A. R. Kosiorek, T. A. Le, C. J. Maddison, M. Igl, F. Wood, and Y. W. Teh · 2018
Later among the works it cites.
On the convergence of adam and beyond
S. J. Reddi, S. Kale, and S. Kumar · 2018
Later among the works it cites.
Generative models of visually grounded imagination
R. Vedantam, I. Fischer, J. Huang, and K. Murphy · 2018
Later among the works it cites.
Multimodal generative models for scalable weakly-supervised learning
M. Wu and N. Goodman · 2018
Later among the works it cites.
Few-shot unsupervised image-to-image translation, 2019
M.-Y. Liu, X. Huang, A. Mallya, T. Karras, T. Aila, J. Lehtinen, and J. Kautz · 2019
Closest in time.
Disentangling disentanglement in variational autoencoders
E. Mathieu, T. Rainforth, N. Siddharth, and Y. W. Teh · 2019
Closest in time.
Latent translation: Crossing modalities by bridging generative models
Y. Tian and J. Engel · 2019
Closest in time.
Learning factorized multimodal representations
Y. H. Tsai, P. P. Liang, A. A. Bagherzade, L.-P. Morency, and R. Salakhutdinov · 2019
Closest in time.
Doubly reparameterized gradient estimators for monte carlo objectives
G. Tucker, D. Lawson, S. Gu, and C. J. Maddison · 2019
Closest in time.