Fetching the paper…
Reading the bibliography…
Multimodal generative models should be able to learn a meaningful latent representation that enables a coherent joint generation of all modalities (e.g., images and text).
Y. LeCun, L. Bottou, Y. Bengio, P. Haffner, Gradient-based learning applied to document recognition, Proceedings of the IEEE 86 (11) (1998) 2278–2324
1998
Earlier work this paper cites.
G. E. Hinton, Training products of experts by minimizing contrastive divergence, Neural Computation 14 (8) (2002) 1771–1800
2002
Earlier work this paper cites.
M. Welling, Product of experts, Scholarpedia 2 (10) (2007) 3879
2007
Earlier work this paper cites.
J. Ngiam, A. Khosla, M. Kim, J. Nam, H. Lee, A. Y. Ng, Multimodal deep learning, in: Proceedings of the 28th International Conference on International Conference on Machine Learning (ICML), 2011, p. 689–696
2011
Earlier work this paper cites.
Y. Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, A. Y. Ng, Reading digits in natural images with unsupervised feature learning, in: NIPS Workshop on Deep Learning and Unsupervised Feature Learning 2011, 2011
2011
Earlier work this paper cites.
C. Wah, S. Branson, P. Welinder, P. Perona, S. Belongie, The Caltech-UCSD Birds-200-2011 Dataset, Tech. Rep. CNS-TR-2011-001, California Institute of Technology (2011)
2011
Earlier work this paper cites.
D. P. Kingma, M. Welling, Auto-encoding variational Bayes, in: International Conference on Learning Representations (ICLR), 2014
2014
Earlier work this paper cites.
C. Silberer, M. Lapata, Learning grounded meaning representations with autoencoders, in: Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (ACL), Vol. 1, Association for Computational Linguistics, 2014, pp. 721–732
2014
Earlier work this paper cites.
D. P. Kingma, S. Mohamed, D. J. Rezende, M. Welling, Semi-supervised learning with deep generative models, in: Advances in Neural Information Processing Systems (NeurIPS), 2014, pp. 3581–3589
2014
Earlier work this paper cites.
D. J. Rezende, S. Mohamed, D. Wierstra, Stochastic backpropagation and approximate inference in deep generative models, in: Proceedings of the 31st International Conference on Machine Learning (ICML), Vol. 32(2) of Proceedings of Machine Learning Research, PMLR, 2014, pp. 1278–1286
2014
Cited alongside, same era.
K. Sohn, H. Lee, X. Yan, Learning structured output representation using deep conditional generative models, in: Advances in Neural Information Processing Systems (NeurIPS), 2015, pp. 3483–3491
2015
Cited alongside, same era.
doi:10.1016/J.TRAC.2015.09.005
M. Vinaixa, E. L. Schymanski, S. Neumann, M. Navarro, R. M. Salek, O. Yanes, Mass spectral databases for LC/MS- and GC/MS-based metabolomics: State of the field and future prospects , TrAC Trends in Analytical Chemistry 78 (2016) 23–35 · 2015
Cited alongside, same era.
D. P. Kingma, J. Ba, Adam: A method for stochastic optimization, in: International Conference on Learning Representations (ICLR), 2015
2015
Cited alongside, same era.
doi:10.1186/s13321-017-0207-1
E. L. Schymanski, C. Ruttkies, M. Krauss, C. Brouard, T. Kind, K. Dührkop, F. Allen, A. Vaniya, D. Verdegem, S. Böcker, J. Rousu, H. Shen, H. Tsugawa, T. Sajed, O. Fiehn, B. Ghesquière, S. Neumann, Critical assessment of small molecule identification 2016: automated methods , Journal of Cheminformatics 9 (1) (Mar. 2017) · 2017
Later among the works it cites.
M. Wu, N. Goodman, Multimodal generative models for scalable weakly-supervised learning, in: Advances in Neural Information Processing Systems 31 (NIPS), 2018, pp. 5575–5585
2018
Later among the works it cites.
D. Massiceti, P. K. Dokania, N. Siddharth, P. H. S. Torr, Visual dialogue without vision or dialogue, in: NeurIPS Workshop on Critiquing and Correcting Trends in Machine Learning, 2018 · 2018
Later among the works it cites.
O. Caglayan, P. Madhyastha, L. Specia, L. Barrault, Probing the need for visual context in multimodal machine translation, in: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL HLT), Vol. 1, 2019, pp. 4159–4170
2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
Y. Burda, R. Grosse, R. Salakhutdinov, Importance weighted autoencoders, in: International Conference on Learning Representations (ICLR), 2016 · 2016
Cited alongside, same era.
K. He, X. Zhang, S. Ren, J. Sun, Deep residual learning for image recognition, in: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 770–778
2016
Cited alongside, same era.
M. Suzuki, K. Nakayama, Y. Matsuo, Joint Multimodal Learning with Deep Generative Models, in: International Conference on Learning Representations Workshop (ICLR) Workshop Track, 2017
2017
Cited alongside, same era.
P. Bojanowski, E. Grave, A. Joulin, T. Mikolov, Enriching word vectors with subword information, Transactions of the Association for Computational Linguistics 5 (2017) 135–146
2017
Cited alongside, same era.
Later among the works it cites.
2019
Later among the works it cites.
Y. Shi, S. N, B. Paige, P. Torr, Variational mixture-of-experts autoencoders for multi-modal deep generative models, in: Advances in Neural Information Processing Systems (NeurIPS), 2019, pp. 15718–15729
2019
Later among the works it cites.
2019
Later among the works it cites.
J. E. van Engelen, H. H. Hoos, A survey on semi-supervised learning, Machine Learning 109 (2) (2020) 373–440
2020
Later among the works it cites.
R. N. Yadav, A. Sardana, V. P. Namboodiri, R. M. Hegde, Bridged variational autoencoders for joint modeling of images and attributes, 2020 IEEE Winter Conference on Applications of Computer Vision (WACV) (2020) 1468–1476
2020
Later among the works it cites.