Fetching the paper…
Reading the bibliography…
In the rapidly advancing field of multi-modal machine learning (MMML), the convergence of multiple data modalities has the potential to reshape various applications.
“StackGAN++: Realistic Image Synthesis with Stacked Generative Adversarial Networks,”
Zhang, H., Xu, T., Li, H., Zhang, S., Wang, X., Huang, X., and Metaxas, D. N., 2019, · 1962
Earlier work this paper cites.
“Some mathematical notes on three-mode factor analysis,”
Tucker, L. R., 1966, · 1966
Earlier work this paper cites.
“Design Prototypes: A Knowledge Representation Schema for Design,”
Gero, J. S., 1990, · 1990
Earlier work this paper cites.
“The importance of drawing in the mechanical design process,”
Ullman, D. G., Wood, S., and Craig, D., 1990, · 1990
Earlier work this paper cites.
“Drawings and the design process: A review of protocol studies in design and other disciplines and related research in cognitive psychology,”
Purcell, A. T., and Gero, J. S., 1998, · 1998
Earlier work this paper cites.
“Combining labeled and unlabeled data with co-training,”
Blum, A., and Mitchell, T., 1998, · 1998
Earlier work this paper cites.
“Image-to-word transformation based on dividing,”
Mori, Y., Takahashi, H., and Oka, R., 1999, · 1999
Earlier work this paper cites.
Product design and development
Ulrich, K. T., and Eppinger, S. D., 2000, · 2000
Earlier work this paper cites.
“Separating style and content with bilinear models,”
Tenenbaum, J. B., and Freeman, W. T., 2000, · 2000
Earlier work this paper cites.
Product Design: Techniques in Reverse Engineering and New Product Development
Wood, K., and Otto, K., 2001, · 2001
Earlier work this paper cites.
“Unsupervised improvement of visual detectors using co-training,”
Levin, A., Viola, P., and Freund, Y., 2003, · 2003
Earlier work this paper cites.
“A study of prototypes, design activity, and design outcome,”
Yang, M. C., 2005, · 2005
Earlier work this paper cites.
“Learning visual representations using images with captions,”
Quattoni, A., Collins, M., and Darrell, T., 2007, · 2007
Earlier work this paper cites.
Improved Statistical Machine Translation for Resource-Poor Languages Using Related Resource-Rich Languages
Nakov, P., and Ng, H. T., 2009, · 2009
Earlier work this paper cites.
“How Uncertainty Helps Sketch Interpretation in a Design Task,”
Tseng, W. S., and Ball, L. J., 2011, · 2010
Earlier work this paper cites.
“Connecting modalities: Semi-supervised segmentation and annotation of images using unaligned text corpora,”
Socher, R., and Fei-Fei, L., 2010, · 2010
Earlier work this paper cites.
Visual Information in Semantic Representation
Feng, Y., and Lapata, M., 2010, · 2010
Earlier work this paper cites.
“Every picture tells a story: Generating sentences from images,”
Farhadi, A., Hejrati, M., Sadeghi, M. A., Young, P., Rashtchian, C., Hockenmaier, J., and Forsyth, D., 2010, · 2010
Earlier work this paper cites.
“Multimodal Deep Learning,”
Ngiam, J., Khosla, A., Kim, M., Nam, J., Lee, H., and Ng, A. Y., 2011, · 2011
Earlier work this paper cites.
“Im2Text: Describing Images Using 1 Million Captioned Photographs,”
Ordonez, V., Kulkarni, G., and Berg, T. L., 2011, · 2011
Earlier work this paper cites.
“Representation Learning: A Review and New Perspectives,”
Bengio, Y., Courville, A., and Vincent, P., 2012, · 2012
Earlier work this paper cites.
“Computer-aided design versus sketching: An exploratory case study,”
Veisz, D., Namouz, E. Z., Joshi, S., and Summers, J. D., 2012, · 2012
Earlier work this paper cites.
“Multimodal Learning with Deep Boltzmann Machines,”
Srivastava, N., and Salakhutdinov, R. R., 2012, · 2012
Earlier work this paper cites.
Distributional Semantics in Technicolor
Bruni, E., Boleda, G., Baroni, M., and Tran, N.-K., 2012, · 2012
Earlier work this paper cites.
“Multi-View Learning in the Presence of View Disagreement,”
Christoudias, C. M., Urtasun, R., and Darrell, T., 2012, · 2012
Earlier work this paper cites.
“AVA: A large-scale database for aesthetic visual analysis,”
Murray, N., Marchesotti, L., and Perronnin, F., 2012, · 2012
Earlier work this paper cites.
“Impact of Product Design Representation on Customer Judgment,”
Reid, T. N., MacDonald, E. F., and Du, P., 2013, · 2013
Earlier work this paper cites.
“DeViSE: A Deep Visual-Semantic Embedding Model,”
Frome, A., Corrado, G. S., Shlens, J., Bengio, S., Dean, J., Ranzato, M., and Mikolov, T., 2013, · 2013
Earlier work this paper cites.
Deep Canonical Correlation Analysis, 5
Andrew, G., Arora, R., Bilmes, J., and Livescu, K., 2013, · 2013
Earlier work this paper cites.
Learning Deep Structured Semantic Models for Web Search using Clickthrough Data, 10
Huang, P.-S., He, X., Gao, J., Deng, L., Acero, A., and Heck, L., 2013, · 2013
Earlier work this paper cites.
“Multiple kernel learning for emotion recognition in the wild,”
Sikka, K., Dykstra, K., Sathyanarayana, S., Littlewort, G., and Bartlett, M., 2013, · 2013
Earlier work this paper cites.
“Distributed Representations of Words and Phrases and their Compositionality,”
Mikolov, T., Chen, K., Corrado, G. S., Dean, J., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J., 2013, · 2013
Earlier work this paper cites.
“Zero-Shot Learning Through Cross-Modal Transfer,”
Socher, R., Ganjoo, M., Sridhar, H., Bastani, O., Manning, C. D., and Ng, A. Y., 2013, · 2013
Earlier work this paper cites.
“Framing Image Description as a Ranking Task: Data, Models and Evaluation Metrics,”
Hodosh, M., Young, P., and Hockenmaier, J., 2013, · 2013
Earlier work this paper cites.
“Data-intensive evaluation of design creativity using novelty, value, and surprise,”
Grace, K., Maher, M. L., Fisher, D., and Brady, K., 2014, · 2014
Earlier work this paper cites.
“Media and Representations in Product Design Education,”
BABAPOUR, M., ORNAS, V. H. A., REXFELT, O., and RAHE, U., 2014, · 2014
Earlier work this paper cites.
“Learning Grounded Meaning Representations with Autoencoders,”
Silberer, C., and Lapata, M., 2014, · 2014
Earlier work this paper cites.
“Cross-modal retrieval with correspondence autoencoder,”
Feng, F., Wang, X., and Li, R., 2014, · 2014
Earlier work this paper cites.
“Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models,”
Kiros, R., Salakhutdinov, R., and Zemel, R. S., 2014, · 2014
Earlier work this paper cites.
“Deep Visual-Semantic Alignments for Generating Image Descriptions,”
Karpathy, A., and Fei-Fei, L., 2014, · 2014
Earlier work this paper cites.
“Deep Fragment Embeddings for Bidirectional Image Sentence Mapping,”
Karpathy, A., Joulin, A., and Fei-Fei, L., 2014, · 2014
Earlier work this paper cites.
“Majority Vote of Diverse Classifiers for Late Fusion,”
Morvant, E., Habrard, A., and Ayache, S., 2014, · 2014
Earlier work this paper cites.
Graves, A., Wayne, G., and Danihelka, I., 2014, · 2014
Earlier work this paper cites.
“Generative Adversarial Networks,”
Tolstikhin, I., Bousquet, O., Schölkopf, B., Thierbach, K., Bazin, P. L., de Back, W., Gavriilidis, F., Kirilina, E., Jäger, C., Morawski, M., Geyer, S., Weiskopf, N., Scherf, N., Goodfellow, I. J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., Bengio, Y., Musk, E., Neuralink, Hjortsø, M. A., Wolenski, P., Ruder, S., Grathwohl, W., Chen, R. T. Q., Bettencourt, J., Sutskever, I., Duvenaud, D., and Doersch, C., 2014, · 2014
Earlier work this paper cites.
“Conditional Generative Adversarial Nets,”
Mirza, M., and Osindero, S., 2014, · 2014
Earlier work this paper cites.
“Show and Tell: A Neural Image Caption Generator,”
Vinyals, O., Toshev, A., Bengio, S., and Erhan, D., 2014, · 2014
Earlier work this paper cites.
“Grounded Compositional Semantics for Finding and Describing Images with Sentences,”
Socher, R., Karpathy, A., Le, Q. V., Manning, C. D., and Ng, A. Y., 2014, · 2014
Earlier work this paper cites.
“Microsoft COCO: Common Objects in Context,”
Lin, T. Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L., 2014, · 2014
Earlier work this paper cites.
“What are you talking about? Text-to-image coreference,”
Kong, C., Lin, D., Bansal, M., Urtasun, R., and Fidler, S., 2014, · 2014
Earlier work this paper cites.
“Connections Between the Design Tool, Design Attributes, and User Preferences in Early Stage Design,”
Häggman, A., Tsai, G., Elsen, C., Honda, T., and Yang, M. C., 2015, · 2015
Earlier work this paper cites.
“Representing analogies to influence fixation and creativity: A study comparing computer-aided design, photographs, and sketches,”
Atilola, O., and Linsey, J., 2015, · 2015
Earlier work this paper cites.
“Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models,”
Plummer, B. A., Wang, L., Cervantes, C. M., Caicedo, J. C., Hockenmaier, J., and Lazebnik, S., 2015, · 2015
Earlier work this paper cites.
“Neural Machine Translation by Jointly Learning to Align and Translate,”
Bahdanau, D., Cho, K., and Bengio, Y., 2014, · 2015
Earlier work this paper cites.
“Ask, Attend and Answer: Exploring Question-Guided Spatial Attention for Visual Question Answering,”
Xu, H., and Saenko, K., 2015, · 2015
Earlier work this paper cites.
“Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks,”
Ren, S., He, K., Girshick, R., and Sun, J., 2015, · 2015
Earlier work this paper cites.
“Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN),”
Mao, J., Xu, W., Yang, Y., Wang, J., Huang, Z., and Yuille, A., 2014, · 2015
Earlier work this paper cites.
“A Distributed Representation Based Query Expansion Approach for Image Captioning,”
Yagcioglu, S., Erdem, E., Erdem, A., and Çakici, R., 2015, · 2015
Earlier work this paper cites.
“Deep Unsupervised Learning using Nonequilibrium Thermodynamics,”
Sohl-Dickstein, J., Weiss, E. A., Maheswaranathan, N., and Ganguli, S., 2015, · 2015
Earlier work this paper cites.
“The Long-Short Story of Movie Description,”
Rohrbach, A., Rohrbach, M., and Schiele, B., 2015, · 2015
Earlier work this paper cites.
“Learning Visual Features from Large Weakly Supervised Data,”
Joulin, A., van Der Maaten, L., Jabri, A., and Vasilache, N., 2015, · 2015
Earlier work this paper cites.
“Grounding Semantics in Olfactory Perception,”
Kiela, D., Bulat, L., and Clark, S., 2015, · 2015
Earlier work this paper cites.
“Fast R-CNN,”
Girshick, R., 2015, · 2015
Earlier work this paper cites.
“Language Models for Image Captioning: The Quirks and What Works,”
Devlin, J., Cheng, H., Fang, H., Gupta, S., Deng, L., He, X., Zweig, G., and Mitchell, M., 2015, · 2015
Earlier work this paper cites.
“Jointly Modeling Deep Video and Compositional Text to Bridge Vision and Language in a Unified Framework,”
Xu, R., Xiong, C., Chen, W., and Corso, J. J., 2015, · 2015
Earlier work this paper cites.
“YFCC100M: The New Data in Multimedia Research,”
Thomee, B., Shamma, D. A., Friedland, G., Elizalde, B., Ni, K., Poland, D., Borth, D., and Li, L.-J., 2015, · 2015
Earlier work this paper cites.
“Effects of 3D CAD applications on the design creativity of students with different representational abilities,”
Chang, Y. S., Chien, Y. H., Lin, H. C., Chen, M. Y., and Hsieh, H. H., 2016, · 2016
Earlier work this paper cites.
“The effects of representation on idea generation and design fixation: A study comparing sketches and function trees,”
Atilola, O., Tomko, M., and Linsey, J. S., 2016, · 2016
Earlier work this paper cites.
“An Assessment of the Effectiveness of Sketch Representations in Early Stage Digital Design,”
Hannibal, C., Brown, A., and Knight, M., 2016, · 2016
Earlier work this paper cites.
“Bridge Correlational Neural Networks for Multilingual Multimodal Representation Learning,”
Rajendran, J., Khapra, M. M., Chandar, S., and Ravindran, B., 2015, · 2016
Earlier work this paper cites.
“Deep multimodal fusion for persuasiveness prediction,”
Nojavanasghari, B., Gopinath, D., Koushik, J., Baltrušaitis, T., and Morency, L. P., 2016, · 2016
Earlier work this paper cites.
“Black Holes and White Rabbits: Metaphor Identification with Visual Features,”
Shutova, E., Kiela, D., and Maillard, J., 2016, · 2016
Earlier work this paper cites.
“Deep visual-semantic hashing for cross-modal retrieval,”
Cao, Y., Long, M., Wang, J., Yang, Q., and Yuy, P. S., 2016, · 2016
Earlier work this paper cites.
“Compact Bilinear Pooling,”
Gao, Y., Beijbom, O., Zhang, N., and Darrell, T., 2015, · 2016
Earlier work this paper cites.
“Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding,”
Fukui, A., Park, D. H., Yang, D., Rohrbach, A., Darrell, T., and Rohrbach, M., 2016, · 2016
Earlier work this paper cites.
“Visual7W: Grounded question answering in images,”
Zhu, Y., Groth, O., Bernstein, M., and Fei-Fei, L., 2016, · 2016
Earlier work this paper cites.
“Where To Look: Focus Regions for Visual Question Answering,”
Shih, K. J., Singh, S., and Hoiem, D., 2015, · 2016
Earlier work this paper cites.
“Generating Images from Captions with Attention,”
Mansimov, E., Parisotto, E., Ba, J. L., and Salakhutdinov, R., 2015, · 2016
Earlier work this paper cites.
“Hierarchical Question-Image Co-Attention for Visual Question Answering,”
Lu, J., Yang, J., Batra, D., and Parikh, D., 2016, · 2016
Earlier work this paper cites.
“Stacked Attention Networks for Image Question Answering,”
Yang, Z., He, X., Gao, J., Deng, L., and Smola, A., 2015, · 2016
Earlier work this paper cites.
“Dynamic Memory Networks for Visual and Textual Question Answering,”
Xiong, C., Merity, S., and Socher, R., 2016, · 2016
Earlier work this paper cites.
“Multimodal Residual Learning for Visual QA,”
Kim, J. H., Lee, S. W., Kwak, D., Heo, M. O., Kim, J., Ha, J. W., and Zhang, B. T., 2016, · 2016
Earlier work this paper cites.
“Image Question Answering using Convolutional Neural Network with Dynamic Parameter Prediction,”
Noh, H., Seo, P. H., and Han, B., 2015, · 2016
Earlier work this paper cites.
“Generative Adversarial Text to Image Synthesis,”
Reed, S., Akata, Z., Yan, X., Logeswaran, L., Schiele, B., and Lee, H., 2016, · 2016
Earlier work this paper cites.
“StackGAN: Text to Photo-realistic Image Synthesis with Stacked Generative Adversarial Networks,”
Zhang, H., Xu, T., Li, H., Zhang, S., Wang, X., Huang, X., and Metaxas, D., 2016, · 2016
Earlier work this paper cites.
“Learning What and Where to Draw,”
Reed, S., Akata, Z., Mohan, S., Tenka, S., Schiele, B., and Lee, H., 2016, · 2016
Earlier work this paper cites.
“Generation and Comprehension of Unambiguous Object Descriptions,”
Mao, J., Huang, J., Toshev, A., Camburu, O., Yuille, A., and Murphy, K., 2015, · 2016
Earlier work this paper cites.
Density estimation using Real NVP, 7
Dinh, L., Sohl-Dickstein Google, J., Samy, B., and Google Brain, B., 2016, · 2016
Earlier work this paper cites.
“Improved Techniques for Training GANs,”
Salimans, T., Goodfellow, I., Zaremba, W., Cheung, V., Radford, A., and Chen, X., 2016, · 2016
Earlier work this paper cites.
“Learning Deep Representations of Fine-grained Visual Descriptions,”
Reed, S., Akata, Z., Lee, H., and Schiele, B., 2016, · 2016
Cited alongside, same era.
“Deep Compositional Captioning: Describing Novel Object Categories without Paired Training Data,”
Hendricks, L. A., Venugopalan, S., Rohrbach, M., Mooney, R., Saenko, K., and Darrell, T., 2015, · 2016
Cited alongside, same era.
“VisualWord2Vec (Vis-W2V): Learning Visually Grounded Word Embeddings Using Abstract Scenes,”
Kottur, S., Vedantam, R., Moura, J. M. F., and Parikh, D., 2016, · 2016
Cited alongside, same era.
“3D-R2N2: A unified approach for single and multi-view 3D object reconstruction,”
Choy, C. B., Xu, D., Gwak, J. Y., Chen, K., and Savarese, S., 2016, · 2016
Cited alongside, same era.
“Learning a Probabilistic Latent Space of Object Shapes via 3D Generative-Adversarial Modeling,”
Wu, J., Zhang, C., Xue, T., Freeman, W. T., and Tenenbaum, J. B., 2016, · 2016
Cited alongside, same era.
“Unsupervised Primitive Discovery for Improved 3D Generative Modeling,”
Khan, S. H., Guo, Y., Hayat, M., and Barnes, N., 2019, · 2019
Later among the works it cites.
“The situated function-behavior-structure co-design model,”
Gero, J., and Milovanovic, J., 2019, · 2019
Later among the works it cites.
“Multi-level 3D CNN for learning multi-scale spatial features,”
Ghadai, S., Lee, X. Y., Balu, A., Sarkar, S., and Krishnamurthy, A., 2019, · 2019
Later among the works it cites.
“Influence of Design Representation on Effectiveness of Idea Generation,”
McKoy, F. L., Vargas-Hernández, N., Summers, J. D., and Shah, J. J., 2020, · 2020
Later among the works it cites.
“Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-Training,”
Li, G., Duan, N., Fang, Y., Gong, M., and Jiang, D., 2020, · 2020
Later among the works it cites.
“Contrastive Learning of Medical Visual Representations from Paired Images and Text,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations,”
Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L. J., Shamma, D. A., Bernstein, M. S., and Fei-Fei, L., 2016, · 2016
Cited alongside, same era.
“How It Is Made Matters: Distinguishing Traits of Designs Created by Sketches, Prototypes, and CAD,”
Tsai, G., and Yang, M. C., 2017, · 2017
Cited alongside, same era.
“Deep Multimodal Representation Learning from Temporal Data,”
Yang, X., Ramesh, P., Chitta, R., Madhvanath, S., Bernal, E. A., and Luo, J., 2017, · 2017
Cited alongside, same era.
“Neural Architecture Search with Reinforcement Learning,”
Zoph, B., and Le, Q. V., 2016, · 2017
Cited alongside, same era.
“Tensor Fusion Network for Multimodal Sentiment Analysis,”
Zadeh, A., Chen, M., Cambria, E., Poria, S., and Morency, L. P., 2017, · 2017
Cited alongside, same era.
“Multi-modal Factorized Bilinear Pooling with Co-Attention Learning for Visual Question Answering,”
Yu, Z., Yu, J., Fan, J., and Tao, D., 2017, · 2017
Cited alongside, same era.
“Beyond Bilinear: Generalized Multimodal Factorized High-order Pooling for Visual Question Answering,”
Yu, Z., Yu, J., Xiang, C., Fan, J., and Tao, D., 2017, · 2017
Cited alongside, same era.
Zhang, Y., Jiang, H., Miura, Y., Manning, C. D., and Langlotz, C. P., 2020, · 2020
Later among the works it cites.
“A Review of Convolutional Neural Networks,”
Ajit, A., Acharya, K., and Samanta, A., 2020, · 2020
Later among the works it cites.
“An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,”
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N., 2020, · 2020
Later among the works it cites.
“Denoising Diffusion Probabilistic Models,”
Ho, J., Jain, A., and Abbeel, P., 2020, · 2020
Later among the works it cites.
“Denoising Diffusion Implicit Models,”
Song, J., Meng, C., and Ermon, S., 2020, · 2020
Later among the works it cites.
“Score-Based Generative Modeling through Stochastic Differential Equations,”
Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B., 2020, · 2020
Later among the works it cites.
“VirTex: Learning Visual Representations from Textual Annotations,”
Desai, K., and Johnson, J., 2020, · 2020
Later among the works it cites.
“Learning Visual Representations with Caption Annotations,”
Bulent Sariyildiz, M., Perez, J., Larlus, D., Sariyildiz, M. B., Perez, J., and Larlus, D., 2020, · 2020
Later among the works it cites.
“Image Captioning through Image Transformer,”
He, S., Liao, W., Tavakoli, H. R., Yang, M., Rosenhahn, B., and Pugeault, N., 2020, · 2020
Later among the works it cites.
“DF-GAN: A Simple and Effective Baseline for Text-to-Image Synthesis,”
Tao, M., Tang, H., Wu, F., Jing, X.-Y., Bao, B.-K., and Xu, C., 2020, · 2020
Later among the works it cites.
“Rank3DGAN: Semantic mesh generation using relative attributes,”
Saquil, Y., Xu, Q. C., Yang, Y. L., and Hall, P., 2020, · 2020
Later among the works it cites.
“A multimodal deep learning framework for predicting drug–drug interaction events,”
Deng, Y., Xu, X., Qiu, Y., Xia, J., Zhang, W., and Liu, S., 2020, · 2020
Later among the works it cites.
“Explainable Explainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI,”
Barredo Arrieta, A., Díaz-Rodríguez, N., Del Ser, J., Bennetot, A., Tabik, S., Barbado, A., Garcia, S., Gil-Lopez, S., Molina, D., Benjamins, R., Chatila, R., and Herrera, F., 2020, · 2020
Later among the works it cites.
“PcDGAN: A Continuous Conditional Diverse Generative Adversarial Network For Inverse Design,”
Nobari, A. H., Chen, W., and Ahmed, F., 2021, · 2021
Later among the works it cites.
“Guiding data-driven design ideation by knowledge distance,”
Luo, J., Sarica, S., and Wood, K. L., 2021, · 2021
Later among the works it cites.
“A Digital Twin-Driven Method for Product Performance Evaluation Based on Intelligent Psycho-Physiological Analysis,”
Feng, Y., Li, M., Lou, S., Zheng, H., Gao, Y., and Tan, J., 2021, · 2021
Later among the works it cites.
“Range-GAN: Range-Constrained Generative Adversarial Network for Conditioned Design Synthesis,”
Nobari, A. H., Chen, W., and Ahmed, F., 2021, · 2021
Later among the works it cites.
“Diffusion Models Beat GANs on Image Synthesis,”
Dhariwal, P., and Nichol, A., 2021, · 2021
Later among the works it cites.
“GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models,”
Nichol, A., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., McGrew, B., Sutskever, I., and Chen, M., 2021, · 2021
Later among the works it cites.
“DiffusionCLIP: Text-Guided Diffusion Models for Robust Image Manipulation,”
Kim, G., Kwon, T., and Ye, J. C., 2021, · 2021
Later among the works it cites.
“Multimodal Fusion with BERT and Attention Mechanism for Fake News Detection,”
Duc Tuan, N. M., and Quang Nhat Minh, P., 2021, · 2021
Later among the works it cites.
“Learning Transferable Visual Models From Natural Language Supervision,”
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I., 2021, · 2021
Later among the works it cites.
“Using DeepGCN to identify the autism spectrum disorder from multi-site resting-state data,”
Cao, M., Yang, M., Qin, C., Zhu, X., Chen, Y., Wang, J., and Liu, T., 2021, · 2021
Later among the works it cites.
“High-Resolution Image Synthesis with Latent Diffusion Models,”
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B., 2021, · 2021
Later among the works it cites.
“CLIP-Forge: Towards Zero-Shot Text-to-Shape Generation,”
Sanghi, A., Chu, H., Lambourne, J. G., Wang, Y., Cheng, C.-Y., Fumero, M., and Malekshan, K. R., 2021, · 2021
Later among the works it cites.
“A Survey of Convolutional Neural Networks: Analysis, Applications, and Prospects,”
Li, Z., Liu, F., Yang, W., Peng, S., and Zhou, J., 2021, · 2021
Later among the works it cites.
“AMMUS : A Survey of Transformer-based Pretrained Models in Natural Language Processing,”
Kalyan, K. S., Rajasekharan, A., and Sangeetha, S., 2021, · 2021
Later among the works it cites.
“Score-based Generative Modeling in Latent Space,”
Vahdat, A., Kreis, K., and Kautz, J., 2021, · 2021
Later among the works it cites.
“Diffusion Probabilistic Models for 3D Point Cloud Generation,”
Luo, S., and Hu, W., 2021, · 2021
Later among the works it cites.
“3D Shape Generation and Completion through Point-Voxel Diffusion,”
Zhou, L., Du, Y., and Wu, J., 2021, · 2021
Later among the works it cites.
“CogView: Mastering Text-to-Image Generation via Transformers,”
Ding, M., Yang, Z., Hong, W., Zheng, W., Zhou, C., Yin, D., Lin, J., Zou, X., Shao, Z., Yang, H., and Tang, J., 2021, · 2021
Later among the works it cites.
“3D Shape Generation with Grid-based Implicit Functions,”
Ibing, M., Lim, I., and Kobbelt, L., 2021, · 2021
Later among the works it cites.
“StyleGAN-NADA: CLIP-Guided Domain Adaptation of Image Generators,”
Gal, R., Patashnik, O., Maron, H., Bermano, A. H., Chechik, G., and Cohen-Or, D., 2021, · 2021
Later among the works it cites.
“Image-Based CLIP-Guided Essence Transfer,”
Chefer, H., Benaim, S., Paiss, R., and Wolf, L., 2021, · 2021
Later among the works it cites.
“Zero-Shot Text-to-Image Generation,”
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I., 2021, · 2021
Later among the works it cites.
“Vector-quantized Image Modeling with Improved VQGAN,”
Yu, J., Li, X., Koh, J. Y., Zhang, H., Pang, R., Qin, J., Ku, A., Xu, Y., Baldridge, J., and Wu, Y., 2021, · 2021
Later among the works it cites.
“CLIPDraw: Exploring Text-to-Drawing Synthesis through Language-Image Encoders,”
Frans, K., Soros, L. B., and Witkowski, O., 2021, · 2021
Later among the works it cites.
“MeshMVS: Multi-View Stereo Guided Mesh Reconstruction,”
Shrestha, R., Fan, Z., Su, Q., Dai, Z., Zhu, S., and Tan, P., 2020, · 2021
Later among the works it cites.
“Text2Mesh: Text-Driven Neural Stylization for Meshes,”
Michel, O., Bar-On, R., Liu, R., Benaim, S., and Hanocka, R., 2021, · 2021
Later among the works it cites.
“ClipMatrix: Text-controlled Creation of 3D Textured Meshes,”
Jetchev, N., 2021, · 2021
Later among the works it cites.
“Deeptake: Prediction of driver takeover behavior using multimodal data,”
Pakdamanian, E., Sheng, S., Baee, S., Heo, S., Kraus, S., and Feng, L., 2021, · 2021
Later among the works it cites.
“Parkinson’s Disease Detection Using CNN Architectures with Transfer Learning,”
Jahan, N., Nesa, A., and Layek, M. A., 2021, · 2021
Later among the works it cites.
“Explainable AI: A Review of Machine Learning Interpretability Methods,”
Linardatos, P., Papastefanopoulos, V., and Kotsiantis, S., 2021, · 2021
Later among the works it cites.
“Assessing Machine Learnability of Image and Graph Representations for Drone Performance Prediction,”
Song, B., McComb, C., and Ahmed, F., 2022, · 2022
Later among the works it cites.
“Deep Multi-modal Fusion of Image and Non-image Data in Disease Diagnosis and Prognosis: A Review,”
Cui, C., Yang, H., Wang, Y., Zhao, S., Asad, Z., Coburn, L. A., Wilson, K. T., Landman, B. A., and Huo, Y., 2022, · 2022
Later among the works it cites.
“Deep-Learning Methods of Cross-Modal Tasks for Conceptual Design of Product Shapes: A Review,”
Li, X., Wang, Y., and Sha, Z., 2022, · 2022
Later among the works it cites.
“Hey, ai! can you see what i see? multimodal transfer learning-based design metrics prediction for sketches with text descriptions,”
Song, B., Miller, S., and Ahmed, F., 2022, · 2022
Later among the works it cites.
“Leveraging End-User Data for Enhanced Design Concept Evaluation: A Multimodal Deep Regression Model,”
Yuan, C., Marion, T., and Moghaddam, M., 2022, · 2022
Later among the works it cites.
“IC3D: Image-Conditioned 3D Diffusion for Shape Generation,”
Sbrolli, C., Cudrano, P., Frosi, M., and Matteucci, M., 2022, · 2022
Later among the works it cites.
“Concise and Effective Network for 3D Human Modeling From Orthogonal Silhouettes,”
Liu, B., Liu, X., Yang, Z., and Wang, C. C. L., 2022, · 2022
Later among the works it cites.
Hadamard Product for Low-rank Bilinear Pooling, 7
Kim, J.-H., On, K.-W., Lim, W., Kim, J., Ha, J.-W., and Zhang, B.-T., 2022, · 2022
Later among the works it cites.
“Deep Learning for Technical Document Classification,”
Jiang, S., Hu, J., Magee, C. L., and Luo, J., 2022, · 2022
Later among the works it cites.
data2vec: A general framework for self-supervised learning in speech, vision and language
Baevski, A., Hsu, W.-N., Xu, Q., Babu, A., Gu, J., and Auli, M., 2022, · 2022
Later among the works it cites.
“End-to-End Transformer Based Model for Image Captioning,”
Wang, Y., Xu, J., and Sun, Y., 2022, · 2022
Later among the works it cites.
“LION: Latent Point Diffusion Models for 3D Shape Generation,”
Zeng, X., Vahdat, A., Williams, F., Gojcic, Z., Litany, O., Fidler, S., and Kreis, K., 2022, · 2022
Later among the works it cites.
“Classifier-Free Diffusion Guidance,”
Ho, J., and Salimans, T., 2022, · 2022
Later among the works it cites.
“Point-E: A System for Generating 3D Point Clouds from Complex Prompts,”
Nichol, A., Jun, H., Dhariwal, P., Mishkin, P., and Chen, M., 2022, · 2022
Later among the works it cites.
“Hierarchical Text-Conditional Image Generation with CLIP Latents,”
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M., 2022, · 2022
Later among the works it cites.
“Scaling Autoregressive Models for Content-Rich Text-to-Image Generation,”
Yu, J., Xu, Y., Koh, J. Y., Luong, T., Baid, G., Wang, Z., Vasudevan, V., Ku, A., Yang, Y., Ayan, B. K., Hutchinson, B., Han, W., Parekh, Z., Li, X., Zhang, H., Baldridge, J., and Wu, Y., 2022, · 2022
Later among the works it cites.
“Flow-based gan for 3d point cloud generation from a single image,”
Wei, Y., Vosselman, G., and Yang, M. Y., 2022, · 2022
Later among the works it cites.
“Explaining transformer-based image captioning models: An empirical analysis,”
Cornia, M., Baraldi, L., and Cucchiara, R., 2022, · 2022
Later among the works it cites.
“VQGAN-CLIP: Open Domain Image Generation and Editing with Natural Language Guidance,”
Crowson, K., Biderman, S., Kornis, D., Stander, D., Hallahan, E., Castricato, L., Raff, E., and Allen Hamilton, B., 2022, · 2022
Later among the works it cites.
“Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding,”
Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E., Ghasemipour, S. K. S., Ayan, B. K., Mahdavi, S. S., Lopes, R. G., Salimans, T., Ho, J., Fleet, D. J., and Norouzi, M., 2022, · 2022
Later among the works it cites.
“A Predictive and Generative Design Approach for Three-Dimensional Mesh Shapes Using Target-Embedding Variational Autoencoder,”
Li, X., Xie, C., and Sha, Z., 2022, · 2022
Later among the works it cites.
“XDGAN: Multi-Modal 3D Shape Generation in 2D Space,”
Alhaija, H. A., Dirik, A., Knörig, A., Fidler, S., and Shugrina, M., 2022, · 2022
Later among the works it cites.
“ShapeCrafter: A Recursive Text-Conditioned 3D Shape Generation Model,”
Fu, R., Zhan, X., Chen, Y., Ritchie, D., and Sridhar, S., 2022, · 2022
Later among the works it cites.
“Tetrahedral Diffusion Models for 3D Shape Generation,”
Kalischek, N., Peters, T., Wegner, J. D., and Schindler, K., 2022, · 2022
Later among the works it cites.
“Pre-train, Self-train, Distill: A simple recipe for Supersizing 3D Reconstruction,”
Alwala, K. V., Gupta, A., and Tulsiani, S., 2022, · 2022
Later among the works it cites.
“ISS: Image as Stepping Stone for Text-Guided 3D Shape Generation,”
Liu, Z., Dai, P., Li, R., Qi, X., and Fu, C.-W., 2022, · 2022
Later among the works it cites.
“3D-LDM: Neural Implicit 3D Shape Generation with Latent Diffusion Models,”
Nam, G., Khlifi, M., Rodriguez, A., Tono, A., Zhou, L., and Guerrero, P., 2022, · 2022
Later among the works it cites.
“Cross-Modal 3D Shape Generation and Manipulation,”
Cheng, Z., Chai, M., Ren, J., Lee, H.-Y., Olszewski, K., Huang, Z., Maji, S., and Tulyakov, S., 2022, · 2022
Later among the works it cites.
“Hybrid contrastive learning of tri-modal representation for multimodal sentiment analysis,”
Mai, S., Zeng, Y., Zheng, S., and Hu, H., 2022, · 2022
Later among the works it cites.
Multimodal fake news detection via clip-guided learning
Zhou, Y., Ying, Q., Qian, Z., Li, S., and Zhang, X., 2022, · 2022
Later among the works it cites.
“Enabling multi-modal search for inspirational design stimuli using deep learning,”
Kwon, E., Huang, F., and Goucher-Lambert, K., 2022, · 2022
Later among the works it cites.
“Deep Learning for Free-Hand Sketch: A Survey,”
Xu, P., Hospedales, T. M., Yin, Q., Song, Y.-Z., Xiang, T., and Wang, L., 2022, · 2022
Later among the works it cites.
“Research on the Design Strategy of Healing Products for Anxious Users during COVID-19,”
Wu, F., Lin, Y. C., and Lu, P., 2022, · 2022
Later among the works it cites.
“Biologically Inspired Design Concept Generation Using Generative Pre-Trained Transformers,”
Zhu, Q., Zhang, X., and Luo, J., 2023, · 2023
Closest in time.
“Generative Transformers for Design Concept Generation,”
Zhu, Q., and Luo, J., 2023, · 2023
Closest in time.
“ATTENTION-ENHANCED MULTIMODAL LEARNING FOR CONCEPTUAL DESIGN EVALUATIONS,”
Song, B., Associate, P., Miller, S., and Ahmed, F., 2023, · 2023
Closest in time.
“DDE-GAN: Integrating a Data-Driven Design Evaluator Into Generative Adversarial Networks for Desirable and Diverse Concept Generation,”
Yuan, C., Marion, T., and Moghaddam, M., 2023, · 2023
Closest in time.
“Beyond Statistical Similarity: Rethinking Metrics for Deep Generative Models in Engineering Design,”
Regenwetter, L., Srivastava, A., Gutfreund, D., and Ahmed, F., 2023, · 2023
Closest in time.
“StyleCLIP: Text-Driven Manipulation of StyleGAN Imagery,”
Patashnik, O., Wu, Z., Shechtman, E., Cohen-Or, D., and Lischinski, D., 2021, · 2074
Closest in time.