Fetching the paper…
Reading the bibliography…
Deep learning methods have revolutionized speech recognition, image recognition, and natural language processing since 2010.
L. Tucker, “Some mathematical notes on three-mode factor analy,”
1966
Earlier work this paper cites.
D. Rumelhart, G. Hinton, and R. Williams, “Learning representations by back-propagating errors,”
1986
Earlier work this paper cites.
N. Ström, “Speaker adaptation by modeling the speaker variation in a continuous speech recognition system,” in
1996
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,”
1997
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,”
1998
Earlier work this paper cites.
B. Maison, C. Neti, and A. Senior, “Audio-visual speaker recognition for video broadcast news: Some fusion techniques,” in
1999
Earlier work this paper cites.
S. Dupont and J. Luettin, “Audio-visual speech modeling for continuous speech recognition,”
2000
Earlier work this paper cites.
J. Tenenbaum and W. Freeman, “Separating style and content with bilinear models,”
2000
Earlier work this paper cites.
Y. Bengio, R. Ducharme, P. Vincent, and C. Jauvin, “A neural probabilistic language model,”
2003
Earlier work this paper cites.
Z. Wu, L. Cai, and H. Meng, “Multi-level fusion of audio and visual features for speaker identification,” in
2005
Earlier work this paper cites.
G. Hinton and R. Salakhutdinov, “Reducing the dimensionality of data with neural networks,”
2006
Earlier work this paper cites.
M. Cookea, J. Barker, S. Cunningham, and X. Shao, “An audio-visual corpus for speech perception and automatic speech recognition,”
2006
Earlier work this paper cites.
M.-E. Nilsback and A. Zisserman, “A visual vocabulary for flower classification,” in
2006
Earlier work this paper cites.
L. Lathauwer, “Decompositions of a higher-order tensor in block terms—part II: Definitions and uniqueness,”
2008
Earlier work this paper cites.
Y. Bengio, “Learning deep architectures for AI,”
2009
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A large-scale hierarchical image database,” in
2009
Earlier work this paper cites.
D. Yu, L. Deng, and G. Dahl, “Roles of pre-training and fine-tuning in context-dependent DBN-HMMs for real-world speech recognition,”
2010
Earlier work this paper cites.
L. Deng, M. Seltzer, D. Yu, A. Acero, A. Mohamed, and G. Hinton, “Binary coding of speech spectrograms using a deep autoencoder,”
2010
Earlier work this paper cites.
P. Welinder, S. Branson, T. Mita, C. Wah, F. Schroff, S. Belongie, and P. Perona, “Caltech-UCSD birds 200,” Tech. Rep. CNS-TR-2010-001, California Institute of Technology, 2010
2010
Earlier work this paper cites.
L. Deng, “An overview of deep-structured learning for information processing,” in
2011
Earlier work this paper cites.
D. Yu, L. Deng, F. Seide, and G. Li, “Discriminative pre-training of deep nerual networks,” in
2011
Earlier work this paper cites.
G. Dahl, D. Yu, and L. Deng, “Large-vocabulry continuous speech recognition with context-dependent DBN-HMMs,” in
2011
Earlier work this paper cites.
F. Seide, L. Gang, and Y. Dong, “Conversational speech transcription using context-dependent deep neural networks,” in
2011
Earlier work this paper cites.
N. Dehak, P. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, “Front-end factor analysis for speaker verification,”
2011
Earlier work this paper cites.
J. Ngiam, A. Khosla, M. Kim, J. Nam, H. Lee, and A. Ng, “Multimodal deep learning,” in
2011
Earlier work this paper cites.
G. Dahl, D. Yu, L. Deng, and A. Acero, “Context-dependent pre-trained deep neural networks for large-vocabulary speech recognition,”
2012
Earlier work this paper cites.
G. Hinton, L. Deng, Y. Dong, G. Dahl, A.-R. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. Sainath, and B. Kingsbury, “Deep neural networks for acoustic modeling in speech recognition,”
2012
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. Hinton, “ImageNet classification with deep convolutional neural networks,” in
2012
Earlier work this paper cites.
N. Srivastava and R. Salakhutdinov, “Multimodal learning with deep boltzmann machines,” in
2012
Earlier work this paper cites.
E. Bruni, G. Boleda, M. Baroni, and N.-K. Tran, “Distributional semantics in technicolor,” in
2012
Earlier work this paper cites.
M. Charikar, K. Chen, and M. Farach-Colton, “Finding frequent items in data streams,” in
2012
Earlier work this paper cites.
L. Deng, J. Li, J.-T. Huang, K. Yao, D. Yu, F. Seide, M. Seltzer, G. Zweig, X. He, J. Williams, Y. Gong, and A. Acero, “Recent advances in deep learning for speech research at Microsoft,” in
2013
Earlier work this paper cites.
L. Deng, G. Hinton, and B. Kingsbury, “New types of deep neural network learning for speech recognition and related applications: An overview,”
2013
Earlier work this paper cites.
Y. Bengio, A. Courville, and P. Vincent, “Representation learning: A review and new perspectives,”
2013
Earlier work this paper cites.
P.-S. Huang, X. He, G. J., L. Deng, A. Acero, and L. Heck, “Learning deep structured semantic models for web search using clickthrough data,” in
2013
Earlier work this paper cites.
T. Mikolov, K. Chen, G. Corrado, and J. Dean, “Efficient estimation of word representations in vector space,” in
2013
Earlier work this paper cites.
T. Mikolov, I. Sutskever, K. Chen, G. Corrado, and J. Dean, “Distributed representations of words and phrases and their compositionality,” in
2013
Earlier work this paper cites.
T. Mikolov, W. Yih, and G. Zweig, “Linguistic regularities in continuous space word representations,” in
2013
Earlier work this paper cites.
M. Karafiát, L. Burget, P. Matějka, O. Glembek, and J. Černocký, “iVector-based discriminative adaptation for automatic speech recognition,” in
2013
Earlier work this paper cites.
G. Saon, H. Soltau, D. Nahamoo, and M. Picheny, “Speaker adaptation of neural network acoustic models using i-vectors,” in
2013
Earlier work this paper cites.
O. Abdel-Hamid and H. Jiang, “Fast speaker adaptation of hybrid NN/HMM model for speech recognition based on discriminative learning of speaker code,” in
2013
Earlier work this paper cites.
J. Portêlo, A. Abad, B. Raj, and I. Trancoso, “Secure binary embeddings of front-end factor analysis for privacy preserving speaker verification,” in
2013
Earlier work this paper cites.
R. Socher, M. Ganjoo, H. Sridhar, O. Bastani, C. Manning, and A. Ng, “Zero-shot learning through cross-modal transfer,” in
2013
Earlier work this paper cites.
A. Frome, G. Corrado, J. Shlens, S. Bengio, J. Dean, M. Ranzato, and T. Mikolov, “DeViSE: A deep visual-semantic embedding model,” in
2013
Earlier work this paper cites.
N. Pham and R. Pagh, “Fast and scalable polynomial kernels via explicit feature maps,” in
2013
Earlier work this paper cites.
A. Graves, “Generating sequences with recurrent neural networks,” in
2013
Earlier work this paper cites.
L. Deng and Y. Dong, “Deep Learning: Methods and Applications,”
2014
Earlier work this paper cites.
R. Girshick, J. Donahue, T. Darrell, and J. Malik, “Rich feature hierarchies for accurate object detection and semantic segmentation,” in
2014
Earlier work this paper cites.
I. Sutskever, O. Vinyals, and Q. Le, “Sequence to sequence learning with neural networks,” in
2014
Earlier work this paper cites.
S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg, “Referitgame: Referring to objects in photographs of natural scenes,” in
2014
Earlier work this paper cites.
Y. Shen, X. He, J. Gao, L. Deng, and G. Mesnil, “Learning semantic representations using convolutional neural networks for web search,” in
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
J. Pennington, R. Socher, and C. Manning, “Glove: Global vectors for word representation,” in
2014
Earlier work this paper cites.
W.-T. Yih, X. He, and C. Meek, “Semantic parsing for single-relation question answering,” in
2014
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in
2014
Earlier work this paper cites.
A. Senior and I. Lopez-Moreno, “Improving DNN speaker independence with i-vector inputs,” in
2014
Earlier work this paper cites.
E. Variani, X. Lei, E. McDermott, I. Moreno, and J. Gonzalez-Dominguez, “Deep neural networks for small footprint text- dependent speaker verification,” in
2014
Earlier work this paper cites.
C. Silberer and M. Lapata, “Learning grounded meaning representations with autoencoders,” in
2014
Earlier work this paper cites.
A. Karpathy, A. Joulin, and F.-F. Li, “Deep fragment embeddings for bidirectional image sentence mapping,” in
2014
Earlier work this paper cites.
A. Graves, G. Wayne, and I. Danihelka, “Neural turing machines,” in
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
M. Mirza and S. Osindero, “Conditional generative adversarial nets,” in
2014
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, L. Bourdev, R. Girshick, J. Hays, P. Perona, D. Ramanan, C. Zitnick, and P. Dollár, “Microsoft COCO: Common objects in context,” in
2014
Earlier work this paper cites.
M. Malinowski and M. Fritz, “A multi-world approach to question answering about real-world scenes based on yncertain input,” in
2014
Earlier work this paper cites.
J. Schmidhuber, “Deep learning in neural networks: An overview,”
2015
Earlier work this paper cites.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,”
2015
Earlier work this paper cites.
Springer, 2015
D. Yu and L. Deng, · 2015
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in
2015
Earlier work this paper cites.
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in
2015
Earlier work this paper cites.
R. Girshick, “Fast R-CNN,” in
2015
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards real-time object detection with region proposal networks,” in
2015
Earlier work this paper cites.
G. Mesnil, Y. Dauphin, K. Yao, Y. Bengio, L. Deng, D. Hakkani-Tur, X. He, L. Heck, G. Tur, D. Yu, and G. Zweig, “Using recurrent neural networks for slot filling in spoken language understanding,”
2015
Earlier work this paper cites.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in
2015
Earlier work this paper cites.
M.-T. Luong, H. Pham, and C. Manning, “Effective approaches to attention-based neural machine translation,” in
2015
Earlier work this paper cites.
B. Plummer, L. Wang, C. Cervantes, J. Caicedo, J. Hockenmaier, and L. S., “Flickr30k entities: Collecting region-to phrase correspondences for richer image-to-sentence models,” in
2015
Earlier work this paper cites.
A. Karpathy and L. Fei-Fei, “Deep visual-semantic alignments for generating image descriptions,” in
2015
Earlier work this paper cites.
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, “Show and tell: A neural image caption generator,” in
2015
Earlier work this paper cites.
D. Geman, S. Geman, N. Hallonquist, and L. Younes, “Visual Turing test for computer vision systems,” in
2015
Earlier work this paper cites.
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batral, C. Zitnick, and D. Parikh, “VQA: Visual question answering,” in
2015
Earlier work this paper cites.
L. Yu, E. Park, A. Berg, and T. Berg, “Visual Madlibs: Fill in the blank description generation and question answering,” in
2015
Earlier work this paper cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. Berg, and F.-F. Li, “ImageNet large scale visual recognition challenge,”
2015
Earlier work this paper cites.
A. Elkahky, Y. Song, and X. He, “A multi-view deep learning approach for cross domain user modeling in recommendation systems,” in
2015
Earlier work this paper cites.
X. Liu, J. Gao, X. He, L. Deng, K. Duh, and Y.-Y. Wang, “Representation learning using multi-task deep neural networks for semantic classification and information retrieval,” in
2015
Earlier work this paper cites.
W.-T. Yih, M.-W. Chang, X. He, and J. Gao, “Semantic parsing via staged query graph generation: Question answering with knowledge base,” in
2015
Earlier work this paper cites.
R. Kiros, Y. Zhu, R. Salakhutdinov, R. Zemel, A. Torralba, R. Urtasun, and S. Fidler, “Skip-thought vectors,” in
2015
Earlier work this paper cites.
S. Yella and A. Stolcke, “A comparison of neural network feature transforms for speaker diarization,” in
2015
Earlier work this paper cites.
Z. Wu, P. Swietojanski, C. Veaux, S. Renals, and S. King, “A study of speaker adaptation for DNN-based speech synthesis,” in
2015
Earlier work this paper cites.
H. Fang, S. Gupta, F. Iandola, R. K. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. C. Platt,
2015
Earlier work this paper cites.
A. Lazaridou, N. Pham, and M. Baroni, “Combining language and vision with a multimodal skip-gram model,” in
2015
Earlier work this paper cites.
J. Ba, K. Swersky, S. Fidler, and R. Salakhutdinov, “Predicting deep zero-shot convolutional neural networks using textual descriptions,” in
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in
2015
Cited alongside, same era.
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhutdinov, R. Zemel, and Y. Bengio, “Show, attend and tell: Neural image caption generation with visual attention,” in
2015
Cited alongside, same era.
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, “Show and tell: A neural image caption generator,” in
2015
Cited alongside, same era.
X. Chen and C. Lawrence Zitnick, “Mind’s eye: A recurrent visual representation for image caption generation,” in
2015
Cited alongside, same era.
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhudinov, R. Zemel, and Y. Bengio, “Show, attend and tell: Neural image caption generation with visual attention,” in
2015
Cited alongside, same era.
T. Afouras, J. Chung, and A. Zisserman, “The conversation: Deep audio-visual speech enhancement,” in
2018
Later among the works it cites.
Q. Wang, C. Downey, L. Wan, P. Mansfield, and I. Lopez Moreno, “Speaker diarization with LSTM,” in
2018
Later among the works it cites.
M. Sarma, P. Ghahremani, D. Povey, N. Goel, K. Sarma, and N. Dehak, “Emotion identification from raw speech signals using DNNs,” in
2018
Later among the works it cites.
Y. Zhu, T. Ko, D. Snyder, B. Mak, and D. Povey, “Self-attentive speaker embeddings for text-independent speaker verification,” in
2018
Later among the works it cites.
K.-H. Lee, X. Chen, G. Hua, H. Hu, and X. He, “Stacked cross attention for image-text matching,” in
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. Gregor, I. Danihelka, A. Graves, D. Rezende, and D. Wierstra, “DRAW: A recurrent neural network for image generation,” in
2015
Cited alongside, same era.
E. Denton, S. Chintala, A. Szlam, and R. Fergus, “Deep generative image models using a laplacian pyramid of adversarial networks,” in
2015
Cited alongside, same era.
M. Malinowski, M. Rohrbach, and M. Fritz, “Ask your neurons: A neural-based approach to answering questions about images,” in
2015
Cited alongside, same era.
M. Ren, R. Kiros, and R. Zemel, “Exploring models and data for image question answering,” in
2015
Cited alongside, same era.
Google, “Freebase data dumps,” in
2015
Cited alongside, same era.
The MIT Press, 2016
I. Goodfellow, Y. Bengio, and A. Courville, · 2016
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in
2016
Cited alongside, same era.
Y.-H. Tsai, P. Liang, A. Zadeh, L.-P. Morency, and R. Salakhutdinov, “Learning factorized multimodal representations,” in
2018
Later among the works it cites.
V. Vielzeuf, A. Lechervy, S. Pateux, and F. Jurie, “CentralNet: A multilayer approach for multimodal fusion,” in
2018
Later among the works it cites.
C. Liu, B. Zoph, M. Neumann, J. Shlens, W. Hua, L.-J. Li, F.-F. Li, A. Yuille, J. Huang, and K. Murphy, “Progressive neural architecture search,” in
2018
Later among the works it cites.
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang, “Bottom-up and top-down attention for image captioning and visual question answering,” in
2018
Later among the works it cites.
P. Lu, H. Li, W. Zhang, J. Wang, and X. Wang, “Co-attending free-form regions and detections with multi-modal multiplicative feature embedding for visual question answering,” in
2018
Later among the works it cites.
H. Fan and J. Zhou, “Stacked latent attention for multimodal reasoning,” in
2018
Later among the works it cites.
Z. Yu, J. Yu, C. Xiang, J. Fan, and D. Tao, “Beyond bilinear: Generalized multimodal factorized high-order pooling for visual question answering,”
2018
Later among the works it cites.
Z. Liu, Y. Shen, V. Lakshminarasimhan, P. Liang, A. Zadeh, and L.-P. Morency, “Efficient low-rank multimodal fusion with modality-specific factors,” in
2018
Later among the works it cites.
J.-H. Kim, J. Jun, and B.-T. Zhang, “Bilinear attention networks,” in
2018
Later among the works it cites.
2018
Later among the works it cites.
Z. Zhang, Y. Xie, and L. Yang, “Photographic text-to-image synthesis with a hierarchically-nested adversarial network,” in
2018
Later among the works it cites.
T. Karras, T. Aila, S. Laine, and J. Lehtinen, “Progressive growing of GANs for improved quality, stability, and variation,” in
2018
Later among the works it cites.
J. Johnson, A. Gupta, and F.-F. Li, “Image generation from scene graphs,” in
2018
Later among the works it cites.
S. Hong, D. Yang, J. Choi, and H. Lee, “Inferring semantic layout for hierarchical text-to-image synthesis,” in
2018
Later among the works it cites.
S. Nam, Y. Kim, and S. Kim, “Text-adaptive generative adversarial networks: Manipulating images with natural language,” in
2018
Later among the works it cites.
S. Sharma, D. Suhubdy, V. Michalski, S. Kahou, and Y. Bengio, “ChatPainter: Improving text to image generation using dialogue,” in
2018
Later among the works it cites.
P. Cascante-Bonilla, X. Yin, V. Ordonez, and S. Feng, “Chat-crowd: A dialog-based platform for visual layout composition,” in
2018
Later among the works it cites.
Y. Li, M. Min, D. Shen, D. Carlson, and L. Carin, “Video generation from text,” in
2018
Later among the works it cites.
Q. Wu, P. Wang, C. Shen, I. Reid, and A. van den Hengel, “Are you talking to me? Reasoned visual dialog generation through adversarial learning,” in
2018
Later among the works it cites.
P. Wang, Q. Wu, C. Shen, A. van den Hengel, and A. Dick, “FVQA: Fact-based visual question answering,”
2018
Later among the works it cites.
A. Agrawal, D. Batra, D. Parikh, and A. Kembhavi, “Don’t just assume; Look and answer: Overcoming priors for visual question answering,” in
2018
Later among the works it cites.
S. Ramakrishnan, A. Agrawal, and S. Lee, “Overcoming language priors in visual question answering with adversarial regularization,” in
2018
Later among the works it cites.
Y. Zhang, J. Hare, and A. Prügel-Bennett, “Learning to count objects in natural images for visual question answering,” in
2018
Later among the works it cites.
D. Gurari, Q. Li, A. Stangl, A. Guo, C. Lin, K. Grauman, J. Luo, and J. Bigham, “VizWiz grand challenge: Answering visual questions from blind people,” in
2018
Later among the works it cites.
I. Misra, R. Girshick, R. Fergus, M. Hebert, A. Gupta, and L. van der Maaten, “Learning by asking questions,” in
2018
Later among the works it cites.
E. Perez, F. Strub, H. de Vries, V. Dumoulin, and A. Courville, “FiLM: Visual reasoning with a general conditioning layer,” in
2018
Later among the works it cites.
R. Hu, J. Andreas, T. Darrell, and K. Saenko, “Explainable neural computation via stack neural module networks,” in
2018
Later among the works it cites.
D. Mascharka, P. Tran, R. Soklaski, and A. Majumdar, “Transparency by design: Closing the gap between performance and interpretability in visual reasoning,” in
2018
Later among the works it cites.
D. Hudson and C. Manning, “Compositional attention networks for machine reasoning,” in
2018
Later among the works it cites.
K. Yi, J. Wu, C. Gan, A. Torralba, P. Kohli, and J. Tenenbaum, “Neural-symbolic VQA: Disentangling reasoning from vision and language understanding,” in
2018
Later among the works it cites.
R. Vedantam, K. Desai, S. Lee, M. Rohrbach, D. Batra, and D. Parikh, “Probabilistic neural-symbolic models for interpretable visual question answering,” in
2018
Later among the works it cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in
2019
Closest in time.
ACM and Morgan & Claypool Publishers, 2019
S. Bengio, L. Deng, L. Morency, and B. Schuller, · 2019
Closest in time.
J. Chung, B.-J. Lee, and I. Han, “Who said that?: Audio-visual speaker diarisation of real-world meetings,” in
2019
Closest in time.
J. Wu, Y. Xu, S.-X. Zhang, L.-W. Chen, M. Yu, L. Xie, and D. Yu, “Time domain audio visual speech separation,” in
2019
Closest in time.
2019
Closest in time.
G. Sun, C. Zhang, and P. Woodland, “Speaker diarisation using 2D self-attentive combination of embeddings,” in
2019
Closest in time.
A. Nautsch, A. Jiménez, A. Treiber, J. Kolberg, C. Jasserand, E. Kindt, H. Delgado, M. Todisco, M. Hmani, A. Mtibaa, M. Abdelraheem, A. Abad, F. Teixeira, D. Matrouf, M. Gomez-Barrero, G. Petrovska-Delacrétaz, Chollet, N. Evans, T. Schneider, J.-F. Bonastre, B. Raj, I. Trancoso, and C. Busch, “Preserving privacy in speaker and speech characterisation,”
2019
Closest in time.
P. Bachman, R. Hjelm, and W. Buchwalter, “Learning representations by maximizing mutual information across views,” in
2019
Closest in time.
H. Wu, J. Mao, Y. Zhang, Y. Jiang, L. Li, W. Sun, and W.-Y. Ma, “Unified visual-semantic embeddings: Bridging vision and language with structured meaning representations,” in
2019
Closest in time.
T. Gupta, A. Schwing, and D. Hoiem, “ViCo: Word embeddings from visual co-occurrences,” in
2019
Closest in time.
D.-K. Nguyen and T. Okatani, “Multi-task learning of hierarchical vision-language representation,” in
2019
Closest in time.
2019
Closest in time.
2019
Closest in time.
2019
Closest in time.
C. Sun, A. Myers, C. Vondrick, K. Murphy, and C. Schmid, “VideoBERT: A joint model for video and language representation learning,” in
2019
Closest in time.
C. Alberti, J. Ling, M. Collins, and D. Reitter, “Fusion of detected objects in text for visual question answering,” in
2019
Closest in time.
H. Tan and B. Mohit, “LXMERT: Learning cross-modality encoder representations from transformers,” in
2019
Closest in time.
J. Lu, D. Batra, D. Parikh, and S. Lee, “ViLBERT: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,” in
2019
Closest in time.
2019
Closest in time.
X. Liu, P. He, W. Chen, and J. Gao, “Multi-task deep neural networks for natural language understanding,” in
2019
Closest in time.
A. Anastasopoulos, S. Kumar, and H. Liao, “Neural language modeling with visual features,” in
2019
Closest in time.
J.-M. Pérez-Rúa, V. Vielzeuf, S. Pateux, M. Baccouche, and F. Jurie, “MFAS: Multimodal fusion architecture search,” in
2019
Closest in time.
J.-M. Pérez-Rúa, M. Baccouche, and S. Pateux, “Efficient progressive neural architecture search,” in
2019
Closest in time.
W. Li, P. Zhang, L. Zhang, Q. Huang, X. He, S. Lyu, and J. Gao, “Object-driven text-to-image synthesis via adversarial training,” in
2019
Closest in time.
A. Osman and W. Samek, “DRAU: Dual recurrent attention units for visual question answering,”
2019
Closest in time.
H. Ben-younes, R. Cadene, N. Thome, and M. Cord, “BLOCK: Bilinear superdiagonal fusion for visual question answering and visual relationship detection,” in
2019
Closest in time.
X. Wang, Q. Huang, A. Celikyilmaz, J. Gao, D. Shen, Y.-F. Wang, W. Wang, and L. Zhang, “Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation,” in
2019
Closest in time.
A. Deshpande, J. Aneja, L. Wang, A. G. Schwing, and D. Forsyth, “Fast, diverse and accurate image captioning guided by part-of-speech,” in
2019
Closest in time.
H. Zhang, T. Xu, H. Li, S. Zhang, X. Wang, X. Huang, and D. Metaxas, “StackGAN++: Realistic image synthesis with stacked generative adversarial networks,”
2019
Closest in time.
M. Zhu, P. Pan, W. Chen, and Y. Yang, “DM-GAN: Dynamic memory generative adversarial networks for text-to-image synthesis,” in
2019
Closest in time.
M. Cha, Y. Gwon, and H. Kung, “Adversarial learning of semantic relevance in text to image synthesis,” in
2019
Closest in time.
X. Chen, M. Rohrbach, and D. Parikh, “Cycle-consistency for robust visual question answering,” in
2019
Closest in time.
T. Qiao, J. Zhang, D. Xu, and D. Tao, “MirrorGAN: Learning text-to-image generation by redescription,” in
2019
Closest in time.
B. Zhao, L. Meng, W. Yin, and L. Sigal, “Image generation from layout,” in
2019
Closest in time.
T. Hinz, S. Heinrich, and S. Wermter, “Generating multiple objects at spatially distinct locations,” in
2019
Closest in time.
Z. He, W. Zuo, M. Kan, S. Shan, and X. Chen, “AttGAN: Facial attribute editing by only changing what you want,”
2019
Closest in time.
Q. Lao, M. Havaei, A. Pesaranghader, F. Dutil, L. Jorio, and T. Fevens, “Dual adversarial inference for text-to-image synthesis,” in
2019
Closest in time.
F. Tan, S. Feng, and V. Ordonez, “Text2Scene: Generating compositional scenes from textual descriptions,” in
2019
Closest in time.
A. El-Nouby, S. Sharma, H. Schulz, D. Hjelm, L. Asri, S. Kahou, Y. Bengio, and G. Taylor, “Tell, draw, and repeat: Generating and modifying images based on continual linguistic instruction,” in
2019
Closest in time.
Y. Chen, Z. Gan, Y. Li, J. Liu, and J. Gao, “Sequential attention GAN for interactive image editing via dialogue,” in
2019
Closest in time.
J.-H. Kim, N. Kitaev, X. Chen, M. Rohrbach, B.-T. Zhang, Y. Tian, D. Batra, and D. Parikh, “CoDraw: Collaborative drawing as a testbed for grounded goal-driven communication,” in
2019
Closest in time.
Y. Li, Z. Gan, Y. Shen, J. Liu, Y. Cheng, Y. Wu, L. Carin, D. Carlson, and J. Gao, “StoryGAN: A sequential conditional GAN for story visualization,” in
2019
Closest in time.
Y. Balaji, M. Min, B. Bai, R. Chellappa, and H. Graf, “Conditional GAN with discriminative filter generation for text-to-video synthesis,” in
2019
Closest in time.
Y. Goyal, T. Khot, A. Agrawal, D. Summers-Stay, D. Batra, and D. Parikh, “Making the V in VQA matter: Elevating the role of image understanding in visual question answering,”
2019
Closest in time.
K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi, “OK-VQA: A visual question answering benchmark requiring external knowledge,” in
2019
Closest in time.
D. Hudson and C. Manning, “GQA: A new dataset for real-world visual reasoning and compositional question answering,” in
2019
Closest in time.
R. Cadene, C. Dancette, H. Ben-younes, M. Cord, and D. Parikh, “RUBi: Reducing unimodal biases in visual question answering,” in
2019
Closest in time.
A. Singh, V. Natarajan, M. Shah, Y. Jiang, X. Chen, D. Batra, D. Parikh, and M. Rohrbach, “Towards vqa models that can read,” in
2019
Closest in time.
2019
Closest in time.
R. Cadene, H. Ben-younes, M. Cord, and N. Thome, “MUREL: Multimodal relational reasoning for visual question answering,” in
2019
Closest in time.
J. Mao, C. Gan, P. Kohli, J. Tenenbaum, and J. Wu, “The neuro-symbolic concept learner: Interpreting scenes, words, and sentences from natural supervision,” in
2019
Closest in time.
2019
Closest in time.
I. Chaturvedi, R. Satapathy, S. Cavallari, and E. Cambria, “Fuzzy commonsense reasoning for multimodal sentiment analysis,”
2019
Closest in time.
2020
Closest in time.
D. Golub, R. Martín-Martín, A. El-Kishky, and S. Savarese, “Leveraging pretrained image classifiers for language-based segmentation,” in
2020
Closest in time.