Fetching the paper…
Reading the bibliography…
Interest in Artificial Intelligence (AI) and its applications has seen unprecedented growth in the last few years.
Visual entailment: A novel task for fine-grained image understanding
Xie, N., Lai, F., Doran, D., and Kadav, A. (2019) · 1901
Earlier work this paper cites.
GQA: A new dataset for compositional question answering over real-world images
Hudson, D. A., and Manning, C. D. (2019) · 1902
Earlier work this paper cites.
Self-supervised visual feature learning with deep neural networks: A survey
Jing, L., and Tian, Y. (2019) · 1902
Earlier work this paper cites.
Cycle-consistency for robust visual question answering
Shah, M., Chen, X., Rohrbach, M., and Parikh, D. (2019a) · 1902
Earlier work this paper cites.
Probabilistic neural-symbolic models for interpretable visual question answering
Vedantam, R., Desai, K., Lee, S., Rohrbach, M., Batra, D., and Parikh, D. (2019) · 1902
Earlier work this paper cites.
Visual semantic information pursuit: A survey
Liu, D., Bober, M., and Kittler, J. (2019) · 1903
Earlier work this paper cites.
Recent advances in natural language inference: A survey of benchmarks, resources, and approaches
Storks, S., Gao, Q., and Chai, J. Y. (2019) · 1904
Earlier work this paper cites.
Learning to compose and reason with language tree structures for visual grounding
Hong, R., Liu, D., Mo, X., He, X., and Zhang, H. (2019) · 1906
Earlier work this paper cites.
Zhu, D., Mogadala, A., and Klakow, D. (2019) · 1912
Earlier work this paper cites.
Doubly-attentive decoder for multi-modal neural machine translation
Calixto, I., Liu, Q., and Campbell, N. (2017) · 1924
Earlier work this paper cites.
Learning distributions over logical forms for referring expression generation
FitzGerald, N., Artzi, Y., and Zettlemoyer, L. (2013) · 1925
Earlier work this paper cites.
Phrase localization and visual relationship detection with comprehensive image-language cues
Plummer, B. A., Mallya, A., Cervantes, C. M., Hockenmaier, J., and Lazebnik, S. (2017a) · 1937
Earlier work this paper cites.
Raven’s progressive matrices: A review and critical evaluation
Burke, H. R. (1958) · 1958
Earlier work this paper cites.
StackGAN++: Realistic image synthesis with stacked generative adversarial networks
Zhang, H., Xu, T., Li, H., Zhang, S., Wang, X., Huang, X., and Metaxas, D. N. (2019) · 1962
Earlier work this paper cites.
Improving lstm-based video description with linguistic knowledge mined from text
Venugopalan, S., Hendricks, L. A., Mooney, R. J., and Saenko, K. (2016) · 1966
Earlier work this paper cites.
Eliza—a computer program for the study of natural language communication between man and machine
Weizenbaum, J. (1966) · 1966
Earlier work this paper cites.
A statistical approach to machine translation
Brown, P. F., Cocke, J., Della Pietra, S. A., Della Pietra, V. J., Jelinek, F., Lafferty, J. D., Mercer, R. L., and Roossin, P. S. (1990) · 1990
Earlier work this paper cites.
Convolutional networks for images, speech, and time series
LeCun, Y., Bengio, Y., et al. (1995) · 1995
Earlier work this paper cites.
Long Short-Term Memory
Hochreiter, S., and Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
MUREL: multimodal relational reasoning for visual question answering
Cadène, R., Ben-younes, H., Cord, M., and Thome, N. (2019) · 1998
Earlier work this paper cites.
Introduction to reinforcement learning
Sutton, R. S., Barto, A. G., et al. (1998) · 1998
Earlier work this paper cites.
Experiments with open-domain textual question answering
Harabagiu, S. M., Pasca, M. A., and Maiorano, S. J. (2000) · 2000
Earlier work this paper cites.
IR evaluation methods for retrieving highly relevant documents
Järvelin, K., and Kekäläinen, J. (2000) · 2000
Earlier work this paper cites.
Building natural language generation systems
Reiter, E., and Dale, R. (2000) · 2000
Earlier work this paper cites.
Ranking and retrieval of image sequences from multiple paragraph queries
Kim, G., Moon, S., and Sigal, L. (2015) · 2001
Earlier work this paper cites.
Vision based navigation for an unmanned aerial vehicle
Sinopoli, B., Micheli, M., Donato, G., and Koo, T.-J. (2001) · 2001
Earlier work this paper cites.
A user attention model for video summarization
Ma, Y.-F., Lu, L., Zhang, H.-J., and Li, M. (2002) · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J. (2002) · 2002
Earlier work this paper cites.
Matching words and pictures
Barnard, K., Duygulu, P., Forsyth, D. A., de Freitas, N., Blei, D. M., and Jordan, M. I. (2003) · 2003
Earlier work this paper cites.
Entailment, intensionality and text understanding
Condoravdi, C., Crouch, D., De Paiva, V., Stolle, R., and Bobrow, D. G. (2003) · 2003
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y. (2004) · 2004
Earlier work this paper cites.
METEOR: An automatic metric for mt evaluation with improved correlation with human judgments
Banerjee, S., and Lavie, A. (2005) · 2005
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020a) · 2005
Earlier work this paper cites.
Walk the talk: Connecting language, knowledge, and action in route instructions
MacMahon, M., Stankiewicz, B., and Kuipers, B. (2006) · 2006
Earlier work this paper cites.
Advances in open domain question answering
Strzalkowski, T., and Harabagiu, S. (2006) · 2006
Earlier work this paper cites.
Neural language generation: Formulation, methods, and evaluation
Garbacea, C., and Mei, Q. (2020) · 2007
Earlier work this paper cites.
Integrating language, vision and action for human robot dialog systems
Rickert, M., Foster, M. E., Giuliani, M., By, T., Panin, G., and Knoll, A. (2007) · 2007
Earlier work this paper cites.
A review and comparison of measures for automatic video surveillance systems
Baumann, A., Boltz, M., Ebling, J., Koenig, M., Loos, H., Merkel, M., Niem, W., Warzelhan, J., and Yu, J. (2008) · 2008
Earlier work this paper cites.
The graph neural network model
Scarselli, F., Gori, M., Tsoi, A. C., Hagenbuchner, M., and Monfardini, G. (2008) · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. (2009) · 2009
Earlier work this paper cites.
Vision based mav navigation in unknown and unstructured environments
Blösch, M., Weiss, S., Scaramuzza, D., and Siegwart, R. (2010) · 2010
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Bottou, L. (2010) · 2010
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
Farhadi, A., Hejrati, S. M. M., Sadeghi, M. A., Young, P., Rashtchian, C., Hockenmaier, J., and Forsyth, D. A. (2010) · 2010
Earlier work this paper cites.
A game-theoretic approach to generating spatial descriptions
Golland, D., Liang, P., and Klein, D. (2010) · 2010
Earlier work this paper cites.
Introduction to information retrieval
Manning, C., Raghavan, P., and Schütze, H. (2010) · 2010
Earlier work this paper cites.
Learning to follow navigational directions
Vogel, A., and Jurafsky, D. (2010) · 2010
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
Chen, D. L., and Dolan, W. B. (2011) · 2011
Earlier work this paper cites.
Composing simple image descriptions using web-scale n-grams
Li, S., Kulkarni, G., Berg, T. L., Berg, A. C., and Choi, Y. (2011) · 2011
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
Ordonez, V., Kulkarni, G., and Berg, T. L. (2011) · 2011
Earlier work this paper cites.
Corpus-guided sentence generation of natural images
Yang, Y., Teo, C. L., Daumé III, H., and Aloimonos, Y. (2011) · 2011
Earlier work this paper cites.
Computational generation of referring expressions: A survey
Krahmer, E., and Van Deemter, K. (2012) · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E. (2012) · 2012
Earlier work this paper cites.
Midge: Generating image descriptions from computer vision detections
Mitchell, M., Han, X., Dodge, J., Mensch, A., Goyal, A., Berg, A., Yamaguchi, K., Berg, T., Stratos, K., and Daumé III, H. (2012) · 2012
Earlier work this paper cites.
Improving video activity recognition using object recognition and text mining.
Motwani, T. S., and Mooney, R. J. (2012) · 2012
Earlier work this paper cites.
A database for fine grained activity detection of cooking activities
Rohrbach, M., Amin, S., Andriluka, M., and Schiele, B. (2012) · 2012
Earlier work this paper cites.
Script data for attribute-based recognition of composite activities
Rohrbach, M., Regneri, M., Andriluka, M., Amin, S., Pinkal, M., and Schiele, B. (2012) · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, T., and Hinton, G. (2012) · 2012
Earlier work this paper cites.
A thousand frames in just a few words: Lingual description of videos through latent topics and sparse object stitching
Das, P., Xu, C., Doell, R. F., and Corso, J. J. (2013) · 2013
Earlier work this paper cites.
Image description using visual dependency representations
Elliott, D., and Keller, F. (2013) · 2013
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
Frome, A., Corrado, G. S., Shlens, J., Bengio, S., Dean, J., Mikolov, T., et al. (2013) · 2013
Earlier work this paper cites.
Framing image description as a ranking task: Data, models and evaluation metrics
Hodosh, M., Young, P., and Hockenmaier, J. (2013) · 2013
Earlier work this paper cites.
Generating natural-language video descriptions using text-mined knowledge
Krishnamoorthy, N., Malkarnenkar, G., Mooney, R., Saenko, K., and Guadarrama, S. (2013) · 2013
Earlier work this paper cites.
Babytalk: Understanding and generating simple image descriptions
Kulkarni, G., Premraj, V., Ordonez, V., Dhar, S., Li, S., Choi, Y., Berg, A. C., and Berg, T. L. (2013) · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J. (2013) · 2013
Earlier work this paper cites.
Generating expressions that refer to visible objects
Mitchell, M., Van Deemter, K., and Reiter, E. (2013) · 2013
Earlier work this paper cites.
Grounding action descriptions in videos
Regneri, M., Rohrbach, M., Wetzel, D., Thater, S., Schiele, B., and Pinkal, M. (2013) · 2013
Earlier work this paper cites.
Translating video content to natural language descriptions
Rohrbach, M., Qiu, W., Titov, I., Thater, S., Pinkal, M., and Schiele, B. (2013) · 2013
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Cho, K., van Merrienboer, B., Gülçehre, Ç., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Chung, J., Gülçehre, Ç., Cho, K., and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
Kazemzadeh, S., Ordonez, V., Matten, M., and Berg, T. (2014) · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P., and Welling, M. (2014) · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L. (2014) · 2014
Earlier work this paper cites.
A multi-world approach to question answering about real-world scenes based on uncertain input
Malinowski, M., and Fritz, M. (2014) · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Pennington, J., Socher, R., and Manning, C. (2014) · 2014
Earlier work this paper cites.
Coherent multi-sentence video description with variable level of detail
Rohrbach, A., Rohrbach, M., Qiu, W., Friedrich, A., Pinkal, M., and Schiele, B. (2014) · 2014
Earlier work this paper cites.
Meaning in interaction: An introduction to pragmatics
Thomas, J. A. (2014) · 2014
Earlier work this paper cites.
Integrating language and vision to generate natural language descriptions of videos in the wild
Thomason, J., Venugopalan, S., Guadarrama, S., Saenko, K., and Mooney, R. (2014) · 2014
Earlier work this paper cites.
C3d: generic features for video analysis
Tran, D., Bourdev, L. D., Fergus, R., Torresani, L., and Paluri, M. (2014) · 2014
Earlier work this paper cites.
Joint video and text parsing for understanding events and answering queries
Tu, K., Meng, M., Lee, M. W., Choe, T. E., and Zhu, S. C. (2014) · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Young, P., Lai, A., Hodosh, M., and Hockenmaier, J. (2014) · 2014
Earlier work this paper cites.
Edge Boxes: Locating object proposals from edges
Zitnick, C. L., and Dollár, P. (2014) · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Lawrence Zitnick, C., and Parikh, D. (2015) · 2015
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y. (2015) · 2015
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Bowman, S. R., Angeli, G., Potts, C., and Manning, C. D. (2015) · 2015
Earlier work this paper cites.
Mind’s eye: A recurrent visual representation for image caption generation
Chen, X., and Lawrence Zitnick, C. (2015) · 2015
Earlier work this paper cites.
A survey on the application of recurrent neural networks to statistical language modeling
De Mulder, W., Bethard, S., and Moens, M.-F. (2015) · 2015
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
Donahue, J., Anne Hendricks, L., Guadarrama, S., Rohrbach, M., Venugopalan, S., Saenko, K., and Darrell, T. (2015) · 2015
Earlier work this paper cites.
Multi-language image description with neural sequence models
Elliott, D., Frank, S., and Hasler, E. (2015) · 2015
Earlier work this paper cites.
From captions to visual concepts and back
Fang, H., Gupta, S., Iandola, F., Srivastava, R. K., Deng, L., Dollár, P., Gao, J., He, X., Mitchell, M., Platt, J. C., et al. (2015) · 2015
Earlier work this paper cites.
A survey of current datasets for vision and language research
Ferraro, F., Mostafazadeh, N., Huang, T. K., Vanderwende, L., Devlin, J., Galley, M., and Mitchell, M. (2015) · 2015
Earlier work this paper cites.
Are you talking to a machine? dataset and methods for multilingual image question
Gao, H., Mao, J., Zhou, J., Huang, Z., Wang, L., and Xu, W. (2015) · 2015
Earlier work this paper cites.
Visual turing test for computer vision systems
Geman, D., Geman, S., Hallonquist, N., and Younes, L. (2015) · 2015
Earlier work this paper cites.
Guiding the long-short term memory model for image caption generation
Jia, X., Gavves, E., Fernando, B., and Tuytelaars, T. (2015) · 2015
Earlier work this paper cites.
Jin, J., Fu, K., Cui, R., Sha, F., and Zhang, C. (2015) · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Karpathy, A., and Fei-Fei, L. (2015) · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P., and Ba, J. (2015) · 2015
Earlier work this paper cites.
Deep learning
LeCun, Y., Bengio, Y., and Hinton, G. (2015) · 2015
Earlier work this paper cites.
Multi-task deep visual-semantic embedding for video thumbnail selection
Liu, W., Mei, T., Zhang, Y., Che, C., and Luo, J. (2015) · 2015
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation
Long, J., Shelhamer, E., and Darrell, T. (2015) · 2015
Earlier work this paper cites.
Ask your neurons: A neural-based approach to answering questions about images
Malinowski, M., Rohrbach, M., and Fritz, M. (2015) · 2015
Earlier work this paper cites.
Deep captioning with multimodal recurrent neural networks (m-rnn)
Mao, J., Xu, W., Yang, Y., Wang, J., and Yuille, A. L. (2015) · 2015
Earlier work this paper cites.
Polylingual multimodal learning
Mogadala, A. (2015) · 2015
Earlier work this paper cites.
Expressing an image stream with a sequence of natural sentences
Park, C. C., and Kim, G. (2015) · 2015
Earlier work this paper cites.
A dataset for movie description
Rohrbach, A., Rohrbach, M., Tandon, N., and Schiele, B. (2015) · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K., and Zisserman, A. (2015) · 2015
Earlier work this paper cites.
Unsupervised learning of video representations using LSTMs
Srivastava, N., Mansimov, E., and Salakhutdinov, R. (2015) · 2015
Earlier work this paper cites.
End-to-end memory networks
Sukhbaatar, S., Weston, J., Fergus, R., et al. (2015) · 2015
Earlier work this paper cites.
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S. E., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A. (2015) · 2015
Earlier work this paper cites.
Book2movie: Aligning video scenes with book chapters
Tapaswi, M., Bäuml, M., and Stiefelhagen, R. (2015) · 2015
Earlier work this paper cites.
Using descriptive video services to create a large data source for video annotation research
Torabi, A., Pal, C. J., Larochelle, H., and Courville, A. C. (2015) · 2015
Earlier work this paper cites.
CIDEr: Consensus-based image description evaluation
Vedantam, R., Zitnick, C. L., and Parikh, D. (2015) · 2015
Earlier work this paper cites.
Translating videos to natural language using deep recurrent neural networks
Venugopalan, S., Xu, H., Donahue, J., Rohrbach, M., Mooney, R. J., and Saenko, K. (2015b) · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator
Vinyals, O., Toshev, A., Bengio, S., and Erhan, D. (2015) · 2015
Earlier work this paper cites.
Video description generation incorporating spatio-temporal features and a soft-attention mechanism
Yao, L., Torabi, A., Cho, K., Ballas, N., Pal, C. J., Larochelle, H., and Courville, A. C. (2015) · 2015
Earlier work this paper cites.
Visual Madlibs: Fill in the blank image generation and question answering
Yu, L., Park, E., Berg, A. C., and Berg, T. L. (2015) · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Zhu, Y., Kiros, R., Zemel, R. S., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S. (2015) · 2015
Earlier work this paper cites.
Sort story: Sorting jumbled images and captions into stories
Agrawal, H., Chandrasekaran, A., Batra, D., Parikh, D., and Bansal, M. (2016) · 2016
Earlier work this paper cites.
SPICE: Semantic propositional image caption evaluation
Anderson, P., Fernando, B., Johnson, M., and Gould, S. (2016) · 2016
Earlier work this paper cites.
Learning to compose neural networks for question answering
Andreas, J., Rohrbach, M., Darrell, T., and Klein, D. (2016a) · 2016
Earlier work this paper cites.
Automatic description generation from images: A survey of models, datasets, and evaluation measures.
Bernardi, R., Cakici, R., Elliott, D., Erdem, A., Erdem, E., Ikizler-Cinbis, N., Keller, F., Muscat, A., and Plank, B. (2016) · 2016
Earlier work this paper cites.
Deep visual-semantic hashing for cross-modal retrieval
Cao, Y., Long, M., Wang, J., Yang, Q., and Yu, P. S. (2016) · 2016
Earlier work this paper cites.
Evaluating prerequisite qualities for learning end-to-end dialog systems
Dodge, J., Gane, A., Zhang, X., Bordes, A., Chopra, S., Miller, A. H., Szlam, A., and Weston, J. (2016) · 2016
Earlier work this paper cites.
Multi30k: Multilingual english-german image descriptions
Elliott, D., Frank, S., Sima’an, K., and Specia, L. (2016) · 2016
Earlier work this paper cites.
Multimodal compact bilinear pooling for visual question answering and visual grounding
Fukui, A., Park, D. H., Yang, D., Rohrbach, A., Darrell, T., and Rohrbach, M. (2016) · 2016
Earlier work this paper cites.
Image style transfer using convolutional neural networks
Gatys, L. A., Ecker, A. S., and Bethge, M. (2016) · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Earlier work this paper cites.
Deep Compositional Captioning: Describing novel object categories without paired training data
Hendricks, L. A., Venugopalan, S., Rohrbach, M., Mooney, R., Saenko, K., and Darrell, T. (2016) · 2016
Earlier work this paper cites.
Multimodal pivots for image caption translation
Hitschler, J., Schamoni, S., and Riezler, S. (2016) · 2016
Earlier work this paper cites.
Natural language object retrieval
Hu, R., Xu, H., Rohrbach, M., Feng, J., Saenko, K., and Darrell, T. (2016) · 2016
Earlier work this paper cites.
Attention-based multimodal neural machine translation
Huang, P.-Y., Liu, F., Shiang, S.-R., Oh, J., and Dyer, C. (2016) · 2016
Earlier work this paper cites.
Visual storytelling
Huang, T. K., Ferraro, F., Mostafazadeh, N., Misra, I., Agrawal, A., Devlin, J., Girshick, R. B., He, X., Kohli, P., Batra, D., Zitnick, C. L., Parikh, D., Vanderwende, L., Galley, M., and Mitchell, M. (2016) · 2016
Earlier work this paper cites.
Revisiting visual question answering baselines
Jabri, A., Joulin, A., and Van Der Maaten, L. (2016) · 2016
Earlier work this paper cites.
Describing videos using multi-modal fusion
Jin, Q., Chen, J., Chen, S., Xiong, Y., and Hauptmann, A. (2016) · 2016
Earlier work this paper cites.
Densecap: Fully convolutional localization networks for dense captioning
Johnson, J., Karpathy, A., and Fei-Fei, L. (2016) · 2016
Earlier work this paper cites.
Deep reinforcement learning for dialogue generation
Li, J., Monroe, W., Ritter, A., Jurafsky, D., Galley, M., and Gao, J. (2016) · 2016
Earlier work this paper cites.
Generating images from captions with attention
Mansimov, E., Parisotto, E., Ba, L. J., and Salakhutdinov, R. (2016) · 2016
Cited alongside, same era.
Generation and comprehension of unambiguous object descriptions
Mao, J., Huang, J., Toshev, A., Camburu, O., Yuille, A. L., and Murphy, K. (2016) · 2016
Cited alongside, same era.
Senticap: Generating image descriptions with sentiments
Mathews, A. P., Xie, L., and He, X. (2016) · 2016
Cited alongside, same era.
Cross-lingual image caption generation.
Miyazaki, T., and Shimizu, N. (2016) · 2016
Cited alongside, same era.
Modeling context between objects for referring expression understanding
Nagaraja, V. K., Morariu, V. I., and Davis, L. S. (2016) · 2016
Cited alongside, same era.
Abstractive text summarization using sequence-to-sequence rnns and beyond
Nallapati, R., Zhou, B., dos Santos, C. N., Gülçehre, Ç., and Xiang, B. (2016) · 2016
Discovering connotations as labels for weakly supervised image-sentence data
Mogadala, A., Kanuparthi, B., Rettinger, A., and Sure-Vetter, Y. (2018b) · 2018
Later among the works it cites.
Text-adaptive generative adversarial networks: Manipulating images with natural language
Nam, S., Kim, Y., and Kim, S. J. (2018) · 2018
Later among the works it cites.
Film: Visual reasoning with a general conditioning layer
Perez, E., Strub, F., De Vries, H., Dumoulin, V., and Courville, A. (2018) · 2018
Later among the works it cites.
Measuring abstract reasoning in neural networks
Santoro, A., Hill, F., Barrett, D., Morcos, A., and Lillicrap, T. (2018) · 2018
Later among the works it cites.
A dataset and reranking method for multimodal MT of user-generated image captions
Schamoni, S., Hitschler, J., and Riezler, S. (2018) · 2018
Later among the works it cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Unsupervised representation learning with deep convolutional generative adversarial networks
Radford, A., Metz, L., and Chintala, S. (2016) · 2016
Cited alongside, same era.
Multimodal video description
Ramanishka, V., Das, A., Park, D. H., Venugopalan, S., Hendricks, L. A., Rohrbach, M., and Saenko, K. (2016) · 2016
Cited alongside, same era.
Generative adversarial text to image synthesis
Reed, S. E., Akata, Z., Yan, X., Logeswaran, L., Schiele, B., and Lee, H. (2016b) · 2016
Cited alongside, same era.
Improved techniques for training GANs
Salimans, T., Goodfellow, I., Zaremba, W., Cheung, V., Radford, A., and Chen, X. (2016) · 2016
Cited alongside, same era.
A shared task on multimodal machine translation and crosslingual image description.
Specia, L., Frank, S., Sima’an, K., and Elliott, D. (2016) · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z. (2016) · 2016
Cited alongside, same era.
Sharma, P., Ding, N., Goodman, S., and Soricut, R. (2018) · 2018
Later among the works it cites.
Image and video captioning with augmented neural architectures
Shetty, R., Tavakoli, H. R., and Laaksonen, J. (2018) · 2018
Later among the works it cites.
Visual reasoning with multi-hop feature modulation
Strub, F., Seurin, M., Perez, E., De Vries, H., Mary, J., Preux, P., and CourvilleOlivier Pietquin, A. (2018) · 2018
Later among the works it cites.
Object referring in videos with language and human gaze
Vasudevan, A. B., Dai, D., and Gool, L. V. (2018) · 2018
Later among the works it cites.
Grounded textual entailment
Vu, H., Greco, C., Erofeeva, A., Jafaritazehjan, S., Linders, G., Tanti, M., Testoni, A., Bernardi, R., and Gatt, A. (2018) · 2018
Later among the works it cites.
Reconstruction network for video captioning
Wang, B., Ma, L., Zhang, W., and Liu, W. (2018) · 2018
Later among the works it cites.
Show, reward and tell: Automatic generation of narrative paragraph from photo stream by adversarial training
Wang, J., Fu, J., Tang, J., Li, Z., and Mei, T. (2018) · 2018
Later among the works it cites.
Object counts! bringing explicit detections back into image captioning
Wang, J., Madhyastha, P. S., and Specia, L. (2018) · 2018
Later among the works it cites.
Look before you leap: Bridging model-free and model-based reinforcement learning for planned-ahead vision-and-language navigation
Wang, X., Xiong, W., Wang, H., and Yang Wang, W. (2018) · 2018
Later among the works it cites.
Incorporating background knowledge into video description generation
Whitehead, S., Ji, H., Bansal, M., Chang, S.-F., and Voss, C. (2018) · 2018
Later among the works it cites.
Image captioning and visual question answering based on attributes and external knowledge
Wu, Q., Shen, C., Wang, P., Dick, A. R., and van den Hengel, A. (2018) · 2018
Later among the works it cites.
Are you talking to me? Reasoned visual dialog generation through adversarial learning
Wu, Q., Wang, P., Shen, C., Reid, I. D., and van den Hengel, A. (2018) · 2018
Later among the works it cites.
AttnGAN: Fine-grained text to image generation with attentional generative adversarial networks
Xu, T., Zhang, P., Huang, Q., Zhang, H., Gan, Z., Huang, X., and He, X. (2018) · 2018
Later among the works it cites.
A dataset and architecture for visual reasoning with a working memory
Yang, G. R., Ganichev, I., Wang, X., Shlens, J., and Sussillo, D. (2018) · 2018
Later among the works it cites.
Cascaded mutual modulation for visual reasoning
Yao, Y., Xu, J., Wang, F., and Xu, B. (2018) · 2018
Later among the works it cites.
Neural-Symbolic VQA: Disentangling reasoning from vision and language understanding
Yi, K., Wu, J., Gan, C., Torralba, A., Kohli, P., and Tenenbaum, J. (2018) · 2018
Later among the works it cites.
MAttNet: Modular attention network for referring expression comprehension
Yu, L., Lin, Z., Shen, X., Yang, J., Lu, X., Bansal, M., and Berg, T. L. (2018) · 2018
Later among the works it cites.
Grounding referring expressions in images by variational context
Zhang, H., Niu, Y., and Chang, S.-F. (2018) · 2018
Later among the works it cites.
Learning to count objects in natural images for visual question answering
Zhang, Y., Hare, J. S., and Prügel-Bennett, A. (2018) · 2018
Later among the works it cites.
Photographic text-to-image synthesis with a hierarchically-nested adversarial network
Zhang, Z., Xie, Y., and Yang, L. (2018) · 2018
Later among the works it cites.
A multi-task learning approach for image captioning
Zhao, W., Wang, B., Ye, J., Yang, M., Zhao, Z., Luo, R., and Qiao, Y. (2018) · 2018
Later among the works it cites.
A visual attention grounding neural model for multimodal machine translation
Zhou, M., Cheng, R., Lee, Y. J., and Yu, Z. (2018c) · 2018
Later among the works it cites.
Parallel Attention: A unified framework for visual object discovery through dialogs and queries
Zhuang, B., Wu, Q., Shen, C., Reid, I. D., and van den Hengel, A. (2018) · 2018
Later among the works it cites.
Spatial knowledge distillation to aid visual reasoning
Aditya, S., Saha, R., Yang, Y., and Baral, C. (2019) · 2019
Later among the works it cites.
nocaps: novel object captioning at scale
Agrawal, H., Anderson, P., Desai, K., Wang, Y., Chen, X., Jain, R., Johnson, M., Batra, D., Parikh, D., and Lee, S. (2019) · 2019
Later among the works it cites.
Audio visual scene-aware dialog
Alamri, H., Cartillier, V., Das, A., Wang, J., Cherian, A., Essa, I., Batra, D., Marks, T. K., Hori, C., Anderson, P., Lee, S., and Parikh, D. (2019a) · 2019
Later among the works it cites.
Fusion of detected objects in text for visual question answering
Alberti, C., Ling, J., Collins, M., and Reitter, D. (2019) · 2019
Later among the works it cites.
Multimodal machine learning: A survey and taxonomy
Baltrušaitis, T., Ahuja, C., and Morency, L.-P. (2019) · 2019
Later among the works it cites.
Probing the need for visual context in multimodal machine translation
Caglayan, O., Madhyastha, P., Specia, L., and Barrault, L. (2019) · 2019
Later among the works it cites.
Latent variable model for multi-modal translation
Calixto, I., Rios, M., and Aziz, W. (2019) · 2019
Later among the works it cites.
Touchdown: Natural language navigation and spatial reasoning in visual street environments
Chen, H., Suhr, A., Misra, D., Snavely, N., and Artzi, Y. (2019) · 2019
Later among the works it cites.
Cross-lingual language model pretraining
Conneau, A., and Lample, G. (2019) · 2019
Later among the works it cites.
Show, control and tell: A framework for generating controllable and grounded captions
Cornia, M., Baraldi, L., and Cucchiara, R. (2019) · 2019
Later among the works it cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M., Lee, K., and Toutanova, K. (2019) · 2019
Later among the works it cites.
Perceptual pyramid adversarial networks for text-to-image synthesis
Gao, L., Chen, D., Song, J., Xu, X., Zhang, D., and Shen, H. T. (2019) · 2019
Later among the works it cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Goyal, Y., Khot, T., Agrawal, A., Summers-Stay, D., Batra, D., and Parikh, D. (2019) · 2019
Later among the works it cites.
Image-question-answer synergistic network for visual dialog
Guo, D., Xu, C., and Tao, D. (2019) · 2019
Later among the works it cites.
Mscap: Multi-style image captioning with unpaired stylized text
Guo, L., Liu, J., Yao, P., Li, J., and Lu, H. (2019) · 2019
Later among the works it cites.
It is not about the journey; it is about the destination: Following soft paths under question-guidance for visual reasoning
Haurilet, M., Roitberg, A., and Stiefelhagen, R. (2019) · 2019
Later among the works it cites.
Generating multiple objects at spatially distinct locations
Hinz, T., Heinrich, S., and Wermter, S. (2019) · 2019
Later among the works it cites.
End-to-end audio visual scene-aware dialog using multimodal attention-based video features
Hori, C., Alamri, H., Wang, J., Wichern, G., Hori, T., Cherian, A., Marks, T. K., Cartillier, V., Lopes, R. G., Das, A., et al. (2019) · 2019
Later among the works it cites.
A comprehensive survey of deep learning for image captioning
Hossain, M., Sohel, F., Shiratuddin, M. F., and Laga, H. (2019) · 2019
Later among the works it cites.
Language-conditioned graph networks for relational reasoning
Hu, R., Rohrbach, A., Darrell, T., and Saenko, K. (2019) · 2019
Later among the works it cites.
Attention on attention for image captioning
Huang, L., Wang, W., Chen, J., and Wei, X. (2019) · 2019
Later among the works it cites.
Hierarchically structured reinforcement learning for topically coherent visual story generation
Huang, Q., Gan, Z., Çelikyilmaz, A., Wu, D. O., Wang, J., and He, X. (2019) · 2019
Later among the works it cites.
Barack’s wife hillary: Using knowledge graphs for fact-aware language modeling
IV, R. L. L., Liu, N. F., Peters, M. E., Gardner, M., and Singh, S. (2019) · 2019
Later among the works it cites.
Challenges and prospects in vision and language research
Kafle, K., Shrestha, R., and Kanan, C. (2019) · 2019
Later among the works it cites.
Reflective decoding network for image captioning
Ke, L., Pei, W., Li, R., Shen, X., and Tai, Y. (2019a) · 2019
Later among the works it cites.
Tactical rewind: Self-correction via backtracking in vision-and-language navigation
Ke, L., Li, X., Bisk, Y., Holtzman, A., Gan, Z., Liu, J., Gao, J., Choi, Y., and Srinivasa, S. S. (2019b) · 2019
Later among the works it cites.
Dense relational captioning: Triple-stream networks for relationship-based captioning
Kim, D., Choi, J., Oh, T., and Kweon, I. S. (2019) · 2019
Later among the works it cites.
CLEVR-dialog: A diagnostic dataset for multi-round reasoning in visual dialog
Kottur, S., Moura, J. M., Parikh, D., Batra, D., and Rohrbach, M. (2019) · 2019
Later among the works it cites.
Knowledge-driven encode, retrieve, paraphrase for medical image report generation
Li, C. Y., Liang, X., Hu, Z., and Xing, E. P. (2019) · 2019
Later among the works it cites.
Residual attention-based LSTM for video captioning
Li, X., Zhou, Z., Chen, L., and Gao, L. (2019) · 2019
Later among the works it cites.
Storygan: A sequential conditional GAN for story visualization
Li, Y., Gan, Z., Shen, Y., Liu, J., Cheng, Y., Wu, Y., Carin, L., Carlson, D. E., and Gao, J. (2019b) · 2019
Later among the works it cites.
Clevr-ref+: Diagnosing visual reasoning with referring expressions
Liu, R., Liu, C., Bai, Y., and Yuille, A. L. (2019) · 2019
Later among the works it cites.
ViLBERT: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Lu, J., Batra, D., Parikh, D., and Lee, S. (2019) · 2019
Later among the works it cites.
Self-monitoring navigation agent via auxiliary progress estimation
Ma, C., Lu, J., Wu, Z., AlRegib, G., Kira, Z., Socher, R., and Xiong, C. (2019a) · 2019
Later among the works it cites.
The regretful agent: Heuristic-aided navigation through progress estimation
Ma, C., Wu, Z., AlRegib, G., Xiong, C., and Kira, Z. (2019b) · 2019
Later among the works it cites.
Explicit bias discovery in visual question answering models
Manjunatha, V., Saini, N., and Davis, L. S. (2019) · 2019
Later among the works it cites.
The neuro-symbolic concept learner: Interpreting scenes, words, and sentences from natural supervision
Mao, J., Gan, C., Kohli, P., Tenenbaum, J. B., and Wu, J. (2019) · 2019
Later among the works it cites.
OK-VQA: A visual question answering benchmark requiring external knowledge
Marino, K., Rastegari, M., Farhadi, A., and Mottaghi, R. (2019) · 2019
Later among the works it cites.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Miech, A., Zhukov, D., Alayrac, J., Tapaswi, M., Laptev, I., and Sivic, J. (2019) · 2019
Later among the works it cites.
Joint processing of language and visual data for better automated understanding (dagstuhl seminar 19021)
Moens, M., Specia, L., and Tuytelaars, T. (2019) · 2019
Later among the works it cites.
Multi-task learning of hierarchical vision-language representation
Nguyen, D.-K., and Okatani, T. (2019) · 2019
Later among the works it cites.
Vision-based navigation with language-based assistance via imitation learning with indirect intervention
Nguyen, K., Dey, D., Brockett, C., and Dolan, B. (2019) · 2019
Later among the works it cites.
Recursive visual attention in visual dialog
Niu, Y., Zhang, H., Zhang, M., Zhang, J., Lu, Z., and Wen, J.-R. (2019) · 2019
Later among the works it cites.
Mirrorgan: Learning text-to-image generation by redescription
Qiao, T., Zhang, J., Xu, D., and Tao, D. (2019) · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. (2019) · 2019
Later among the works it cites.
Explain yourself! leveraging language models for commonsense reasoning
Rajani, N. F., McCann, B., Xiong, C., and Socher, R. (2019) · 2019
Later among the works it cites.
A simple baseline for audio-visual scene-aware dialog
Schwartz, I., Schwing, A. G., and Hazan, T. (2019) · 2019
Later among the works it cites.
KVQA: knowledge-aware visual question answering
Shah, S., Mishra, A., Yadati, N., and Talukdar, P. P. (2019b) · 2019
Later among the works it cites.
Beyond task success: A closer look at jointly learning to see, ask, and guesswhat
Shekhar, R., Venkatesh, A., Baumgärtner, T., Bruni, E., Plank, B., Bernardi, R., and Fernández, R. (2019) · 2019
Later among the works it cites.
Explainable and explicit visual reasoning over scene graphs
Shi, J., Zhang, H., and Li, J. (2019) · 2019
Later among the works it cites.
Engaging image captioning via personality
Shuster, K., Humeau, S., Hu, H., Bordes, A., and Weston, J. (2019) · 2019
Later among the works it cites.
Towards vqa models that can read
Singh, A., Natarajan, V., Shah, M., Jiang, Y., Chen, X., Batra, D., Parikh, D., and Rohrbach, M. (2019) · 2019
Later among the works it cites.
Unsupervised multi-modal neural machine translation
Su, Y., Fan, K., Bach, N., Kuo, C. J., and Huang, F. (2019) · 2019
Later among the works it cites.
A corpus for reasoning about natural language grounded in photographs
Suhr, A., Zhou, S., Zhang, A., Zhang, I., Bai, H., and Artzi, Y. (2019) · 2019
Later among the works it cites.
VideoBERT: A joint model for video and language representation learning
Sun, C., Myers, A., Vondrick, C., Murphy, K., and Schmid, C. (2019) · 2019
Later among the works it cites.
Lxmert: Learning cross-modality encoder representations from transformers
Tan, H., and Bansal, M. (2019) · 2019
Later among the works it cites.
Learning to navigate unseen environments: Back translation with environmental dropout
Tan, H., Yu, L., and Bansal, M. (2019) · 2019
Later among the works it cites.
Efficientnet: Rethinking model scaling for convolutional neural networks
Tan, M., and Le, Q. (2019) · 2019
Later among the works it cites.
COIN: A large-scale dataset for comprehensive instructional video analysis
Tang, Y., Ding, D., Rao, Y., Zheng, Y., Zhang, D., Zhao, L., Lu, J., and Zhou, J. (2019) · 2019
Later among the works it cites.
Vision-and-dialog navigation
Thomason, J., Murray, M., Cakmak, M., and Zettlemoyer, L. (2019) · 2019
Later among the works it cites.
Learning two-branch neural networks for image-text matching tasks
Wang, L., Li, Y., Huang, J., and Lazebnik, S. (2019) · 2019
Later among the works it cites.
Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation
Wang, X., Huang, Q., Çelikyilmaz, A., Gao, J., Shen, D., Wang, Y., Wang, W. Y., and Zhang, L. (2019a) · 2019
Later among the works it cites.
Vatex: A large-scale, high-quality multilingual dataset for video-and-language research
Wang, X., Wu, J., Chen, J., Li, L., Wang, Y., and Wang, W. Y. (2019b) · 2019
Later among the works it cites.
Unified visual-semantic embeddings: Bridging vision and language with structured meaning representations
Wu, H., Mao, J., Zhang, Y., Jiang, Y., Li, L., Sun, W., and Ma, W.-Y. (2019) · 2019
Later among the works it cites.
Joint event detection and description in continuous video streams
Xu, H., Li, B., Ramanishka, V., Sigal, L., and Saenko, K. (2019) · 2019
Later among the works it cites.
Cross-modal relationship inference for grounding referring expressions
Yang, S., Li, G., and Yu, Y. (2019) · 2019
Later among the works it cites.
From recognition to cognition: Visual commonsense reasoning
Zellers, R., Bisk, Y., Farhadi, A., and Choi, Y. (2019) · 2019
Later among the works it cites.
RAVEN: A dataset for relational and analogical visual reasoning
Zhang, C., Gao, F., Jia, B., Zhu, Y., and Zhu, S. (2019) · 2019
Later among the works it cites.
Reasoning visual dialogs with structural and partial observations
Zheng, Z., Wang, W., Qi, S., and Zhu, S.-C. (2019) · 2019
Later among the works it cites.
Grounded video description
Zhou, L., Kalantidis, Y., Chen, X., Corso, J. J., and Rohrbach, M. (2019) · 2019
Later among the works it cites.
Improving image captioning by leveraging knowledge graphs
Zhou, Y., Sun, Y., and Honavar, V. G. (2019) · 2019
Later among the works it cites.
Video description: A survey of methods, datasets, and evaluation metrics
Aafaq, N., Mian, A., Liu, W., Gilani, S. Z., and Shah, M. (2020) · 2020
Later among the works it cites.
Referit3d: Neural listeners for fine-grained 3D object identification in real-world scenes
Achlioptas, P., Abdelreheem, A., Xia, F., Elhoseiny, M., and Guibas, L. (2020) · 2020
Later among the works it cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., et al. (2020b) · 2020
Later among the works it cites.
Behind the scene: Revealing the secrets of pre-trained vision-and-language models
Cao, J., Gan, Z., Cheng, Y., Yu, L., Chen, Y.-C., and Liu, J. (2020) · 2020
Later among the works it cites.
End-to-end object detection with transformers
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S. (2020) · 2020
Later among the works it cites.
Scanrefer: 3d object localization in RGB-D scans using natural language
Chen, D. Z., Chang, A. X., and Nießner, M. (2020) · 2020
Later among the works it cites.
Uniter: Universal image-text representation learning
Chen, Y.-C., Li, L., Yu, L., Kholy, A. E., Ahmed, F., Gan, Z., Cheng, Y., and Liu, J. (2020) · 2020
Later among the works it cites.
Just Ask: An interactive learning framework for vision and language navigation
Chi, T., Shen, M., Eric, M., Kim, S., and Hakkani-Tür, D. (2020) · 2020
Later among the works it cites.
Meshed-memory transformer for image captioning
Cornia, M., Stefanini, M., Baraldi, L., and Cucchiara, R. (2020) · 2020
Later among the works it cites.
Show, tell, and polish: Ruminant decoding for image captioning
Guo, L., Liu, J., Lu, S., and Lu, H. (2020) · 2020
Later among the works it cites.
Overcoming language priors in vqa via decomposed linguistic representations
Jing, C., Wu, Y., Zhang, X., Jia, Y., and Wu, Q. (2020) · 2020
Later among the works it cites.
MULE: multimodal universal language embedding
Kim, D., Saito, K., Saenko, K., Sclaroff, S., and Plummer, B. A. (2020) · 2020
Later among the works it cites.
TVQA+: spatio-temporal grounding for video question answering
Lei, J., Yu, L., Berg, T. L., and Bansal, M. (2020) · 2020
Later among the works it cites.
ManiGAN: Text-guided image manipulation
Li, B., Qi, X., Lukasiewicz, T., and Torr, P. H. (2020) · 2020
Later among the works it cites.
Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training
Li, G., Duan, N., Fang, Y., Gong, M., and Jiang, D. (2020) · 2020
Later among the works it cites.
Video storytelling: Textual summaries for events
Li, J., Wong, Y., Zhao, Q., and Kankanhalli, M. S. (2020) · 2020
Later among the works it cites.
Oscar: Object-semantics aligned pre-training for vision-language tasks
Li, X., Yin, X., Li, C., Hu, X., Zhang, P., Zhang, L., Wang, L., Hu, H., Dong, L., Wei, F., Choi, Y., and Gao, J. (2020) · 2020
Later among the works it cites.
Federated learning for vision-and-language grounding problems
Liu, F., Wu, X., Ge, S., Fan, W., and Zou, Y. (2020) · 2020
Later among the works it cites.
VIOLIN: A large-scale dataset for video-and-language inference
Liu, J., Chen, W., Cheng, Y., Gan, Z., Yu, L., Yang, Y., and Liu, J. (2020) · 2020
Later among the works it cites.
VisualCOMET: Reasoning about the dynamic context of a still image
Park, J. S., Bhagavatula, C., Mottaghi, R., Farhadi, A., and Choi, Y. (2020) · 2020
Later among the works it cites.
Show, edit and tell: A framework for editing image captions
Sammani, F., and Melas-Kyriazi, L. (2020) · 2020
Later among the works it cites.
BLEURT: learning robust metrics for text generation
Sellam, T., Das, D., and Parikh, A. P. (2020) · 2020
Later among the works it cites.
Alfred: A benchmark for interpreting grounded instructions for everyday tasks
Shridhar, M., Thomason, J., Gordon, D., Bisk, Y., Han, W., Mottaghi, R., Zettlemoyer, L., and Fox, D. (2020) · 2020
Later among the works it cites.
VL-BERT: pre-training of generic visual-linguistic representations
Su, W., Zhu, X., Cao, Y., Li, B., Lu, L., Wei, F., and Dai, J. (2020) · 2020
Later among the works it cites.
Temporally grounding language queries in videos by contextual boundary-aware prediction
Wang, J., Ma, L., and Jiang, W. (2020) · 2020
Later among the works it cites.
Bertscore: Evaluating text generation with BERT
Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., and Artzi, Y. (2020) · 2020
Later among the works it cites.
Object relational graph with teacher-recommended learning for video captioning
Zhang, Z., Shi, Y., Yuan, C., Li, B., Wang, P., Hu, W., and Zha, Z. (2020) · 2020
Later among the works it cites.
Unified vision-language pre-training for image captioning and VQA
Zhou, L., Palangi, H., Zhang, L., Hu, H., Corso, J. J., and Gao, J. (2020) · 2020
Later among the works it cites.
Deep learning for AI
Bengio, Y., LeCun, Y., and Hinton, G. E. (2021) · 2021
Later among the works it cites.
Towards better adversarial synthesis of human images from text
Briq, R., Kochar, P., and Gall, J. (2021) · 2021
Later among the works it cites.
Fusion models for improved image captioning
Kalimuthu, M., Mogadala, A., Mosbach, M., and Klakow, D. (2020) · 2021
Later among the works it cites.
Refer-it-in-rgbd: A bottom-up approach for 3d visual grounding in rgbd images
Liu, H., Lin, A., Han, X., Yang, L., Yu, Y., and Cui, S. (2021) · 2021
Later among the works it cites.
Languagerefer: Spatial-language model for 3d visual grounding
Roh, J., Desingh, K., Farhadi, A., and Fox, D. (2021) · 2021
Later among the works it cites.
Vinvl: Revisiting visual representations in vision-language models
Zhang, P., Li, X., Hu, X., Yang, J., Zhang, L., Wang, L., Choi, Y., and Gao, J. (2021) · 2021
Later among the works it cites.
Video captioning with attention-based lstm and semantic consistency
Gao, L., Guo, Z., Zhang, H., Xu, X., and Shen, H. T. (2017) · 2055
Later among the works it cites.
Show, Attend and Tell: Neural image caption generation with visual attention
Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A. C., Salakhutdinov, R., Zemel, R. S., and Bengio, Y. (2015a) · 2057
Later among the works it cites.
Embodied question answering
Das, A., Datta, S., Gkioxari, G., Lee, S., Parikh, D., and Batra, D. (2018a) · 2063
Later among the works it cites.