Fetching the paper…
Reading the bibliography…
Recent methods for visual question answering rely on large-scale annotated datasets.
Good question! Statistical ranking for question generation
Michael Heilman and Noah A Smith · 2010
Earlier work this paper cites.
The IWSLT 2011 evaluation campaign on automatic talk translation
Marcello Federico, Sebastian Stüker, Luisa Bentivogli, Michael Paul, Mauro Cettolo, Teresa Herrmann, Jan Niehues, and Giovanni Moretti · 2012
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
VQA: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Microsoft COCO captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick · 2015
Earlier work this paper cites.
Exploring models and data for image question answering
Mengye Ren, Ryan Kiros, and Richard Zemel · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Multimodal compact bilinear pooling for visual question answering and visual grounding
Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach · 2016
Earlier work this paper cites.
Gaussian error linear units (GELUs)
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
Visual Genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, Michael Bernstein, and Li Fei-Fei · 2016
Earlier work this paper cites.
Hierarchical question-image co-attention for visual question answering
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh · 2016
Earlier work this paper cites.
Generating natural questions about an image
Nasrin Mostafazadeh, Ishan Misra, Jacob Devlin, Margaret Mitchell, Xiaodong He, and Lucy Vanderwende · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
MovieQA: Understanding stories in movies through question-answering
Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba, Raquel Urtasun, and Sanja Fidler · 2016
Earlier work this paper cites.
Zero-shot visual question answering
Damien Teney and Anton van den Hengel · 2016
Earlier work this paper cites.
Bidirectional recurrent neural network with attention mechanism for punctuation restoration
Ottokar Tilk and Tanel Alumäe · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al · 2016
Earlier work this paper cites.
Dynamic memory networks for visual and textual question answering
Caiming Xiong, Stephen Merity, and Richard Socher · 2016
Earlier work this paper cites.
Ask, attend and answer: Exploring question-guided spatial attention for visual question answering
Huijuan Xu and Kate Saenko · 2016
Earlier work this paper cites.
Stacked attention networks for image question answering
Zichao Yang, Xiaodong He, Jianfeng Gao, Li Deng, and Alex Smola · 2016
Earlier work this paper cites.
https://github.com/ottokart/punctuator2 , 2017
Punctuator · 2017
Earlier work this paper cites.
MUTAN: Multimodal tucker fusion for visual question answering
Hedi Ben-Younes, Rémi Cadene, Matthieu Cord, and Nicolas Thome · 2017
Earlier work this paper cites.
Learning to ask: Neural question generation for reading comprehension
Xinya Du, Junru Shao, and Claire Cardie · 2017
Earlier work this paper cites.
Making the V in VQA matter: : Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Earlier work this paper cites.
TGIF-QA: Toward spatio-temporal reasoning in visual question answering
Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim · 2017
Earlier work this paper cites.
Deepstory: Video story qa by deep embedded memory networks
Kyung-Min Kim, Min-Oh Heo, Seong-Ho Choi, and Byoung-Tak Zhang · 2017
Earlier work this paper cites.
MarioQA: Answering questions by watching gameplay videos
Jonghwan Mun, Paul Hongsuck Seo, Ilchae Jung, and Bohyung Han · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion
Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang · 2017
Earlier work this paper cites.
Video question answering via attribute-augmented attention network learning
Yunan Ye, Zhou Zhao, Yimeng Li, Long Chen, Jun Xiao, and Yueting Zhuang · 2017
Earlier work this paper cites.
Leveraging video descriptions to learn video question answering
Kuo-Hao Zeng, Tseng-Hung Chen, Ching-Yao Chuang, Yuan-Hong Liao, Juan Carlos Niebles, and Min Sun · 2017
Earlier work this paper cites.
Video question answering via hierarchical spatio-temporal attention networks
Zhou Zhao, Qifan Yang, Deng Cai, Xiaofei He, Yueting Zhuang, Zhou Zhao, Qifan Yang, Deng Cai, Xiaofei He, and Yueting Zhuang · 2017
Earlier work this paper cites.
Neural question generation from text: A preliminary study
Qingyu Zhou, Nan Yang, Furu Wei, Chuanqi Tan, Hangbo Bao, and Ming Zhou · 2017
Earlier work this paper cites.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang · 2018
Cited alongside, same era.
Motion-appearance co-memory networks for video question answering
Jiyang Gao, Runzhou Ge, Kan Chen, and Ram Nevatia · 2018
Cited alongside, same era.
Learning answer embeddings for visual question answering
Hexiang Hu, Wei-Lun Chao, and Fei Sha · 2018
Cited alongside, same era.
TVQA: Localized, compositional video question answering
Jie Lei, Licheng Yu, Mohit Bansal, and Tamara L Berg · 2018
Cited alongside, same era.
Visual question generation as dual task of visual question answering
Yikang Li, Nan Duan, Bolei Zhou, Xiao Chu, Wanli Ouyang, Xiaogang Wang, and Ming Zhou · 2018
Cited alongside, same era.
Conceptual Captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Multi-modal transformer for video retrieval
Valentin Gabeur, Chen Sun, Karteek Alahari, and Cordelia Schmid · 2020
Closest in time.
KnowIT VQA: Answering knowledge-based questions about videos
Noa Garcia, Mayu Otani, Chenhui Chu, and Yuta Nakashima · 2020
Closest in time.
Location-aware graph convolutional networks for video question answering
Deng Huang, Peihao Chen, Runhao Zeng, Qing Du, Mingkui Tan, and Chuang Gan · 2020
Closest in time.
Pixel-BERT: Aligning image pixels with text by deep multi-modal transformers
Zhicheng Huang, Zhaoyang Zeng, Bei Liu, Dongmei Fu, and Jianlong Fu · 2020
Closest in time.
Divide and conquer: Question-guided spatio-temporal contextual attention for video question answering
Jianwen Jiang, Ziqiang Chen, Haojie Lin, Xibin Zhao, and Yue Gao · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut · 2018
Cited alongside, same era.
Explore multi-step reasoning in video question answering
Xiaomeng Song, Yucheng Shi, Xin Chen, and Yahong Han · 2018
Cited alongside, same era.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy · 2018
Cited alongside, same era.
A better way to attend: Attention with trees for video question answering
Hongyang Xue, Wenqing Chu, Zhou Zhao, and Deng Cai · 2018
Cited alongside, same era.
Teaching machines to ask questions
Kaichun Yao, Libo Zhang, Tiejian Luo, Lili Tao, and Yanjun Wu · 2018
Cited alongside, same era.
Open-ended long-form video question answering via adaptive hierarchical reinforced networks
Zhou Zhao, Zhu Zhang, Shuwen Xiao, Zhou Yu, Jun Yu, Deng Cai, Fei Wu, and Yueting Zhuang · 2018
Cited alongside, same era.
Synthetic QA corpora generation with roundtrip consistency
Chris Alberti, Daniel Andor, Emily Pitler, Jacob Devlin, and Michael Collins · 2019
Cited alongside, same era.
Pin Jiang and Yahong Han · 2020
Closest in time.
Dense-caption matching and frame-selection gating for temporal localization in VideoQA
Hyounghun Kim, Zineng Tang, and Mohit Bansal · 2020
Closest in time.
Modality shifting attention network for multi-modal video question answering
Junyeong Kim, Minuk Ma, Trung Pham, Kyungsu Kim, and Chang D Yoo · 2020
Closest in time.
Hierarchical conditional relation networks for video question answering
Thao Minh Le, Vuong Le, Svetha Venkatesh, and Truyen Tran · 2020
Closest in time.
Neural reasoning, fast and slow, for video question answering
Thao Minh Le, Vuong Le, Svetha Venkatesh, and Truyen Tran · 2020
Closest in time.
TVQA+: Spatio-temporal grounding for video question answering
Jie Lei, Licheng Yu, Tamara L Berg, and Mohit Bansal · 2020
Closest in time.
Unicoder-VL: A universal encoder for vision and language by cross-modal pre-training
Gen Li, Nan Duan, Yuejian Fang, Ming Gong, Daxin Jiang, and Ming Zhou · 2020
Closest in time.
HERO: Hierarchical encoder for video+language omni-representation pre-training
Linjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan, Licheng Yu, and Jingjing Liu · 2020
Closest in time.
Oscar: Object-semantics aligned pre-training for vision-language tasks
Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et al · 2020
Closest in time.
Transformer-based end-to-end question generation
Luis Enrico Lopez, Diane Kathryn Cruz, Jan Christian Blaise Cruz, and Charibeth Cheng · 2020
Closest in time.
12-in-1: Multi-task vision and language representation learning
Jiasen Lu, Vedanuj Goswami, Marcus Rohrbach, Devi Parikh, and Stefan Lee · 2020
Closest in time.
UniViLM: A unified video and language pre-training model for multimodal understanding and generation
Huaishao Luo, Lei Ji, Botian Shi, Haoyang Huang, Nan Duan, Tianrui Li, Xilin Chen, and Ming Zhou · 2020
Closest in time.
End-to-end learning of visual representations from uncurated instructional videos
Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman · 2020
Closest in time.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Closest in time.
VQA with no questions-answers training
Ben-Zion Vatashsky and Shimon Ullman · 2020
Closest in time.
Long video question answering: A matching-guided attention model
Weining Wang, Yan Huang, and Liang Wang · 2020
Closest in time.
On modality bias in the TVQA dataset
Thomas Winterbottom, Sarah Xiao, Alistair McLean, and Noura Al Moubayed · 2020
Closest in time.
BERT representations for video question answering
Zekun Yang, Noa Garcia, Chenhui Chu, Mayu Otani, Yuta Nakashima, and Haruo Takemura · 2020
Closest in time.
Open-ended video question answering via multi-modal conditional adversarial networks
Zhou Zhao, Shuwen Xiao, Zehan Song, Chujie Lu, Jun Xiao, and Yueting Zhuang · 2020
Closest in time.
Unified vision-language pre-training for image captioning and VQA
Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason J Corso, and Jianfeng Gao · 2020
Closest in time.
ActBERT: Learning global-local video-text representations
Linchao Zhu and Yi Yang · 2020
Closest in time.
Multichannel attention refinement for video question answering
Yueting Zhuang, Dejing Xu, Xin Yan, Wenzhuo Cheng, Zhou Zhao, Shiliang Pu, and Jun Xiao · 2020
Closest in time.
Noise estimation using density estimation for self-supervised multimodal learning
Elad Amrani, Rami Ben-Ari, Daniel Rotman, and Alex Bronstein · 2021
Closest in time.
iPerceive: Applying common-sense reasoning to multi-modal dense video captioning and video question answering
Aman Chadha, Gurneet Arora, and Navpreet Kaloty · 2021
Closest in time.
DramaQA: Character-centered video story understanding with hierarchical qa
Seongho Choi, Kyoung-Woon On, Yu-Jung Heo, Ahjeong Seo, Youwon Jang, Seungchan Lee, Minsu Lee, and Byoung-Tak Zhang · 2021
Closest in time.
VirTex: Learning visual representations from textual annotations
Karan Desai and Justin Johnson · 2021
Closest in time.
Less is more: Clipbert for video-and-language learning via sparse sampling
Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L Berg, Mohit Bansal, and Jingjing Liu · 2021
Closest in time.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Closest in time.
Look before you speak: Visually contextualized utterances
Paul Hongsuck Seo, Arsha Nagrani, and Cordelia Schmid · 2021
Closest in time.