Fetching the paper…
Reading the bibliography…
Existing video understanding datasets mostly focus on human interactions, with little attention being paid to the "in the wild" settings, where the videos are recorded outdoors.
Temporal localization of moments in video collections with natural language
Victor Escorcia, Mattia Soldan, Josef Sivic, Bernard Ghanem, and Bryan Russell. 2019 · 1907
Earlier work this paper cites.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Multitask learning: A knowledge-based source of inductive bias
Richard Caruana. 1993 · 1993
Earlier work this paper cites.
Text chunking using transformation-based learning
Lance Ramshaw and Mitch Marcus. 1995 · 1995
Earlier work this paper cites.
The OpenCV Library
G. Bradski. 2000 · 2000
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Heterogeneous memory enhanced multimodal attention model for video question answering
Chenyou Fan, Xiaofan Zhang, Shu Zhang, Wensheng Wang, Chi Zhang, and Heng Huang. 2019 · 2007
Earlier work this paper cites.
A unified architecture for natural language processing: deep neural networks with multitask learning
Ronan Collobert and Jason Weston. 2008 · 2008
Earlier work this paper cites.
The PASCAL visual object classes (VOC) challenge
Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. 2010 · 2010
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
David Chen and William Dolan. 2011 · 2011
Earlier work this paper cites.
New types of deep neural network learning for speech recognition and related applications: An overview
Li Deng, Geoffrey Hinton, and Brian Kingsbury. 2013 · 2013
Earlier work this paper cites.
A multi-world approach to question answering about real-world scenes based on uncertain input
Mateusz Malinowski and Mario Fritz. 2014 · 2014
Earlier work this paper cites.
Fast R-CNN
Ross B. Girshick. 2015 · 2015
Earlier work this paper cites.
Exploring models and data for image question answering
Mengye Ren, Ryan Kiros, and Richard S. Zemel. 2015 · 2015
Earlier work this paper cites.
Physical causality of action verbs in grounded language understanding
Qiaozi Gao, Malcolm Doering, Shaohua Yang, and Joyce Yue Chai. 2016 · 2016
Earlier work this paper cites.
PororoQA: Cartoon video series dataset for story understanding
K Kim, C Nan, MO Heo, SH Choi, and BT Zhang. 2016 · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
MovieQA: Understanding stories in movies through question-answering
Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba, Raquel Urtasun, and Sanja Fidler. 2016 · 2016
Earlier work this paper cites.
VideoMCC: a new benchmark for video comprehension
Du Tran, Maksim Bolonkin, Manohar Paluri, and Lorenzo Torresani. 2016 · 2016
Earlier work this paper cites.
Yin and yang: Balancing and answering binary visual questions
Peng Zhang, Yash Goyal, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2016 · 2016
Earlier work this paper cites.
Visual7W: Grounded question answering in images
Yuke Zhu, Oliver Groth, Michael S. Bernstein, and Li Fei-Fei. 2016 · 2016
Earlier work this paper cites.
Quo vadis, action recognition? A new model and the kinetics dataset
João Carreira and Andrew Zisserman. 2017 · 2017
Earlier work this paper cites.
TALL: temporal activity localization via language query
Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia. 2017 · 2017
Earlier work this paper cites.
Localizing moments in video with natural language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan C. Russell. 2017 · 2017
Earlier work this paper cites.
TGIF-QA: toward spatio-temporal reasoning in visual question answering
Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim. 2017 · 2017
Cited alongside, same era.
Dense-captioning events in videos
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles. 2017 · 2017
Cited alongside, same era.
A dataset and exploration of models for understanding video data through fill-in-the-blank question-answering
Tegan Maharaj, Nicolas Ballas, Anna Rohrbach, Aaron C. Courville, and Christopher Joseph Pal. 2017 · 2017
Cited alongside, same era.
MarioQA: Answering questions by watching gameplay videos
Jonghwan Mun, Paul Hongsuck Seo, Ilchae Jung, and Bohyung Han. 2017 · 2017
Cited alongside, same era.
Video question answering via gradually refined attention over appearance and motion
Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang. 2017 · 2017
Cited alongside, same era.
From recognition to cognition: Visual commonsense reasoning
Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019 · 2019
Later among the works it cites.
Frame augmented alternating attention network for video question answering
Wenqiao Zhang, Siliang Tang, Yanpeng Cao, Shiliang Pu, Fei Wu, and Yueting Zhuang. 2019 · 2019
Later among the works it cites.
LifeQA: A real-life dataset for video question answering
Santiago Castro, Mahmoud Azab, Jonathan Stroud, Cristina Noujaim, Ruoyao Wang, Jia Deng, and Rada Mihalcea. 2020 · 2020
Later among the works it cites.
UNITER: Universal image-text representation learning
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020 · 2020
Later among the works it cites.
TutorialVQA: Question answering dataset for tutorial videos
Anthony Colas, Seokhwan Kim, Franck Dernoncourt, Siddhesh Gupte, Zhe Wang, and Doo Soon Kim. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Unifying the video and question attentions for open-ended video question answering
Hongyang Xue, Zhou Zhao, and Deng Cai. 2017 · 2017
Cited alongside, same era.
Video question answering via attribute-augmented attention network learning
Yunan Ye, Zhou Zhao, Yimeng Li, Long Chen, Jun Xiao, and Yueting Zhuang. 2017 · 2017
Cited alongside, same era.
Leveraging video descriptions to learn video question answering
Kuo-Hao Zeng, Tseng-Hung Chen, Ching-Yao Chuang, Yuan-Hong Liao, Juan Carlos Niebles, and Min Sun. 2017 · 2017
Cited alongside, same era.
Video question answering via hierarchical spatio-temporal attention networks
Zhou Zhao, Qifan Yang, Deng Cai, Xiaofei He, and Yueting Zhuang. 2017 · 2017
Cited alongside, same era.
Uncovering the temporal context for video question answering
Linchao Zhu, Zhongwen Xu, Yi Yang, and Alexander G Hauptmann. 2017 · 2017
Cited alongside, same era.
Motion-appearance co-memory networks for video question answering
Jiyang Gao, Runzhou Ge, Kan Chen, and Ram Nevatia. 2018a · 2018
Cited alongside, same era.
What action causes this? towards naive physical action-effect prediction
Qiaozi Gao, Shaohua Yang, Joyce Yue Chai, and Lucy Vanderwende. 2018b · 2018
Cited alongside, same era.
Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C. Wallace. 2020 · 2020
Later among the works it cites.
KnowIT VQA: answering knowledge-based questions about videos
Noa Garcia, Mayu Otani, Chenhui Chu, and Yuta Nakashima. 2020 · 2020
Later among the works it cites.
Location-aware graph convolutional networks for video question answering
Deng Huang, Peihao Chen, Runhao Zeng, Qing Du, Mingkui Tan, and Chuang Gan. 2020 · 2020
Later among the works it cites.
Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness?
Alon Jacovi and Yoav Goldberg. 2020 · 2020
Later among the works it cites.
Divide and conquer: Question-guided spatio-temporal contextual attention for video question answering
Jianwen Jiang, Ziqiang Chen, Haojie Lin, Xibin Zhao, and Yue Gao. 2020 · 2020
Later among the works it cites.
Uncertainty-aware self-training for few-shot text classification
Subhabrata Mukherjee and Ahmed Awadallah. 2020 · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Later among the works it cites.
Neural semantic parsing in low-resource settings with back-translation and meta-learning
Yibo Sun, Duyu Tang, Nan Duan, Yeyun Gong, Xiaocheng Feng, Bing Qin, and Daxin Jiang. 2020 · 2020
Later among the works it cites.
Long-term video question answering via multimodal hierarchical memory attentive networks
Ting Yu, Jun Yu, Zhou Yu, Qingming Huang, and Qi Tian. 2020 · 2020
Later among the works it cites.
DramaQA: Character-centered video story understanding with hierarchical qa
Seongho Choi, Kyoung-Woon On, Yu-Jung Heo, Ahjeong Seo, Youwon Jang, Minsu Lee, and Byoung-Tak Zhang. 2021 · 2021
Later among the works it cites.
Ego4D: Around the world in 3,000 hours of egocentric video
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. 2021 · 2021
Later among the works it cites.
KaggleDBQA: Realistic evaluation of text-to-SQL parsers
Chia-Hsuan Lee, Oleksandr Polozov, and Matthew Richardson. 2021 · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021 · 2021
Later among the works it cites.
Next-qa: Next phase of question-answering to explaining temporal actions
Junbin Xiao, Xindi Shang, Angela Yao, and Tat-Seng Chua. 2021 · 2021
Later among the works it cites.
MERLOT: Multimodal neural script knowledge models
Rowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu, Jae Sung Park, Jize Cao, Ali Farhadi, and Yejin Choi. 2021 · 2021
Later among the works it cites.
FIBER: Fill-in-the-blanks as a challenging video understanding evaluation framework
Santiago Castro, Ruoyao Wang, Pingxuan Huang, Ian Stewart, Oana Ignat, Nan Liu, Jonathan Stroud, and Rada Mihalcea. 2022 · 2022
Closest in time.
DeepStory: Video story QA by deep embedded memory networks
Kyung-Min Kim, Min-Oh Heo, Seong-Ho Choi, and Byoung-Tak Zhang. 2017 · 2022
Closest in time.
xGQA: Cross-lingual visual question answering
Jonas Pfeiffer, Gregor Geigle, Aishwarya Kamath, Jan-Martin Steitz, Stefan Roth, Ivan Vulić, and Iryna Gurevych. 2022 · 2022
Closest in time.
Video question answering: Datasets, algorithms and challenges
Yaoyao Zhong, Wei Ji, Junbin Xiao, Yicong Li, Weihong Deng, and Tat-Seng Chua. 2022 · 2022
Closest in time.