Fetching the paper…
Reading the bibliography…
It is challenging for models to understand complex, multimodal content such as television clips, and this is in part because video-language models often rely on single-modality reasoning and lack interpretability.
Visual entailment: A novel task for fine-grained image understanding
Ning Xie, Farley Lai, Derek Doran, and Asim Kadav. 2019 · 1901
Earlier work this paper cites.
Tvqa+: Spatio-temporal grounding for video question answering
Jie Lei, Licheng Yu, Tamara L Berg, and Mohit Bansal. 2019 · 1904
Earlier work this paper cites.
Multimodal logical inference system for visual-textual entailment
Riko Suzuki, Hitomi Yanaka, Masashi Yoshikawa, Koji Mineshima, and Daisuke Bekki. 2019 · 1906
Earlier work this paper cites.
Explainable deep learning for video recognition tasks: A framework & recommendations
Liam Hiley, Alun Preece, and Yulia Hicks. 2019 · 1909
Earlier work this paper cites.
Logical self-defense
Ralph H. Johnson and J. Anthony Blair. 1977 · 1977
Earlier work this paper cites.
e-snli-ve: Corrected visual-textual entailment with natural language explanations
Virginie Do, Oana-Maria Camburu, Zeynep Akata, and Thomas Lukasiewicz. 2020 · 2004
Earlier work this paper cites.
Hero: Hierarchical encoder for video+ language omni-representation pre-training
Linjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan, Licheng Yu, and Jingjing Liu. 2020 · 2005
Earlier work this paper cites.
Mahsan Nourani, Chiradeep Roy, Tahrima Rahman, Eric D Ragan, Nicholas Ruozzi, and Vibhav Gogate. 2020 · 2005
Earlier work this paper cites.
Ms marco: A human generated machine reading comprehension dataset
Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, et al. 2016 · 2016
Earlier work this paper cites.
Movieqa: Understanding stories in movies through question-answering
Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba, Raquel Urtasun, and Sanja Fidler. 2016 · 2016
Earlier work this paper cites.
Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim. 2017 · 2017
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel R Bowman. 2017 · 2017
Earlier work this paper cites.
Tvqa: Localized, compositional video question answering
Jie Lei, Licheng Yu, Mohit Bansal, and Tamara L Berg. 2018 · 2018
Earlier work this paper cites.
Open-ended long-form video question answering via adaptive hierarchical reinforced networks
Zhou Zhao, Zhu Zhang, Shuwen Xiao, Zhou Yu, Jun Yu, Deng Cai, Fei Wu, and Yueting Zhuang. 2018 · 2018
Earlier work this paper cites.
Are we asking the right questions in movieqa?
Bhavan Jasani, Rohit Girdhar, and Deva Ramanan. 2019 · 2019
Earlier work this paper cites.
Explainable activity recognition in videos
Chiradeep Roy, Mahesh Shanbhag, Mahsan Nourani, Tahrima Rahman, Samia Kabir, Vibhav Gogate, Nicholas Ruozzi, and Eric D Ragan. 2019 · 2019
Cited alongside, same era.
Explainable video action reasoning via prior knowledge and state transitions
Tao Zhuo, Zhiyong Cheng, Peng Zhang, Yongkang Wong, and Mohan Kankanhalli. 2019 · 2019
Cited alongside, same era.
Violin: A large-scale dataset for video-and-language inference
Jingzhou Liu, Wenhu Chen, Yu Cheng, Zhe Gan, Licheng Yu, Yiming Yang, and Jingjing Liu. 2020 · 2020
Cited alongside, same era.
Open-ended video question answering via multi-modal conditional adversarial networks
Zhou Zhao, Shuwen Xiao, Zehan Song, Chujie Lu, Jun Xiao, and Yueting Zhuang. 2020 · 2020
Cited alongside, same era.
Unified vision-language pre-training for image captioning and vqa
Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason Corso, and Jianfeng Gao. 2020 · 2020
Cited alongside, same era.
Dynamic multistep reasoning based on video scene graph for video question answering
Jianguo Mao, Wenbin Jiang, Xiangdong Wang, Zhifan Feng, Yajuan Lyu, Hong Liu, and Yong Zhu. 2022 · 2022
Later among the works it cites.
Entailment tree explanations via iterative retrieval-generation reasoner
Danilo Neves Ribeiro, Shen Wang, Xiaofei Ma, Rui Dong, Xiaokai Wei, Henghui Zhu, Xinchi Chen, Peng Xu, Zhiheng Huang, Andrew Arnold, and Dan Roth. 2022 · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Later among the works it cites.
Are vision-language transformers learning multimodal representations? a probing perspective
Emmanuelle Salin, Badreddine Farah, Stéphane Ayache, and Benoit Favre. 2022 · 2022
Later among the works it cites.
Long-form video-language pre-training with multimodal temporal contrastive learning
Yuchong Sun, Hongwei Xue, Ruihua Song, Bei Liu, Huan Yang, and Jianlong Fu. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yeyun Zou and Qiyu Xie. 2020 · 2020
Cited alongside, same era.
Evaluation of text generation: A survey
Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2021 · 2021
Cited alongside, same era.
Explainable video entailment with grounded visual evidence
Junwen Chen and Yu Kong. 2021 · 2021
Cited alongside, same era.
Explaining answers with entailment trees
Bhavana Dalvi, Peter Jansen, Oyvind Tafjord, Zhengnan Xie, Hannah Smith, Leighanna Pipatanangkura, and Peter Clark. 2021 · 2021
Cited alongside, same era.
Understanding by understanding not: Modeling negation in language models
Arian Hosseini, Siva Reddy, Dzmitry Bahdanau, R Devon Hjelm, Alessandro Sordoni, and Aaron Courville. 2021 · 2021
Cited alongside, same era.
Video question-answering techniques, benchmark datasets and evaluation metrics leveraging video captioning: a comprehensive survey
Khushboo Khurana and Umesh Deshpande. 2021 · 2021
Cited alongside, same era.
Vilt: Vision-and-language transformer without convolution or region supervision
Wonjae Kim, Bokyung Son, and Ildoo Kim. 2021 · 2021
Cited alongside, same era.
Fine-grained visual entailment
Christopher Thomas, Yipeng Zhang, and Shih-Fu Chang. 2022 · 2022
Later among the works it cites.
Zero-shot video question answering via frozen bidirectional language models
Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid. 2022 · 2022
Later among the works it cites.
Video question answering: Datasets, algorithms and challenges
Yaoyao Zhong, Junbin Xiao, Wei Ji, Yicong Li, Weihong Deng, and Tat-Seng Chua. 2022 · 2022
Later among the works it cites.
Large language models are temporal and causal reasoners for video question answering
Dohwan Ko, Ji Soo Lee, Wooyoung Kang, Byungseok Roh, and Hyunwoo J Kim. 2023 · 2023
Later among the works it cites.
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023 · 2023
Later among the works it cites.
Automated evaluation of written discourse coherence using gpt-4
Ben Naismith, Phoebe Mulcaire, and Jill Burstein. 2023 · 2023
Later among the works it cites.
Revealing the illusion of joint multimodal understanding in videoqa models
Ishaan Singh Rawal, Shantanu Jaiswal, Basura Fernando, and Cheston Tan. 2023 · 2023
Later among the works it cites.
Self-chained image-language model for video localization and question answering
Shoubin Yu, Jaemin Cho, Prateek Yadav, and Mohit Bansal. 2023 · 2023
Later among the works it cites.
Explainability for large language models: A survey
Haiyan Zhao, Hanjie Chen, Fan Yang, Ninghao Liu, Huiqi Deng, Hengyi Cai, Shuaiqiang Wang, Dawei Yin, and Mengnan Du. 2023 · 2023
Later among the works it cites.
Entailer: Answering questions with faithful and truthful chains of reasoning
Oyvind Tafjord, Bhavana Dalvi Mishra, and Peter Clark. 2022 · 2093
Closest in time.