Fetching the paper…
Reading the bibliography…
Causal video question answering (QA) has garnered increasing interest, yet existing datasets often lack depth in causal reasoning.
The illusion of life : Disney animation
Frank Thomas and Ollie Johnston · 1981
Earlier work this paper cites.
How is children’s learning from television distinctive? exploiting the medium methodologically
Laurene K Meringoff, Martha M Vibbert, Cynthia A Char, David E Fernie, Gail S Banker, and Howard Gardner · 1983
Earlier work this paper cites.
Identification and ratings of caricatures: Implications for mental representations of faces
Gillian Rhodes, Susan Brennan, and Susan Carey · 1987
Earlier work this paper cites.
Caricature and face recognition
Robert Mauro and Michael Kubovy · 1992
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
Functions of animation in comprehension and learning
Wolfgang Schnotz and Thorsten Rasch · 2008
Earlier work this paper cites.
Youtube2text: Recognizing and describing arbitrary activities using semantic hierarchies and zero-shot recognition
Sergio Guadarrama, Niveda Krishnamoorthy, Girish Malkarnenkar, Subhashini Venugopalan, Raymond Mooney, Trevor Darrell, and Kate Saenko · 2013
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Data association for semantic world modeling from partial views
Lawson LS Wong, Leslie Pack Kaelbling, and Tomás Lozano-Pérez · 2015
Earlier work this paper cites.
Spice: Semantic propositional image caption evaluation
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould · 2016
Earlier work this paper cites.
Human action recognition without human
Yun He, Soma Shirakabe, Yutaka Satoh, and Hirokatsu Kataoka · 2016
Earlier work this paper cites.
Does animation facilitate better learning in primary education? a comparative study of three different subjects
Mairaru Shreesha and Sanjay Kumar Tyagi · 2016
Earlier work this paper cites.
Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim · 2017
Earlier work this paper cites.
A dataset and exploration of models for understanding video data through fill-in-the-blank question-answering
Tegan Maharaj, Nicolas Ballas, Anna Rohrbach, Aaron Courville, and Christopher Pal · 2017
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion
Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang · 2017
Cited alongside, same era.
Unifying the video and question attentions for open-ended video question answering
Hongyang Xue, Zhou Zhao, and Deng Cai · 2017
Cited alongside, same era.
Peer learning with concept cartoons enhance critical thinking and performance in secondary school economics
Khoo Yin Yin and Robert Fitzgerald · 2017
Cited alongside, same era.
Leveraging video descriptions to learn video question answering
Kuo-Hao Zeng, Tseng-Hung Chen, Ching-Yao Chuang, Yuan-Hong Liao, Juan Carlos Niebles, and Min Sun · 2017
Cited alongside, same era.
Uncovering the temporal context for video question answering
Linchao Zhu, Zhongwen Xu, Yi Yang, and Alexander G Hauptmann · 2017
Cited alongside, same era.
Agqa: A benchmark for compositional spatio-temporal reasoning
Madeleine Grunde-McLaughlin, Ranjay Krishna, and Maneesh Agrawala · 2021
Later among the works it cites.
Clipcap: Clip prefix for image captioning
Ron Mokady, Amir Hertz, and Amit H Bermano · 2021
Later among the works it cites.
STAR: A benchmark for situated reasoning in real-world videos
Bo Wu, Shoubin Yu, Zhenfang Chen, Joshua B Tenenbaum, and Chuang Gan · 2021
Later among the works it cites.
Next-qa: Next phase of question-answering to explaining temporal actions
Junbin Xiao, Xindi Shang, Angela Yao, and Tat-Seng Chua · 2021
Later among the works it cites.
From representation to reasoning: Towards both evidence and commonsense reasoning for video question-answering
Jiangtong Li, Li Niu, and Liqing Zhang · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jiyang Gao, Runzhou Ge, Kan Chen, and Ram Nevatia · 2018
Cited alongside, same era.
Open-ended long-form video question answering via adaptive hierarchical reinforced networks
Zhou Zhao, Zhu Zhang, Shuwen Xiao, Zhou Yu, Jun Yu, Deng Cai, Fei Wu, and Yueting Zhuang · 2018
Cited alongside, same era.
Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models
Andrei Barbu, David Mayo, Julian Alverio, William Luo, Christopher Wang, Dan Gutfreund, Josh Tenenbaum, and Boris Katz · 2019
Cited alongside, same era.
Heterogeneous memory enhanced multimodal attention model for video question answering
Chenyou Fan, Xiaofan Zhang, Shu Zhang, Wensheng Wang, Chi Zhang, and Heng Huang · 2019
Cited alongside, same era.
Video question answering with spatio-temporal reasoning
Yunseok Jang, Yale Song, Chris Dongjoo Kim, Youngjae Yu, Youngjin Kim, and Gunhee Kim · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych · 2019
Cited alongside, same era.
Clevrer-humans: Describing physical and causal events the human way
Jiayuan Mao, Xuelin Yang, Xikun Zhang, Noah Goodman, and Jiajun Wu · 2022
Later among the works it cites.
Human mesh recovery from multiple shots
Georgios Pavlakos, Jitendra Malik, and Angjoo Kanazawa · 2022
Later among the works it cites.
Video question answering: Datasets, algorithms and challenges
Yaoyao Zhong, Junbin Xiao, Wei Ji, Yicong Li, Weihong Deng, and Tat-Seng Chua · 2022
Later among the works it cites.
Open-vocabulary universal image segmentation with maskclip
Zheng Ding, Jieke Wang, and Zhuowen Tu · 2023
Later among the works it cites.
Mist: Multi-modal iterative spatial-temporal transformer for long-form video question answering
Difei Gao, Luowei Zhou, Lei Ji, Linchao Zhu, Yi Yang, and Mike Zheng Shou · 2023
Later among the works it cites.
Caricaturing shapes in visual memory
Subin Han, Zekun Sun, and Chaz Firestone · 2023
Later among the works it cites.
team OpenAI · 2023
Later among the works it cites.
The use of concept cartoons in overcoming the misconception in electricity concepts
Laı Chin Siong, Yunn Ong, Fatin Aliah Phang, and Jaysuman Pusppanathan · 2023
Later among the works it cites.
Alpha-clip: A clip model focusing on wherever you want
Zeyi Sun, Ye Fang, Tong Wu, Pan Zhang, Yuhang Zang, Shu Kong, Yuanjun Xiong, Dahua Lin, and Jiaqi Wang · 2023
Later among the works it cites.
Video-LLaMA: An instruction-tuned audio-visual language model for video understanding
Hang Zhang, Xin Li, and Lidong Bing · 2023
Later among the works it cites.
Explore spurious correlations at the concept level in language models for text classification
Yuhang Zhou, Paiheng Xu, Xiaoyu Liu, Bang An, Wei Ai, and Furong Huang · 2023
Later among the works it cites.
Hello gpt-4o
OpenAI · 2024
Closest in time.