Fetching the paper…
Reading the bibliography…
Large language models such as GPT-3 have demonstrated an impressive capability to adapt to new tasks without requiring task-specific training data.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Movie script summarization as graph-based scene extraction
Philip John Gorinski and Mirella Lapata · 2015
Earlier work this paper cites.
Movieqa: Understanding stories in movies through question-answering
Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba, Raquel Urtasun, and Sanja Fidler · 2016
Earlier work this paper cites.
Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim · 2017
Earlier work this paper cites.
Deepstory: video story qa by deep embedded memory networks
Kyung-Min Kim, Min-Oh Heo, Seong-Ho Choi, and Byoung-Tak Zhang · 2017
Earlier work this paper cites.
A read-write memory network for movie story understanding
Seil Na, Sangho Lee, Jisung Kim, and Gunhee Kim · 2017
Earlier work this paper cites.
Movie Description
Anna Rohrbach, Atousa Torabi, Marcus Rohrbach, Niket Tandon, Christopher Pal, Hugo Larochelle, Aaron Courville, and Bernt Schiele · 2017
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion
Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang · 2017
Earlier work this paper cites.
Leveraging video descriptions to learn video question answering
Kuo-Hao Zeng, Tseng-Hung Chen, Ching-Yao Chuang, Yuan-Hong Liao, Juan Carlos Niebles, and Min Sun · 2017
Earlier work this paper cites.
Video question answering via hierarchical dual-level attention network learning
Zhou Zhao, Jinghao Lin, Xinghua Jiang, Deng Cai, Xiaofei He, and Yueting Zhuang · 2017
Earlier work this paper cites.
Motion-appearance co-memory networks for video question answering
Jiyang Gao, Runzhou Ge, Kan Chen, and Ram Nevatia · 2018
Earlier work this paper cites.
Tvqa: Localized, compositional video question answering
Jie Lei, Licheng Yu, Mohit Bansal, and Tamara L Berg · 2018
Earlier work this paper cites.
Heterogeneous memory enhanced multimodal attention model for video question answering
Chenyou Fan, Xiaofan Zhang, Shu Zhang, Wensheng Wang, Chi Zhang, and Heng Huang · 2019
Cited alongside, same era.
Are we asking the right questions in movieqa?
Bhavan Jasani, Rohit Girdhar, and Deva Ramanan · 2019
Cited alongside, same era.
Tvqa+: Spatio-temporal grounding for video question answering
Jie Lei, Licheng Yu, Tamara L Berg, and Mohit Bansal · 2019
Cited alongside, same era.
A2a: Attention to attention reasoning for movie question answering
Chao-Ning Liu, Ding-Jie Chen, Hwann-Tzong Chen, and Tyng-Luh Liu · 2019
Cited alongside, same era.
Movie plot analysis via turning point identification
Pinelopi Papalampidi, Frank Keller, and Mirella Lapata · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Violet: End-to-end video-language transformers with masked visual-token modeling
Tsu-Jui Fu, Linjie Li, Zhe Gan, Kevin Lin, William Yang Wang, Lijuan Wang, and Zicheng Liu · 2021
Later among the works it cites.
Self-supervised pre-training and contrastive representation learning for multiple-choice video qa
Seonhoon Kim, Seohyeong Jeong, Eunbyul Kim, Inho Kang, and Nojun Kwak · 2021
Later among the works it cites.
Transformer-based screenplay summarization using augmented learning representation with dialogue information
Myungji Lee, Hong-Seok Kwon, Jaehun Shin, WonKee Lee, Baikjin Jung, and Jong-Hyeok Lee · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
Star: A benchmark for situated reasoning in real-world videos
Bo Wu, Shoubin Yu, Zhenfang Chen, Joshua B Tenenbaum, and Chuang Gan · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
DramaQA: character-centered video story understanding with hierarchical qa
Seongho Choi, Kyoung-Woon On, Yu-Jung Heo, Ahjeong Seo, Youwon Jang, Seungchan Lee, Minsu Lee, and Byoung-Tak Zhang · 2020
Cited alongside, same era.
Dual hierarchical temporal convolutional network with qa-aware dynamic normalization for video story question answering
Fei Liu, Jing Liu, Xinxin Zhu, Richang Hong, and Hanqing Lu · 2020
Cited alongside, same era.
Screenplay summarization using latent narrative structure
Pinelopi Papalampidi, Frank Keller, Lea Frermann, and Mirella Lapata · 2020
Cited alongside, same era.
Pegasus: Pre-training with extracted gap-sentences for abstractive summarization
Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter Liu · 2020
Cited alongside, same era.
Dramaqa: Character-centered video story understanding with hierarchical qa
Seongho Choi, Kyoung-Woon On, Yu-Jung Heo, Ahjeong Seo, Youwon Jang, Minsu Lee, and Byoung-Tak Zhang · 2021
Cited alongside, same era.
Progressive attention memory network for movie story question answering
Junyeong Kim, Minuk Ma, Kyungsu Kim, Sungjin Kim, and Chang D Yoo
Cited in the paper.
Later among the works it cites.
Next-qa: Next phase of question-answering to explaining temporal actions
Junbin Xiao, Xindi Shang, Angela Yao, and Tat-Seng Chua · 2021
Later among the works it cites.
Merlot: Multimodal neural script knowledge models
Rowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu, Jae Sung Park, Jize Cao, Ali Farhadi, and Yejin Choi · 2021
Later among the works it cites.
Z-code++: A pre-trained language model optimized for abstractive summarization
Pengcheng He, Baolin Peng, Liyang Lu, Songhe Wang, Jie Mei, Yang Liu, Ruochen Xu, Hany Hassan Awadalla, Yu Shi, Chenguang Zhu, Wayne Xiong, Michael Zeng, Jianfeng Gao, and Xuedong Huang · 2022
Later among the works it cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi · 2022
Later among the works it cites.
Merlot reserve: Neural script knowledge through vision and language and sound
Rowan Zellers, Jiasen Lu, Ximing Lu, Youngjae Yu, Yanpeng Zhao, Mohammadreza Salehi, Aditya Kusupati, Jack Hessel, Ali Farhadi, and Yejin Choi · 2022
Later among the works it cites.
Socratic models: Composing zero-shot multimodal reasoning with language
Andy Zeng, Adrian Wong, Stefan Welker, Krzysztof Choromanski, Federico Tombari, Aveek Purohit, Michael S Ryoo, Vikas Sindhwani, Johnny Lee, Vincent Vanhoucke, et al · 2022
Later among the works it cites.