Fetching the paper…
Reading the bibliography…
This paper considers the problem of Multi-Hop Video Question Answering (MH-VidQA) in long-form egocentric videos.
Bleu: a method for automatic evaluation of machine translation
Papineni, K.; Roukos, S.; Ward, T.; and Zhu, W.-J. 2002 · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y. 2004 · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Banerjee, S.; and Lavie, A. 2005 · 2005
Earlier work this paper cites.
MovieQA: Understanding Stories in Movies through Question-Answering
Tapaswi, M.; Zhu, Y.; Stiefelhagen, R.; Torralba, A.; Urtasun, R.; and Fidler, S. 2016 · 2016
Earlier work this paper cites.
How2: A Large-scale Dataset for Multimodal Language Understanding
Sanabria, R.; Caglayan, O.; Palaskar, S.; Elliott, D.; Barrault, L.; Specia, L.; and Metze, F. 2018 · 2018
Earlier work this paper cites.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Yang, Z.; Qi, P.; Zhang, S.; Bengio, Y.; Cohen, W. W.; Salakhutdinov, R.; and Manning, C. D. 2018 · 2018
Earlier work this paper cites.
Are we asking the right questions in MovieQA?
Jasani, B.; Girdhar, R.; and Ramanan, D. 2019 · 2019
Earlier work this paper cites.
HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips
Miech, A.; Zhukov, D.; Alayrac, J.-B.; Tapaswi, M.; Laptev, I.; and Sivic, J. 2019 · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Reimers, N.; and Gurevych, I. 2019 · 2019
Earlier work this paper cites.
ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering
Yu, Z.; Xu, D.; Yu, J.; Yu, T.; Zhao, Z.; Zhuang, Y.; and Tao, D. 2019 · 2019
Earlier work this paper cites.
Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps
Ho, X.; Nguyen, A.-K. D.; Sugawara, S.; and Aizawa, A. 2020 · 2020
Earlier work this paper cites.
spaCy: Industrial-strength natural language processing in python
Honnibal, M.; Montani, I.; Van Landeghem, S.; Boyd, A.; et al. 2020 · 2020
Earlier work this paper cites.
Action genome: Actions as compositions of spatio-temporal scene graphs
Ji, J.; Krishna, R.; Fei-Fei, L.; and Niebles, J. C. 2020 · 2020
Earlier work this paper cites.
AGQA: A Benchmark for Compositional Spatio-Temporal Reasoning
Grunde-McLaughlin, M.; Krishna, R.; and Agrawala, M. 2021 · 2021
Earlier work this paper cites.
STAR: A Benchmark for Situated Reasoning in Real-World Videos
Wu, B.; Yu, S.; Chen, Z.; Tenenbaum, J. B.; and Gan, C. 2021 · 2021
Earlier work this paper cites.
NExT-QA:Next Phase of Question-Answering to Explaining Temporal Actions
Xiao, J.; Shang, X.; Yao, A.; and Chua, T.-S. 2021 · 2021
Earlier work this paper cites.
Answering Complex Open-Domain Questions with Multi-Hop Dense Retrieval
Xiong, W.; Li, X.; Iyer, S.; Du, J.; Lewis, P.; Wang, W. Y.; Mehdad, Y.; Yih, S.; Riedel, S.; Kiela, D.; et al. 2021 · 2021
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B.; Donahue, J.; Luc, P.; Miech, A.; Barr, I.; Hasson, Y.; Lenc, K.; Mensch, A.; Millican, K.; Reynolds, M.; et al. 2022 · 2022
Cited alongside, same era.
Where did i leave my keys?-episodic-memory-based question answering on egocentric videos
Bärmann, L.; and Waibel, A. 2022 · 2022
Cited alongside, same era.
Ego4D: Around the World in 3,000 Hours of Egocentric Video
Grauman, K.; Westbury, A.; Byrne, E.; Chavis, Z.; Furnari, A.; Girdhar, R.; Hamburger, J.; Jiang, H.; Liu, M.; Liu, X.; Martin, M.; Nagarajan, T.; Radosavovic, I.; Ramakrishnan, S. K.; Ryan, F.; Sharma, J.; Wray, M.; Xu, M.; Xu, E. Z.; Zhao, C.; Bansal, S.; Batra, D.; Cartillier, V.; Crane, S.; Do, T.; Doulaty, M.; Erapalli, A.; Feichtenhofer, C.; Fragomeni, A.; Fu, Q.; Gebreselasie, A.; Gonzalez, C.; Hillis, J.; Huang, X.; Huang, Y.; Jia, W.; Khoo, W.; Kolar, J.; Kottur, S.; Kumar, A.; Landini, F.; Li, C.; Li, Y.; Li, Z.; Mangalam, K.; Modhugu, R.; Munro, J.; Murrell, T.; Nishiyasu, T.; Price, W.; Puentes, P. R.; Ramazanova, M.; Sari, L.; Somasundaram, K.; Southerland, A.; Sugano, Y.; Tao, R.; Vo, M.; Wang, Y.; Wu, X.; Yagi, T.; Zhao, Z.; Zhu, Y.; Arbelaez, P.; Crandall, D.; Damen, D.; Farinella, G. M.; Fuegen, C.; Ghanem, B.; Ithapu, V. K.; Jawahar, C. V.; Joo, H.; Kitani, K.; Li, H.; Newcombe, R.; Oliva, A.; Park, H. S.; Rehg, J. M.; Sato, Y.; Shi, J.; Shou, M. Z.; Torralba, A.; Torresani, L.; Yan, M.; and Malik, J. 2022 · 2022
Llama 3 Model Card
AI@Meta. 2024 · 2024
Closest in time.
Ataallah, K.; Shen, X.; Abdelrahman, E.; Sleiman, E.; Zhu, D.; Ding, J.; and Elhoseiny, M. 2024 · 2024
Closest in time.
Grounded Question-Answering in Long Egocentric Videos
Di, S.; and Xie, W. 2024 · 2024
Closest in time.
Jiang, A. Q.; Sablayrolles, A.; Roux, A.; Mensch, A.; Savary, B.; Bamford, C.; Chaplot, D. S.; Casas, D. d. l.; Hanna, E. B.; Bressand, F.; et al. 2024 · 2024
Closest in time.
Lisa: Reasoning segmentation via large language model
Lai, X.; Tian, Z.; Chen, Y.; Li, Y.; Yuan, Y.; Liu, S.; and Jia, J. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
MuSiQue: Multihop Questions via Single-hop Question Composition
Trivedi, H.; Balasubramanian, N.; Khot, T.; and Sabharwal, A. 2022 · 2022
Cited alongside, same era.
Internvideo: General video foundation models via generative and discriminative learning
Wang, Y.; Li, K.; Li, Y.; He, Y.; Huang, B.; Zhao, Z.; Zhang, H.; Xu, J.; Liu, Y.; Wang, Z.; et al. 2022 · 2022
Cited alongside, same era.
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023 · 2023
Cited alongside, same era.
Can Large Language Models Be an Alternative to Human Evaluations?
Chiang, C.-H.; and Lee, H.-y. 2023 · 2023
Cited alongside, same era.
Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality
Chiang, W.-L.; Li, Z.; Lin, Z.; Sheng, Y.; Wu, Z.; Zhang, H.; Zheng, L.; Zhuang, S.; Zhuang, Y.; Gonzalez, J. E.; Stoica, I.; and Xing, E. P. 2023 · 2023
Cited alongside, same era.
Univtg: Towards unified video-language temporal grounding
Lin, K. Q.; Zhang, P.; Chen, J.; Pramanick, S.; Gao, D.; Wang, A. J.; Yan, R.; and Shou, M. Z. 2023 · 2023
Cited alongside, same era.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Lu, P.; Bansal, H.; Xia, T.; Liu, J.; Li, C.; Hajishirzi, H.; Cheng, H.; Chang, K.-W.; Galley, M.; and Gao, J. 2023 · 2023
Cited alongside, same era.
EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding
Mangalam, K.; Akshulakov, R.; and Malik, J. 2023 · 2023
Cited alongside, same era.
Liu, H.; Li, C.; Li, Y.; Li, B.; Zhang, Y.; Shen, S.; and Lee, Y. J. 2024 · 2024
Closest in time.
SnAG: Scalable and Accurate Video Grounding
Mu, F.; Mo, S.; and Li, Y. 2024 · 2024
Closest in time.
Perception test: A diagnostic benchmark for multimodal video models
Patraucean, V.; Smaira, L.; Gupta, A.; Recasens, A.; Markeeva, L.; Banarse, D.; Koppula, S.; Malinowski, M.; Yang, Y.; Doersch, C.; et al. 2024 · 2024
Closest in time.
Momentor: Advancing video large language model with fine-grained temporal reasoning
Qian, L.; Li, J.; Wu, Y.; Ye, Y.; Fei, H.; Chua, T.-S.; Zhuang, Y.; and Tang, S. 2024 · 2024
Closest in time.
Timechat: A time-sensitive multimodal large language model for long video understanding
Ren, S.; Yao, L.; Li, S.; Sun, X.; and Hou, L. 2024 · 2024
Closest in time.
Action Scene Graphs for Long-Form Understanding of Egocentric Videos
Rodin, I.; Furnari, A.; Min, K.; Tripathi, S.; and Farinella, G. M. 2024 · 2024
Closest in time.
Can i trust your answer? visually grounded video question answering
Xiao, J.; Yao, A.; Li, Y.; and Chua, T.-S. 2024 · 2024
Closest in time.
VISA: Reasoning Video Object Segmentation via Large Language Models
Yan, C.; Wang, H.; Yan, S.; Jiang, X.; Hu, Y.; Kang, G.; Xie, W.; and Gavves, E. 2024 · 2024
Closest in time.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Yue, X.; Ni, Y.; Zhang, K.; Zheng, T.; Liu, R.; Zhang, G.; Stevens, S.; Jiang, D.; Ren, W.; Sun, Y.; et al. 2024 · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E.; et al. 2024 · 2024
Closest in time.