Fetching the paper…
Reading the bibliography…
Benchmarks for large multimodal language models (MLMs) now serve to simultaneously assess the general capabilities of models instead of evaluating for a specific capability.
Prefixspan: Mining sequential patterns efficiently by prefix-projected pattern growth
Jiawei Han, Jian Pei, Behzad Mortazavi-Asl, Helen Pinto, Qiming Chen, Umeshwar Dayal, and Meichun Hsu · 2001
Earlier work this paper cites.
Bide: Efficient mining of frequent closed sequences
Jianyong Wang and Jiawei Han · 2004
Earlier work this paper cites.
Indoor segmentation and support inference from rgbd images
Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus · 2012
Earlier work this paper cites.
Blinkdb: queries with bounded errors and bounded response times on very large data
Sameer Agarwal, Barzan Mozafari, Aurojit Panda, Henry Milner, Samuel Madden, and Ion Stoica · 2013
Earlier work this paper cites.
String similarity joins: An experimental evaluation
Yu Jiang, Guoliang Li, Jianhua Feng, and Wen-Syan Li · 2014
Earlier work this paper cites.
Deep compositional question answering with neural module networks
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein · 2016
Earlier work this paper cites.
An empirical evaluation of set similarity join techniques
Willi Mann, Nikolaus Augsten, and Panagiotis Bouros · 2016
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
Gunnar A Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta · 2016
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2017
Earlier work this paper cites.
Vega-lite: A grammar of interactive graphics
Arvind Satyanarayan, Dominik Moritz, Kanit Wongsuphasawat, and Jeffrey Heer · 2017
Earlier work this paper cites.
Tallyqa: Answering complex counting questions
Manoj Acharya, Kushal Kafle, and Christopher Kanan · 2018
Earlier work this paper cites.
Blender - a 3d modelling and rendering package, 2018
Blender Online Community · 2018
Earlier work this paper cites.
Daniel Kang, Peter Bailis, and Matei Zaharia · 2018
Earlier work this paper cites.
Tvqa: Localized, compositional video question answering
Jie Lei, Licheng Yu, Mohit Bansal, and Tamara L. Berg · 2018
Earlier work this paper cites.
Altair: Interactive statistical visualizations for python
Jacob VanderPlas, Brian Granger, Jeffrey Heer, Dominik Moritz, Kanit Wongsuphasawat, Arvind Satyanarayan, Eitan Lees, Ilia Timofeev, Ben Welsh, and Scott Sievert · 2018
Earlier work this paper cites.
Mistique: A system to store and query model intermediates for model diagnosis
Manasi Vartak, Joana M F. da Trindade, Samuel Madden, and Matei Zaharia · 2018
Earlier work this paper cites.
Gradio: Hassle-free sharing and testing of ml models in the wild
Abubakar Abid, Ali Abdalla, Ali Abid, Dawood Khan, Abdulrahman Alfozan, and James Zou · 2019
Earlier work this paper cites.
Lvis: A dataset for large vocabulary instance segmentation
Agrim Gupta, Piotr Dollar, and Ross Girshick · 2019
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Drew A Hudson and Christopher D Manning · 2019
Earlier work this paper cites.
Model cards for model reporting
Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Earlier work this paper cites.
Efficient prefdiv algorithms for effective top-k result diversification
Xiaoyu Ge and Panos K Chrysanthis · 2020
Earlier work this paper cites.
Action genome: Actions as compositions of spatio-temporal scene graphs
Jingwei Ji, Ranjay Krishna, Li Fei-Fei, and Juan Carlos Niebles · 2020
Earlier work this paper cites.
Agqa: A benchmark for compositional spatio-temporal reasoning
Madeleine Grunde-McLaughlin, Ranjay Krishna, and Maneesh Agrawala · 2021
Earlier work this paper cites.
Deepeverest: Accelerating declarative top-k queries for deep neural network interpretation
Dong He, Maureen Daum, Walter Cai, and Magdalena Balazinska · 2021
Earlier work this paper cites.
Mind your outliers! investigating the negative impact of outliers on active learning for visual question answering
Siddharth Karamcheti, Ranjay Krishna, Li Fei-Fei, and Christopher D Manning · 2021
Earlier work this paper cites.
Dynabench: Rethinking benchmarking in NLP
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, and Adina Williams · 2021
Earlier work this paper cites.
Pervasive label errors in test sets destabilize machine learning benchmarks
Curtis G. Northcutt, Anish Athalye, and Jonas Mueller · 2021
Earlier work this paper cites.
Simvqa: Exploring simulated environments for visual question answering
Paola Cascante-Bonilla, Hui Wu, Letao Wang, Rogerio S Feris, and Vicente Ordonez · 2022
Earlier work this paper cites.
Measuring compositional consistency for video question answering
Mona Gandhi, Mustafa Omer Gul, Eva Prakash, Madeleine Grunde-McLaughlin, Ranjay Krishna, and Maneesh Agrawala · 2022
Earlier work this paper cites.
Data mining: concepts and techniques
Jiawei Han, Jian Pei, and Hanghang Tong · 2022
Earlier work this paper cites.
Solving diversity-aware maximum inner product search efficiently and effectively
Kohei Hirata, Daichi Amagata, Sumio Fujita, and Takahiro Hara · 2022
Earlier work this paper cites.
3db: A framework for debugging computer vision models
Guillaume Leclerc, Hadi Salman, Andrew Ilyas, Sai Vemprala, Logan Engstrom, Vibhav Vineet, Kai Xiao, Pengchuan Zhang, Shibani Santurkar, Greg Yang, et al · 2022
Earlier work this paper cites.
Openflamingo: An open-source framework for training large autoregressive vision-language models, 2023
Anas Awadalla, Irena Gao, Josh Gardner, Jack Hessel, Yusuf Hanafy, Wanrong Zhu, Kalyani Marathe, Yonatan Bitton, Samir Gadre, Shiori Sagawa, Jenia Jitsev, Simon Kornblith, Pang Wei Koh, Gabriel Ilharco, Mitchell Wortsman, and Ludwig Schmidt · 2023
Cited alongside, same era.
Qwen-vl: A frontier large vision-language model with versatile abilities
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou · 2023
Cited alongside, same era.
Lasagna: Layered score distillation for disentangled object relighting
Dina Bashkirova, Arijit Ray, Rupayan Mallick, Sarah Adel Bargal, Jianming Zhang, Ranjay Krishna, and Kate Saenko · 2023
Cited alongside, same era.
Misar: A multimodal instructional system with augmented reality, 2023
Jing Bi, Nguyen Manh Nguyen, Ali Vosoughi, and Chenliang Xu · 2023
Cited alongside, same era.
Video-bench: A comprehensive benchmark and toolkit for evaluating video-based large language models, 2023
Munan Ning, Bin Zhu, Yujia Xie, Bin Lin, Jiaxi Cui, Lu Yuan, Dongdong Chen, and Li Yuan · 2023
Later among the works it cites.
Gpt-4v(ision) system card
OpenAI · 2023
Later among the works it cites.
Kosmos-2: Grounding multimodal large language models to the world, 2023
Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, and Furu Wei · 2023
Later among the works it cites.
Unival: Unified model for image, video, audio and language tasks, 2023
Mustafa Shukor, Corentin Dancette, Alexandre Rame, and Matthieu Cord · 2023
Later among the works it cites.
Fine-grained audio-visual joint representations for multimodal large language models, 2023
Guangzhi Sun, Wenyi Yu, Changli Tang, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun Ma, and Chao Zhang · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rizhao Cai, Zirui Song, Dayan Guan, Zhenhao Chen, Xing Luo, Chenyu Yi, and Alex Kot · 2023
Cited alongside, same era.
Videollm: Modeling video sequence with large language models, 2023
Guo Chen, Yin-Dong Zheng, Jiahao Wang, Jilan Xu, Yifei Huang, Junting Pan, Yi Wang, Yali Wang, Yu Qiao, Tong Lu, and Limin Wang · 2023
Cited alongside, same era.
Grounding-prompter: Prompting llm with multimodal information for temporal sentence grounding in long videos, 2023
Houlun Chen, Xin Wang, Hong Chen, Zihan Song, Jia Jia, and Wenwu Zhu · 2023
Cited alongside, same era.
Pali-3 vision language models: Smaller, faster, stronger, 2023
Xi Chen, Xiao Wang, Lucas Beyer, Alexander Kolesnikov, Jialin Wu, Paul Voigtlaender, Basil Mustafa, Sebastian Goodman, Ibrahim Alabdulmohsin, Piotr Padlewski, Daniel Salz, Xi Xiong, Daniel Vlasic, Filip Pavetic, Keran Rong, Tianli Yu, Daniel Keysers, Xiaohua Zhai, and Radu Soricut · 2023
Cited alongside, same era.
Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, Bin Li, Ping Luo, Tong Lu, Yu Qiao, and Jifeng Dai · 2023
Cited alongside, same era.
Opencompass: A universal evaluation platform for foundation models
OpenCompass Contributors · 2023
Cited alongside, same era.
Holistic analysis of hallucination in gpt-4v(ision): Bias and interference challenges, 2023
Chenhang Cui, Yiyang Zhou, Xinyu Yang, Shirley Wu, Linjun Zhang, James Zou, and Huaxiu Yao · 2023
Cited alongside, same era.
Objaverse: A universe of annotated 3d objects
Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi · 2023
Cited alongside, same era.
Yunlong Tang, Jinrui Zhang, Xiangchen Wang, Teng Wang, and Feng Zheng · 2023
Later among the works it cites.
Chatvideo: A tracklet-centric multimodal and versatile video understanding system, 2023
Junke Wang, Dongdong Chen, Chong Luo, Xiyang Dai, Lu Yuan, Zuxuan Wu, and Yu-Gang Jiang · 2023
Later among the works it cites.
Internvid: A large-scale video-text dataset for multimodal understanding and generation
Yi Wang, Yinan He, Yizhuo Li, Kunchang Li, Jiashuo Yu, Xin Jian Ma, Xinyuan Chen, Yaohui Wang, Ping Luo, Ziwei Liu, Yali Wang, Limin Wang, and Y. Qiao · 2023
Later among the works it cites.
Gpt4video: A unified multimodal large language model for lnstruction-followed understanding and safety-aware generation, 2023
Zhanyu Wang, Longyue Wang, Zhen Zhao, Minghao Wu, Chenyang Lyu, Huayang Li, Deng Cai, Luping Zhou, Shuming Shi, and Zhaopeng Tu · 2023
Later among the works it cites.
A simple llm framework for long-range video question-answering
Ce Zhang, Taixi Lu, Md Mohaiminul Islam, Ziyang Wang, Shoubin Yu, Mohit Bansal, and Gedas Bertasius · 2023
Later among the works it cites.
Video-llama: An instruction-tuned audio-visual language model for video understanding
Hang Zhang, Xin Li, and Lidong Bing · 2023
Later among the works it cites.
Movqa: A benchmark of versatile question-answering for long-form movie understanding
Hongjie Zhang, Yi Liu, Lu Dong, Yifei Huang, Zhen-Hua Ling, Yali Wang, Limin Wang, and Yu Qiao · 2023
Later among the works it cites.
M3exam: A multilingual, multimodal, multilevel benchmark for examining large language models
Wenxuan Zhang, Sharifah Mahani Aljunied, Chang Gao, Yew Ken Chia, and Lidong Bing · 2023
Later among the works it cites.
Instructblip: Towards general-purpose vision-language models with instruction tuning
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale N Fung, and Steven Hoi · 2024
Closest in time.
Objaverse-xl: A universe of 10m+ 3d objects
Matt Deitke, Ruoshi Liu, Matthew Wallingford, Huong Ngo, Oscar Michel, Aditya Kusupati, Alan Fan, Christian Laforte, Vikram Voleti, Samir Yitzhak Gadre, et al · 2024
Closest in time.
Muffin or chihuahua? challenging large vision-language models with multipanel vqa
Yue Fan, Jing Gu, Kaiwen Zhou, Qianqi Yan, Shan Jiang, Ching-Chen Kuo, Xinze Guan, and Xin Eric Wang · 2024
Closest in time.
Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis
Chaoyou Fu, Yuhan Dai, Yondong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, Peixian Chen, Yanwei Li, Shaohui Lin, Sirui Zhao, Ke Li, Tong Xu, Xiawu Zheng, Enhong Chen, Rongrong Ji, and Xing Sun · 2024
Closest in time.
Blink: Multimodal large language models can see but not perceive
Xingyu Fu, Yushi Hu, Bangzheng Li, Yu Feng, Haoyu Wang, Xudong Lin, Dan Roth, Noah A Smith, Wei-Chiu Ma, and Ranjay Krishna · 2024
Closest in time.
Mllm-bench: Evaluating multimodal llms with per-sample criteria, 2024
Wentao Ge, Shunian Chen, Guiming Hardy Chen, Zhihong Chen, Junying Chen, Shuo Yan, Chenghao Zhu, Ziyue Lin, Wenya Xie, Xinyi Zhang, Yichen Chai, Xiaoyu Liu, Nuo Chen, Dingjie Song, Xidong Wang, Anningzhe Gao, Zhiyi Zhang, Jianquan Li, Xiang Wan, and Benyou Wang · 2024
Closest in time.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2024
Closest in time.
Prismer: A vision-language model with multi-task experts, 2024
Shikun Liu, Linxi Fan, Edward Johns, Zhiding Yu, Chaowei Xiao, and Anima Anandkumar · 2024
Closest in time.
Tempcompass: Do video llms really understand videos?
Yuanxin Liu, Shicheng Li, Yi Liu, Yuxiang Wang, Shuhuai Ren, Lei Li, Sishuo Chen, Xu Sun, and Lu Hou · 2024
Closest in time.
Object 3dit: Language-guided 3d-aware image editing
Oscar Michel, Anand Bhattad, Eli VanderBilt, Ranjay Krishna, Aniruddha Kembhavi, and Tanmay Gupta · 2024
Closest in time.
Lifelong benchmarks: Efficient model evaluation in an era of rapid progress
Ameya Prabhu, Vishaal Udandarao, Philip Torr, Matthias Bethge, Adel Bibi, and Samuel Albanie · 2024
Closest in time.
Lance: Stress-testing visual models by generating language-guided counterfactual images
Viraj Prabhu, Sriram Yenamandra, Prithvijit Chattopadhyay, and Judy Hoffman · 2024
Closest in time.
Cinepile: A long video question answering dataset and benchmark
Ruchit Rawal, Khalid Saifullah, Ronen Basri, David Jacobs, Gowthami Somepalli, and Tom Goldstein · 2024
Closest in time.
Generative multimodal models are in-context learners, 2024
Quan Sun, Yufeng Cui, Xiaosong Zhang, Fan Zhang, Qiying Yu, Zhengxiong Luo, Yueze Wang, Yongming Rao, Jingjing Liu, Tiejun Huang, and Xinlong Wang · 2024
Closest in time.
Emu: Generative pretraining in multimodality, 2024
Quan Sun, Qiying Yu, Yufeng Cui, Fan Zhang, Xiaosong Zhang, Yueze Wang, Hongcheng Gao, Jingjing Liu, Tiejun Huang, and Xinlong Wang · 2024
Closest in time.
Eyes wide shut? exploring the visual shortcomings of multimodal llms, 2024
Shengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma, Yann LeCun, and Saining Xie · 2024
Closest in time.
Can you rely on your model evaluation? improving model evaluation with synthetic test data
Boris van Breugel, Nabeel Seedat, Fergus Imrie, and Mihaela van der Schaar · 2024
Closest in time.
Internvideo2: Scaling video foundation models for multimodal video understanding
Yi Wang, Kunchang Li, Xinhao Li, Jiashuo Yu, Yinan He, Guo Chen, Baoqi Pei, Rongkun Zheng, Jilan Xu, Zun Wang, Yansong Shi, Tianxiang Jiang, Songze Li, Hongjie Zhang, Yifei Huang, Yu Qiao, Yali Wang, and Limin Wang · 2024
Closest in time.
Have seen me before? automating dataset updates towards reliable and timely evaluation
Jiahao Ying, Yixin Cao, Bo Wang, Wei Tang, Yizhe Yang, and Shuicheng Yan · 2024
Closest in time.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, Cong Wei, Botao Yu, Ruibin Yuan, Renliang Sun, Ming Yin, Boyuan Zheng, Zhenzhu Yang, Yibo Liu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen · 2024
Closest in time.
Benchmarking large multimodal models against common corruptions, 2024
Jiawei Zhang, Tianyu Pang, Chao Du, Yi Ren, Bo Li, and Min Lin · 2024
Closest in time.