Fetching the paper…
Reading the bibliography…
This paper presents GRASP, a novel benchmark to evaluate the language grounding and physical understanding capabilities of video-based multimodal large language models (LLMs).
Object permanence in five-month-old infants
Renée Baillargeon, Elizabeth S. Spelke, and Stanley Wasserman · 1985
Earlier work this paper cites.
Object permanence in 3½- and 4½-month-old infants
Renée Baillargeon · 1987
Earlier work this paper cites.
Infants’ sensitivity to effects of gravity on visible object motion
In K. Kim and Elizabeth S. Spelke · 1992
Earlier work this paper cites.
Origins of knowledge
Elizabeth S. Spelke, Karen Breinlinger, Janet Macomber, and Kristen Jacobson · 1992
Earlier work this paper cites.
Early knowledge of object motion: continuity and inertia
Elizabeth S. Spelke, Gary Katz, Susan E. Purcell, Sheryl M. Ehrlich, and Karen Breinlinger · 1994
Earlier work this paper cites.
Physical reasoning in infancy
Renee Baillargeon · 1995
Earlier work this paper cites.
Perception and understanding of effects of gravity and inertia on object motion
In K. Kim and Elizabeth S. Spelke · 1999
Earlier work this paper cites.
Object individuation and physical reasoning in infancy: An integrative account
Renée Baillargeon, Maayan Stavans, Di Wu, Yael Gertner, Peipei Setoh, Audrey K. Kittredge, and Amélie Bernard · 2012
Earlier work this paper cites.
SemEval-2012 task 7: Choice of plausible alternatives: An evaluation of commonsense causal reasoning
Andrew Gordon, Zornitsa Kozareva, and Melissa Roemmele · 2012
Earlier work this paper cites.
Midge: Generating image descriptions from computer vision detections
Margaret Mitchell, Jesse Dodge, Amit Goyal, Kota Yamaguchi, Karl Stratos, Xufeng Han, Alyssa Mensch, Alex Berg, Tamara Berg, and Hal Daumé III · 2012
Earlier work this paper cites.
VQA: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Learning to poke by poking: Experiential learning of intuitive physics
Pulkit Agrawal, Ashvin V. Nair, Pieter Abbeel, Jitendra Malik, and Sergey Levine · 2016
Earlier work this paper cites.
Interaction networks for learning about objects, relations and physics
Peter Battaglia, Razvan Pascanu, Matthew Lai, Danilo Jimenez Rezende, and Koray kavukcuoglu · 2016
Earlier work this paper cites.
Learning physical intuition of block towers by example
Adam Lerer, Sam Gross, and Rob Fergus · 2016
Earlier work this paper cites.
CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick · 2017
Cited alongside, same era.
Visual interaction networks: Learning a physics simulator from video
Nicholas Watters, Andrea Tacchetti, Théophane Weber, Razvan Pascanu, Peter Battaglia, and Daniel Zoran · 2017
Cited alongside, same era.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sünderhauf, Ian Reid, Stephen Gould, and Anton van den Hengel · 2018
Cited alongside, same era.
Embodied question answering
Abhishek Das, Samyak Datta, Georgia Gkioxari, Stefan Lee, Devi Parikh, and Dhruv Batra · 2018
Cited alongside, same era.
From recognition to cognition: Visual commonsense reasoning
Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Cited alongside, same era.
Intphys 2019: A benchmark for visual intuitive physics understanding
Ronan Riochet, Mario Ynocente Castro, Mathieu Bernard, Adam Lerer, Rob Fergus, Véronique Izard, and Emmanuel Dupoux · 2022
Later among the works it cites.
Benchmarking progress to infant-level physical reasoning in AI
Luca Weihs, Amanda Yuile, Renée Baillargeon, Cynthia Fisher, Gary Marcus, Roozbeh Mottaghi, and Aniruddha Kembhavi · 2022
Later among the works it cites.
Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing · 2023
Closest in time.
Imagebind: One embedding space to bind them all
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra · 2023
Closest in time.
Vtimellm: Empower llm to grasp video moments, 2023
Bin Huang, Xin Wang, Hong Chen, Zihan Song, and Wenwu Zhu · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yonatan Bisk, Ari Holtzman, Jesse Thomason, Jacob Andreas, Yoshua Bengio, Joyce Chai, Mirella Lapata, Angeliki Lazaridou, Jonathan May, Aleksandr Nisnevich, Nicolas Pinto, and Joseph Turian · 2020
Cited alongside, same era.
CLEVRER: Collision events for video representation and reasoning
Kexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli, Jiajun Wu, Antonio Torralba, and Joshua B. Tenenbaum · 2020
Cited alongside, same era.
A benchmark for modeling violation-of-expectation in physical reasoning across event categories
Arijit Dasgupta, Jiafei Duan, Marcelo H Ang Jr, Yi Lin, Su-hua Wang, Renée Baillargeon, and Cheston Tan · 2021
Cited alongside, same era.
Avoe: A synthetic 3d dataset on understanding violation of expectation for artificial cognition
Arijit Dasgupta, Jiafei Duan, Marcelo H Ang Jr, and Cheston Tan · 2021
Cited alongside, same era.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karen Simonyan · 2022
Cited alongside, same era.
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi · 2023
Closest in time.
Videochat: Chat-centric video understanding
KunChang Li, Yinan He, Yi Wang, Yizhuo Li, Wenhai Wang, Ping Luo, Yali Wang, Limin Wang, and Yu Qiao · 2023
Closest in time.
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2023
Closest in time.
Video-ChatGPT: Towards detailed video understanding via large vision and language models
Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan · 2023
Closest in time.
A comprehensive overview of large language models, 2023
Humza Naveed, Asad Ullah Khan, Shi Qiu, Muhammad Saqib, Saeed Anwar, Muhammad Usman, Naveed Akhtar, Nick Barnes, and Ajmal Mian · 2023
Closest in time.
PandaGPT: One model to instruction-follow them all
Yixuan Su, Tian Lan, Huayang Li, Jialu Xu, Yan Wang, and Deng Cai · 2023
Closest in time.
Video-LLaMa: An instruction-tuned audio-vidual language model for video understanding
Hang Zhang, Xin Li, and Lidong Bing · 2023
Closest in time.
MiniGPT-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny · 2023
Closest in time.