Fetching the paper…
Reading the bibliography…
Recent advancements in video-based large language models (Video LLMs) have witnessed the emergence of diverse capabilities to reason and interpret dynamic visual content.
Conceptual precursors to language
Susan J Hespos and Elizabeth S Spelke · 2004
Earlier work this paper cites.
A framework for the semi-automatic testing of video games
Alfredo Nantes, Ross Brown, and Frederic Maire · 2008
Earlier work this paper cites.
Simulation as an engine of physical scene understanding
Peter W Battaglia, Jessica B Hamrick, and Joshua B Tenenbaum · 2013
Earlier work this paper cites.
Learning visual predictive models of physics for playing billiards
Katerina Fragkiadaki, Pulkit Agrawal, Sergey Levine, and Jitendra Malik · 2015
Earlier work this paper cites.
Learning to poke by poking: Experiential learning of intuitive physics
Pulkit Agrawal, Ashvin V Nair, Pieter Abbeel, Jitendra Malik, and Sergey Levine · 2016
Earlier work this paper cites.
Interaction networks for learning about objects, relations and physics
Peter Battaglia, Razvan Pascanu, Matthew Lai, Danilo Jimenez Rezende, et al · 2016
Earlier work this paper cites.
A compositional object-based approach to learning physical dynamics
Michael B Chang, Tomer Ullman, Antonio Torralba, and Joshua B Tenenbaum · 2016
Earlier work this paper cites.
Unsupervised learning for physical interaction through video prediction
Chelsea Finn, Ian Goodfellow, and Sergey Levine · 2016
Earlier work this paper cites.
Learning physical intuition of block towers by example
Adam Lerer, Sam Gross, and Rob Fergus · 2016
Earlier work this paper cites.
“what happens if…” learning to predict the effect of forces in images
Roozbeh Mottaghi, Mohammad Rastegari, Abhinav Gupta, and Ali Farhadi · 2016
Earlier work this paper cites.
Se3-pose-nets: Structured deep dynamics models for visuomotor planning and control
Arunkumar Byravan, Felix Leeb, Franziska Meier, and Dieter Fox · 2017
Earlier work this paper cites.
Self-supervised visual planning with temporal skip connections
Frederik Ebert, Chelsea Finn, Alex X Lee, and Sergey Levine · 2017
Earlier work this paper cites.
Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Visual interaction networks: Learning a physics simulator from video
Nicholas Watters, Daniel Zoran, Theophane Weber, Peter Battaglia, Razvan Pascanu, and Andrea Tacchetti · 2017
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion
Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang · 2017
Earlier work this paper cites.
Tvqa: Localized, compositional video question answering
Jie Lei, Licheng Yu, Mohit Bansal, and Tamara L Berg · 2018
Earlier work this paper cites.
Virtualhome: Simulating household activities via programs
Xavier Puig, Kevin Ra, Marko Boben, Jiaman Li, Tingwu Wang, Sanja Fidler, and Antonio Torralba · 2018
Earlier work this paper cites.
Differentiable physics and stable modes for tool-use and manipulation planning
Marc A Toussaint, Kelsey Rebecca Allen, Kevin A Smith, and Joshua B Tenenbaum · 2018
Earlier work this paper cites.
Interpretable intuitive physics model
Tian Ye, Xiaolong Wang, James Davidson, and Abhinav Gupta · 2018
Earlier work this paper cites.
On the utility of learning about humans for human-ai coordination
Micah Carroll, Rohin Shah, Mark K Ho, Tom Griffiths, Sanjit Seshia, Pieter Abbeel, and Anca Dragan · 2019
Earlier work this paper cites.
Tool macgyvering: Tool construction using geometric reasoning
Lakshmi Nair, Jonathan Balloch, and Sonia Chernova · 2019
Earlier work this paper cites.
Activitynet-qa: A dataset for understanding complex web videos via question answering
Zhou Yu, Dejing Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, and Dacheng Tao · 2019
Earlier work this paper cites.
Wuji: Automatic online combat game testing using evolutionary deep reinforcement learning
Yan Zheng, Xiaofei Xie, Ting Su, Lei Ma, Jianye Hao, Zhaopeng Meng, Yang Liu, Ruimin Shen, Yingfeng Chen, and Changjie Fan · 2019
Earlier work this paper cites.
Craft: A benchmark for causal reasoning about forces and interactions
Tayfun Ates, M Samil Atesoglu, Cagatay Yigit, Ilker Kesen, Mert Kobas, Erkut Erdem, Aykut Erdem, Tilbe Goksun, and Deniz Yuret · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Hero: Hierarchical encoder for video+ language omni-representation pre-training
Linjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan, Licheng Yu, and Jingjing Liu · 2020
Earlier work this paper cites.
Using deep convolutional neural networks to detect rendered glitches in video games
Carlos Ling, Konrad Tollmar, and Linus Gisslén · 2020
Earlier work this paper cites.
A video game testing method utilizing deep learning
Mohammad Reza Taesiri, Moslem Habibi, and Mohammad Amin Fazli · 2020
Earlier work this paper cites.
Clevrer: Collision events for video representation and reasoning
Kexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli, Jiajun Wu, Antonio Torralba, and Joshua B Tenenbaum · 2020
Earlier work this paper cites.
Cophy: Counterfactual learning of physical dynamics
Fabien Baradel, Natalia Neverova, Julien Mille, Greg Mori, and Christian Wolf · 2021
Earlier work this paper cites.
Physion: Evaluating physical prediction from vision in humans and machines
Daniel Bear, Elias Wang, Damian Mrowca, Felix Jedidja Binder, Hsiao-Yu Tung, RT Pramod, Cameron Holdaway, Sirui Tao, Kevin A Smith, Fan-Yun Sun, et al · 2021
Earlier work this paper cites.
On pursuit of designing multi-modal transformer for video grounding
Meng Cao, Long Chen, Mike Zheng Shou, Can Zhang, and Yuexian Zou · 2021
Cited alongside, same era.
Glib: towards automated test oracle for graphically-rich applications
Ke Chen, Yufei Li, Yingfeng Chen, Changjie Fan, Zhipeng Hu, and Wei Yang · 2021
Cited alongside, same era.
Cater: A diagnostic dataset for compositional actions & temporal reasoning
Rohit Girdhar and Deva Ramanan · 2021
Cited alongside, same era.
Video game industry market analysis: Approaches that resulted in industry success and high demand
Savelii Pashkov · 2021
Cited alongside, same era.
Watch-and-help: A challenge for social perception and human-ai collaboration
Xavier Puig, Tianmin Shu, Shuang Li, Zilin Wang, Yuan-Hong Liao, Joshua B Tenenbaum, Sanja Fidler, and Antonio Torralba · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Weak supervision for label efficient visual bug detection
Farrukh Rahman · 2023
Later among the works it cites.
A deep reinforcement learning technique for bug detection in video games
Geeta Rani, Upasana Pandey, Aniket Anil Wagde, and Vijaypal Singh Dhaka · 2023
Later among the works it cites.
Avalon’s game of thoughts: Battle against deception through recursive contemplation
Shenzhi Wang, Chang Liu, Zilong Zheng, Siyuan Qi, Shuo Chen, Qisen Yang, Andrew Zhao, Chaofei Wang, Shiji Song, and Gao Huang · 2023
Later among the works it cites.
Concept-aware video captioning: Describing videos with effective prior information
Bang Yang, Meng Cao, and Yuexian Zou · 2023
Later among the works it cites.
Qilin-med: Multi-stage knowledge injection advanced medical large language model
Qichen Ye, Junling Liu, Dading Chong, Peilin Zhou, Yining Hua, Fenglin Liu, Meng Cao, Ziming Wang, Xuxin Cheng, Zhu Lei, et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Next-qa: Next phase of question-answering to explaining temporal actions
Junbin Xiao, Xindi Shang, Angela Yao, and Tat-Seng Chua · 2021
Cited alongside, same era.
Rr-net: Relation reasoning for end-to-end human-object interaction detection
Dongming Yang, Yuexian Zou, Can Zhang, Meng Cao, and Jie Chen · 2021
Cited alongside, same era.
A survey on in-context learning
Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Jingyuan Ma, Rui Li, Heming Xia, Jingjing Xu, Zhiyong Wu, Tianyu Liu, et al · 2022
Cited alongside, same era.
Video referring expression comprehension via transformer with content-aware query
Ji Jiang, Meng Cao, Tengtao Song, and Yuexian Zou · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
Chess as a testbed for language model state tracking
Shubham Toshniwal, Sam Wiseman, Karen Livescu, and Kevin Gimpel · 2022
Cited alongside, same era.
Later among the works it cites.
Video-llama: An instruction-tuned audio-visual language model for video understanding
Hang Zhang, Xin Li, and Lidong Bing · 2023
Later among the works it cites.
Claude 3.5 sonnet, 2024
Anthropic · 2024
Closest in time.
Videophy: Evaluating physical commonsense for video generation
Hritik Bansal, Zongyu Lin, Tianyi Xie, Zeshun Zong, Michal Yarom, Yonatan Bitton, Chenfanfu Jiang, Yizhou Sun, Kai-Wei Chang, and Aditya Grover · 2024
Closest in time.
S-agents: Self-organizing agents in open-ended environments
Jiaqi Chen, Yuxian Jiang, Jiachen Lu, and Li Zhang · 2024
Closest in time.
How to continually adapt text-to-image diffusion models for flexible customization?
Jiahua Dong, Wenqi Liang, Hongliu Li, Duzhen Zhang, Meng Cao, Henghui Ding, Salman Khan, and Fahad Shahbaz Khan · 2024
Closest in time.
Towards event-oriented long video understanding
Yifan Du, Kun Zhou, Yuqi Huo, Yifan Li, Wayne Xin Zhao, Haoyu Lu, Zijia Zhao, Bingning Wang, Weipeng Chen, and Ji-Rong Wen · 2024
Closest in time.
Chessgpt: Bridging policy learning and language modeling
Xidong Feng, Yicheng Luo, Ziyan Wang, Hongrui Tang, Mengyue Yang, Kun Shao, David Mguni, Yali Du, and Jun Wang · 2024
Closest in time.
Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis
Chaoyou Fu, Yuhan Dai, Yondong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, et al · 2024
Closest in time.
Vila: On pre-training for visual language models
Ji Lin, Hongxu Yin, Wei Ping, Pavlo Molchanov, Mohammad Shoeybi, and Song Han · 2024
Closest in time.
Textual inversion and self-supervised refinement for radiology report generation
Yuanjiang Luo, Hongxiang Li, Xuan Wu, Meng Cao, Xiaoshuang Huang, Zhihong Zhu, Peixi Liao, Hu Chen, and Yi Zhang · 2024
Closest in time.
Towards world simulator: Crafting physical commonsense-based benchmark for video generation
Fanqing Meng, Jiaqi Liao, Xinyu Tan, Wenqi Shao, Quanfeng Lu, Kaipeng Zhang, Yu Cheng, Dianqi Li, Yu Qiao, and Ping Luo · 2024
Closest in time.
Hello gpt-4o
OpenAI · 2024
Closest in time.
Perception test: A diagnostic benchmark for multimodal video models
Viorica Patraucean, Lucas Smaira, Ankush Gupta, Adria Recasens, Larisa Markeeva, Dylan Banarse, Skanda Koppula, Mateusz Malinowski, Yi Yang, Carl Doersch, et al · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean-baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, et al · 2024
Closest in time.
Swarmbrain: Embodied agent for real-time strategy game starcraft ii via large language models
Xiao Shao, Weifu Jiang, Fei Zuo, and Mengqing Liu · 2024
Closest in time.
Moviechat: From dense token to sparse memory for long video understanding
Enxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang, Haoyang Zhou, Feiyang Wu, Haozhe Chi, Xun Guo, Tian Ye, Yanting Zhang, et al · 2024
Closest in time.
Videogamebunny: Towards vision assistants for video games
Mohammad Reza Taesiri and Cor-Paul Bezemer · 2024
Closest in time.
Muse: Mamba is efficient multi-scale learner for text-video retrieval
Haoran Tang, Meng Cao, Jinfa Huang, Ruyang Liu, Peng Jin, Ge Li, and Xiaodan Liang · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al · 2024
Closest in time.
Physion++: Evaluating physical scene understanding that requires online inference of different physical properties
Hsiao-Yu Tung, Mingyu Ding, Zhenfang Chen, Daniel Bear, Chuang Gan, Josh Tenenbaum, Dan Yamins, Judith Fan, and Kevin Smith · 2024
Closest in time.
Lvbench: An extreme long video understanding benchmark
Weihan Wang, Zehai He, Wenyi Hong, Yean Cheng, Xiaohan Zhang, Ji Qi, Shiyu Huang, Bin Xu, Yuxiao Dong, Ming Ding, et al · 2024
Closest in time.
Longvideobench: A benchmark for long-context interleaved video-language understanding
Haoning Wu, Dongxu Li, Bei Chen, and Junnan Li · 2024
Closest in time.
A survey on game playing agents and large models: Methods, applications, and challenges
Xinrun Xu, Yuxin Wang, Chaoyi Xu, Ziluo Ding, Jiechuan Jiang, Zhiming Ding, and Börje F Karlsson · 2024
Closest in time.
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, et al · 2024
Closest in time.
Uncertainty-aware sign language video retrieval with probability distribution modeling
Xuan Wu, Hongxiang Li, Yuanjiang Luo, Xuxin Cheng, Xianwei Zhuang, Meng Cao, and Keren Fu · 2025
Closest in time.