Fetching the paper…
Reading the bibliography…
We introduce CameraBench, a large-scale dataset and benchmark designed to assess and improve camera motion understanding.
A grammar of the film: An analysis of film technique
Raymond Spottiswoode · 1969
Earlier work this paper cites.
Motion parallax and absolute distance
Steven H Ferris · 1972
Earlier work this paper cites.
Motion parallax as an independent cue for depth perception
Brian Rogers and Maureen Graham · 1979
Earlier work this paper cites.
Ambiguity in structure from motion: Sphere versus plane
Cornelia Fermüller and Yiannis Aloimonos · 1998
Earlier work this paper cites.
Optimal structure from motion: Local ambiguities and global estimates
Alessandro Chiuso, Roger Brockett, and Stefano Soatto · 2000
Earlier work this paper cites.
The ecological approach to visual perception
James J Gibson · 2003
Earlier work this paper cites.
Monoslam: Real-time single camera slam
Andrew J Davison, Ian D Reid, Nicholas D Molton, and Olivier Stasse · 2007
Earlier work this paper cites.
Cinematic journeys: Film and movement
Dimitris Eleftheriotis · 2010
Earlier work this paper cites.
Understanding noise sensitivity in structure from motion
Kostas Daniilidis and Minas E Spetsakis · 2013
Earlier work this paper cites.
Lsd-slam: Large-scale direct monocular slam
Jakob Engel, Thomas Schöps, and Daniel Cremers · 2014
Earlier work this paper cites.
Structure-from-motion revisited
Johannes L Schonberger and Jan-Michael Frahm · 2016
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Earlier work this paper cites.
Visual slam algorithms: A survey from 2010 to 2016
Takafumi Taketomi, Hideaki Uchiyama, and Sei Ikeda · 2017
Earlier work this paper cites.
Stereo magnification: Learning view synthesis using multiplane images
Tinghui Zhou, Richard Tucker, John Flynn, Graham Fyffe, and Noah Snavely · 2018
Earlier work this paper cites.
Types of camera movements in film explained: Definitive guide, 2020
Kyle Deguzman · 2020
Earlier work this paper cites.
Movienet: A holistic dataset for movie understanding
Qingqiu Huang, Yu Xiong, Anyi Rao, Jiaze Wang, and Dahua Lin · 2020
Earlier work this paper cites.
A unified framework for shot type classification based on subject centric lens
Anyi Rao, Jiaze Wang, Linning Xu, Xuekun Jiang, Qingqiu Huang, Bolei Zhou, and Dahua Lin · 2020
Earlier work this paper cites.
Clipscore: A reference-free evaluation metric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
The anatomy of video editing: A dataset and benchmark suite for ai-assisted video editing
Dawit Mureja Argaw, Fabian Caba Heilbron, Joon-Young Lee, Markus Woodson, and In So Kweon · 2022
Cited alongside, same era.
Winoground: Probing vision and language models for visio-linguistic compositionality
Tristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh, Adina Williams, Douwe Kiela, and Candace Ross · 2022
Cited alongside, same era.
Structure and motion from casual videos
Zhoutong Zhang, Forrester Cole, Zhengqi Li, Michael Rubinstein, Noah Snavely, and William T Freeman · 2022
Cited alongside, same era.
Multimodality helps unimodality: Cross-modal few-shot learning with multimodal models, 2023
Zhiqiu Lin, Samuel Yu, Zhiyi Kuang, Deepak Pathak, and Deva Ramanan · 2023
Cited alongside, same era.
OpenAI · 2023
Cited alongside, same era.
Movie gen: A cast of media foundation models
Adam Polyak, Amit Zohar, Andrew Brown, Andros Tjandra, Animesh Sinha, Ann Lee, Apoorv Vyas, Bowen Shi, Chih-Yao Ma, Ching-Yao Chuang, et al · 2024
Later among the works it cites.
Transnet v2: An effective deep network architecture for fast shot transition detection
Tomás Soucek and Jakub Lokoc · 2024
Later among the works it cites.
Vidcomposition: Can mllms analyze compositions in compiled videos?
Yunlong Tang, Junjia Guo, Hang Hua, Susan Liang, Mingqian Feng, Xinyang Li, Rui Mao, Chao Huang, Jing Bi, Zeliang Zhang, et al · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Languagebind: Extending video-language pretraining to n-modality by language-based semantic alignment, 2023
Bin Zhu, Bin Lin, Munan Ning, Yang Yan, Jiaxi Cui, Wang HongFa, Yatian Pang, Wenhao Jiang, Junwu Zhang, Zongwei Li, Cai Wan Zhang, Zhifeng Li, Wei Liu, and Li Yuan · 2023
Cited alongside, same era.
Auroracap: Efficient, performant video detailed captioning and a new benchmark
Wenhao Chai, Enxin Song, Yilun Du, Chenlin Meng, Vashisht Madhavan, Omer Bar-Tal, Jenq-Neng Hwang, Saining Xie, and Christopher D Manning · 2024
Cited alongside, same era.
Boosting camera motion control for video diffusion transformers
Soon Yau Cheong, Duygu Ceylan, Armin Mustafa, Andrew Gilbert, and Chun-Hao Paul Huang · 2024
Cited alongside, same era.
Et the exceptional trajectories: Text-to-camera-trajectory generation with character awareness
Robin Courant, Nicolas Dufour, Xi Wang, Marc Christie, and Vicky Kalogeiton · 2024
Cited alongside, same era.
Mast3r-sfm: a fully-integrated solution for unconstrained structure-from-motion
Bardienus Duisterhof, Lojze Zust, Philippe Weinzaepfel, Vincent Leroy, Yohann Cabon, and Jerome Revaud · 2024
Cited alongside, same era.
Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives
Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, et al · 2024
Cited alongside, same era.
Cameractrl: Enabling camera control for text-to-video generation
Hao He, Yinghao Xu, Yuwei Guo, Gordon Wetzstein, Bo Dai, Hongsheng Li, and Ceyuan Yang · 2024
Cited alongside, same era.
Zeqi Xiao, Wenqi Ouyang, Yifan Zhou, Shuai Yang, Lei Yang, Jianlou Si, and Xingang Pan · 2024
Later among the works it cites.
Direct-a-video: Customized video generation with user-directed camera movement and object motion
Shiyuan Yang, Liang Hou, Haibin Huang, Chongyang Ma, Pengfei Wan, Di Zhang, Xiaodong Chen, and Jing Liao · 2024
Later among the works it cites.
Latent-reframe: Enabling camera control for video diffusion model without training
Zhenghong Zhou, Jie An, and Jiebo Luo · 2024
Later among the works it cites.
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, et al · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al · 2025
Closest in time.
Wenyi Hong, Yean Cheng, Zhuoyi Yang, Weihan Wang, Lefan Wang, Xiaotao Gu, Shiyu Huang, Yuxiao Dong, and Jie Tang · 2025
Closest in time.
Realcam-i2v: Real-world image-to-video generation with interactive complex camera control
Teng Li, Guangcong Zheng, Rui Jiang, Tao Wu, Yehao Lu, Yining Lin, Xi Li, et al · 2025
Closest in time.
Chatcam: Empowering camera control through conversational ai
Xinhang Liu, Yu-Wing Tai, and Chi-Keung Tang · 2025
Closest in time.
Xincheng Shuai, Henghui Ding, Zhenyuan Qin, Hao Luo, Xingjun Ma, and Dacheng Tao · 2025
Closest in time.
Continuous 3d perception model with persistent state
Qianqian Wang, Yifei Zhang, Aleksander Holynski, Alexei A Efros, and Angjoo Kanazawa · 2025
Closest in time.
Motionbooth: Motion-aware customized text-to-video generation
Jianzong Wu, Xiangtai Li, Yanhong Zeng, Jiangning Zhang, Qianyu Zhou, Yining Li, Yunhai Tong, and Kai Chen · 2025
Closest in time.
Motioncanvas: Cinematic shot design with controllable image-to-video generation
Jinbo Xing, Long Mai, Cusuh Ham, Jiahui Huang, Aniruddha Mahapatra, Chi-Wing Fu, Tien-Tsin Wong, and Feng Liu · 2025
Closest in time.
Liping Yuan, Jiawei Wang, Haomiao Sun, Yuchen Zhang, and Yuan Lin · 2025
Closest in time.
Vidcraft3: Camera, object, and lighting control for image-to-video generation
Sixiao Zheng, Zimian Peng, Yanpeng Zhou, Yi Zhu, Hang Xu, Xiangru Huang, and Yanwei Fu · 2025
Closest in time.