Fetching the paper…
Reading the bibliography…
We propose a novel multimodal video benchmark - the Perception Test - to evaluate the perception and reasoning skills of pre-trained multimodal models (e.g.
The Frostig program for the development of visual perception: Teacher’s guide
Marianne Frostig and David Horne · 1965
Earlier work this paper cites.
Physical reasoning in young infants: Seeking explanations for impossible events
Renée Baillargeon · 1994
Earlier work this paper cites.
Child psychology: A contemporary viewpoint, 5th ed
Eileen Mavis Hetherington, Ross D. Parke, and Virginia Otis Locke · 1999
Earlier work this paper cites.
Developments in young infants’ reasoning about occluded objects
A. Aguiar and R. Baillargeon · 2002
Earlier work this paper cites.
The reliability of the occupational therapy adult perceptual screening test (ot-apst)
Deirdre M Cooke, Kryss McKenna, Jennifer Fleming, and Ross Darnell · 2005
Earlier work this paper cites.
Test of visual perceptual skills
Nancy A Martin and Morrison F Gardner · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Towards a visual turing challenge
Mateusz Malinowski and Mario Fritz · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles · 2015
Earlier work this paper cites.
Visual Turing test for computer vision systems
Donald Geman, Stuart Geman, Neil Hallonquist, and Laurent Younes · 2015
Earlier work this paper cites.
Fully-convolutional siamese networks for object tracking
Luca Bertinetto, Jack Valmadre, João F Henriques, Andrea Vedaldi, and Philip HS Torr · 2016
Earlier work this paper cites.
Visual object tracking performance measures revisited
Luka Čehovin, Aleš Leonardis, and Matej Kristan · 2016
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
Gunnar A. Sigurdsson, Gül Varol, X. Wang, Ali Farhadi, Ivan Laptev, and Abhinav Kumar Gupta · 2016
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Earlier work this paper cites.
The kinetics human action video dataset, 2017
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, Mustafa Suleyman, and Andrew Zisserman · 2017
Earlier work this paper cites.
Tracking by natural language specification
Zhenyang Li, Ran Tao, Efstratios Gavves, Cees G. M. Snoek, and Arnold W. M. Smeulders · 2017
Earlier work this paper cites.
Scaling egocentric vision: The epic-kitchens dataset
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray · 2018
Earlier work this paper cites.
Evaluating visual "common sense" using fine-grained classification and captioning tasks, 2018
Raghav Goyal, Farzaneh Mahdisoltani, Guillaume Berger, Waseem Gharbieh, Ingo Bax, and Roland Memisevic · 2018
Cited alongside, same era.
Got-10k: A large high-diversity benchmark for generic object tracking in the wild
Lianghua Huang, Xin Zhao, and Kaiqi Huang · 2019
Cited alongside, same era.
Video Question Answering with Spatio-Temporal Reasoning
Yunseok Jang, Yale Song, Chris Dongjoo Kim, Youngjae Yu, Youngjin Kim, and Gunhee Kim · 2019
Cited alongside, same era.
Invariant Action Recognition Dataset, 2019
Andrea Tacchetti, Leyla Isik, and Tomaso Poggio · 2019
Cited alongside, same era.
Activitynet-qa: A dataset for understanding complex web videos via question answering
Zhou Yu, Dejing Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, and Dacheng Tao · 2019
Cited alongside, same era.
Self-supervised multimodal versatile networks
Mdetr–modulated detection for end-to-end multi-modal understanding
Aishwarya Kamath, Mannat Singh, Yann LeCun, Ishan Misra, Gabriel Synnaeve, and Nicolas Carion · 2021
Later among the works it cites.
Roses are red, violets are blue… but should vqa expect them to?
Corentin Kervadec, Grigory Antipov, Moez Baccouche, and Christian Wolf · 2021
Later among the works it cites.
Multibench: Multiscale benchmarks for multimodal representation learning
Paul Pu Liang, Yiwei Lyu, Xiang Fan, Zetian Wu, Yun Cheng, Jason Wu, Leslie Yufan Chen, Peter Wu, Michelle A Lee, Yuke Zhu, et al · 2021
Later among the works it cites.
Intphys: A framework and benchmark for visual intuitive physics reasoning
Ronan Riochet, Mario Castro, Mathieu Bernard, Adam Lerer, Rob Fergus, Véronique Izard, and Emmanuel Dupoux · 2021
Later among the works it cites.
Do different tracking tasks require different appearance models?
Zhongdao Wang, Hengshuang Zhao, Ya-Li Li, Shengjin Wang, Philip Torr, and Luca Bertinetto · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jean-Baptiste Alayrac, Adria Recasens, Rosalia Schneider, Relja Arandjelović, Jason Ramapuram, Jeffrey De Fauw, Lucas Smaira, Sander Dieleman, and Andrew Zisserman · 2020
Cited alongside, same era.
Vggsound: A large-scale audio-visual dataset
Honglie Chen, Weidi Xie, Andrea Vedaldi, and Andrew Zisserman · 2020
Cited alongside, same era.
Tao: A large-scale benchmark for tracking any object
Achal Dave, Tarasha Khurana, Pavel Tokmakov, Cordelia Schmid, and Deva Ramanan · 2020
Cited alongside, same era.
Large scale holistic video understanding
Ali Diba, Mohsen Fayyaz, Vivek Sharma, Manohar Paluri, Jürgen Gall, Rainer Stiefelhagen, and Luc Van Gool · 2020
Cited alongside, same era.
CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning
Rohit Girdhar and Deva Ramanan · 2020
Cited alongside, same era.
Hota: A higher order metric for evaluating multi-object tracking
Jonathon Luiten, Aljosa Osep, Patrick Dendorfer, Philip Torr, Andreas Geiger, Laura Leal-Taixé, and Bastian Leibe · 2020
Cited alongside, same era.
Transtrack: Multiple-object tracking with transformer
Peize Sun, Jinkun Cao, Yi Jiang, Rufeng Zhang, Enze Xie, Zehuan Yuan, Changhu Wang, and Ping Luo · 2020
Cited alongside, same era.
Next-qa: Next phase of question-answering to explaining temporal actions
Junbin Xiao, Xindi Shang, Angela Yao, and Tat-Seng Chua · 2021
Later among the works it cites.
Learning spatio-temporal transformer for visual tracking
Bin Yan, Houwen Peng, Jianlong Fu, Dong Wang, and Huchuan Lu · 2021
Later among the works it cites.
Rescaling egocentric vision: Collection, pipeline and challenges for EPIC-KITCHENS-100
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Jian Ma, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray · 2022
Later among the works it cites.
TAP-vid: A benchmark for tracking any point in a video
Carl Doersch, Ankush Gupta, Larisa Markeeva, Adria Recasens Continente, Lucas Smaira, Yusuf Aytar, Joao Carreira, Andrew Zisserman, and Yi Yang · 2022
Later among the works it cites.
Kubric: A scalable dataset generator
Klaus Greff, Francois Belletti, Lucas Beyer, Carl Doersch, Yilun Du, Daniel Duckworth, David J Fleet, Dan Gnanapragasam, Florian Golemo, Charles Herrmann, et al · 2022
Later among the works it cites.
Robustness analysis of video-language models against visual and language perturbations
Madeline Chantry Schiappa, Shruti Vyas, Hamid Palangi, Yogesh S Rawat, and Vibhav Vineet · 2022
Later among the works it cites.
Image as a foreign language: Beit pretraining for all vision and vision-language tasks, 2022
Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhiliang Peng, Qiang Liu, Kriti Aggarwal, Owais Khan Mohammed, Saksham Singhal, Subhojit Som, and Furu Wei · 2022
Later among the works it cites.
Multiview Transformers for Video Recognition
Shen Yan, Xuehan Xiong, Anurag Arnab, Zhichao Lu, Mi Zhang, Chen Sun, and Cordelia Schmid · 2022
Later among the works it cites.
Actionformer: Localizing moments of actions with transformers
Chenlin Zhang, Jianxin Wu, and Yin Li · 2022
Later among the works it cites.
Multi-modal classifiers for open-vocabulary object detection
Prannay Kaul, Weidi Xie, and Andrew Zisserman · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Self-chained image-language model for video localization and question answering
Shoubin Yu, Jaemin Cho, Prateek Yadav, and Mohit Bansal · 2023
Closest in time.