Fetching the paper…
Reading the bibliography…
Recent advances in zero-shot and few-shot classification heavily rely on the success of pre-trained vision-language models (VLMs) such as CLIP.
Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories
Li Fei-Fei, Rob Fergus, and Pietro Perona · 2004
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
A qvga 143 db dynamic range frame-free pwm image sensor with lossless pixel-level video compression and time-domain cds
Christoph Posch, Daniel Matolin, and Rainer Wohlgenannt · 2010
Earlier work this paper cites.
Event-based visual flow
Ryad Benosman, Charles Clercq, Xavier Lagorce, Sio-Hoi Ieng, and Chiara Bartolozzi · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Converting static image datasets to spiking neuromorphic datasets using saccades
Garrick Orchard, Ajinkya Jayawant, Gregory K Cohen, and Nitish Thakor · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Hots: a hierarchy of event-based time-surfaces for pattern recognition
Xavier Lagorce, Garrick Orchard, Francesco Galluppi, Bertram E Shi, and Ryad B Benosman · 2016
Earlier work this paper cites.
Training deep spiking neural networks using backpropagation
Jun Haeng Lee, Tobi Delbruck, and Michael Pfeiffer · 2016
Earlier work this paper cites.
Phased lstm: Accelerating recurrent network training for long or event-based sequences
Daniel Neil, Michael Pfeiffer, and Shih-Chii Liu · 2016
Earlier work this paper cites.
A low power, fully event-based gesture recognition system
Arnon Amir, Brian Taba, David Berg, Timothy Melano, Jeffrey McKinstry, Carmelo Di Nolfo, Tapan Nayak, Alexander Andreopoulos, Guillaume Garreau, Marcela Mendoza, et al · 2017
Earlier work this paper cites.
Event-based, 6-dof camera tracking from photometric depth maps
Guillermo Gallego, Jon EA Lund, Elias Mueggler, Henri Rebecq, Tobi Delbruck, and Davide Scaramuzza · 2017
Earlier work this paper cites.
The event-camera dataset and simulator: Event-based data for pose estimation, visual odometry, and slam
Elias Mueggler, Henri Rebecq, Guillermo Gallego, Tobi Delbruck, and Davide Scaramuzza · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Deep sets
Manzil Zaheer, Satwik Kottur, Siamak Ravanbakhsh, Barnabas Poczos, Russ R Salakhutdinov, and Alexander J Smola · 2017
Earlier work this paper cites.
Ace: An efficient asynchronous corner tracker for event cameras
Ignacio Alzugaray and Margarita Chli · 2018
Earlier work this paper cites.
Spatial and temporal downsampling in event-based visual classification
Gregory Cohen, Saeed Afshar, Garrick Orchard, Jonathan Tapson, Ryad Benosman, and Andre van Schaik · 2018
Earlier work this paper cites.
Asynchronous, photometric feature tracking using events and frames
Daniel Gehrig, Henri Rebecq, Guillermo Gallego, and Davide Scaramuzza · 2018
Earlier work this paper cites.
Event-based vision meets deep learning on steering prediction for self-driving cars
Ana I Maqueda, Antonio Loquercio, Guillermo Gallego, Narciso García, and Davide Scaramuzza · 2018
Earlier work this paper cites.
Esim: an open event camera simulator
Henri Rebecq, Daniel Gehrig, and Davide Scaramuzza · 2018
Earlier work this paper cites.
HATS: Histograms of averaged time surfaces for robust event-based object classification
Amos Sironi, Manuele Brambilla, Nicolas Bourdis, Xavier Lagorce, and Ryad Benosman · 2018
Earlier work this paper cites.
Ev-flownet: Self-supervised optical flow estimation for event-based cameras
Alex Zihao Zhu and Liangzhe Yuan · 2018
Earlier work this paper cites.
End-to-end learning of representations for asynchronous event-based data
Daniel Gehrig, Antonio Loquercio, Konstantinos G Derpanis, and Davide Scaramuzza · 2019
Earlier work this paper cites.
Imagenet-trained cnns are biased towards texture; increasing shape bias improves accuracy and robustness
Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix A Wichmann, and Wieland Brendel · 2019
Earlier work this paper cites.
Space-time event clouds for gesture recognition: From rgb cameras to event cameras
Qinyi Wang, Yexin Zhang, Junsong Yuan, and Yilong Lu · 2019
Cited alongside, same era.
A differentiable recurrent surface for asynchronous event-based data
Marco Cannici, Marco Ciccone, Andrea Romanoni, and Matteo Matteucci · 2020
Cited alongside, same era.
Randaugment: Practical automated data augmentation with a reduced search space
Ekin D Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V Le · 2020
Cited alongside, same era.
A large scale event-based detection dataset for automotive
Pierre De Tournemire, Davide Nitti, Etienne Perot, Davide Migliore, and Amos Sironi · 2020
Cited alongside, same era.
Event-based vision: A survey
Guillermo Gallego, Tobi Delbrück, Garrick Orchard, Chiara Bartolozzi, Brian Taba, Andrea Censi, Stefan Leutenegger, Andrew J Davison, Jörg Conradt, Kostas Daniilidis, et al · 2020
Cited alongside, same era.
Event-based video reconstruction using transformer
Wenming Weng, Yueyi Zhang, and Zhiwei Xiong · 2021
Later among the works it cites.
Eventgan: Leveraging large scale image datasets for event cameras
Alex Zihao Zhu, Ziyun Wang, Kaung Khant, and Kostas Daniilidis · 2021
Later among the works it cites.
Open-vocabulary object detection via vision and language knowledge distillation
Xiuye Gu, Tsung-Yi Lin, Weicheng Kuo, and Yin Cui · 2022
Later among the works it cites.
Prompting visual-language models for efficient video understanding
Chen Ju, Tengda Han, Kunhao Zheng, Ya Zhang, and Weidi Xie · 2022
Later among the works it cites.
Ev-tta: Test-time adaptation for event-based object recognition
Junho Kim, Inwoo Hwang, and Young Min Kim · 2022
Later among the works it cites.
Masked event modeling: Self-supervised pretraining for event cameras
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning to exploit multiple vision modalities by using grafted networks
Yuhuang Hu, Tobi Delbruck, and Shih-Chii Liu · 2020
Cited alongside, same era.
Event-based asynchronous sparse convolutional networks
Nico Messikommer, Daniel Gehrig, Antonio Loquercio, and Davide Scaramuzza · 2020
Cited alongside, same era.
Learning to detect objects with a 1 megapixel event camera
Etienne Perot, Pierre De Tournemire, Davide Nitti, Jonathan Masci, and Amos Sironi · 2020
Cited alongside, same era.
Fast image reconstruction with an event camera
Cedric Scheerlinck, Henri Rebecq, Daniel Gehrig, Nick Barnes, Robert Mahony, and Davide Scaramuzza · 2020
Cited alongside, same era.
Reducing the sim-to-real gap for event cameras
Timo Stoffregen, Cedric Scheerlinck, Davide Scaramuzza, Tom Drummond, Nick Barnes, Lindsay Kleeman, and Robert Mahony · 2020
Cited alongside, same era.
On layer normalization in the transformer architecture
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tieyan Liu · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2021
Cited alongside, same era.
Simon Klenk, David Bonello, Lukas Koestler, and Daniel Cremers · 2022
Later among the works it cites.
Asynchronous spatio-temporal memory network for continuous event-based object detection
Jianing Li, Jia Li, Lin Zhu, Xijie Xiang, Tiejun Huang, and Yonghong Tian · 2022
Later among the works it cites.
Bridging the gap between events and frames through unsupervised domain adaptation
Nico Messikommer, Daniel Gehrig, Mathias Gehrig, and Davide Scaramuzza · 2022
Later among the works it cites.
Denseclip: Language-guided dense prediction with context-aware prompting
Yongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang, Zheng Zhu, Guan Huang, Jie Zhou, and Jiwen Lu · 2022
Later among the works it cites.
Aegnn: Asynchronous event-based graph neural networks
Simon Schaefer, Daniel Gehrig, and Davide Scaramuzza · 2022
Later among the works it cites.
Ess: Learning event-based semantic segmentation from still images
Zhaoning Sun, Nico Messikommer, Daniel Gehrig, and Davide Scaramuzza · 2022
Later among the works it cites.
Robust fine-tuning of zero-shot models
Mitchell Wortsman, Gabriel Ilharco, Jong Wook Kim, Mike Li, Simon Kornblith, Rebecca Roelofs, Raphael Gontijo Lopes, Hannaneh Hajishirzi, Ali Farhadi, Hongseok Namkoong, et al · 2022
Later among the works it cites.
Retriever: Learning content-style representation as a token-level bipartite graph
Dacheng Yin, Xuanchi Ren, Chong Luo, Yuwang Wang, Zhiwei Xiong, and Wenjun Zeng · 2022
Later among the works it cites.
Label-free event-based object recognition via joint learning with image reconstruction from events
Hoonhee Cho, Hyeonseong Kim, Yujeong Chae, and Kuk-Jin Yoon · 2023
Closest in time.
Recurrent vision transformers for object detection with event cameras
Mathias Gehrig and Davide Scaramuzza · 2023
Closest in time.
Imagebind: One embedding space to bind them all
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra · 2023
Closest in time.
Is synthetic data from generative models ready for image recognition?
Ruifei He, Shuyang Sun, Xin Yu, Chuhui Xue, Wenqing Zhang, Philip Torr, Song Bai, and Xiaojuan Qi · 2023
Closest in time.
Scaling language-image pre-training via masking
Yanghao Li, Haoqi Fan, Ronghang Hu, Christoph Feichtenhofer, and Kaiming He · 2023
Closest in time.
Slotformer: Unsupervised visual dynamics simulation with object-centric models
Ziyi Wu, Nikita Dvornik, Klaus Greff, Thomas Kipf, and Animesh Garg · 2023
Closest in time.
Ulip: Learning a unified representation of language, images, and point clouds for 3d understanding
Le Xue, Mingfei Gao, Chen Xing, Roberto Martín-Martín, Jiajun Wu, Caiming Xiong, Ran Xu, Juan Carlos Niebles, and Silvio Savarese · 2023
Closest in time.
Visual-language prompt tuning with knowledge-guided context optimization
Hantao Yao, Rui Zhang, and Changsheng Xu · 2023
Closest in time.
E-clip: Towards label-efficient event-based open-world understanding by clip
Jiazhou Zhou, Xu Zheng, Yuanhuiyi Lyu, and Lin Wang · 2023
Closest in time.
Pointclip v2: Adapting clip for powerful 3d open-world learning
Xiangyang Zhu, Renrui Zhang, Bowei He, Ziyao Zeng, Shanghang Zhang, and Peng Gao · 2023
Closest in time.