Fetching the paper…
Reading the bibliography…
Vision transformers have achieved remarkable progress in vision tasks such as image classification and detection.
Object retrieval with large vocabularies and fast spatial matching
James Philbin, Ondrej Chum, Michael Isard, Josef Sivic, and Andrew Zisserman · 2007
Earlier work this paper cites.
Lost in quantization:Improving particular object retrieval in large scale image databases
James Philbin, Ondrej Chum, Michael Isard, Josef Sivic, and Andrew Zisserman · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Diffusion processes for retrieval revisited
Michael Donoser and Horst Bischof · 2013
Earlier work this paper cites.
Neural codes for image retrieval
Artem Babenko, Anton Slesarev, Alexandr Chigorin, and Victor Lempitsky · 2014
Earlier work this paper cites.
Aggregating Local Deep Features for Image Retrieval
Artem Babenko and Victor Lempitsky · 2015
Earlier work this paper cites.
Hypercolumns for object segmentation and fine-grained localization
Bharath Hariharan, Pablo Arbelaez, Ross Girshick, and Jitendra Malik · 2015
Earlier work this paper cites.
Aggregating local deep features for image retrieval
Victor Lempitsky and Artem Babenko · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
Deep image retrieval: Learning global representations for image search
Albert Gordo, Jon Almazan, Jerome Revaud, and Diane Larlus · 2016
Earlier work this paper cites.
Cross-dimensional weighting for aggregated deep convolutional features
Yannis Kalantidis, Clayton Mellina, and Simon Osindero · 2016
Earlier work this paper cites.
SSD: Single shot multibox detector
Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, and Alexander C Berg · 2016
Earlier work this paper cites.
CNN image retrieval learns from BoW: Unsupervised fine-tuning with hard examples
Filip Radenović, Giorgos Tolias, and Ondřej Chum · 2016
Earlier work this paper cites.
Rethinking atrous convolution for semantic image segmentation
Liang-Chieh Chen, George Papandreou, Florian Schroff, and Hartwig Adam · 2017
Earlier work this paper cites.
What is the best practice for cnns applied to visual instance retrieval?
Jiedong Hao, Jing Dong, Wei Wang, and Tieniu Tan · 2017
Earlier work this paper cites.
Densely connected convolutional networks
Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger · 2017
Earlier work this paper cites.
Efficient diffusion on region manifolds: Recovering small objects with compact cnn representations
Ahmet Iscen, Giorgos Tolias, Yannis Avrithis, Teddy Furon, and Ondrej Chum · 2017
Earlier work this paper cites.
Feature pyramid networks for object detection
Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie · 2017
Earlier work this paper cites.
Large-scale image retrieval with attentive deep local features
Hyeonwoo Noh, Andre Araujo, Jack Sim, Tobias Weyand, and Bohyung Han · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Attention-aware generalized mean pooling for image retrieval
Yinzheng Gu, Chuanpeng Li, and Jinbin Xie · 2018
Earlier work this paper cites.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein · 2018
Earlier work this paper cites.
DetNet: A backbone network for object detection
Zeming Li, Chao Peng, Gang Yu, Xiangyu Zhang, Yangdong Deng, and Jian Sun · 2018
Earlier work this paper cites.
On the loss landscape of a class of deep neural networks with no bad local valleys
Quynh Nguyen, Mahesh Chandra Mukkamala, and Matthias Hein · 2018
Earlier work this paper cites.
Revisiting Oxford and Paris: Large-Scale Image Retrieval Benchmarking
Filip Radenović, Ahmet Iscen, Giorgos Tolias, Yannis Avrithis, and Ondřej Chum · 2018
Cited alongside, same era.
Mobilenetv2: Inverted residuals and linear bottlenecks
Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen · 2018
Cited alongside, same era.
Sparsely connected convolutional networks
Ligeng Zhu, Ruizhi Deng, Zhiwei Deng, Greg Mori, and Ping Tan · 2018
Cited alongside, same era.
Fine-tuning cnn image retrieval with no human annotation
Filip Radenović, Giorgos Tolias, and Ondřej Chum · 2019
Cited alongside, same era.
Local Features and Visual Words Emerge in Activations
Oriane Simeoni, Yannis Avrithis, and Ondréj Chum · 2019
Cited alongside, same era.
Efficient image retrieval via decoupling diffusion into online and offline processing
Fan Yang, Ryota Hinami, Yusuke Matsui, Steven Ly, and Shin’ichi Satoh · 2019
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2021
Later among the works it cites.
Rethinking spatial dimensions of vision transformers
Byeongho Heo, Sangdoo Yun, Dongyoon Han, Sanghyuk Chun, Junsuk Choe, and Seong Joon Oh · 2021
Later among the works it cites.
Token labeling: Training a 85.5% top-1 accuracy vision transformer with 56m parameters on imagenet
Zihang Jiang, Qibin Hou, Li Yuan, Daquan Zhou, Xiaojie Jin, Anran Wang, and Jiashi Feng · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexander Kolesnikov, Alexey Dosovitskiy, Dirk Weissenborn, Georg Heigold, Jakob Uszkoreit, Lucas Beyer, Matthias Minderer, Mostafa Dehghani, Neil Houlsby, Sylvain Gelly, Thomas Unterthiner, and Xiaohua Zhai · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Unifying deep local and global features for image search
Bingyi Cao, André Araujo, and Jack Sim · 2020
Cited alongside, same era.
Densely connected search space for more flexible neural architecture search
Jiemin Fang, Yuzhu Sun, Qian Zhang, Yuan Li, Wenyu Liu, and Xinggang Wang · 2020
Cited alongside, same era.
Training data-efficient image transformers & distillation through attention
Matthieu Cord Hugo Touvron, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2020
Cited alongside, same era.
SOLAR: Second-Order Loss and Attention for Image Retrieval
Tony Ng, Vassileios Balntas, Yurun Tian, and Krystian Mikolajczyk · 2020
Cited alongside, same era.
Efficientdet: Scalable and efficient object detection
Mingxing Tan, Ruoming Pang, and Quoc V. Le · 2020
Cited alongside, same era.
Learning and aggregating deep local descriptors for instance-level recognition
Giorgos Tolias, Tomas Jenicek, , and Ondréj Chum · 2020
Cited alongside, same era.
Yawei Li, Kai Zhang, Jiezhang Cao, Radu Timofte, and Luc Van Gool · 2021
Later among the works it cites.
MST: Masked self-supervised transformer for visual representation
Zhaowen Li, Zhiyang Chen, Fan Yang, Wei Li, Yousong Zhu, Chaoyang Zhao, Rui Deng, Liwei Wu, Rui Zhao, Ming Tang, et al · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Later among the works it cites.
Mobilevit: Light-weight, general-purpose, and mobile-friendly vision transformer
Sachin Mehta and Mohammad Rastegari · 2021
Later among the works it cites.
Do wide and deep networks learn the same things? uncovering how neural network representations vary with width and depth
Thao Nguyen, Maithra Raghu, and Simon Kornblith · 2021
Later among the works it cites.
Conformer: Local features coupling global representations for visual recognition
Zhiliang Peng, Wei Huang, Shanzhi Gu, Lingxi Xie, Yaowei Wang, Jianbin Jiao, and Qixiang Ye · 2021
Later among the works it cites.
Do vision transformers see like convolutional neural networks?
Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang, and Alexey Dosovitskiy · 2021
Later among the works it cites.
Instance-level image retrieval using reranking transformers
Fuwen Tan, Jiangbo Yuan, and Vicente Ordonez · 2021
Later among the works it cites.
Learning deep local features with multiple dynamic attentions for large-scale image retrieval
Hui Wu, Min Wang, Wengang Zhou, and Houqiang Li · 2021
Later among the works it cites.
Cvt: Introducing convolutions to vision transformers
Haiping Wu, Bin Xiao, Noel Codella, Mengchen Liu, Xiyang Dai, Lu Yuan, and Lei Zhang · 2021
Later among the works it cites.
Dolg: Single-stage image retrieval with deep orthogonal fusion of local and global features
Min Yang, Dongliang He, Miao Fan, Baorong Shi, Xuetong Xue, Fu Li, Errui Ding, and Jizhou Huang · 2021
Later among the works it cites.
Incorporating convolution designs into visual transformers
Kun Yuan, Shaopeng Guo, Ziwei Liu, Aojun Zhou, Fengwei Yu, and Wei Wu · 2021
Later among the works it cites.
Tokens-to-token vit: Training vision transformers from scratch on imagenet
Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Zihang Jiang, Francis EH Tay, Jiashi Feng, and Shuicheng Yan · 2021
Later among the works it cites.
Multi-scale vision longformer: A new vision transformer for high-resolution image encoding
Pengchuan Zhang, Xiyang Dai, Jianwei Yang, Bin Xiao, Lu Yuan, Lei Zhang, and Jianfeng Gao · 2021
Later among the works it cites.
Deepvit: Towards deeper vision transformer
Daquan Zhou, Bingyi Kang, Xiaojie Jin, Linjie Yang, Xiaochen Lian, Zihang Jiang, Qibin Hou, and Jiashi Feng · 2021
Later among the works it cites.
Elsa: Enhanced local self-attention for vision transformer
Jingkai Zhou, Pichao Wang, Fan Wang, Qiong Liu, Hao Li, and Rong Jin · 2021
Later among the works it cites.
All the attention you need: Global-local, spatial-channel attention for image retrieval
Chull Hwan Song, Hye Joo Han, and Yannis Avrithis · 2022
Closest in time.
Attentive waveblock: Complementarity-enhanced mutual networks for unsupervised domain adaptation in person re-identification and beyond
Wenhao Wang, Fang Zhao, Shengcai Liao, and Ling Shao · 2022
Closest in time.
Learning Super-Features for Image Retrieval
Weinzaepfel, Philippe and Lucas, Thomas and Larlus, Diane and Kalantidis, Yannis · 2022
Closest in time.