Fetching the paper…
Reading the bibliography…
Transformer, an attention-based encoder-decoder model, has already revolutionized the field of natural language processing (NLP).
Learning multiple layers of features from tiny images
A. Krizhevsky et al · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng et al · 2009
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
V. Ordonez et al · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky et al · 2012
Earlier work this paper cites.
Sequence to sequence learning with neural networks
I. Sutskever et al · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin et al · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
P. Young et al · 2014
Earlier work this paper cites.
Why deep learning works: A manifold disentanglement perspective
P. P. Brahma et al · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
O. Ronneberger et al · 2015
Earlier work this paper cites.
Vqa: Visual question answering
S. Antol et al · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator
O. Vinyals et al · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He et al · 2016
Earlier work this paper cites.
Lstm: A search space odyssey
K. Greff et al · 2016
Earlier work this paper cites.
J. L. Ba et al · 2016
Earlier work this paper cites.
You only look once: Unified, real-time object detection
J. Redmon et al · 2016
Earlier work this paper cites.
Attention is all you need
A. Vaswani et al · 2017
Earlier work this paper cites.
Revisiting unreasonable effectiveness of data in deep learning era
C. Sun et al · 2017
Earlier work this paper cites.
Convolution in convolution for network in network
Y. Pang et al · 2017
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren et al · 2017
Earlier work this paper cites.
Focal loss for dense object detection
T.-Y. Lin et al · 2017
Earlier work this paper cites.
Deformable convolutional networks
J. Dai et al · 2017
Earlier work this paper cites.
Feature pyramid networks for object detection
T.-Y. Lin et al · 2017
Earlier work this paper cites.
Mask r-cnn
K. He et al · 2017
Earlier work this paper cites.
Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs
L.-C. Chen et al · 2017
Earlier work this paper cites.
Pointnet++: Deep hierarchical feature learning on point sets in a metric space
C. R. Qi et al · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna et al · 2017
Earlier work this paper cites.
Aggregated residual transformations for deep neural networks
S. Xie et al · 2017
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Y. Goyal et al · 2017
Earlier work this paper cites.
Improving language understanding by generative pre-training
A. Radford et al · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin et al · 2018
Earlier work this paper cites.
Non-local neural networks
X. Wang et al · 2018
Earlier work this paper cites.
Squeeze-and-excitation networks
J. Hu et al · 2018
Earlier work this paper cites.
Cbam: Convolutional block attention module
S. Woo et al · 2018
Earlier work this paper cites.
Image transformer
N. Parmar et al · 2018
Earlier work this paper cites.
Relation networks for object detection
H. Hu et al · 2018
Earlier work this paper cites.
a 2 a^{2} -nets: Double attention networks
Y. Chen et al · 2018
Earlier work this paper cites.
Self-attention with relative position representations
P. Shaw et al · 2018
Earlier work this paper cites.
Relational inductive biases, deep learning, and graph networks
P. W. Battaglia et al · 2018
Earlier work this paper cites.
Mobilenetv2: Inverted residuals and linear bottlenecks
M. Sandler et al · 2018
Earlier work this paper cites.
Cascade r-cnn: Delving into high quality object detection
Z. Cai and N. Vasconcelos · 2018
Earlier work this paper cites.
Unified perceptual parsing for scene understanding
T. Xiao et al · 2018
Earlier work this paper cites.
Bottom-up and top-down attention for image captioning and visual question answering
P. Anderson et al · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
P. Sharma et al · 2018
Earlier work this paper cites.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
S. Xie et al · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford et al · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Y. Liu et al · 2019
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Z. Yang et al · 2019
Earlier work this paper cites.
Efficientnet: Rethinking model scaling for convolutional neural networks
M. Tan and Q. Le · 2019
Earlier work this paper cites.
Ccnet: Criss-cross attention for semantic segmentation
Z. Huang et al · 2019
Earlier work this paper cites.
Gcnet: Non-local networks meet squeeze-excitation networks and beyond
Y. Cao et al · 2019
Earlier work this paper cites.
Local relation networks for image recognition
H. Hu et al · 2019
Earlier work this paper cites.
Attention augmented convolutional networks
I. Bello et al · 2019
Earlier work this paper cites.
Stand-alone self-attention in vision models
P. Ramachandran et al · 2019
Earlier work this paper cites.
Videobert: A joint model for video and language representation learning
C. Sun et al · 2019
Earlier work this paper cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
J. Lu et al · 2019
Earlier work this paper cites.
Lxmert: Learning cross-modality encoder representations from transformers
H. Tan and M. Bansal · 2019
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language
L. H. Li et al · 2019
Earlier work this paper cites.
Vl-bert: Pre-training of generic visual-linguistic representations
W. Su et al · 2019
Earlier work this paper cites.
Tsm: Temporal shift module for efficient video understanding
J. Lin et al · 2019
Earlier work this paper cites.
Representation degeneration problem in training natural language generation models
J. Gao et al · 2019
Earlier work this paper cites.
Cutmix: Regularization strategy to train strong classifiers with localizable features
S. Yun et al · 2019
Earlier work this paper cites.
Object detection with deep learning: A review
Z.-Q. Zhao et al · 2019
Earlier work this paper cites.
Fcos: Fully convolutional one-stage object detection
Z. Tian et al · 2019
Earlier work this paper cites.
Deep high-resolution representation learning for human pose estimation
K. Sun et al · 2019
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
D. A. Hudson and C. D. Manning · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. B. Brown et al · 2020
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Z. Lan et al · 2020
Earlier work this paper cites.
A survey of the usages of deep learning for natural language processing
D. W. Otter et al · 2020
Earlier work this paper cites.
Attention in natural language processing
A. Galassi et al · 2020
Earlier work this paper cites.
Eca-net: Efficient channel attention for deep convolutional neural networks
Q. Wang et al · 2020
Earlier work this paper cites.
Exploring self-attention for image recognition
H. Zhao et al · 2020
Earlier work this paper cites.
Global and local knowledge-aware attention network for action recognition
Z. Zheng et al · 2020
Earlier work this paper cites.
On the relationship between self-attention and convolutional layers
J.-B. Cordonnier et al · 2020
Earlier work this paper cites.
End-to-end object detection with transformers
N. Carion et al · 2020
Cited alongside, same era.
Efficient transformers: A survey
Y. Tay et al · 2020
Cited alongside, same era.
Visual transformers: Token-based image representation and processing for computer vision
B. Wu et al · 2020
Cited alongside, same era.
Generative pretraining from pixels
M. Chen et al · 2020
Cited alongside, same era.
End-to-end object detection with adaptive clustering transformer
M. Zheng et al · 2020
Cited alongside, same era.
Feature pyramid transformer
D. Zhang et al · 2020
Cited alongside, same era.
Segformer: Simple and efficient design for semantic segmentation with transformers
E. Xie et al · 2021
Closest in time.
End-to-end video instance segmentation with transformers
Y. Wang et al · 2021
Closest in time.
Instances as queries
Y. Fang et al · 2021
Closest in time.
Istr: End-to-end instance segmentation with transformers
J. Hu et al · 2021
Closest in time.
Solq: Segmenting objects by learning queries
B. Dong et al · 2021
Closest in time.
Segmenter: Transformer for semantic segmentation
R. Strudel et al · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Attention-based transformers for instance segmentation of cells in microstructures
T. Prangemeier et al · 2020
Cited alongside, same era.
Uniter: Universal image-text representation learning
Y.-C. Chen et al · 2020
Cited alongside, same era.
Oscar: Object-semantics aligned pre-training for vision-language tasks
X. Li et al · 2020
Cited alongside, same era.
Unified vision-language pre-training for image captioning and vqa
L. Zhou et al · 2020
Cited alongside, same era.
Transferring inductive biases through knowledge distillation
S. Abnar et al · 2020
Cited alongside, same era.
How much position information do convolutional neural networks encode
A. Islam et al · 2020
Cited alongside, same era.
H. Zhao et al · 2021
Closest in time.
Pct: Point cloud transformer
M.-H. Guo et al · 2021
Closest in time.
3d object detection with pointformer
X. Pan et al · 2021
Closest in time.
Voxel transformer for 3d object detection
J. Mao et al · 2021
Closest in time.
An end-to-end transformer model for 3d object detection
I. Misra et al · 2021
Closest in time.
Group-free 3d object detection via transformers
Z. Liu et al · 2021
Closest in time.
Improving 3d object detection with channel-wise transformer
H. Sheng et al · 2021
Closest in time.
Pointr: Diverse point cloud completion with geometry-aware transformers
X. Yu et al · 2021
Closest in time.
Snowflakenet: Point cloud completion by snowflake point deconvolution with skip-transformer
P. Xiang et al · 2021
Closest in time.
Deep point cloud reconstruction
J. H. Choe et al · 2021
Closest in time.
Mvt: Multi-view vision transformer for 3d object recognition
S. Chen et al · 2021
Closest in time.
Multiview detection with shadow transformer (and view-coherent data augmentation)
Y. Hou and L. Zheng · 2021
Closest in time.
Multi-modal fusion transformer for end-to-end autonomous driving
A. Prakash et al · 2021
Closest in time.
Cotr: Correspondence transformer for matching across images
W. Jiang et al · 2021
Closest in time.
Multi-view 3d reconstruction with transformers
D. Wang et al · 2021
Closest in time.
Transformerfusion: Monocular rgb scene reconstruction using transformers
A. Bozic et al · 2021
Closest in time.
Multi-view analysis of unregistered medical images using cross-view transformers
G. v. Tulder et al · 2021
Closest in time.
Multi-view depth estimation using epipolar spatio-temporal networks
X. Long et al · 2021
Closest in time.
Deep relation transformer for diagnosing glaucoma with optical coherence tomography and visual field function
D. Song et al · 2021
Closest in time.
Vilt: Vision-and-language transformer without convolution or region supervision
W. Kim et al · 2021
Closest in time.
Vinvl: Revisiting visual representations in vision-language models
P. Zhang et al · 2021
Closest in time.
Learning transferable visual models from natural language supervision
A. Radford et al · 2021
Closest in time.
Zero-shot text-to-image generation
A. Ramesh et al · 2021
Closest in time.
Scaling up visual and vision-language representation learning with noisy text supervision
C. Jia et al · 2021
Closest in time.
Unit: Multimodal multitask learning with a unified transformer
R. Hu and A. Singh · 2021
Closest in time.
Simvlm: Simple visual language model pretraining with weak supervision
Z. Wang et al · 2021
Closest in time.
Mdetr-modulated detection for end-to-end multi-modal understanding
A. Kamath et al · 2021
Closest in time.
Transvg: End-to-end visual grounding with transformers
J. Deng et al · 2021
Closest in time.
Referring transformer: A one-step approach to multi-task visual grounding
M. Li and L. Sigal · 2021
Closest in time.
Transrefer3d: Entity-and-relation aware transformer for fine-grained 3d visual grounding
D. He et al · 2021
Closest in time.
Attention is not all you need: pure attention loses rank doubly exponentially with depth
Y. Dong et al · 2021
Closest in time.
Crossvit: Cross-attention multi-scale vision transformer for image classification
C.-F. Chen et al · 2021
Closest in time.
What makes for end-to-end object detection?
P. Sun et al · 2021
Closest in time.
Sparse r-cnn: End-to-end object detection with learnable proposals
P. Sun et al · 2021
Closest in time.
Searching the search space of vision transformer
M. Chen et al · 2021
Closest in time.
Not all images are worth 16x16 words: Dynamic transformers for efficient image recognition
Y. Wang et al · 2021
Closest in time.
How do vision transformers work?
N. Park and S. Kim · 2021
Closest in time.
Rethinking and improving relative position encoding for vision transformer
K. Wu et al · 2021
Closest in time.
Position, padding and predictions: A deeper look at position information in cnns
M. A. Islam et al · 2021
Closest in time.
Scaling vision with sparse mixture of experts
C. Riquelme et al · 2021
Closest in time.
A survey on vision transformer
K. Han et al · 2022
Closest in time.
A survey of transformers
T. Lin et al · 2022
Closest in time.
Masked autoencoders are scalable vision learners
K. He et al · 2022
Closest in time.
Anchor detr: Query design for transformer-based detector
Y. Wang et al · 2022
Closest in time.
Y. Liu et al · 2022
Closest in time.
Dn-detr: Accelerate detr training by introducing query denoising
F. Li et al · 2022
Closest in time.
Dino: Detr with improved denoising anchor boxes for end-to-end object detection
H. Zhang et al · 2022
Closest in time.
3dctn: 3d convolution-transformer network for point cloud classification
D. Lu et al · 2022
Closest in time.
Fast point transformer
C. Park et al · 2022
Closest in time.
Embracing single stride 3d object detector with sparse transformer
L. Fan et al · 2022
Closest in time.
Voxel set transformer: A set-to-set approach to 3d object detection from point clouds
C. He et al · 2022
Closest in time.
Point-bert: Pre-training 3d point cloud transformers with masked point modeling
X. Yu et al · 2022
Closest in time.
Masked autoencoders for point cloud self-supervised learning
Y. Pang et al · 2022
Closest in time.
Masked discrimination for self-supervised learning on point clouds
H. Liu et al · 2022
Closest in time.
Monodtr: Monocular 3d object detection with depth-aware transformer
K.-C. Huang et al · 2022
Closest in time.
Monodetr: Depth-aware transformer for monocular 3d object detection
R. Zhang et al · 2022
Closest in time.
Detr3d: 3d object detection from multi-view images via 3d-to-2d queries
Y. Wang et al · 2022
Closest in time.
Transfusion: Robust lidar-camera fusion for 3d object detection with transformers
X. Bai et al · 2022
Closest in time.
Futr3d: A unified sensor fusion framework for 3d detection
X. Chen et al · 2022
Closest in time.
mmformer: Multimodal medical transformer for incomplete multimodal learning of brain tumor segmentation
Y. Zhang et al · 2022
Closest in time.
Data2vec: A general framework for self-supervised learning in speech, vision and language
A. Baevski et al · 2022
Closest in time.
Visual grounding with transformers
Y. Du et al · 2022
Closest in time.
Pseudo-q: Generating pseudo language queries for visual grounding
H. Jiang et al · 2022
Closest in time.
Languagerefer: Spatial-language model for 3d visual grounding
J. Roh et al · 2022
Closest in time.
Multi-view transformer for 3d visual grounding
S. Huang et al · 2022
Closest in time.
Tubedetr: Spatio-temporal video grounding with transformers
A. Yang et al · 2022
Closest in time.
Metaformer is actually what you need for vision
W. Yu et al · 2022
Closest in time.