Fetching the paper…
Reading the bibliography…
The aim of this paper is to study the influence of locality mechanisms in vision transformers.
M. Saleh, Y. Wang, N. Navab, B. Busam, and F. Tombari, “Cloudattention: Efficient multi-scale attention scheme for 3d point cloud learning,” in Proc. IROS . IEEE, 2022, pp. 1986–1992
1992
Earlier work this paper cites.
Y. LeCun, Y. Bengio et al. , “Convolutional networks for images, speech, and time series,” The handbook of brain theory and neural networks , vol. 3361, no. 10, p. 1995, 1995
1995
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A large-scale hierarchical image database,” in Proc. CVPR , 2009, pp. 248–255
2009
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Proc. NeurIPS , 2012, pp. 1097–1105
2012
Earlier work this paper cites.
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proc. CVPR , 2016, pp. 770–778
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
G. Huang, Z. Liu, L. van der Maaten, and K. Q. Weinberger, “Densely connected convolutional networks,” in Proc. CVPR , 2017
2017
Earlier work this paper cites.
M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “MobileNetV2: Inverted residuals and linear bottlenecks,” in Proc. CVPR , 2018, pp. 4510–4520
2018
Earlier work this paper cites.
X. Wang, R. Girshick, A. Gupta, and K. He, “Non-local neural networks,” in Proc. CVPR , 2018, pp. 7794–7803
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in Proc. CVPR , 2018, pp. 7132–7141
2018
Earlier work this paper cites.
A. Howard, M. Sandler, G. Chu, L.-C. Chen, B. Chen, M. Tan, W. Wang, Y. Zhu, R. Pang, V. Vasudevan et al. , “Searching for mobilenetv3,” in Proc. ICCV , 2019, pp. 1314–1324
2019
Earlier work this paper cites.
M. Tan, B. Chen, R. Pang, V. Vasudevan, M. Sandler, A. Howard, and Q. V. Le, “Mnasnet: Platform-aware neural architecture search for mobile,” in Proc. CVPR , 2019, pp. 2820–2828
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in Proc. ECCV . Springer, 2020, pp. 213–229
2020
Cited alongside, same era.
Q. Wang, B. Wu, P. Zhu, P. Li, W. Zuo, and Q. Hu, “Eca-net: Efficient channel attention for deep convolutional neural networks,” in Proc. CVPR , 2020
2020
Cited alongside, same era.
I. Radosavovic, R. P. Kosaraju, R. Girshick, K. He, and P. Dollár, “Designing network design spaces,” in Proc. CVPR , 2020, pp. 10 428–10 436
2020
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Closest in time.
2021
Closest in time.
W. Wang, L. Yao, L. Chen, D. Cai, X. He, and W. Liu, “Crossformer: A versatile vision transformer based on cross-scale attention,” arXiv e-prints , pp. arXiv–2108, 2021
2021
Closest in time.
X. Chu, Z. Tian, Y. Wang, B. Zhang, H. Ren, X. Wei, H. Xia, and C. Shen, “Twins: Revisiting the design of spatial attention in vision transformers,” Advances in Neural Information Processing Systems , vol. 34, 2021
2021
Closest in time.
P. Zhang, X. Dai, J. Yang, B. Xiao, L. Yuan, L. Zhang, and J. Gao, “Multi-scale vision longformer: A new vision transformer for high-resolution image encoding,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 2998–3008
2021
Closest in time.
2022
Closest in time.
C. Shi, Y. Zheng, and A. M. Fey, “Recognition and prediction of surgical gestures and trajectories using transformer models in robot-assisted surgery,” in Proc. IROS . IEEE, 2022, pp. 8017–8024
2022
Closest in time.
S. Lee, E. Yi, J. Lee, J. Yoo, H. Lee, and S. H. Kim, “Fully convolutional transformer with local-global attention,” in Proc. IROS . IEEE, 2022, pp. 552–559
2022
Closest in time.
E. V. Mascaro, S. Ma, H. Ahn, and D. Lee, “Robust human motion forecasting using transformer-based model,” in Proc. IROS . IEEE, 2022, pp. 10 674–10 680
2022
Closest in time.
A. Bucker, L. Figueredo, S. Haddadinl, A. Kapoor, S. Ma, and R. Bonatti, “Reshaping robot trajectories using natural language commands: A study of multi-modal data alignment using transformers,” in Proc. IROS . IEEE, 2022, pp. 978–984
2022
Closest in time.
Y. Li, Y. Fan, X. Xiang, D. Demandolx, R. Ranjan, R. Timofte, and L. Van Gool, “Efficient and explicit modelling of image hierarchies for image restoration,” in Proc. CVPR , 2023, pp. 18 278–18 289
2023
Closest in time.