Fetching the paper…
Reading the bibliography…
In the post-deep learning era, the Transformer architecture has demonstrated its powerful performance across pre-trained big models and various downstream tasks.
R. E. Kalman, “A new approach to linear filtering and prediction problems,”
1960
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,”
1997
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in
2009
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in
2012
Earlier work this paper cites.
M. Lin, Q. Chen, and S. Yan, “Network in network,”
2013
Earlier work this paper cites.
K. Cho, B. van Merriënboer, Ç. Gu̇lçehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase representations using rnn encoder–decoder for statistical machine translation,” in
2014
Earlier work this paper cites.
Y. Deng, P. Luo, C. C. Loy, and X. Tang, “Pedestrian attribute recognition at far distance,” in
2014
Earlier work this paper cites.
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in
2015
Earlier work this paper cites.
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in
2015
Earlier work this paper cites.
L. Zheng, L. Shen, L. Tian, S. Wang, J. Wang, and Q. Tian, “Scalable person re-identification: A benchmark,” in
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in
2016
Earlier work this paper cites.
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot
2016
Earlier work this paper cites.
D. Demner-Fushman, M. D. Kohli, M. B. Rosenman, S. E. Shooshan, L. Rodriguez, S. Antani, G. R. Thoma, and C. J. McDonald, “Preparing a collection of radiology examinations for distribution and retrieval,”
2016
Earlier work this paper cites.
X. Liu, W. Liu, H. Ma, and H. Fu, “Large-scale vehicle re-identification in urban surveillance videos,” in
2016
Earlier work this paper cites.
H. Liu, Y. Tian, Y. Yang, L. Pang, and T. Huang, “Deep relative distance learning: Tell the difference between similar vehicles,” in
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in
2017
Earlier work this paper cites.
X. Liu, H. Zhao, M. Tian, L. Sheng, J. Shao, S. Yi, J. Yan, and X. Wang, “Hydraplus-net: Attentive deep features for pedestrian analysis,” in
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Y. Li, X. Liang, Z. Hu, and E. P. Xing, “Hybrid retrieval-generation reinforced agent for medical image report generation,” in
2018
Earlier work this paper cites.
G. Wang, Y. Yuan, X. Chen, J. Li, and X. Zhou, “Learning discriminative features with multiple granularities for person re-identification,” in
2018
Earlier work this paper cites.
L. Wei, S. Zhang, W. Gao, and Q. Tian, “Person transfer gan to bridge domain gap for person re-identification,” in
2018
Earlier work this paper cites.
J. D. M.-W. C. Kenton and L. K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in
2019
Earlier work this paper cites.
G. Bhat, M. Danelljan, L. V. Gool, and R. Timofte, “Learning discriminative model prediction for tracking,” in
2019
Earlier work this paper cites.
M. Danelljan, L. V. Gool, and R. Timofte, “Probabilistic regression for visual tracking,” in
2019
Earlier work this paper cites.
M. Danelljan, G. Bhat, F. Shahbaz Khan, and M. Felsberg, “Atom: Accurate tracking by overlap maximization,” in
2019
Earlier work this paper cites.
C. Y. Li, X. Liang, Z. Hu, and E. P. Xing, “Knowledge-driven encode, retrieve, paraphrase for medical image report generation,” in
2019
Earlier work this paper cites.
B. He, J. Li, Y. Zhao, and Y. Tian, “Part-regularized near-duplicate vehicle re-identification,” in
2019
Earlier work this paper cites.
K. Zhou, Y. Yang, A. Cavallaro, and T. Xiang, “Omni-scale feature learning for person re-identification,” in
2019
Earlier work this paper cites.
R. Chu, Y. Sun, Y. Li, Z. Liu, C. Zhang, and Y. Wei, “Vehicle re-identification with viewpoint-aware metric learning,” in
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
T. Chen, S. Ding, J. Xie, Y. Yuan, W. Chen, Y. Yang, Z. Ren, and Z. Wang, “Abd-net: Attentive but diverse person re-identification,” in
2019
Earlier work this paper cites.
J. Miao, Y. Wu, P. Liu, Y. Ding, and Y. Yang, “Pose-guided feature alignment for occluded person re-identification,” in
2019
Earlier work this paper cites.
J. Miao, Y. Wu, P. Liu, Y. Ding, and Y. Yang, “Pose-guided feature alignment for occluded person re-identification,” in
2019
Earlier work this paper cites.
H. Luo, Y. Gu, X. Liao, S. Lai, and W. Jiang, “Bag of tricks and a strong baseline for deep person re-identification,” in
2019
Earlier work this paper cites.
Z. Wu, S. Pan, F. Chen, G. Long, C. Zhang, and S. Y. Philip, “A comprehensive survey on graph neural networks,”
2020
Earlier work this paper cites.
M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer, “Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension,” in
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret, “Transformers are rnns: Fast autoregressive transformers with linear attention,” in
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
C. Subakan, M. Ravanelli, S. Cornell, M. Bronzi, and J. Zhong, “Attention is all you need in speech separation,” in
2020
Earlier work this paper cites.
Z. Tan, Y. Yang, J. Wan, G. Guo, and S. Z. Li, “Relation-aware pedestrian attribute recognition with graph convolutional networks,”
2020
Earlier work this paper cites.
J. Wu, H. Liu, J. Jiang, M. Qi, B. Ren, X. Li, and Y. Wang, “Person attribute recognition by sequence contextual relation learning,”
2020
Earlier work this paper cites.
H. Fan, H.-M. Hu, S. Liu, W. Lu, and S. Pu, “Correlation graph convolutional network for pedestrian attribute recognition,”
2020
Earlier work this paper cites.
G. Bhat, M. Danelljan, L. Van Gool, and R. Timofte, “Know your surroundings: Exploiting scene information for object tracking,” in
2020
Earlier work this paper cites.
Z. Chen, Y. Song, T.-H. Chang, and X. Wan, “Generating radiology reports via memory-driven transformer,” in
2020
Earlier work this paper cites.
Y. Zhang, X. Wang, Z. Xu, Q. Yu, A. Yuille, and D. Xu, “When radiology report generation meets knowledge graph,” in
2020
Earlier work this paper cites.
Z. Zhuang, L. Wei, L. Xie, T. Zhang, H. Zhang, H. Wu, H. Ai, and Q. Tian, “Rethinking the distribution gap of person re-identification with camera-based batch normalization,” in
2020
Earlier work this paper cites.
J. Qian, W. Jiang, H. Luo, and H. Yu, “Stripe-based and attribute-aware network: A two-branch deep model for vehicle re-identification,”
2020
Earlier work this paper cites.
X. Jin, C. Lan, W. Zeng, and Z. Chen, “Uncertainty-aware multi-shot knowledge distillation for image-based object re-identification,” in
2020
Earlier work this paper cites.
Z. Zhang, C. Lan, W. Zeng, X. Jin, and Z. Chen, “Relation-aware global attention for person re-identification,” in
2020
Earlier work this paper cites.
X. Jin, C. Lan, W. Zeng, G. Wei, and Z. Chen, “Semantics-aligned representation learning for person re-identification,” in
2020
Earlier work this paper cites.
T.-S. Chen, C.-T. Liu, C.-W. Wu, and S.-Y. Chien, “Orientation-aware vehicle re-identification with semantics-guided part attention network,” in
2020
Earlier work this paper cites.
X. Chen, C. Fu, Y. Zhao, F. Zheng, J. Song, R. Ji, and Y. Yang, “Salience-guided cascaded suppression network for person re-identification,” in
2020
Earlier work this paper cites.
D. Meng, L. Li, X. Liu, Y. Li, S. Yang, Z.-J. Zha, X. Gao, S. Wang, and Q. Huang, “Parsing-based view-aware embedding network for vehicle re-identification,” in
2020
Earlier work this paper cites.
P. Khorramshahi, N. Peri, J.-c. Chen, and R. Chellappa, “The devil is in the details: Self-supervised attention for vehicle re-identification,” in
2020
Earlier work this paper cites.
G. Wang, S. Yang, H. Liu, Z. Wang, Y. Yang, S. Wang, G. Yu, E. Zhou, and J. Sun, “High-order information matters: Learning relation and topology for occluded person re-identification,” in
2020
Earlier work this paper cites.
Z. Sun, X. Nie, X. Xi, and Y. Yin, “Cfvmnet: A multi-branch network for vehicle re-identification based on common field of view,” in
2020
Earlier work this paper cites.
K. Zhu, H. Guo, Z. Liu, M. Tang, and J. Wang, “Identity-guided human semantic parsing for person re-identification,” in
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark
2021
Earlier work this paper cites.
H. Ren, H. Dai, Z. Dai, M. Yang, J. Leskovec, D. Schuurmans, and B. Dai, “Combiner: Full attention transformer with sparse computation cost,” in
2021
Earlier work this paper cites.
A. Gu, K. Goel, and C. R’e, “Efficiently modeling long sequences with structured state spaces,”
2021
Earlier work this paper cites.
A. Gu, I. Johnson, K. Goel, K. Saab, T. Dao, A. Rudra, and C. Ré, “Combining recurrent, convolutional, and continuous-time models with linear state space layers,” in
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
L. Chen, K. Lu, A. Rajeswaran, K. Lee, A. Grover, M. Laskin, P. Abbeel, A. Srinivas, and I. Mordatch, “Decision transformer: Reinforcement learning via sequence modeling,” in
2021
Earlier work this paper cites.
J. Jia, X. Chen, and K. Huang, “Spatial and semantic consistency regularizations for pedestrian attribute recognition,”
2021
Earlier work this paper cites.
Y. Yang, Z. Tan, P. Tiwari, H. M. Pandey, J. Wan, Z. Lei, G. Guo, and S. Z. Li, “Cascaded split-and-aggregate learning with feature recombination for pedestrian attribute recognition,”
2021
Earlier work this paper cites.
N. Wang, W. Zhou, J. Wang, and H. Li, “Transformer meets tracker: Exploiting temporal context for robust visual tracking,” in
2021
Earlier work this paper cites.
X. Chen, J. Yan, Bin Zhu, D. Wang, X. Yang, and H. Lu, “Transformer tracking,” in
2021
Earlier work this paper cites.
B. Chen, P. Li, L. Bai, L. Qiao, Q. Shen, B. Li, W. Gan, W. Wu, and W. Ouyang, “Backbone is all your need: A simplified architecture for visual object tracking,” in
2021
Earlier work this paper cites.
F. Liu, X. Wu, S. Ge, W. Fan, and Y. Zou, “Exploring and distilling posterior and prior knowledge for radiology report generation,” in
2021
Earlier work this paper cites.
F. Liu, C. Yin, X. Wu, S. Ge, P. Zhang, and X. Sun, “Contrastive attention for automatic chest x-ray report generation,” in
2021
Earlier work this paper cites.
F. Liu, S. Ge, and X. Wu, “Competence-based multimodal curriculum learning for medical report generation,” in
2021
Earlier work this paper cites.
S. He, H. Luo, P. Wang, F. Wang, H. Li, and W. Jiang, “Transreid: Transformer-based object re-identification,” in
2021
Earlier work this paper cites.
X. Shu, X. Wang, X. Zang, S. Zhang, Y. Chen, G. Li, and Q. Tian, “Large-scale spatio-temporal person re-identification: Algorithms and benchmark,”
2021
Earlier work this paper cites.
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” in
2021
Earlier work this paper cites.
X. Wang, S. Zheng, R. Yang, A. Zheng, Z. Chen, J. Tang, and B. Luo, “Pedestrian attribute recognition: A survey,”
2022
Earlier work this paper cites.
E. Nguyen, K. Goel, A. Gu, G. Downs, P. Shah, T. Dao, S. Baccus, and C. Ré, “S4nd: Modeling images and videos as multidimensional signals with state spaces,” in
2022
Earlier work this paper cites.
A. Gu, K. Goel, A. Gupta, and C. Ré, “On the parameterization and initialization of diagonal state space models,” in
2022
Earlier work this paper cites.
A. Gupta, A. Gu, and J. Berant, “Diagonal state spaces are as effective as structured state spaces,”
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
M. M. Islam and G. Bertasius, “Long movie clip classification with state-space video models,” in
2022
Earlier work this paper cites.
J. T. Smith, A. Warrington, and S. Linderman, “Simplified state space layers for sequence modeling,” in
2022
Cited alongside, same era.
2022
Cited alongside, same era.
X. Ma, C. Zhou, X. Kong, J. He, L. Gui, G. Neubig, J. May, and L. Zettlemoyer, “Mega: Moving average equipped gated attention,” in
2022
Cited alongside, same era.
D. Y. Fu, T. Dao, K. K. Saab, A. W. Thomas, A. Rudra, and C. Re, “Hungry hungry hippos: Towards language modeling with state space models,” in
2022
Cited alongside, same era.
Z.-Q. Wang, S. Cornell, S. Choi, Y. Lee, B. Kim, and S. Watanabe, “Tf-gridnet: Making time-frequency domain models great again for monaural speaker separation,” in
2022
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
K. Goel, A. Gu, C. Donahue, and C. Ré, “It’s raw! audio generation with state-space models,” in
2022
Cited alongside, same era.
J. Wang, J. N. Yan, A. Gu, and A. M. Rush, “Pretraining without attention,”
2022
Cited alongside, same era.
“Inter-attribute awareness for pedestrian attribute recognition,”
2022
Cited alongside, same era.
L. Chen, J. Song, X. Zhang, and M. Shang, “Mcfl: multi-label contrastive focal loss for deep imbalanced pedestrian attribute recognition,”
2022
Cited alongside, same era.
Z. Tang and J. Huang, “Drformer: Learning dual relations using transformer for pedestrian attribute recognition,”
2022
Cited alongside, same era.
H. Guo, X. Fan, and S. Wang, “Visual attention consistency for human attribute recognition,”
2022
Cited alongside, same era.
J. Jia, N. Gao, F. He, X. Chen, and K. Huang, “Learning disentangled attribute representations for robust pedestrian attribute recognition,” in
2022
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
X. Pei, T. Huang, and C. Xu, “Efficientvmamba: Atrous selective scan for light weight visual mamba,”
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Y. Yang, Z. Xing, and L. Zhu, “Vivim: a video vision mamba for medical video object segmentation,”
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
A. Archit and C. Pape, “Vim-unet: Vision mamba for biomedical segmentation,”
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
H. Hou and F. R. Yu, “Rwkv-ts: Beyond traditional recurrent neural network for time series tasks,”
2024
Closest in time.
Z. Zhu, W. Shao, and D. Jiao, “Tls-rwkv: Real-time online action detection with temporal label smoothing,”
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
A. Ali, I. Zimerman, and L. Wolf, “The hidden attention of mamba models,”
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
T. Ota, “Decision mamba: Reinforcement learning via sequence modeling with selective state spaces,”
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
N. Zubi’c, M. Gehrig, and D. Scaramuzza, “State space models for event cameras,”
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
M. R. Samsami, A. Zholus, J. Rajendran, and S. Chandar, “Mastering memory tasks with world models,”
2024
Closest in time.
F. Liu and Q. Li, “From generalization analysis to optimization designs for state space models,” 2024. [Online]. Available:
2024
Closest in time.
A. Yu, A. Nigmetov, D. Morozov, M. W. Mahoney, and N. B. Erichson, “Robustifying state-space models for long sequences via approximate diagonalization,” in
2024
Closest in time.
E. David, J. Bellot, and S. L. Corff, “Variational quantization for state space models,” 2024. [Online]. Available:
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Z. Shi, “Mambastock: Selective state space model for stock prediction,”
2024
Closest in time.
2024
Closest in time.
M. A. Ahamed and Q. Cheng, “Timemachine: A time series is worth 4 mambas for long-term forecasting,”
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Z. Xu, “Rankmamba, benchmarking mamba’s document ranking performance in the era of transformers,”
2024
Closest in time.
A. S. Sharma, D. Atkinson, and D. Bau, “Locating and editing factual associations in mamba,” 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
J. Yang, Y. Li, J. Zhao, H. Wang, M. Ma, J. Ma, Z. Ren, M. Zhang, X. Xin, Z. Chen
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
X. Wang, S. Wang, C. Tang, L. Zhu, B. Jiang, Y. Tian, and J. Tang, “Event stream-based visual object tracking: A high-resolution benchmark dataset and a novel baseline,” in
2024
Closest in time.
X. Wang, W. Wu, C. Li, Z. Zhao, Z. Chen, Y. Shi, and J. Tang, “Structural information guided multimodal pre-training for vehicle-centric perception,” in
2024
Closest in time.
2024
Closest in time.