Fetching the paper…
Reading the bibliography…
Understanding the real world through point cloud video is a crucial aspect of robotics and autonomous driving systems.
A. Shahroudy, J. Liu, T. Ng, and G. Wang, “NTU RGB+D: A large scale dataset for 3d human activity analysis,” in
2016
Earlier work this paper cites.
C. R. Qi, L. Yi, H. Su, and L. J. Guibas, “Pointnet++: Deep hierarchical feature learning on point sets in a metric space,” in
2017
Earlier work this paper cites.
C. R. Qi, H. Su, K. Mo, and L. J. Guibas, “Pointnet: Deep learning on point sets for 3d classification and segmentation,” in
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
Y. Wang, Y. Sun, Z. Liu, S. E. Sarma, M. M. Bronstein, and J. M. Solomon, “Dynamic graph cnn for learning on point clouds,”
2019
Earlier work this paper cites.
X. Liu, M. Yan, and J. Bohg, “Meteornet: Deep learning on dynamic 3d point cloud sequences,” in
2019
Earlier work this paper cites.
G. Liu, J. Qian, F. Wen, X. Zhu, R. Ying, and P. Liu, “Action recognition based on 3d skeleton and rgb frame fusion,” in
2019
Earlier work this paper cites.
J. Liu, A. Shahroudy, M. Perez, G. Wang, L.-Y. Duan, and A. C. Kot, “Ntu rgb+d 120: A large-scale benchmark for 3d human activity understanding,”
2019
Earlier work this paper cites.
A. Ošep, P. Voigtlaender, M. Weber, J. Luiten, and B. Leibe, “4d generic video object proposals,” in
2020
Earlier work this paper cites.
Y. Wang, Y. Xiao, F. Xiong, W. Jiang, Z. Cao, J. T. Zhou, and J. Yuan, “3dv: 3d dynamic voxel for action recognition in depth video,” in
2020
Earlier work this paper cites.
S. Das, S. Sharma, R. Dai, F. Brémond, and M. Thonnat, “VPN: learning video-pose embedding for activities of daily living,” in
2020
Earlier work this paper cites.
D. Seichter, M. Köhler, B. Lewandowski, T. Wengefeld, and H.-M. Gross, “Efficient rgb-d semantic segmentation for indoor scene analysis,” in
2021
Earlier work this paper cites.
H. Zhao, L. Jiang, J. Jia, P. H. S. Torr, and V. Koltun, “Point transformer,” in
2021
Earlier work this paper cites.
H. Fan, X. Yu, Y. Ding, Y. Yang, and M. Kankanhalli, “Pstnet: Point spatio-temporal convolution on point cloud sequences,” in
2021
Earlier work this paper cites.
J. Liu and D. Xu, “Geometrymotion-net: A strong two-stream baseline for 3d action recognition,”
2021
Earlier work this paper cites.
H. Fan, Y. Yang, and M. Kankanhalli, “Point 4d transformer networks for spatio-temporal modeling in point cloud videos,” in
2021
Earlier work this paper cites.
B. Ni, H. Peng, M. Chen, S. Zhang, G. Meng, J. Fu, S. Xiang, and H. Ling, “Expanding language-image pretrained models for general video recognition,” in
2022
Earlier work this paper cites.
G. Qian, Y. Li, H. Peng, J. Mai, H. Hammoud, M. Elhoseiny, and B. Ghanem, “Pointnext: Revisiting pointnet++ with improved training and scaling strategies,” in
2022
Earlier work this paper cites.
X. Yu, L. Tang, Y. Rao, T. Huang, J. Zhou, and J. Lu, “Point-bert: Pre-training 3d point cloud transformers with masked point modeling,” in
2022
Earlier work this paper cites.
Y. Pang, W. Wang, F. E. H. Tay, W. Liu, Y. Tian, and L. Yuan, “Masked autoencoders for point cloud self-supervised learning,” in
2022
Earlier work this paper cites.
X. Yan, H. Zhan, C. Zheng, J. Gao, R. Zhang, S. Cui, and Z. Li, “Let images give you more: Point cloud cross-modal training for shape analysis,” in
2022
Cited alongside, same era.
H. Fan, X. Yu, Y. Yang, and M. S. Kankanhalli, “Deep hierarchical representation of point cloud videos via spatio-temporal decomposition,”
2022
Cited alongside, same era.
J. Zhong, K. Zhou, Q. Hu, B. Wang, N. Trigoni, and A. Markham, “No pain, big gain: Classify dynamic point cloud sequences with static models by fitting feature-level space-time surfaces,” in
2022
Cited alongside, same era.
J. Liu, J. Guo, and D. Xu, “Geometrymotion-transformer: An end-to-end framework for 3d action recognition,”
2022
Cited alongside, same era.
H. Wen, Y. Liu, J. Huang, B. Duan, and L. Yi, “Point primitive transformer for long-term 4d point cloud video understanding,” in
2022
H. Rasheed, M. U. khattak, M. Maaz, S. Khan, and F. S. Khan, “Finetuned clip models are efficient video learners,” in
2023
Later among the works it cites.
2023
Later among the works it cites.
L. Xue, M. Gao, C. Xing, R. Martín-Martín, J. Wu, C. Xiong, R. Xu, J. C. Niebles, and S. Savarese, “Ulip: Learning a unified representation of language, images, and point clouds for 3d understanding,” in
2023
Later among the works it cites.
J. Chen, Y. Zhang, F. Ma, and Z. Tan, “Eb-lg module for 3d point cloud classification and segmentation,”
2023
Later among the works it cites.
L. Chen, H. Wang, H. Kong, W. Yang, and M. Ren, “Ptc-net: Point-wise transformer with sparse convolution network for place recognition,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Y. Wei, H. Liu, T. Xie, Q. Ke, and Y. Guo, “Spatial-temporal transformer for 3d point cloud sequences,” in
2022
Cited alongside, same era.
X. Chen, W. Liu, X. Liu, Y. Zhang, J. Han, and T. Mei, “MAPLE: masked pseudo-labeling autoencoder for semi-supervised point cloud action recognition,” in
2022
Cited alongside, same era.
R. Zhang, Z. Guo, W. Zhang, K. Li, X. Miao, B. Cui, Y. Qiao, P. Gao, and H. Li, “Pointclip: Point cloud understanding by CLIP,” in
2022
Cited alongside, same era.
J. Liu, J. Guo, and D. Xu, “Apsnet: Toward adaptive point sampling for efficient 3d action recognition,”
2022
Cited alongside, same era.
B. Zhou, P. Wang, J. Wan, Y. Liang, F. Wang, D. Zhang, Z. Lei, H. Li, and R. Jin, “Decoupling and recoupling spatiotemporal representation for rgb-d-based motion recognition,” in
2022
Cited alongside, same era.
L. Yao, S. Liu, C. Li, S. Zou, S. Chen, and D. Guan, “Pa-awcnn: Two-stream parallel attention adaptive weight network for rgb-d action recognition,” in
2022
Cited alongside, same era.
H. Duan, Y. Zhao, K. Chen, D. Lin, and B. Dai, “Revisiting skeleton-based action recognition,” in
2022
Cited alongside, same era.
2023
Later among the works it cites.
Z. Fang, X. Li, X. Li, J. M. Buhmann, C. C. Loy, and M. Liu, “Explore in-context learning for 3d point cloud understanding,”
2023
Later among the works it cites.
X. Wang, W. Zhang, C. Wang, Y. Gao, and M. Liu, “Dynamic dense graph convolutional network for skeleton-based human motion prediction,”
2023
Later among the works it cites.
Z. Shen, X. Sheng, L. Wang, Y. Guo, Q. Liu, and Z. Xi, “Pointcmp: Contrastive mask prediction for self-supervised learning on point cloud videos,” in
2023
Later among the works it cites.
T. Huang, B. Dong, Y. Yang, X. Huang, R. W. Lau, W. Ouyang, and W. Zuo, “Clip2point: Transfer clip to point cloud classification with image-depth pre-training,” in
2023
Later among the works it cites.
T. Yang, Y. Zhu, Y. Xie, A. Zhang, C. Chen, and M. Li, “Aim: Adapting image models for efficient video understanding,” in
2023
Later among the works it cites.
W. Wu, X. Wang, H. Luo, J. Wang, Y. Yang, and W. Ouyang, “Bidirectional cross-modal knowledge exploration for video recognition with pre-trained vision-language models,” in
2023
Later among the works it cites.
S. T. Wasim, M. Naseer, S. Khan, F. S. Khan, and M. Shah, “Vita-clip: Video and text adaptive clip via multimodal prompting,” in
2023
Later among the works it cites.
D. Ahn, S. Kim, H. Hong, and B. Ko, “Star-transformer: A spatio-temporal cross attention transformer for human action recognition,” in
2023
Later among the works it cites.
B. X. Yu, Y. Liu, X. Zhang, S.-h. Zhong, and K. C. Chan, “Mmnet: A model-based multimodal network for human action recognition in rgb-d videos,”
2023
Later among the works it cites.
Z. Fang, X. Li, X. Li, J. M. Buhmann, C. C. Loy, and M. Liu, “Explore in-context learning for 3d point cloud understanding,”
2024
Closest in time.
2024
Closest in time.
J. Wu, X. Li, S. Xu, H. Yuan, H. Ding, Y. Yang, X. Li, J. Zhang, Y. Tong, X. Jiang, B. Ghanem, and D. Tao, “Towards open vocabulary learning: A survey,”
2024
Closest in time.
X. Wang, Q. Cui, C. Chen, and M. Liu, “Gcnext: Towards the unity of graph convolutions for human motion prediction,” in
2024
Closest in time.
X. Wang, Z. Fang, X. Li, X. Li, C. Chen, and M. Liu, “Skeleton-in-context: Unified skeleton sequence modeling with in-context learning,”
2024
Closest in time.