Fetching the paper…
Reading the bibliography…
Point cloud videos can faithfully capture real-world spatial geometries and temporal dynamics, which are essential for enabling intelligent agents to understand the dynamically changing world.
Development of a stereo video see-through hmd for ar systems
Akinari Takagi, Shoichi Yamazaki, Yoshihiro Saito, and Naosato Taniguchi · 2000
Earlier work this paper cites.
Action recognition based on a bag of 3d points
Wanqing Li, Zhengyou Zhang, and Zicheng Liu · 2010
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei · 2014
Earlier work this paper cites.
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
Xception: Deep learning with depthwise separable convolutions
François Chollet · 2017
Earlier work this paper cites.
Fast and furious: Real time end-to-end 3d detection, tracking and motion forecasting with a single convolutional net
Wenjie Luo, Bin Yang, and Raquel Urtasun · 2018
Earlier work this paper cites.
4d spatio-temporal convnets: Minkowski convolutional neural networks
Christopher Choy, JunYoung Gwak, and Silvio Savarese · 2019
Earlier work this paper cites.
Pointrnn: Point recurrent neural network for moving point cloud processing
Hehe Fan and Yi Yang · 2019
Earlier work this paper cites.
Hippo: Recurrent memory with optimal polynomial projections
Albert Gu, Tri Dao, Stefano Ermon, Atri Rudra, and Christopher Ré · 2020
Earlier work this paper cites.
ψ \psi -net: Stacking densely convolutional lstms for sub-cortical brain structure segmentation
Lihao Liu, Xiaowei Hu, Lei Zhu, Chi-Wing Fu, Jing Qin, and Pheng-Ann Heng · 2020
Earlier work this paper cites.
3dv: 3d dynamic voxel for action recognition in depth video
Yancheng Wang, Yang Xiao, Fu Xiong, Wenxiang Jiang, Zhiguo Cao, Joey Tianyi Zhou, and Junsong Yuan · 2020
Earlier work this paper cites.
Dynamic sampling networks for efficient action recognition in videos
Yin-Dong Zheng, Zhaoyang Liu, Tong Lu, and Limin Wang · 2020
Cited alongside, same era.
Combining recurrent, convolutional, and continuous-time models with linear state space layers
Albert Gu, Isys Johnson, Karan Goel, Khaled Saab, Tri Dao, Atri Rudra, and Christopher Ré · 2021
Cited alongside, same era.
Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras
Zachary Teed and Jia Deng · 2021
Cited alongside, same era.
Pointmixer: Mlp-mixer for point cloud understanding
Jaesung Choe, Chunghyun Park, Francois Rameau, Jaesik Park, and In So Kweon · 2022
Cited alongside, same era.
Point spatio-temporal transformer networks for point cloud video modeling
Hehe Fan, Yi Yang, and Mohan Kankanhalli · 2022
Cited alongside, same era.
Efficiently modeling long sequences with structured state spaces
Spherical frustum sparse convolution network for lidar point cloud semantic segmentation
Yu Zheng, Guangming Wang, Jiuming Liu, Marc Pollefeys, and Hesheng Wang · 2023
Later among the works it cites.
Simba: Mamba augmented u-shiftgcn for skeletal action recognition in videos
Soumyabrata Chaudhuri and Saumik Bhattacharya · 2024
Closest in time.
Video mamba suite: State space model as a versatile alternative for video understanding
Guo Chen, Yifei Huang, Jilan Xu, Baoqi Pei, Zhe Chen, Zhiqi Li, Jiahao Wang, Kunchang Li, Tong Lu, and Limin Wang · 2024
Closest in time.
3dsflabelling: Boosting 3d scene flow estimation by pseudo auto-labelling
Chaokang Jiang, Guangming Wang, Jiuming Liu, Hesheng Wang, Zhuang Ma, Zhenqiang Liu, Zhujin Liang, Yi Shan, and Dalong Du · 2024
Closest in time.
X4d-sceneformer: Enhanced scene understanding on 4d point cloud videos through cross-modal knowledge transfer
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Albert Gu, Karan Goel, and Christopher Ré · 2022
Cited alongside, same era.
Hoi4d: A 4d egocentric dataset for category-level human-object interaction
Yunze Liu, Yun Liu, Che Jiang, Kangbo Lyu, Weikang Wan, Hao Shen, Boqiang Liang, Zhoujie Fu, He Wang, and Li Yi · 2022
Cited alongside, same era.
Point primitive transformer for long-term 4d point cloud video understanding
Hao Wen, Yunze Liu, Jingwei Huang, Bo Duan, and Li Yi · 2022
Cited alongside, same era.
Long-term visual simultaneous localization and mapping: Using a bayesian persistence filter-based global map prediction
Tianchen Deng, Hongle Xie, Jingchuan Wang, and Weidong Chen · 2023
Cited alongside, same era.
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao · 2023
Cited alongside, same era.
Complete-to-partial 4d distillation for self-supervised point cloud sequence representation learning
Zhuoyang Zhang, Yuhao Dong, Yunze Liu, and Li Yi · 2023
Cited alongside, same era.
Point 4d transformer networks for spatio-temporal modeling in point cloud videos
Hehe Fan, Yi Yang, and Mohan Kankanhalli
Cited in the paper.
Linglin Jing, Ying Xue, Xu Yan, Chaoda Zheng, Dong Wang, Ruimao Zhang, Zhigang Wang, Hui Fang, Bin Zhao, and Zhen Li · 2024
Closest in time.
Videomamba: State space model for efficient video understanding
Kunchang Li, Xinhao Li, Yi Wang, Yinan He, Yali Wang, Limin Wang, and Yu Qiao · 2024
Closest in time.
Pointmamba: A simple state space model for point cloud analysis
Dingkang Liang, Xin Zhou, Xinyu Wang, Xingkui Zhu, Wei Xu, Zhikang Zou, Xiaoqing Ye, and Xiang Bai · 2024
Closest in time.
Ssm meets video diffusion models: Efficient video generation with structured state spaces
Yuta Oshima, Shohei Taniguchi, Masahiro Suzuki, and Yutaka Matsuo · 2024
Closest in time.
Segmamba: Long-range sequential modeling mamba for 3d medical image segmentation
Zhaohu Xing, Tian Ye, Yijun Yang, Guang Liu, and Lei Zhu · 2024
Closest in time.
Vivim: a video vision mamba for medical video object segmentation
Yijun Yang, Zhaohu Xing, and Lei Zhu · 2024
Closest in time.