Fetching the paper…
Reading the bibliography…
Video Internet of Things (VIoT) has shown full potential in collecting an unprecedented volume of video data.
View adaptive neural networks for high performance skeleton-based human action recognition
Zhang, P.; Lan, C.; Xing, J.; Zeng, W.; Xue, J.; and Zheng, N. 2019 · 1978
Earlier work this paper cites.
Yolov4: Optimal speed and accuracy of object detection
Bochkovskiy, A.; Wang, C.-Y.; and Liao, H.-Y. M. 2020 · 2004
Earlier work this paper cites.
Internet of things: Objectives and scientific challenges
Ma, H.-D. 2011 · 2011
Earlier work this paper cites.
Robust view transformation model for gait recognition
Zheng, S.; Zhang, J.; Huang, K.; He, R.; and Tan, T. 2011 · 2011
Earlier work this paper cites.
3D convolutional neural networks for human action recognition
Ji, S.; Xu, W.; Yang, M.; and Yu, K. 2012 · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A.; Sutskever, I.; and Hinton, G. E. 2012 · 2012
Earlier work this paper cites.
Deep learning face representation by joint identification-verification
Sun, Y.; Chen, Y.; Wang, X.; and Tang, X. 2014 · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Caba Heilbron, F.; Escorcia, V.; Ghanem, B.; and Carlos Niebles, J. 2015 · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K.; and Zisserman, A. 2015 · 2015
Earlier work this paper cites.
Scalable person re-identification: A benchmark
Zheng, L.; Shen, L.; Tian, L.; Wang, S.; Wang, J.; and Tian, Q. 2015 · 2015
Earlier work this paper cites.
Ms-celeb-1m: A dataset and benchmark for large-scale face recognition
Guo, Y.; Zhang, L.; Hu, Y.; He, X.; and Gao, J. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Carreira, J.; and Zisserman, A. 2017 · 2017
Earlier work this paper cites.
Crowd Scene Understanding from Video: A Survey
Grant, J. M.; and Flynn, P. J. 2017 · 2017
Earlier work this paper cites.
From Benedict Cumberbatch to Sherlock Holmes: Character Identification in TV series without a Script
Nagrani, A.; and Zisserman, A. 2017 · 2017
Earlier work this paper cites.
Places: A 10 million Image Database for Scene Recognition
Zhou, B.; Lapedriza, A.; Khosla, A.; Oliva, A.; and Torralba, A. 2017 · 2017
Earlier work this paper cites.
Squeeze-and-excitation networks
Hu, J.; Shen, L.; and Sun, G. 2018 · 2018
Earlier work this paper cites.
Real-world anomaly detection in surveillance videos
Sultani, W.; Chen, C.; and Shah, M. 2018 · 2018
Earlier work this paper cites.
Towards end-to-end license plate detection and recognition: A large dataset and baseline
Xu, Z.; Yang, W.; Meng, A.; Lu, N.; Huang, H.; Ying, C.; and Huang, L. 2018 · 2018
Earlier work this paper cites.
Pedestrian alignment network for large-scale person re-identification
Zheng, Z.; Zheng, L.; and Yang, Y. 2018 · 2018
Earlier work this paper cites.
Arcface: Additive angular margin loss for deep face recognition
Deng, J.; Guo, J.; Xue, N.; and Zafeiriou, S. 2019 · 2019
Earlier work this paper cites.
Locality-constrained spatial transformer network for video crowd counting
Fang, Y.; Zhan, B.; Cai, W.; Gao, S.; and Hu, B. 2019 · 2019
Earlier work this paper cites.
Exploring background-bias for anomaly detection in surveillance videos
Liu, K.; and Ma, H. 2019 · 2019
Earlier work this paper cites.
An autoencoder and LSTM-based traffic flow prediction method
Wei, W.; Wu, H.; and Ma, H. 2019 · 2019
Earlier work this paper cites.
A review of machine learning and IoT in smart transportation
Zantalis, F.; Koulouras, G.; Karabetsos, S.; and Kandris, D. 2019 · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 2020
Cited alongside, same era.
Internet of video things: Next-generation IoT with visual sensors
Chen, C. W. 2020 · 2020
Cited alongside, same era.
Enhancing anomaly detection in surveillance videos with transfer learning from action recognition
Liu, K.; Zhu, M.; Fu, H.; Ma, H.; and Chua, T.-S. 2020 · 2020
Cited alongside, same era.
Global structure graph guided fine-grained vehicle recognition
Wang, C.; Fu, H.; and Ma, H. 2020 · 2020
Cited alongside, same era.
Empowering things with intelligence: a survey of the progress, challenges, and opportunities in artificial intelligence of things
ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings
Hao, S.; Liu, T.; Wang, Z.; and Hu, Z. 2023 · 2023
Closest in time.
Fastreid: A pytorch toolbox for general instance re-identification
He, L.; Liao, X.; Liu, W.; Liu, X.; Cheng, P.; and Mei, T. 2023 · 2023
Closest in time.
Segment anything is not always perfect: An investigation of sam on different real-world applications
Ji, W.; Li, J.; Bi, Q.; Li, W.; and Cheng, L. 2023 · 2023
Closest in time.
Guide Your Agent with Adaptive Multimodal Rewards
Kim, C.; Seo, Y.; Liu, H.; Lee, L.; Shin, J.; Lee, H.; and Lee, K. 2023 · 2023
Closest in time.
Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A. C.; Lo, W.-Y.; Dollár, P.; and Girshick, R. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhang, J.; and Tao, D. 2020 · 2020
Cited alongside, same era.
Exploring simple siamese representation learning
Chen, X.; and He, K. 2021 · 2021
Cited alongside, same era.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N. 2021 · 2021
Cited alongside, same era.
Paddleseg: A high-efficient development toolkit for image segmentation
Liu, Y.; Chu, L.; Chen, G.; Wu, Z.; Chen, Z.; Lai, B.; and Hao, Y. 2021 · 2021
Cited alongside, same era.
Rethinking counting and localization in crowds: A purely point-based framework
Song, Q.; Wang, C.; Jiang, Z.; Wang, Y.; Tai, Y.; Wang, C.; Li, J.; Huang, F.; and Wu, Y. 2021 · 2021
Cited alongside, same era.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Wang, B.; and Komatsuzaki, A. 2021 · 2021
Cited alongside, same era.
Deep face recognition: A survey
Wang, M.; and Deng, W. 2021 · 2021
Cited alongside, same era.
Closest in time.
Learning Prompt-Enhanced Context Features for Weakly-Supervised Video Anomaly Detection
Pu, Y.; Wu, X.; and Wang, S. 2023 · 2023
Closest in time.
Toolformer: Language models can teach themselves to use tools
Schick, T.; Dwivedi-Yu, J.; Dessì, R.; Raileanu, R.; Lomeli, M.; Zettlemoyer, L.; Cancedda, N.; and Scialom, T. 2023 · 2023
Closest in time.
Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface
Shen, Y.; Song, K.; Tan, X.; Li, D.; Lu, W.; and Zhuang, Y. 2023 · 2023
Closest in time.
Vipergpt: Visual inference via python execution for reasoning
Surís, D.; Menon, S.; and Vondrick, C. 2023 · 2023
Closest in time.
Can sam segment anything? when sam meets camouflaged object detection
Tang, L.; Xiao, H.; and Li, B. 2023 · 2023
Closest in time.
Image as a foreign language: BEiT pretraining for vision and vision-language tasks
Wang, W.; Bao, H.; Dong, L.; Bjorck, J.; Peng, Z.; Liu, Q.; Aggarwal, K.; Mohammed, O. K.; Singhal, S.; Som, S.; and Wei, F. 2023 · 2023
Closest in time.
Visual chatgpt: Talking, drawing and editing with visual foundation models
Wu, C.; Yin, S.; Qi, W.; Wang, X.; Tang, Z.; and Duan, N. 2023 · 2023
Closest in time.
YuNet: A Tiny Millisecond-level Face Detector
Wu, W.; Peng, H.; and Yu, S. 2023 · 2023
Closest in time.
React: Synergizing reasoning and acting in language models
Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K.; and Cao, Y. 2023 · 2023
Closest in time.
Recognize Anything: A Strong Image Tagging Model
Zhang, Y.; Huang, X.; Ma, J.; Li, Z.; Luo, Z.; Xie, Y.; Qin, Y.; Luo, T.; Li, Y.; Liu, S.; et al. 2023 · 2023
Closest in time.
A survey of large language models
Zhao, W. X.; Zhou, K.; Li, J.; Tang, T.; Wang, X.; Hou, Y.; Min, Y.; Zhang, B.; Zhang, J.; Dong, Z.; et al. 2023 · 2023
Closest in time.
Judging LLM-as-a-judge with MT-Bench and Chatbot Arena
Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E.; et al. 2023 · 2023
Closest in time.
Zhou, T.; Zhang, Y.; Zhou, Y.; Wu, Y.; and Gong, C. 2023 · 2023
Closest in time.
Video generation models as world simulators
Brooks, T.; Peebles, B.; Holmes, C.; DePue, W.; Guo, Y.; Jing, L.; Schnurr, D.; Taylor, J.; Luhman, T.; Luhman, E.; Ng, C.; Wang, R.; and Ramesh, A. 2024 · 2024
Closest in time.
CLOVA: A closed-loop visual assistant with tool usage and update
Gao, Z.; Du, Y.; Zhang, X.; Ma, X.; Han, W.; Zhu, S.-C.; and Li, Q. 2024 · 2024
Closest in time.
Disentangled counterfactual learning for physical audiovisual commonsense reasoning
Lv, C.; Zhang, S.; Tian, Y.; Qi, M.; and Ma, H. 2024 · 2024
Closest in time.
Weakly-Supervised Temporal Action Localization by Inferring Salient Snippet-Feature
Yun, W.; Qi, M.; Wang, C.; and Ma, H. 2024 · 2024
Closest in time.
Internet of video things in 2030: A world with many cameras
Mohan, A.; Gauen, K.; Lu, Y.-H.; Li, W. W.; and Chen, X. 2017 · 2030
Closest in time.