Fetching the paper…
Reading the bibliography…
This paper offers an insightful examination of how currently top-trending AI technologies, i.e., generative artificial intelligence (Generative AI) and large language models (LLMs), are reshaping the field of video technology, including video generation, understanding, and streaming.
T. Wiegand et al. , “Overview of the h. 264/avc video coding standard,” IEEE Transactions on circuits and systems for video technology , vol. 13, no. 7, pp. 560–576, 2003
2003
Earlier work this paper cites.
I. Sodagar, “The mpeg-dash standard for multimedia streaming over the internet,” IEEE multimedia , vol. 18, no. 4, pp. 62–67, 2011
2011
Earlier work this paper cites.
G. J. Sullivan et al. , “Overview of the high efficiency video coding(hevc) standard,” IEEE Transactions on circuits and systems for video technology , vol. 22, no. 12, pp. 1649–1668, 2012
2012
Earlier work this paper cites.
O. Ronneberger et al. , “U-net: Convolutional networks for biomedical image segmentation,” in Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, October 5-9, 2015, Proceedings, Part III 18 . Springer, 2015, pp. 234–241
2015
Earlier work this paper cites.
C. Vondrick et al. , “Generating videos with scene dynamics,” Advances in neural information processing systems , vol. 29, 2016
2016
Earlier work this paper cites.
A. Van den Oord et al. , “Conditional image generation with pixelcnn decoders,” Advances in neural information processing systems , vol. 29, 2016
2016
Earlier work this paper cites.
J. van der Hooft et al. , “HTTP/2-Based Adaptive Streaming of HEVC Video Over 4G/LTE Networks,” IEEE Communications Letters , vol. 20, no. 11, pp. 2177–2180, 2016
2016
Earlier work this paper cites.
N. Kalchbrenner et al. , “Video pixel networks,” in International Conference on Machine Learning . PMLR, 2017, pp. 1771–1779
2017
Earlier work this paper cites.
A. Vaswani et al. , “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
W. Kay et al. , “The kinetics human action video dataset,” arXiv preprint arXiv:1705.06950 , 2017
2017
Earlier work this paper cites.
Huawei, “Cloud vr-oriented bearer network white paper,” Huawei iLab VR Technology White Paper , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
E. Denton et al. , “Stochastic video generation with a learned prior,” in International conference on machine learning . PMLR, 2018, pp. 1174–1183
2018
Earlier work this paper cites.
M. Oliu et al. , “Folded recurrent neural networks for future video prediction,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 716–731
2018
Earlier work this paper cites.
N. Aafaq et al. , “Video description: A survey of methods, datasets, and evaluation metrics,” ACM Computing Surveys (CSUR) , vol. 52, no. 6, pp. 1–37, 2019
2019
Earlier work this paper cites.
C. Chan et al. , “Everybody dance now,” in Proceedings of the IEEE/CVF international conference on computer vision , 2019, pp. 5933–5942
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
L. Castrejon et al. , “Improved conditional vrnns for video prediction,” in Proceedings of the IEEE/CVF international conference on computer vision , 2019, pp. 7608–7617
2019
Earlier work this paper cites.
R. Bhagwatkar et al. , “A review of video generation approaches,” in 2020 International Conference on Power, Instrumentation, Control and Computing (PICC) . IEEE, 2020, pp. 1–5
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
J. van der Hooft et al. , “From capturing to rendering: Volumetric media delivery with six degrees of freedom,” IEEE Communications Magazine , vol. 58, no. 10, pp. 49–55, 2020
2020
Earlier work this paper cites.
M. Saito et al. , “Train sparsely, generate densely: Memory-efficient unsupervised training of high-resolution temporal gan,” International Journal of Computer Vision , vol. 128, no. 10-11, pp. 2586–2606, 2020
2020
Earlier work this paper cites.
A. Radford et al. , “Learning transferable visual models from natural language supervision,” in International conference on machine learning . PMLR, 2021, pp. 8748–8763
2021
Earlier work this paper cites.
F.-Y. Chao et al. , “Transformer-based long-term viewport prediction in 360° video: Scanpath is all you need.” in MMSP , 2021, pp. 1–6
2021
Earlier work this paper cites.
M. Ding et al. , “Cogview: Mastering text-to-image generation via transformers,” Advances in Neural Information Processing Systems , vol. 34, pp. 19 822–19 835, 2021
2021
Earlier work this paper cites.
M. Bain et al. , “Frozen in time: A joint video and image encoder for end-to-end retrieval,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 1728–1738
2021
Earlier work this paper cites.
Y. Kasten et al. , “Layered neural atlases for consistent video editing,” ACM Transactions on Graphics (TOG) , vol. 40, no. 6, pp. 1–12, 2021
2021
Earlier work this paper cites.
Z. Liu et al. , “Point cloud video streaming: Challenges and solutions,” IEEE Network , vol. 35, no. 5, pp. 202–209, 2021
2021
Earlier work this paper cites.
J. Guo et al. , “A video-quality driven strategy in short video streaming,” in Proceedings of the 24th International ACM Conference on Modeling, Analysis and Simulation of Wireless and Mobile Systems , 2021, pp. 221–228
2021
Earlier work this paper cites.
C. Kattadige et al. , “Videotrain: A generative adversarial framework for synthetic video traffic generation,” in 2021 IEEE 22nd International Symposium on a World of Wireless, Mobile and Multimedia Networks (WoWMoM) . IEEE, 2021, pp. 209–218
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
L. Wang et al. , “Knowledge distillation and student-teacher learning for visual intelligence: A review and new outlooks,” IEEE transactions on pattern analysis and machine intelligence , vol. 44, no. 6, pp. 3048–3068, 2021
2021
Earlier work this paper cites.
N. Aldausari et al. , “Video generative adversarial networks: a review,” ACM Computing Surveys (CSUR) , vol. 55, no. 2, pp. 1–25, 2022
2022
Earlier work this paper cites.
J. Ho et al. , “Video diffusion models,” arXiv preprint arXiv:2204.03458 , 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
J. Dave et al. , “Hierarchical language modeling for dense video captioning,” in Inventive Computation and Information Technologies: Proceedings of ICICIT 2021 . Springer, 2022, pp. 421–431
2022
Cited alongside, same era.
M. Yuksekgonul et al. , “When and why vision-language models behave like bags-of-words, and what to do about it?” in The Eleventh International Conference on Learning Representations , 2022
2022
Cited alongside, same era.
T. Azmin et al. , “Bandwidth prediction in 5g mobile networks using informer,” in 2022 13th International Conference on Network of the Future (NoF) . IEEE, 2022, pp. 1–9
2022
Cited alongside, same era.
S. Van Damme et al. , “Machine learning based content-agnostic viewport prediction for 360-degree video,” ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM) , vol. 18, no. 2, pp. 1–24, 2022
2022
Cited alongside, same era.
Y. Zhao et al. , “Learning video representations from large language models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 6586–6597
2023
Later among the works it cites.
K. Ma et al. , “Llavilo: Boosting video moment retrieval via adapter-based multimodal modeling,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 2798–2803
2023
Later among the works it cites.
Z. Shao et al. , “Prompting large language models with answer heuristics for knowledge-based visual question answering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 14 974–14 983
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Cheng et al. , “Gaze estimation using transformer,” in 2022 26th International Conference on Pattern Recognition (ICPR) . IEEE, 2022, pp. 3341–3347
2022
Cited alongside, same era.
J. Xiang et al. , “Mimt: Masked image modeling transformer for video compression,” in The Eleventh International Conference on Learning Representations , 2022
2022
Cited alongside, same era.
F. Mentzer et al. , “Vct: A video compression transformer,” arXiv preprint arXiv:2206.07307 , 2022
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
O. Bar-Tal et al. , “Text2live: Text-driven layered image and video editing,” in European conference on computer vision . Springer, 2022, pp. 707–723
2022
Cited alongside, same era.
2022
Cited alongside, same era.
V. Voleti et al. , “Mcvd-masked conditional video diffusion for prediction, generation, and interpolation,” Advances in Neural Information Processing Systems , vol. 35, pp. 23 371–23 385, 2022
2022
Cited alongside, same era.
H. J. Singh et al. , “Visual questions answering developments, applications, datasets and opportunities: A state-of-the-art survey,” in 2023 International Conference on Sustainable Computing and Data Communication Systems (ICSCDS) . IEEE, 2023, pp. 778–785
2023
Later among the works it cites.
J. Guo et al. , “From images to textual prompts: Zero-shot visual question answering with frozen large language models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 10 867–10 877
2023
Later among the works it cites.
A. Salaberria et al. , “Image captioning for effective use of language models in knowledge-based visual question answering,” Expert Systems with Applications , vol. 212, p. 118669, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Z. Hu et al. , “Reveal: Retrieval-augmented visual-language pre-training with multi-source multimodal knowledge memory,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 23 369–23 379
2023
Later among the works it cites.
M. Gao et al. , “Deep learning for video object segmentation: a review,” Artificial Intelligence Review , vol. 56, no. 1, pp. 457–531, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
J. Li et al. , “Spherical convolution empowered viewport prediction in 360 video multicast with limited fov feedback,” ACM Transactions on Multimedia Computing, Communications and Applications , vol. 19, no. 1, pp. 1–23, 2023
2023
Later among the works it cites.
Z. Ni et al. , “Human-object interaction prediction in videos through gaze following,” Computer Vision and Image Understanding , p. 103741, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
N. Ruiz et al. , “Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 22 500–22 510
2023
Later among the works it cites.
A. Kirillov et al. , “Segment anything,” arXiv preprint arXiv:2304.02643 , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
G. Kim et al. , “Diffusion video autoencoders: Toward temporally consistent face video editing via disentangled video encoding,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 6091–6100
2023
Later among the works it cites.
2023
Later among the works it cites.
G. A. S. Surek et al. , “Video-based human activity recognition using deep learning approaches,” Sensors , vol. 23, no. 14, p. 6384, 2023
2023
Later among the works it cites.
M. G. Morshed et al. , “Human action recognition: A taxonomy-based survey, updates, and opportunities,” Sensors , vol. 23, no. 4, p. 2182, 2023
2023
Later among the works it cites.
C. Zhang et al. , “Large language models for human-robot interaction: A review,” Biomimetic Intelligence and Robotics , p. 100131, 2023
2023
Later among the works it cites.
H. Kaneko et al. , “Toward pioneering sensors and features using large language models in human activity recognition,” in Adjunct Proceedings of the 2023 ACM International Joint Conference on Pervasive and Ubiquitous Computing & the 2023 ACM International Symposium on Wearable Computing , 2023, pp. 475–479
2023
Later among the works it cites.
2023
Later among the works it cites.
W. Wu et al. , “Bidirectional cross-modal knowledge exploration for video recognition with pre-trained vision-language models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 6620–6630
2023
Later among the works it cites.
K. Li et al. , “Videochat: Chat-centric video understanding,” arXiv preprint arXiv:2305.06355 , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
D. Shah et al. , “Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action,” in Conference on Robot Learning . PMLR, 2023, pp. 492–504
2023
Later among the works it cites.
I. Singh et al. , “Progprompt: Generating situated robot task plans using large language models,” in 2023 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2023, pp. 11 523–11 530
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
N. Renella et al. , “Towards automated video game commentary using generative ai,” 2023
2023
Later among the works it cites.
S. Angarano et al. , “Generative adversarial super-resolution at the edge with knowledge distillation,” Engineering Applications of Artificial Intelligence , vol. 123, p. 106407, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
J. Zhu et al. , “A good student is cooperative and reliable: Cnn-transformer collaborative learning for semantic segmentation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 11 720–11 730
2023
Later among the works it cites.