Fetching the paper…
Reading the bibliography…
Video represents the majority of internet traffic today, driving a continual race between the generation of higher quality content, transmission of larger file sizes, and the development of network infrastructure.
R. A. Bradley and M. E. Terry, “Rank analysis of incomplete block designs: I. The method of paired comparisons,” Biometrika , vol. 39, no. 3/4, pp. 324–345, 1952
1952
Earlier work this paper cites.
K. Brandenburg, “MP3 and AAC explained,” in Audio Engineering Society Conference: 17th International Conference: High-Quality Audio Coding . Audio Engineering Society, 1999
1999
Earlier work this paper cites.
E. Cosatto and H. P. Graf, “Photo-realistic talking-heads from image samples,” IEEE Transactions on Multimedia , 2000
2000
Earlier work this paper cites.
T. Wiegand, G. J. Sullivan, G. Bjontegaard, and A. Luthra, “Overview of the H. 264/AVC video coding standard,” IEEE Transactions on circuits and systems for video technology , vol. 13, no. 7, pp. 560–576, 2003
2003
Earlier work this paper cites.
A. Verma, L. V. Subramaniam, N. Rajput, C. Neti, and T. A. Faruquie, “Animating expressive faces across languages,” IEEE Transactions on Multimedia , vol. 6, no. 6, pp. 791–800, 2004
2004
Earlier work this paper cites.
K.-H. Choi and J.-N. Hwang, “Automatic creation of a talking head from a video sequence,” IEEE Transactions on Multimedia , 2005
2005
Earlier work this paper cites.
K.-T. Chen, C.-C. Wu, Y.-C. Chang, and C.-L. Lei, “A crowdsourceable QoE evaluation framework for multimedia content,” in Proceedings of the 17th ACM international conference on Multimedia , 2009
2009
Earlier work this paper cites.
G. Fanelli, J. Gall, H. Romsdorfer, T. Weise, and L. Van Gool, “A 3-d audio-visual corpus of affective communication,” IEEE Transactions on Multimedia , vol. 12, no. 6, pp. 591–598, 2010
2010
Earlier work this paper cites.
R. D. Luce, Individual choice behavior: A theoretical analysis . Courier Corporation, 2012
2012
Earlier work this paper cites.
Y. Chen, K. Wu, and Q. Zhang, “From QoS to QoE: A tutorial on video quality assessment,” IEEE Communications Surveys & Tutorials , vol. 17, no. 2, pp. 1126–1165, 2014
2014
Earlier work this paper cites.
E. Mansimov, E. Parisotto, L. J. Ba, and R. Salakhutdinov, “Generating images from captions with attention,” in ICLR , 2016
2016
Earlier work this paper cites.
S. Suwajanakorn, S. M. Seitz, and I. Kemelmacher-Shlizerman, “Synthesizing Obama: learning lip sync from audio,” ACM Transactions on Graphics (ToG) , vol. 36, no. 4, pp. 1–13, 2017
2017
Earlier work this paper cites.
S. Taylor, T. Kim, Y. Yue, M. Mahler, J. Krahe, A. G. Rodriguez, J. Hodgins, and I. Matthews, “A deep learning approach for generalized speech animation,” ACM Transactions on Graphics (TOG) , vol. 36, no. 4, pp. 1–11, 2017
2017
Earlier work this paper cites.
T. Karras, T. Aila, S. Laine, A. Herva, and J. Lehtinen, “Audio-driven facial animation by joint end-to-end learning of pose and emotion,” ACM Transactions on Graphics (TOG) , vol. 36, no. 4, pp. 1–12, 2017
2017
Earlier work this paper cites.
Y. Chen, D. Murherjee, J. Han, A. Grange, Y. Xu, Z. Liu, S. Parker, C. Chen, H. Su, U. Joshi et al. , “An overview of core coding tools in the AV1 video codec,” in 2018 Picture Coding Symposium (PCS) . IEEE, 2018, pp. 41–45
2018
Earlier work this paper cites.
Y. Li, M. Min, D. Shen, D. Carlson, and L. Carin, “Video generation from text,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 32, no. 1, 2018
2018
Earlier work this paper cites.
X. Yang, T. Zhang, and C. Xu, “Text2video: An end-to-end learning framework for expressing text with videos,” IEEE Transactions on Multimedia , vol. 20, no. 9, pp. 2360–2370, 2018
2018
Earlier work this paper cites.
T.-C. Wang, M.-Y. Liu, J.-Y. Zhu, N. Yakovenko, A. Tao, J. Kautz, and B. Catanzaro, “Video-to-Video Synthesis,” in NeurIPS , 2018
2018
Earlier work this paper cites.
Y. Zhou, Z. Xu, C. Landreth, E. Kalogerakis, S. Maji, and K. Singh, “Visemenet: Audio-driven animator-centric speech animation,” ACM Transactions on Graphics (TOG) , vol. 37, no. 4, pp. 1–10, 2018
2018
Earlier work this paper cites.
M. Yuan and Y. Peng, “CKD: Cross-task knowledge distillation for text-to-image synthesis,” IEEE Transactions on Multimedia , vol. 22, no. 8, pp. 1955–1968, 2019
2019
Earlier work this paper cites.
A. Siarohin, S. Lathuilière, S. Tulyakov, E. Ricci, and N. Sebe, “First order motion model for image animation,” Advances in Neural Information Processing Systems , vol. 32, pp. 7137–7147, 2019
2019
Cited alongside, same era.
O. Fried, A. Tewari, M. Zollhöfer, A. Finkelstein, E. Shechtman, D. B. Goldman, K. Genova, Z. Jin, C. Theobalt, and M. Agrawala, “Text-based editing of talking-head video,” ACM Transactions on Graphics (TOG) , vol. 38, no. 4, pp. 1–14, 2019
2019
Cited alongside, same era.
C. Coupé, Y. M. Oh, D. Dediu, and F. Pellegrino, “Different languages, similar encoding efficiency: Comparable information rates across the human communicative niche,” Science advances , p. eaaw2594, 2019
2019
Cited alongside, same era.
D. Cudeiro, T. Bolkart, C. Laidlaw, A. Ranjan, and M. J. Black, “Capture, learning, and synthesis of 3D speaking styles,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 10 101–10 111
2019
C. Vaccari and A. Chadwick, “Deepfakes and disinformation: Exploring the impact of synthetic political video on deception, uncertainty, and trust in news,” Social Media+ Society , vol. 6, no. 1, p. 2056305120903408, 2020
2020
Later among the works it cites.
J. Thies, M. Elgharib, A. Tewari, C. Theobalt, and M. Nießner, “Neural voice puppetry: Audio-driven facial reenactment,” in European Conference on Computer Vision . Springer, 2020, pp. 716–731
2020
Later among the works it cites.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations,” Advances in Neural Information Processing Systems , vol. 33, 2020
2020
Later among the works it cites.
Cisco, ”Cisco visual networking index: global mobile data traffic forecast update, 2017–2022.” , accessed 2021. [Online]. Available: https://s3.amazonaws.com/media.mediapost.com/uploads/CiscoForecast.pdf
2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
L. Chen, H. Zheng, R. K. Maddox, Z. Duan, and C. Xu, “Sound to visual: Hierarchical cross-modal talking face video generation,” in IEEE Computer Society Conference on Computer Vision and Pattern Recognition workshops , 2019
2019
Cited alongside, same era.
A. Jamaludin, J. S. Chung, and A. Zisserman, “You said that?: Synthesising talking faces from audio,” International Journal of Computer Vision , vol. 127, no. 11, pp. 1767–1779, 2019
2019
Cited alongside, same era.
P. KR, R. Mukhopadhyay, J. Philip, A. Jha, V. Namboodiri, and C. Jawahar, “Towards automatic face-to-face translation,” in Proceedings of the 27th ACM International Conference on Multimedia , 2019
2019
Cited alongside, same era.
N. Pandey, A. Pal et al. , “Impact of digital surge during Covid-19 pandemic: A viewpoint on research and practice,” International Journal of Information Management , vol. 55, p. 102171, 2020
2020
Cited alongside, same era.
M. Candela, V. Luconi, and A. Vecchio, “Impact of the covid-19 pandemic on the internet latency: A large-scale study,” Computer Networks , vol. 182, p. 107495, 2020
2020
Cited alongside, same era.
G. S. Ford, “Covid-19 and broadband speeds: A multi-country analysis,” Available at SSRN 3689044 , 2020
2020
Cited alongside, same era.
D. Kim, D. Joo, and J. Kim, “TiVGAN: Text to Image to Video Generation With Step-by-Step Evolutionary Generator,” IEEE Access , vol. 8, pp. 153 113–153 122, 2020
2020
Cited alongside, same era.
D. Weissenborn, J. Uszkoreit, and O. Täckström, “Scaling Autoregressive Video Models,” in ICLR , 2020
2020
Cited alongside, same era.
Cisco, ”Cisco Annual Internet Report (2018–2023) White Paper” , accessed 2021. [Online]. Available: https://www.cisco.com/c/en/us/solutions/collateral/executive-perspectives/annual-internet-report/white-paper-c11-741490.html
2021
Closest in time.
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever, “Zero-Shot Text-to-Image Generation,” in Proceedings of the 38th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, M. Meila and T. Zhang, Eds., vol. 139. PMLR, 18–24 Jul 2021, pp. 8821–8831. [Online]. Available: https://proceedings.mlr.press/v139/ramesh21a.html
2021
Closest in time.
S. E. Eskimez, Y. Zhang, and Z. Duan, “Speech driven talking face generation from a single image and an emotion condition,” IEEE Transactions on Multimedia , 2021
2021
Closest in time.
L. Yu, H. Xie, and Y. Zhang, “Multimodal learning for temporally coherent talking face generation with articulator synergy,” IEEE Transactions on Multimedia , 2021
2021
Closest in time.
T.-C. Wang, A. Mallya, and M.-Y. Liu, “One-shot free-view neural talking-head synthesis for video conferencing,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 10 039–10 049
2021
Closest in time.
R. Prabhakar, S. Chandak, C. Chiu, R. Liang, H. Nguyen, K. Tatwawadi, and T. Weissman, “Reducing latency and bandwidth for video streaming using keypoint extraction and digital puppetry,” in 2021 Data Compression Conference (DCC) . IEEE, 2021, pp. 360–360
2021
Closest in time.
Resemble, RESEMBLE.AI: Create AI Voices that sound real. , accessed 2021. [Online]. Available: https://www.resemble.ai
2021
Closest in time.
Gzip, Gzip , accessed 2021. [Online]. Available: https://www.gzip.org
2021
Closest in time.
Bzip2, Bzip2 , accessed 2021. [Online]. Available: http://www.bzip.org
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
Google, Google Text-to-Speech. , accessed 2021. [Online]. Available: https://cloud.google.com/text-to-speech
2021
Closest in time.
M. Azure, Microsoft Text-to-Speech. , accessed 2021. [Online]. Available: https://azure.microsoft.com/en-in/services/cognitive-services/text-to-speech/
2021
Closest in time.
Descript, Descript: Ultra-realistic voice cloning. , accessed 2021. [Online]. Available: https://www.descript.com/overdub
2021
Closest in time.
Google, Google Speech-to-Text. , accessed 2021. [Online]. Available: https://cloud.google.com/speech-to-text
2021
Closest in time.