Fetching the paper…
Reading the bibliography…
We propose VLOGGER, a method for audio-driven human video generation from a single input image of a person, which builds on the success of recent generative diffusion models.
Narvekar, N.D., Karam, L.J.: A no-reference perceptual image sharpness metric based on a cumulative probability of blur detection. In: International Workshop on Quality of Multimedia Experience. pp. 87–91. IEEE (2009)
2009
Earlier work this paper cites.
Paysan, P., Knothe, R., Amberg, B., Romdhani, S., Vetter, T.: A 3d face model for pose and illumination invariant face recognition. In: International Conference on Advanced Video and Signal Based Surveillance. pp. 296–301. Ieee (2009)
2009
Earlier work this paper cites.
Narvekar, N.D., Karam, L.J.: A no-reference image blur metric based on the cumulative probability of blur detection (cpbd). Image Processing 20
2011
Earlier work this paper cites.
Mittal, A., Soundararajan, R., Bovik, A.C.: Making a “completely blind” image quality analyzer. Signal Processing Letters 20
2012
Earlier work this paper cites.
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., Bengio, Y.: Generative adversarial nets. In: NeurIPS (2014)
2014
Earlier work this paper cites.
Bogo, F., Kanazawa, A., Lassner, C., Gehler, P., Romero, J., Black, M.J.: Keep it SMPL: Automatic estimation of 3d human pose and shape from a single image. In: ECCV (2016)
2016
Earlier work this paper cites.
Chung, J.S., Zisserman, A.: Out of time: automated lip sync in the wild. In: ACCV-W (2016)
2016
Earlier work this paper cites.
Thies, J., Zollhofer, M., Stamminger, M., Theobalt, C., Nießner, M.: Face2face: Real-time face capture and reenactment of rgb videos. In: CVPR. pp. 2387–2395 (2016)
2016
Earlier work this paper cites.
Chung, J.S., Jamaludin, A., Zisserman, A.: You said that? (2017)
2017
Earlier work this paper cites.
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., Hochreiter, S.: Gans trained by a two time-scale update rule converge to a local nash equilibrium. NeurIPS 30
2017
Earlier work this paper cites.
Li, T., Bolkart, T., Black, M.J., Li, H., Romero, J.: Learning a model of facial shape and expression from 4d scans. TOG 36
2017
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I.: Attention is all you need. NeurIPS (2017)
2017
Earlier work this paper cites.
Joo, H., Simon, T., Sheikh, Y.: Total capture: A 3d deformation model for tracking faces, hands, and bodies. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 8320–8329 (2018)
2018
Earlier work this paper cites.
Wiles, O., Koepke, A., Zisserman, A.: X2face: A network for controlling face generation using images, audio, and pose codes. In: ECCV. pp. 670–686 (2018)
2018
Earlier work this paper cites.
Chen, L., Maddox, R.K., Duan, Z., Xu, C.: Hierarchical cross-modal talking face generation with dynamic pixel-wise loss. In: CVPR. pp. 7832–7841 (2019)
2019
Earlier work this paper cites.
Cudeiro, D., Bolkart, T., Laidlaw, C., Ranjan, A., Black, M.J.: Capture, learning, and synthesis of 3d speaking styles. In: CVPR (2019)
2019
Earlier work this paper cites.
Deng, J., Guo, J., Xue, N., Zafeiriou, S.: Arcface: Additive angular margin loss for deep face recognition. In: CVPR. pp. 4690–4699 (2019)
2019
Earlier work this paper cites.
Jamaludin, A., Chung, J.S., Zisserman, A.: You said that?: Synthesising talking faces from audio. IJCV 127
2019
Earlier work this paper cites.
Nirkin, Y., Keller, Y., Hassner, T.: Fsgan: Subject agnostic face swapping and reenactment. In: ICCV. pp. 7184–7193 (2019)
2019
Earlier work this paper cites.
Pavlakos, G., Choutas, V., Ghorbani, N., Bolkart, T., Osman, A.A., Tzionas, D., Black, M.J.: Expressive body capture: 3d hands, face, and body from a single image. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 10975–10985 (2019)
2019
Earlier work this paper cites.
Pumarola, A., Agudo, A., Martinez, A., Sanfeliu, A., Moreno-Noguer, F.: Ganimation: One-shot anatomically consistent facial animation. In: IJCV (2019)
2019
Earlier work this paper cites.
Text-to-speech - google cloud. https://cloud.google.com/text-to-speech (2019)
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Zhou, H., Liu, Y., Liu, Z., Luo, P., Wang, X.: Talking face generation by adversarially disentangled audio-visual representation. In: AAAI (2019)
2019
Earlier work this paper cites.
Ha, S., Kersner, M., Kim, B., Seo, S., Kim, D.: Marionette: Few-shot face reenactment preserving identity of unseen targets. In: AAAI (2020)
2020
Earlier work this paper cites.
Karras, T., Laine, S., Aittala, M., Hellsten, J., Lehtinen, J., Aila, T.: Analyzing and improving the image quality of stylegan. In: CVPR. pp. 8110–8119 (2020)
2020
Earlier work this paper cites.
Mildenhall, B., Srinivasan, P.P., Tancik, M., Barron, J.T., Ramamoorthi, R., Ng, R.: Nerf: Representing scenes as neural radiance fields for view synthesis. In: ECCV (2020)
2020
Earlier work this paper cites.
Prajwal, K., Mukhopadhyay, R., Namboodiri, V.P., Jawahar, C.: A lip sync expert is all you need for speech to lip generation in the wild. In: ACM International Conference on Multimedia. pp. 484–492 (2020)
2020
Earlier work this paper cites.
Vougioukas, K., Petridis, S., Pantic, M.: Realistic speech-driven facial animation with gans. IJCV 128
2020
Earlier work this paper cites.
Xu, H., Bazavan, E.G., Zanfir, A., Freeman, B., Sukthankar, R., Sminchisescu, C.: GHUM & GHUML: Generative 3D human shape and articulated pose models. CVPR (2020)
2020
Earlier work this paper cites.
Zhou, Y., Han, X., Shechtman, E., Echevarria, J., Kalogerakis, E., Li, D.: Makelttalk: speaker-aware talking-head animation. ACM Transactions On Graphics (TOG) 39
2020
Earlier work this paper cites.
Adam, M., Wessel, M., Benlian, A.: Ai-based chatbots in customer service and their effects on user compliance. Electronic Markets 31
2021
Earlier work this paper cites.
Dhariwal, P., Nichol, A.: Diffusion models beat gans on image synthesis. NeurIPS 34
2021
Earlier work this paper cites.
Doukas, M.C., Zafeiriou, S., Sharmanska, V.: Headgan: One-shot neural head synthesis and editing. In: ICCV. pp. 14398–14407 (October 2021)
2021
Earlier work this paper cites.
Guo, Y., Chen, K., Liang, S., Liu, Y.J., Bao, H., Zhang, J.: Ad-nerf: Audio driven neural radiance fields for talking head synthesis. In: ICCV. pp. 5784–5794 (2021)
2021
Earlier work this paper cites.
Lu, Y., Chai, J., Cao, X.: Live Speech Portraits: Real-time photorealistic talking-head animation. ACM Transactions on Graphics 40
2021
Cited alongside, same era.
Moussawi, S., Koufaris, M., Benbunan-Fich, R.: How perceptions of intelligence and anthropomorphism affect adoption of personal intelligent agents. Electronic Markets 31
2021
Cited alongside, same era.
Ren, Y., Li, G., Chen, Y., Li, T.H., Liu, S.: Pirenderer: Controllable portrait image generation via semantic neural rendering. In: ICCV. pp. 13759–13768 (2021)
2021
Cited alongside, same era.
Richard, A., Zollhöfer, M., Wen, Y., De la Torre, F., Sheikh, Y.: Meshtalk: 3d face animation from speech using cross-modality disentanglement. In: ICCV. pp. 1173–1182 (2021)
2021
Cited alongside, same era.
Roesler, E., Manzey, D., Onnasch, L.: A meta-analysis on the effectiveness of anthropomorphism in human-robot interaction. Science Robotics 6
Bard: Bard: A large language model from google ai (2023), https://blog.google/technology/ai/bard-google-ai-search-updates/
2023
Later among the works it cites.
2023
Later among the works it cites.
Bounareli, S., Tzelepis, C., Argyriou, V., Patras, I., Tzimiropoulos, G.: Hyperreenact: one-shot reenactment via jointly learning to refine and retarget faces. In: ICCV. pp. 7149–7159 (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
Seeger, A.M., Pfeiffer, J., Heinzl, A.: Texting with humanlike conversational agents: Designing for anthropomorphism. Journal of the Association for Information Systems 22
2021
Cited alongside, same era.
Suzhen, W., Lincheng, L., Yu, D., Changjie, F., Xin, Y.: Audio2head: Audio-driven one-shot talking-head generation with natural head motion. In: IJCAI (2021)
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Wang, T.C., Mallya, A., Liu, M.Y.: One-shot free-view neural talking-head synthesis for video conferencing. In: CVPR (2021)
2021
Cited alongside, same era.
Wang, T.C., Mallya, A., Liu, M.Y.: One-shot free-view neural talking-head synthesis for video conferencing. In: CVPR. pp. 10039–10049 (2021)
2021
Cited alongside, same era.
Yi, X., Zhou, Y., Xu, F.: Transpose: Real-time 3d human translation and pose estimation with six inertial sensors. ACM Transactions on Graphics (TOG) 40
2021
Cited alongside, same era.
Zhang, Z., Li, L., Ding, Y., Fan, C.: Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset. In: CVPR. pp. 3661–3670 (2021)
2021
Cited alongside, same era.
Gao, Y., Zhou, Y., Wang, J., Li, X., Ming, X., Lu, Y.: High-fidelity and freely controllable talking head video generation. In: CVPR. pp. 5609–5619 (2023)
2023
Later among the works it cites.
Guan, J., Zhang, Z., Zhou, H., Hu, T., Wang, K., He, D., Feng, H., Liu, J., Ding, E., Liu, Z., et al.: Stylesync: High-fidelity generalized and personalized lip sync in style-based generator. In: CVPR. pp. 1505–1515 (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Liu, Y., Lin, L., Yu, F., Zhou, C., Li, Y.: Moda: Mapping-once audio-driven portrait animation with dual attentions. In: ICCV (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
Manyika, J.: An overview of bard: an early experiment with generative ai. AI. Google Static Documents (2023)
2023
Later among the works it cites.
OpenAI: Gpt-4 technical report (2023)
2023
Later among the works it cites.
Pizzi, G., Vannucci, V., Mazzoli, V., Donvito, R.: I, chatbot! the impact of anthropomorphism and gaze direction on willingness to disclose personal information and behavioral intentions. Psychology & Marketing 40
2023
Later among the works it cites.
Shen, K., Guo, C., Kaufmann, M., Zarate, J.J., Valentin, J., Song, J., Hilliges, O.: X-avatar: Expressive human avatars. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 16911–16921 (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Thambiraja, B., Habibie, I., Aliakbarian, S., Cosker, D., Theobalt, C., Thies, J.: Imitator: Personalized speech-driven 3d facial animation. In: ICCV. pp. 20621–20631 (2023)
2023
Later among the works it cites.
Wang, J., Qian, X., Zhang, M., Tan, R.T., Li, H.: Seeing what you said: Talking face generation guided by a lip reading expert. In: CVPR. pp. 14653–14662 (2023)
2023
Later among the works it cites.
Wang, T., Li, L., Lin, K., Lin, C.C., Yang, Z., Zhang, H., Liu, Z., Wang, L.: Disco: Disentangled control for referring human dance generation in real world. arXiv e-prints pp. arXiv–2307 (2023)
2023
Later among the works it cites.
Wu, X., Hu, P., Wu, Y., Lyu, X., Cao, Y.P., Shan, Y., Yang, W., Sun, Z., Qi, X.: Speech2lip: High-fidelity speech to lip generation by learning from a short video. In: ICCV. pp. 22168–22177 (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
Xing, J., Xia, M., Zhang, Y., Cun, X., Wang, J., Wong, T.T.: Codetalker: Speech-driven 3d facial animation with discrete motion prior. In: CVPR. pp. 12780–12790 (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Yi, H., Liang, H., Liu, Y., Cao, Q., Wen, Y., Bolkart, T., Tao, D., Black, M.J.: Generating holistic 3d human motion from speech. In: CVPR. pp. 469–480 (June 2023)
2023
Later among the works it cites.
Yu, Z., Yin, Z., Zhou, D., Wang, D., Wong, F., Wang, B.: Talking head generation with probabilistic audio-to-visual diffusion priors. In: ICCV. pp. 7645–7655 (2023)
2023
Later among the works it cites.
Zhang, L., Rao, A., Agrawala, M.: Adding conditional control to text-to-image diffusion models (2023)
2023
Later among the works it cites.
Zhang, W., Cun, X., Wang, X., Zhang, Y., Shen, X., Guo, Y., Shan, Y., Wang, F.: Sadtalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation. In: CVPR. pp. 8652–8661 (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
Zhou, Q., Li, B., Han, L., Jou, M.: Talking to a bot or a wall? how chatbots vs. human agents affect anticipated communication quality. Computers in Human Behavior 143
2023
Later among the works it cites.
2024
Closest in time.
Bazavan, E.G., Zanfir, A., Szente, T.A., Zanfir, M., Alldieck, T., Sminchisescu, C.: Sphear: Spherical head registration for complete statistical 3d modeling. 3DV (2024)
2024
Closest in time.
Sora. https://openai.com/sora (2024)
2024
Closest in time.