Fetching the paper…
Reading the bibliography…
Conversation is an essential component of virtual avatar activities in the metaverse.
H. H. Clark, Using language . Cambridge university press, 1996
1996
Earlier work this paper cites.
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image quality assessment: from error visibility to structural similarity,” IEEE Trans. Image Process. (TIP) , vol. 13, no. 4, pp. 600–612, 2004
2004
Earlier work this paper cites.
J.-F. Bonastre, F. Wils, and S. Meignier, “Alize, a free toolkit for speaker recognition,” in IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , 2005, pp. I–737
2005
Earlier work this paper cites.
N. Narvekar and L. Karam, “A no-reference perceptual image sharpness metric based on a cumulative probability of blur detection,” in International Workshop on Quality of Multimedia Experience , 2009, pp. 87–91
2009
Earlier work this paper cites.
A. Hore and D. Ziou, “Image quality metrics: Psnr vs. ssim,” in International Conference on Pattern Recognition (ICPR) , 2010, pp. 2366–2369
2010
Earlier work this paper cites.
C. Danescu-Niculescu-Mizil and L. Lee, “Chameleons in imagined conversations: A new approach to understanding coordination of linguistic style in dialogs,” in Workshop on Cognitive Modeling and Computational Linguistics , 2011, pp. 76–87
2011
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz et al. , “The kaldi speech recognition toolkit,” in IEEE 2011 workshop on automatic speech recognition and understanding (ASRU) , no. CONF, 2011
2011
Earlier work this paper cites.
J. Xu, L. Xing, A. Perkis, and Y. Jiang, “On the properties of mean opinion scores for quality of experience management,” in IEEE International Symposium on Multimedia , 2011, pp. 500–505
2011
Earlier work this paper cites.
X. Anguera, S. Bozonnet, N. Evans, C. Fredouille, G. Friedland, and O. Vinyals, “Speaker diarization: A review of recent research,” IEEE Transactions on audio, speech, and language processing , vol. 20, no. 2, pp. 356–370, 2012
2012
Earlier work this paper cites.
R. Lowe, N. Pow, I. Serban, and J. Pineau, “The Ubuntu dialogue corpus: A large dataset for research in unstructured multi-turn dialogue systems,” in Annual Meeting of the Special Interest Group on Discourse and Dialogue , 2015, pp. 285–294
2015
Earlier work this paper cites.
T. Giannakopoulos, “pyaudioanalysis: An open-source python library for audio signal analysis,” PloS one , vol. 10, no. 12, 2015
2015
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in International Conference on Learning and Representation (ICLR) , 2015
2015
Earlier work this paper cites.
A. Kumar, O. Irsoy, P. Ondruska, M. Iyyer, J. Bradbury, I. Gulrajani, V. Zhong, R. Paulus, and R. Socher, “Ask me anything: Dynamic memory networks for natural language processing,” in International conference on machine learning (ICML) , 2016, pp. 1378–1387
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
P. Lison and J. Tiedemann, “OpenSubtitles2016: Extracting large parallel corpora from movie and TV subtitles,” in LREC , 2016, pp. 923–929
2016
Earlier work this paper cites.
D. Amodei, S. Ananthanarayanan, R. Anubhai, J. Bai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, Q. Cheng, G. Chen et al. , “Deep speech 2: End-to-end speech recognition in english and mandarin,” in International Conference on Machine Learning (ICML) , 2016, pp. 173–182
2016
Earlier work this paper cites.
J. S. Chung and A. Zisserman, “Out of time: automated lip sync in the wild,” in Workshop on Multi-view Lip-reading, ACCV , 2016
2016
Earlier work this paper cites.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, Q. Le, Y. Agiomyrgiannakis, R. Clark, and R. A. Saurous, “Tacotron: Towards End-to-End Speech Synthesis,” in Proc. Interspeech 2017 , 2017, pp. 4006–4010
2017
Earlier work this paper cites.
L. Hu, S. Saito, L. Wei, K. Nagano, J. Seo, J. Fursund, I. Sadeghi, C. Sun, Y.-C. Chen, and H. Li, “Avatar digitization from a single image for real-time rendering,” ACM Trans. Graph. (TOG) , vol. 36, no. 6, 2017
2017
Earlier work this paper cites.
H. X. Pham, S. Cheung, and V. Pavlovic, “Speech-driven 3d facial animation with implicit emotional awareness: A deep learning approach,” in CVPRW , 2017, pp. 2328–2336
2017
Earlier work this paper cites.
S. Suwajanakorn, S. M. Seitz, and I. Kemelmacher-Shlizerman, “Synthesizing obama: learning lip sync from audio,” ACM Transactions on Graphics (TOG) , vol. 36, no. 4, pp. 1–13, 2017
2017
Cited alongside, same era.
T. Karras, T. Aila, S. Laine, A. Herva, and J. Lehtinen, “Audio-driven facial animation by joint end-to-end learning of pose and emotion,” ACM Transactions on Graphics (TOG) , vol. 36, no. 4, pp. 1–12, 2017
2017
Cited alongside, same era.
L. Luo, J. Xu, J. Lin, Q. Zeng, and X. Sun, “An auto-encoder matching model for learning utterance-level semantic dependency in dialogue generation,” in Conference on Empirical Methods in Natural Language Processing , 2018, pp. 702–707
2018
Cited alongside, same era.
Y. Wang, C. Liu, M. Huang, and L. Nie, “Learning to ask questions in open-domain conversational systems with typed decoders,” in Annual Meeting of the Association for Computational Linguistics , 2018, pp. 2193–2203
2018
Cited alongside, same era.
H. Bredin, R. Yin, J. M. Coria, G. Gelly, P. Korshunov, M. Lavechin, D. Fustes, H. Titeux, W. Bouaziz, and M.-P. Gill, “Pyannote. audio: neural building blocks for speaker diarization,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020, pp. 7124–7128
2020
Later among the works it cites.
K. Schwarz, Y. Liao, M. Niemeyer, and A. Geiger, “Graf: Generative radiance fields for 3d-aware image synthesis,” in Advances in Neural Information Processing Systems (NeurIPS) , 2020, pp. 20 154–20 166
2020
Later among the works it cites.
K. R. Prajwal, R. Mukhopadhyay, V. P. Namboodiri, and C. Jawahar, “A lip sync expert is all you need for speech to lip generation in the wild,” in ACM International Conference on Multimedia (ACM MM) , 2020, p. 484–492
2020
Later among the works it cites.
K. Vougioukas, S. Petridis, and M. Pantic, “Realistic speech-driven facial animation with gans,” International Journal of Computer Vision (IJCV) , vol. 128, no. 5, pp. 1398–1413, 2020
2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P. Hömke, J. Holler, and S. C. Levinson, “Eye blinks are perceived as communicative signals in human face-to-face interaction,” PloS one , vol. 13, no. 12, p. e0208030, 2018
2018
Cited alongside, same era.
L. Jiang, J. Zhang, B. Deng, H. Li, and L. Liu, “3d face reconstruction with geometry details from a single image,” IEEE Trans. Image Process. (TIP) , vol. 27, no. 10, pp. 4756–4770, 2018
2018
Cited alongside, same era.
P.-A. Broux, F. Desnous, A. Larcher, S. Petitrenaud, J. Carrive, and S. Meignier, “S4d: Speaker diarization toolkit in python,” in Interspeech , 2018
2018
Cited alongside, same era.
L. Chen, Z. Li, R. K. Maddox, Z. Duan, and C. Xu, “Lip movements generation at a glance,” in European Conference on Computer Vision (ECCV) , 2018, pp. 520–535
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
D. Cudeiro, T. Bolkart, C. Laidlaw, A. Ranjan, and M. J. Black, “Capture, learning, and synthesis of 3d speaking styles,” in 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2019, pp. 10 093–10 103
2019
Cited alongside, same era.
L. Chen, R. K. Maddox, Z. Duan, and C. Xu, “Hierarchical cross-modal talking face generation with dynamic pixel-wise loss,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2019, pp. 7824–7833
2019
Cited alongside, same era.
Later among the works it cites.
Y. Guo, K. Chen, S. Liang, Y. Liu, H. Bao, and J. Zhang, “Ad-nerf: Audio driven neural radiance fields for talking head synthesis,” in IEEE/CVF International Conference on Computer Vision (ICCV) , 2021
2021
Later among the works it cites.
G. Gafni, J. Thies, M. Zollhöfer, and M. Nießner, “Dynamic neural radiance fields for monocular 4d facial avatar reconstruction,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2021, pp. 8649–8658
2021
Later among the works it cites.
Y. Guo, L. Cai, and J. Zhang, “3d face from X: learning face shape from diverse sources,” IEEE Trans. Image Process. (TIP) , vol. 30, pp. 3815–3827, 2021
2021
Later among the works it cites.
A. Richard, M. Zollhöfer, Y. Wen, F. D. la Torre, and Y. Sheikh, “Meshtalk: 3d face animation from speech using cross-modality disentanglement,” in IEEE/CVF International Conference on Computer Vision (ICCV) , 2021, pp. 1153–1162
2021
Later among the works it cites.
Z. Zhang, L. Li, Y. Ding, and C. Fan, “Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2021, pp. 3661–3670
2021
Later among the works it cites.
M. Niemeyer and A. Geiger, “Giraffe: Representing scenes as compositional generative neural feature fields,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2021, pp. 11 453–11 464
2021
Later among the works it cites.
S. Peng, Y. Zhang, Y. Xu, Q. Wang, Q. Shuai, H. Bao, and X. Zhou, “Neural body: Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2021, pp. 9054–9063
2021
Later among the works it cites.
2021
Later among the works it cites.
K. Park, U. Sinha, J. T. Barron, S. Bouaziz, D. B. Goldman, S. M. Seitz, and R. Martin-Brualla, “Nerfies: Deformable neural radiance fields,” in IEEE/CVF International Conference on Computer Vision (ICCV) , 2021, pp. 5865–5874
2021
Later among the works it cites.
Y. Lu, J. Chai, and X. Cao, “Live Speech Portraits: Real-time photorealistic talking-head animation,” ACM Transactions on Graphics (TOG) , vol. 40, no. 6, 2021
2021
Later among the works it cites.
H. Zhou, Y. Sun, W. Wu, C. C. Loy, X. Wang, and Z. Liu, “Pose-controllable talking face generation by implicitly modularized audio-visual representation,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2022
Closest in time.
W. Yang, Z. Chen, C. Chen, G. Chen, and K. K. Wong, “Deep face video inpainting via UV mapping,” IEEE Trans. Image Process. (TIP) , vol. 32, pp. 1145–1157, 2023
2023
Closest in time.