Fetching the paper…
Reading the bibliography…
We introduce a new conversation head generation benchmark for synthesizing behaviors of a single interlocutor in a face-to-face conversation.
A. Kendon, “Movement coordination in social interaction: Some examples described,” Acta psychologica , vol. 32, pp. 101–125, 1970
1970
Earlier work this paper cites.
V. Blanz and T. Vetter, “A morphable model for the synthesis of 3d faces,” in Proceedings of the 26th annual conference on Computer graphics and interactive techniques , 1999, pp. 187–194
1999
Earlier work this paper cites.
J. Parker and E. Coiera, “Improving clinical communication: a view from psychology,” Journal of the American Medical Informatics Association , vol. 7, no. 5, pp. 453–461, 2000
2000
Earlier work this paper cites.
J. M. Honeycutt and S. G. Ford, “Mental imagery and intrapersonal communication: A review of research on imagined interactions (iis) and current developments,” Annals of the International Communication Association , vol. 25, no. 1, pp. 315–345, 2001
2001
Earlier work this paper cites.
C. R. Berger, “Interpersonal communication: Theoretical perspectives, future prospects.” Journal of communication , 2005
2005
Earlier work this paper cites.
K. Robertson, “Active listening: more than just paying attention,” Australian family physician , vol. 34, no. 12, 2005
2005
Earlier work this paper cites.
R. Maatman, J. Gratch, and S. Marsella, “Natural behavior of a listening agent,” in International workshop on intelligent virtual agents . Springer, 2005, pp. 25–36
2005
Earlier work this paper cites.
D. Heylen, E. Bevacqua, M. Tellier, and C. Pelachaud, “Searching for prototypical facial feedback signals,” in International Workshop on Intelligent Virtual Agents . Springer, 2007, pp. 147–153
2007
Earlier work this paper cites.
D. McNaughton, D. Hamlin, J. McCarthy, D. Head-Reeves, and M. Schreiner, “Learning to listen: Teaching an active listening strategy to preservice education professionals,” Topics in Early Childhood Special Education , vol. 27, no. 4, pp. 223–231, 2008
2008
Earlier work this paper cites.
M. Gillies, X. Pan, M. Slater, and J. Shawe-Taylor, “Responsive listening behavior,” Computer animation and virtual worlds , vol. 19, no. 5, pp. 579–589, 2008
2008
Earlier work this paper cites.
M. Tomasello, Origins of human communication . MIT press, 2010
2010
Earlier work this paper cites.
G. J. Stephens, L. J. Silbert, and U. Hasson, “Speaker–listener neural coupling underlies successful communication,” Proceedings of the National Academy of Sciences , vol. 107, no. 32, pp. 14 425–14 430, 2010
2010
Earlier work this paper cites.
G. McKeown, M. Valstar, R. Cowie, M. Pantic, and M. Schroder, “The semaine database: Annotated multimodal records of emotionally colored conversations between a person and a limited agent,” IEEE transactions on affective computing , vol. 3, no. 1, pp. 5–17, 2011
2011
Earlier work this paper cites.
D. Heylen, E. Bevacqua, C. Pelachaud, I. Poggi, J. Gratch, and M. Schröder, “Generating listening behaviour,” in Emotion-oriented systems . Springer, 2011, pp. 321–347
2011
Earlier work this paper cites.
H. Bunt, J. Alexandersson, J.-W. Choe, A. C. Fang, K. Hasida, V. Petukhova, A. Popescu-Belis, and D. Traum, “Iso 24617-2: A semantically-based standard for dialogue annotation,” UNIVERSITY OF SOUTHERN CALIFORNIA LOS ANGELES, Tech. Rep., 2012
2012
Earlier work this paper cites.
M. Rost and J. Wilson, Active listening . Routledge, 2013
2013
Earlier work this paper cites.
S. Petridis, B. Martinez, and M. Pantic, “The mahnob laughter database,” Image and Vision Computing , vol. 31, no. 2, pp. 186–202, 2013
2013
Cited alongside, same era.
D. W. Stacks and M. B. Salwen, An integrated approach to communication theory and research . Routledge, 2014
2014
Cited alongside, same era.
H. Buschmeier, Z. Malisz, J. Skubisz, M. Wlodarczak, I. Wachsmuth, S. Kopp, and P. Wagner, “Alico: A multimodal corpus for the study of active listening,” in LREC 2014, Ninth International Conference on Language Resources and Evaluation, 26-31 May, Reykjavik, Iceland , 2014, pp. 3638–3643
2014
Cited alongside, same era.
C. Dong, C. C. Loy, K. He, and X. Tang, “Image super-resolution using deep convolutional networks,” IEEE transactions on pattern analysis and machine intelligence , vol. 38, no. 2, pp. 295–307, 2015
2015
Cited alongside, same era.
W. Bao, W.-S. Lai, X. Zhang, Z. Gao, and M.-H. Yang, “Memc-net: Motion estimation and motion compensation driven neural network for video interpolation and enhancement,” IEEE transactions on pattern analysis and machine intelligence , vol. 43, no. 3, pp. 933–948, 2019
2019
Later among the works it cites.
R. Szeto, X. Sun, K. Lu, and J. J. Corso, “A temporally-aware interpolation network for video frame inpainting,” IEEE transactions on pattern analysis and machine intelligence , vol. 42, no. 5, pp. 1053–1068, 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
G. Melis, T. Kočiskỳ, and P. Blunsom, “Mogrifier lstm,” arXiv preprint arXiv:1909.01792 , 2019
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. S. Chung and A. Zisserman, “Out of time: automated lip sync in the wild,” in Asian conference on computer vision . Springer, 2016, pp. 251–263
2016
Cited alongside, same era.
J. S. Chung, A. Jamaludin, and A. Zisserman, “You said that?” arXiv preprint arXiv:1705.02966 , 2017
2017
Cited alongside, same era.
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, “Gans trained by a two time-scale update rule converge to a local nash equilibrium,” Advances in neural information processing systems , vol. 30, 2017
2017
Cited alongside, same era.
2018
Cited alongside, same era.
A. Bansal, S. Ma, D. Ramanan, and Y. Sheikh, “Recycle-gan: Unsupervised video retargeting,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 119–135
2018
Cited alongside, same era.
A. Tewari, M. Zollhoefer, F. Bernard, P. Garrido, H. Kim, P. Perez, and C. Theobalt, “High-fidelity monocular face reconstruction based on an unsupervised model-based face autoencoder,” IEEE transactions on pattern analysis and machine intelligence , vol. 42, no. 2, pp. 357–370, 2018
2018
Cited alongside, same era.
J. Booth, A. Roussos, E. Ververas, E. Antonakos, S. Ploumpis, Y. Panagakis, and S. Zafeiriou, “3d reconstruction of “in-the-wild” faces in images and videos,” IEEE transactions on pattern analysis and machine intelligence , vol. 40, no. 11, pp. 2638–2652, 2018
2018
Cited alongside, same era.
H. Kim, P. Garrido, A. Tewari, W. Xu, J. Thies, M. Niessner, P. Pérez, C. Richardt, M. Zollhöfer, and C. Theobalt, “Deep video portraits,” ACM Transactions on Graphics (TOG) , vol. 37, no. 4, pp. 1–14, 2018
2018
Cited alongside, same era.
K. Prajwal, R. Mukhopadhyay, V. P. Namboodiri, and C. Jawahar, “A lip sync expert is all you need for speech to lip generation in the wild,” in Proceedings of the 28th ACM International Conference on Multimedia , 2020, pp. 484–492
2020
Later among the works it cites.
K. Wang, Q. Wu, L. Song, Z. Yang, W. Wu, C. Qian, R. He, Y. Qiao, and C. C. Loy, “Mead: A large-scale audio-visual dataset for emotional talking-face generation,” in European Conference on Computer Vision . Springer, 2020, pp. 700–717
2020
Later among the works it cites.
P. Yi, Z. Wang, K. Jiang, J. Jiang, T. Lu, and J. Ma, “A progressive fusion generative adversarial network for realistic and consistent video super-resolution,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 44, no. 5, pp. 2264–2280, 2020
2020
Later among the works it cites.
H. Zhu, M.-D. Luo, R. Wang, A.-H. Zheng, and R. He, “Deep audio-visual learning: A survey,” International Journal of Automation and Computing , pp. 1–26, 2021
2021
Later among the works it cites.
C. Zhang, Y. Zhao, Y. Huang, M. Zeng, S. Ni, M. Budagavi, and X. Guo, “Facial: Synthesizing dynamic talking face with implicit attribute learning,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 3867–3876
2021
Later among the works it cites.
L. Yu, H. Xie, and Y. Zhang, “Multimodal learning for temporally coherent talking face generation with articulator synergy,” IEEE Transactions on Multimedia , vol. 24, pp. 2950–2962, 2021
2021
Later among the works it cites.
S. E. Eskimez, Y. Zhang, and Z. Duan, “Speech driven talking face generation from a single image and an emotion condition,” IEEE Transactions on Multimedia , vol. 24, pp. 3480–3490, 2021
2021
Later among the works it cites.
C. Oertel, P. Jonell, D. Kontogiorgos, K. F. Mora, J.-M. Odobez, and J. Gustafson, “Towards an engagement-aware attentive artificial listener for multi-party interactions,” Frontiers in Robotics and AI , p. 189, 2021
2021
Later among the works it cites.
Y. Ren, G. Li, Y. Chen, T. H. Li, and S. Liu, “Pirenderer: Controllable portrait image generation via semantic neural rendering,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 13 759–13 768
2021
Later among the works it cites.
H. Chen, Y. Jin, K. Xu, Y. Chen, and C. Zhu, “Multiframe-to-multiframe network for video denoising,” IEEE Transactions on Multimedia , vol. 24, pp. 2164–2178, 2021
2021
Later among the works it cites.
A. Richard, M. Zollhöfer, Y. Wen, F. De la Torre, and Y. Sheikh, “Meshtalk: 3d face animation from speech using cross-modality disentanglement,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 1173–1182
2021
Later among the works it cites.
M. C. Doukas, E. Ververas, V. Sharmanska, and S. Zafeiriou, “Free-headgan: Neural talking head synthesis with explicit gaze control,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2023
2023
Closest in time.