Fetching the paper…
Reading the bibliography…
We address the problem of generating diverse 3D human motions from textual descriptions.
Hill, I.: Natural language versus computer language. In: Designing for Human-Computer Communication (1983)
1983
Earlier work this paper cites.
Ionescu, C., Li, F., Sminchisescu, C.: Latent structured models for human pose estimation. In: International Conference on Computer Vision (ICCV) (2011)
2011
Earlier work this paper cites.
Mikolov, T., Sutskever, I., Chen, K., Corrado, G.S., Dean, J.: Distributed representations of words and phrases and their compositionality. In: Neural Information Processing Systems (NeurIPS) (2013)
2013
Earlier work this paper cites.
Ionescu, C., Papava, D., Olaru, V., Sminchisescu, C.: Human3.6M: Large scale datasets and predictive methods for 3D human sensing in natural environments. Transactions on Pattern Analysis and Machine Intelligence (TPAMI) (2014)
2014
Earlier work this paper cites.
Kingma, D.P., Welling, M.: Auto-encoding variational bayes. In: International Conference on Learning Representations (ICLR) (2014)
2014
Earlier work this paper cites.
Terlemez, O., Ulbrich, S., Mandery, C., Do, M., Vahrenkamp, N., Asfour, T.: Master motor map (MMM) — framework and toolkit for capturing, representing, and reproducing human motion on humanoid robots. In: International Conference on Humanoid Robots (2014)
2014
Earlier work this paper cites.
Gao, T., Dontcheva, M., Adar, E., Liu, Z., Karahalios, K.G.: DataTone: Managing ambiguity in natural language interfaces for data visualization. In: ACM Symposium on User Interface Software & Technology (2015)
2015
Earlier work this paper cites.
Kingma, D.P., Ba, J.: Adam: A method for stochastic optimization. In: International Conference on Learning Representations (ICLR) (2015)
2015
Earlier work this paper cites.
Loper, M., Mahmood, N., Romero, J., Pons-Moll, G., Black, M.J.: SMPL: A skinned multi-person linear model. ACM Transactions on Graphics (TOG) (2015)
2015
Earlier work this paper cites.
Mandery, C., Terlemez, O., Do, M., Vahrenkamp, N., Asfour, T.: The kit whole-body human motion database. In: International Conference on Advanced Robotics (ICAR) (2015)
2015
Earlier work this paper cites.
Holden, D., Saito, J., Komura, T.: A deep learning framework for character motion synthesis and editing. ACM Transactions on Graphics (TOG) (2016)
2016
Earlier work this paper cites.
Plappert, M., Mandery, C., Asfour, T.: The KIT motion-language dataset. Big Data (2016)
2016
Earlier work this paper cites.
Xu, J., Mei, T., Yao, T., Rui, Y.: MSR-VTT: A large video description dataset for bridging video and language. In: Computer Vision and Pattern Recognition (CVPR) (2016)
2016
Earlier work this paper cites.
Habibie, I., Holden, D., Schwarz, J., Yearsley, J., Komura, T.: A recurrent variational autoencoder for human motion synthesis. In: British Machine Vision Conference (BMVC) (2017)
2017
Earlier work this paper cites.
Karras, T., Aila, T., Laine, S., Herva, A., Lehtinen, J.: Audio-driven facial animation by joint end-to-end learning of pose and emotion. ACM Transactions on Graphics (TOG) (2017)
2017
Earlier work this paper cites.
Martinez, J., Black, M.J., Romero, J.: On human motion prediction using recurrent neural networks. In: Computer Vision and Pattern Recognition (CVPR) (2017)
2017
Earlier work this paper cites.
Romero, J., Tzionas, D., Black, M.J.: Embodied hands: Modeling and capturing hands and bodies together. ACM Transactions on Graphics (TOG) (2017)
2017
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., Polosukhin, I.: Attention is all you need. In: Neural Information Processing Systems (NeurIPS) (2017)
2017
Earlier work this paper cites.
Ahn, H., Ha, T., Choi, Y., Yoo, H., Oh, S.: Text2Action: Generative adversarial synthesis from language to action. In: International Conference on Robotics and Automation (ICRA) (2018)
2018
Earlier work this paper cites.
Barsoum, E., Kender, J., Liu, Z.: HP-GAN: Probabilistic 3D human motion prediction via GAN. In: Computer Vision and Pattern Recognition Workshops (CVPRW) (2018)
2018
Earlier work this paper cites.
Lin, A.S., Wu, L., Corona, R., Tai, K., Huang, Q., Mooney, R.J.: Generating animated videos of human activities from natural language descriptions. Visually Grounded Interaction and Language (ViGIL) NeurIPS Workshop (2018)
2018
Earlier work this paper cites.
Lin, X., Amer, M.: Human motion modeling using DVGANs. arXiv preprint arXiv:1804.10652 (2018)
2018
Cited alongside, same era.
Pavllo, D., Grangier, D., Auli, M.: QuaterNet: A quaternion-based recurrent model for human motion. In: British Machine Vision Conference (BMVC) (2018)
2018
Cited alongside, same era.
Plappert, M., Mandery, C., Asfour, T.: Learning a bidirectional mapping between human whole-body motion and natural language using deep recurrent neural networks. Robotics Auton. Syst. (2018)
2018
Cited alongside, same era.
Yamada, T., Matsunaga, H., Ogata, T.: Paired recurrent autoencoders for bidirectional translation between robot actions and linguistic descriptions. Robotics and Automation Letters (2018)
2018
Cited alongside, same era.
Ahuja, C., Morency, L.P.: Language2Pose: Natural language grounded pose forecasting. In: International Conference on 3D Vision (3DV) (2019)
Henter, G.E., Alexanderson, S., Beskow, J.: MoGlow: Probabilistic and controllable motion synthesis using normalising flows. ACM Transactions on Graphics (TOG) (2020)
2020
Later among the works it cites.
2020
Later among the works it cites.
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Scao, T.L., Gugger, S., Drame, M., Lhoest, Q., Rush, A.M.: Transformers: State-of-the-art natural language processing. In: Empirical Methods in Natural Language Processing: System Demonstrations (2020)
2020
Later among the works it cites.
Yuan, Y., Kitani, K.: Dlow: Diversifying latent flows for diverse human motion prediction. In: European Conference on Computer Vision (ECCV) (2020)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
Aksan, E., Kaufmann, M., Hilliges, O.: Structured prediction helps 3D human motion modelling. In: International Conference on Computer Vision (ICCV) (2019)
2019
Cited alongside, same era.
Cudeiro, D., Bolkart, T., Laidlaw, C., Ranjan, A., Black, M.: Capture, learning, and synthesis of 3D speaking styles. In: Computer Vision and Pattern Recognition (CVPR) (2019)
2019
Cited alongside, same era.
Devlin, J., Chang, M.W., Lee, K., Toutanova, K.: BERT: Pre-training of deep bidirectional transformers for language understanding. In: North American Chapter of the Association for Computational Linguistics (NAACL) (2019)
2019
Cited alongside, same era.
Falcon et al., W.: Pytorch lightning. GitHub. Note: https://github.com/PyTorchLightning/pytorch-lightning (2019)
2019
Cited alongside, same era.
Ginosar, S., Bar, A., Kohavi, G., Chan, C., Owens, A., Malik, J.: Learning individual styles of conversational gesture. In: Computer Vision and Pattern Recognition (CVPR) (2019)
2019
Cited alongside, same era.
Lee, H.Y., Yang, X., Liu, M.Y., Wang, T.C., Lu, Y.D., Yang, M.H., Kautz, J.: Dancing to music. In: Neural Information Processing Systems (NeurIPS) (2019)
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2020
Later among the works it cites.
Zanfir, A., Bazavan, E.G., Xu, H., Freeman, W.T., Sukthankar, R., Sminchisescu, C.: Weakly supervised 3D human pose and shape reconstruction with normalizing flows. In: European Conference on Computer Vision (ECCV) (2020)
2020
Later among the works it cites.
2020
Later among the works it cites.
Zhao, R., Su, H., Ji, Q.: Bayesian adversarial human motion synthesis. In: Computer Vision and Pattern Recognition (CVPR) (2020)
2020
Later among the works it cites.
Bain, M., Nagrani, A., Varol, G., Zisserman, A.: Frozen in time: A joint video and image encoder for end-to-end retrieval. In: International Conference on Computer Vision (ICCV) (2021)
2021
Later among the works it cites.
Bain, M., Nagrani, A., Varol, G., Zisserman, A.: Frozen in time: A joint video and image encoder for end-to-end retrieval. In: International Conference on Computer Vision (ICCV) (2021)
2021
Later among the works it cites.
Bhattacharya, U., Childs, E., Rewkowski, N., Manocha, D.: Speech2AffectiveGestures: Synthesizing Co-Speech Gestures with Generative Adversarial Affective Expression Learning (2021)
2021
Later among the works it cites.
2021
Later among the works it cites.
Ghosh, A., Cheema, N., Oguz, C., Theobalt, C., Slusallek, P.: Synthesis of compositional animations from textual descriptions. In: International Conference on Computer Vision (ICCV) (2021)
2021
Later among the works it cites.
Li, R., Yang, S., Ross, D.A., Kanazawa, A.: AI choreographer: Music conditioned 3D dance generation with AIST++. In: International Conference on Computer Vision (ICCV) (2021)
2021
Later among the works it cites.
Petrovich, M., Black, M.J., Varol, G.: Action-conditioned 3D human motion synthesis with transformer VAE. In: International Conference on Computer Vision (ICCV) (2021)
2021
Later among the works it cites.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., Sutskever, I.: Learning transferable visual models from natural language supervision. In: International Conference on Machine Learning (ICML) (2021)
2021
Later among the works it cites.
Richard, A., Zollhöfer, M., Wen, Y., de la Torre, F., Sheikh, Y.: Meshtalk: 3D face animation from speech using cross-modality disentanglement. In: International Conference on Computer Vision (ICCV) (2021)
2021
Later among the works it cites.
Saunders, B., Camgoz, N.C., Bowden, R.: Mixed SIGNals: Sign language production via a mixture of motion primitives. In: International Conference on Computer Vision (ICCV) (2021)
2021
Later among the works it cites.
2021
Later among the works it cites.
Zhang, Y., Black, M.J., Tang, S.: We are more than our joints: Predicting how 3D bodies move. In: Computer Vision and Pattern Recognition (CVPR) (2021)
2021
Later among the works it cites.
Fan, Y., Lin, Z., Saito, J., Wang, W., Komura, T.: Faceformer: Speech-driven 3D facial animation with transformers. In: Computer Vision and Pattern Recognition (CVPR) (2022)
2022
Closest in time.
Yang, J., Li, C., Zhang, P., Xiao, B., Liu, C., Yuan, L., Gao, J.: Unified contrastive learning in image-text-label space. In: Computer Vision and Pattern Recognition (CVPR) (2022)
2022
Closest in time.