Fetching the paper…
Reading the bibliography…
We propose CLIP-Actor, a text-driven motion recommendation and neural mesh stylization system for human mesh animation.
Blanz, V., Vetter, T.: A morphable model for the synthesis of 3d faces. In: SIGGRAPH (1999)
1999
Earlier work this paper cites.
Hanser, E., Kevitt, P.M., Lunney, T.F., Condell, J.: Scenemaker: Intelligent multimodal visualisation of natural language scripts. In: Proceedings of the 20th Irish conference on Artificial intelligence and cognitive science (2009)
2009
Earlier work this paper cites.
Agirre, E., Diab, M., Cer, D., Gonzalez-Agirre, A.: Semeval-2012 task 6: A pilot on semantic textual similarity. In: Proceedings of the First Joint Conference on Lexical and Computational Semantics - Volume 1: Proceedings of the Main Conference and the Shared Task, and Volume 2: Proceedings of the Sixth International Workshop on Semantic Evaluation. p. 385–393 (2012)
2012
Earlier work this paper cites.
Hodosh, M., Young, P., Hockenmaier, J.: Framing image description as a ranking task: Data, models and evaluation metrics. Journal of Artificial Intelligence Research 47
2013
Earlier work this paper cites.
Marelli, M., Menini, S., Baroni, M., Bentivogli, L., Bernardi, R., Zamparelli, R.: A sick cure for the evaluation of compositional distributional semantic models. In: International Conference on Language Resources and Evaluation (LREC) (2014)
2014
Earlier work this paper cites.
Fragkiadaki, K., Levine, S., Felsen, P., Malik, J.: Recurrent network models for human dynamics. IEEE International Conference on Computer Vision (ICCV) (2015)
2015
Earlier work this paper cites.
Loper, M., Mahmood, N., Romero, J., Pons-Moll, G., Black, M.J.: SMPL: A skinned multi-person linear model. ACM Transactions on Graphics (SIGGRAPH Asia) 34
2015
Earlier work this paper cites.
Romero, J., Tzionas, D., Black, M.J.: Embodied hands: Modeling and capturing hands and bodies together. ACM Transactions on Graphics (SIGGRAPH Asia) 36
2017
Earlier work this paper cites.
Zuffi, S., Kanazawa, A., Jacobs, D., Black, M.J.: 3D menagerie: Modeling the 3D shape and pose of animals. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2017)
2017
Earlier work this paper cites.
Ahn, H., Ha, T., Choi, Y., Yoo, H., Oh, S.: Text2action: Generative adversarial synthesis from language to action. In: IEEE International Conference on Robotics and Automation (ICRA) (2018)
2018
Earlier work this paper cites.
Biggs, B., Roddick, T., Fitzgibbon, A., Cipolla, R.: Creatures great and SMAL: Recovering the shape and motion of animals from video. In: Asia Conference on Computer Vision (ACCV) (2018)
2018
Earlier work this paper cites.
Lin, A.S., Wu, L., Corona, R., Tai, K.W.H., Huang, Q., Mooney, R.J.: Generating animated videos of human activities from natural language descriptions. In: Proceedings of the Visually Grounded Interaction and Language Workshop at NeurIPS 2018 (2018)
2018
Earlier work this paper cites.
Plappert, M., Mandery, C., Asfour, T.: Learning a bidirectional mapping between human whole-body motion and natural language using deep recurrent neural networks. Robotics and Autonomous Systems 109
2018
Earlier work this paper cites.
Ahuja, C., Morency, L.P.: Language2pose: Natural language grounded pose forecasting. International Conference on 3D Vision (3DV) (2019)
2019
Earlier work this paper cites.
Bhatnagar, B.L., Tiwari, G., Theobalt, C., Pons-Moll, G.: Multi-garment net: Learning to dress 3d people from images. In: IEEE International Conference on Computer Vision (ICCV) (2019)
2019
Earlier work this paper cites.
Mahmood, N., Ghorbani, N., Troje, N.F., Pons-Moll, G., Black, M.J.: AMASS: Archive of motion capture as surface shapes. In: IEEE International Conference on Computer Vision (ICCV) (2019)
2019
Earlier work this paper cites.
Pavlakos, G., Choutas, V., Ghorbani, N., Bolkart, T., Osman, A.A.A., Tzionas, D., Black, M.J.: Expressive body capture: 3D hands, face, and body from a single image. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2019)
2019
Earlier work this paper cites.
Saito, S., Huang, Z., Natsume, R., Morishima, S., Kanazawa, A., Li, H.: Pifu: Pixel-aligned implicit function for high-resolution clothed human digitization. In: IEEE International Conference on Computer Vision (ICCV) (2019)
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Tsang, C.F., Shugrina, M., Lafleche, J.F., Takikawa, T., Wang, J., Loop, C., Chen, W., Jatavallabhula, K.M., Smith, E., Rozantsev, A., Perel, O., Shen, F., Gao, J., Fidler, S., State, G., Gorski, J., Xiang, T., Li, J., Li, M., Lebaredian, R.: Kaolin (2019)
2019
Earlier work this paper cites.
Yoon, Y., Ko, W.R., Jang, M., Lee, J., Kim, J., Lee, G.: Robots learn social skills: End-to-end learning of co-speech gesture generation for humanoid robots. In: IEEE International Conference on Robotics and Automation (ICRA) (2019)
2019
Earlier work this paper cites.
Zuffi, S., Kanazawa, A., Berger-Wolf, T., Black, M.J.: Three-d safari: Learning to estimate zebra pose, shape, and texture from images “in the wild”. In: IEEE International Conference on Computer Vision (ICCV) (2019)
2019
Earlier work this paper cites.
Bhatnagar, B.L., Sminchisescu, C., Theobalt, C., Pons-Moll, G.: Combining implicit function learning and parametric models for 3d human reconstruction. In: European Conference on Computer Vision (ECCV) (2020)
2020
Earlier work this paper cites.
Biggs, B., Boyne, O., Charles, J., Fitzgibbon, A., Cipolla, R.: Who left the dogs out?: 3D animal reconstruction with expectation maximization in the loop. In: European Conference on Computer Vision (ECCV) (2020)
2020
Cited alongside, same era.
Bozic, A., Palafox, P., Zollöfer, M., Dai, A., Thies, J., Nießner, M.: Neural non-rigid tracking. In: Advances in Neural Information Processing Systems (NeurIPS) (2020)
2020
Cited alongside, same era.
Guo, C., Zuo, X., Wang, S., Zou, S., Sun, Q., Deng, A., Gong, M., Cheng, L.: Action2motion: Conditioned generation of 3d human motions. In: ACM International Conference on Multimedia (MM) (2020)
2020
Cited alongside, same era.
Jiang, J., Xia, G.G., Carlton, D.B., Anderson, C.N., Miyakawa, R.H.: Transformer vae: A hierarchical model for structure-aware and interpretable music representation learning. In: IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) (2020)
2020
Cited alongside, same era.
Palafox, P., Bozic, A., Thies, J., Nießner, M., Dai, A.: Neural parametric models for 3d deformable shapes. In: IEEE International Conference on Computer Vision (ICCV) (2021)
2021
Later among the works it cites.
Patashnik, O., Wu, Z., Shechtman, E., Cohen-Or, D., Lischinski, D.: Styleclip: Text-driven manipulation of stylegan imagery. In: IEEE International Conference on Computer Vision (ICCV) (2021)
2021
Later among the works it cites.
Petrovich, M., Black, M.J., Varol, G.: Action-conditioned 3D human motion synthesis with transformer VAE. In: IEEE International Conference on Computer Vision (ICCV) (2021)
2021
Later among the works it cites.
Punnakkal, A.R., Chandrasekaran, A., Athanasiou, N., Quiros-Ramirez, A., Black, M.J.: BABEL: Bodies, action and behavior with english labels. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2021)
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
Ma, Q., Yang, J., Ranjan, A., Pujades, S., Pons-Moll, G., Tang, S., Black, M.J.: Learning to dress 3d people in generative clothing. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2020)
2020
Cited alongside, same era.
Mildenhall, B., Srinivasan, P.P., Tancik, M., Barron, J.T., Ramamoorthi, R., Ng, R.: Nerf: Representing scenes as neural radiance fields for view synthesis. In: European Conference on Computer Vision (ECCV) (2020)
2020
Cited alongside, same era.
Niemeyer, M., Mescheder, L., Oechsle, M., Geiger, A.: Differentiable volumetric rendering: Learning implicit 3d representations without 3d supervision. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2020)
2020
Cited alongside, same era.
2020
Cited alongside, same era.
Saito, S., Simon, T., Saragih, J., Joo, H.: Pifuhd: Multi-level pixel-aligned implicit function for high-resolution 3d human digitization. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2020)
2020
Cited alongside, same era.
Song, K., Tan, X., Qin, T., Lu, J., Liu, T.Y.: Mpnet: Masked and permuted pre-training for language understanding. Advances in Neural Information Processing Systems (NeurIPS) (2020)
2020
Cited alongside, same era.
Tancik, M., Srinivasan, P.P., Mildenhall, B., Fridovich-Keil, S., Raghavan, N., Singhal, U., Ramamoorthi, R., Barron, J.T., Ng, R.: Fourier features let networks learn high frequency functions in low dimensional domains. Advances in Neural Information Processing Systems (NeurIPS) (2020)
2020
Cited alongside, same era.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., Sutskever, I.: Learning transferable visual models from natural language supervision. In: International Conference on Machine Learning (ICML) (2021)
2021
Later among the works it cites.
Saito, S., Yang, J., Ma, Q., Black, M.J.: SCANimate: Weakly supervised learning of skinned clothed avatar networks. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2021)
2021
Later among the works it cites.
Shree, V., Asfora, B., Zheng, R., Hong, S., Banfi, J., Campbell, M.: Exploiting natural language for efficient risk-aware multi-robot sar planning. IEEE Robotics and Automation Letters 6
2021
Later among the works it cites.
Wang, S., Mihajlovic, M., Ma, Q., Geiger, A., Tang, S.: Metaavatar: Learning animatable clothed human models from few depth images. In: Advances in Neural Information Processing Systems (NeurIPS) (2021)
2021
Later among the works it cites.
Youwang, K., Ji-Yeon, K., Joo, K., Oh, T.H.: Unified 3d mesh recovery of humans and animals by learning animal exercise. In: British Machine Vision Conference (BMVC) (2021)
2021
Later among the works it cites.
Yu, A., Ye, V., Tancik, M., Kanazawa, A.: pixelNeRF: Neural radiance fields from one or few images. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2021)
2021
Later among the works it cites.
Zhang, R., Guo, Z., Zhang, W., Li, K., Miao, X., Cui, B., Qiao, Y., Gao, P., Li, H.: Pointclip: Point cloud understanding by clip. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2021)
2021
Later among the works it cites.
2022
Closest in time.
Gal, R., Patashnik, O., Maron, H., Chechik, G., Cohen-Or, D.: Stylegan-nada: Clip-guided domain adaptation of image generators. In: ACM Transactions on Graphics (SIGGRAPH) (2022)
2022
Closest in time.
Guo, C., Zuo, X., Wang, S., Liu, X., Zou, S., Gong, M., Cheng, L.: Action2video: Generating videos of human 3d actions. International Journal of Computer Vision (IJCV) pp. 1–31 (2022)
2022
Closest in time.
Jain, A., Mildenhall, B., Barron, J.T., Abbeel, P., Poole, B.: Zero-shot text-guided object generation with dream fields. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2022)
2022
Closest in time.
Kim, G., Ye, J.C.: Diffusionclip: Text-guided diffusion models for robust image manipulation. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2022)
2022
Closest in time.
Kwon, G., Ye, J.C.: Clipstyler: Image style transfer with a single text condition. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2022)
2022
Closest in time.
Li, Y., Tang, M., Yang, Y., Huang, Z., Tong, R., Yang, S., Li, Y., Manocha, D.: N-Cloth: Predicting 3D cloth deformation with mesh-based networks. Computer Graphics Forum (Proceedings of Eurographics) pp. 547–558 (2022)
2022
Closest in time.
Michel, O., Bar-On, R., Liu, R., Benaim, S., Hanocka, R.: Text2mesh: Text-driven neural stylization for meshes. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2022)
2022
Closest in time.
Nichol, A., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., McGrew, B., Sutskever, I., Chen, M.: Glide: Towards photorealistic image generation and editing with text-guided diffusion models. In: International Conference on Machine Learning (ICML) (2022)
2022
Closest in time.
Petrovich, M., Black, M.J., Varol, G.: TEMOS: Generating diverse human motions from textual descriptions. In: European Conference on Computer Vision (ECCV) (2022)
2022
Closest in time.
Tevet, G., Gordon, B., Hertz, A., Bermano, A.H., Cohen-Or, D.: Motionclip: Exposing human motion generation to clip space. In: European Conference on Computer Vision (ECCV) (2022)
2022
Closest in time.
Wang, C., Chai, M., He, M., Chen, D., Liao, J.: Clip-nerf: Text-and-image driven manipulation of neural radiance fields. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2022)
2022
Closest in time.