Fetching the paper…
Reading the bibliography…
We introduce MotionCLIP, a 3D human motion auto-encoder featuring a latent embedding that is disentangled, well behaved, and supports highly semantic textual descriptions.
Dimensionality reduction by learning an invariant mapping. In 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’06) , Vol. 2. IEEE, 1735–1742
Raia Hadsell, Sumit Chopra, and Yann LeCun. 2006 · 2006
Earlier work this paper cites.
Nice: Non-linear independent components estimation
Laurent Dinh, David Krueger, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
SMPL: A skinned multi-person linear model
Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J Black. 2015 · 2015
Earlier work this paper cites.
Learning structured output representation using deep conditional generative models
Kihyuk Sohn, Honglak Lee, and Xinchen Yan. 2015 · 2015
Earlier work this paper cites.
Deep unsupervised clustering with gaussian mixture variational autoencoders
Nat Dilokthanakul, Pedro AM Mediano, Marta Garnelo, Matthew CH Lee, Hugh Salimbeni, Kai Arulkumaran, and Murray Shanahan. 2016 · 2016
Earlier work this paper cites.
JALI: an animator-centric viseme model for expressive lip synchronization
Pif Edwards, Chris Landreth, Eugene Fiume, and Karan Singh. 2016 · 2016
Earlier work this paper cites.
Image Style Transfer Using Convolutional Neural Networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
Leon A. Gatys, Alexander S. Ecker, and Matthias Bethge. 2016 · 2016
Earlier work this paper cites.
A deep learning framework for character motion synthesis and editing
Daniel Holden, Jun Saito, and Taku Komura. 2016 · 2016
Earlier work this paper cites.
The KIT motion-language dataset
Matthias Plappert, Christian Mandery, and Tamim Asfour. 2016 · 2016
Earlier work this paper cites.
Fine-grained image classification via combining vision and language. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 5994–6002
Xiangteng He and Yuxin Peng. 2017 · 2017
Earlier work this paper cites.
Arbitrary style transfer in real-time with adaptive instance normalization. In Proceedings of the IEEE international conference on computer vision . 1501–1510
Xun Huang and Serge Belongie. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
A large-scale RGB-D database for arbitrary-view human action recognition. In Proceedings of the 26th ACM international Conference on Multimedia . 1510–1518
Yanli Ji, Feixiang Xu, Yang Yang, Fumin Shen, Heng Tao Shen, and Wei-Shi Zheng. 2018 · 2018
Earlier work this paper cites.
Generating animated videos of human activities from natural language descriptions
Angela S Lin, Lemeng Wu, Rodolfo Corona, Kevin Tai, Qixing Huang, and Raymond J Mooney. 2018 · 2018
Earlier work this paper cites.
Learning a bidirectional mapping between human whole-body motion and natural language using deep recurrent neural networks
Matthias Plappert, Christian Mandery, and Tamim Asfour. 2018 · 2018
Earlier work this paper cites.
Paired recurrent autoencoders for bidirectional translation between robot actions and linguistic descriptions
Tatsuro Yamada, Hiroyuki Matsunaga, and Tetsuya Ogata. 2018 · 2018
Cited alongside, same era.
Language2pose: Natural language grounded pose forecasting. In 2019 International Conference on 3D Vision (3DV) . IEEE, 719–728
Chaitanya Ahuja and Louis-Philippe Morency. 2019 · 2019
Cited alongside, same era.
Ntu rgb+ d 120: A large-scale benchmark for 3d human activity understanding
Jun Liu, Amir Shahroudy, Mauricio Perez, Gang Wang, Ling-Yu Duan, and Alex C Kot. 2019 · 2019
Cited alongside, same era.
AMASS: Archive of Motion Capture as Surface Shapes. In International Conference on Computer Vision . 5442–5451
Naureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll, and Michael J. Black. 2019 · 2019
Cited alongside, same era.
Expressive Body Capture: 3D Hands, Face, and Body from a Single Image. In Proceedings IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) . 10975–10985
CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval
Huaishao Luo, Lei Ji, Ming Zhong, Yang Chen, Wen Lei, Nan Duan, and Tianrui Li. 2021 · 2021
Later among the works it cites.
Text2Mesh: Text-Driven Neural Stylization for Meshes
Oscar Michel, Roi Bar-On, Richard Liu, Sagie Benaim, and Rana Hanocka. 2021 · 2021
Later among the works it cites.
Action-Conditioned 3D Human Motion Synthesis with Transformer VAE. In International Conference on Computer Vision (ICCV) . 10985–10995
Mathis Petrovich, Michael J. Black, and Gül Varol. 2021 · 2021
Later among the works it cites.
BABEL: Bodies, Action and Behavior with English Labels. In Proceedings IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR) . 722–731
Abhinanda R. Punnakkal, Arjun Chandrasekaran, Nikos Athanasiou, Alejandra Quiros-Ramirez, and Michael J. Black. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A. A. Osman, Dimitrios Tzionas, and Michael J. Black. 2019 · 2019
Cited alongside, same era.
Two-stream adaptive graph convolutional networks for skeleton-based action recognition. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 12026–12035
Lei Shi, Yifan Zhang, Jian Cheng, and Hanqing Lu. 2019 · 2019
Cited alongside, same era.
On the continuity of rotation representations in neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 5745–5753
Yi Zhou, Connelly Barnes, Jingwan Lu, Jimei Yang, and Hao Li. 2019 · 2019
Cited alongside, same era.
Unpaired motion style transfer from video to animation
Kfir Aberman, Yijia Weng, Dani Lischinski, Daniel Cohen-Or, and Baoquan Chen. 2020 · 2020
Cited alongside, same era.
A simple framework for contrastive learning of visual representations. In International conference on machine learning . PMLR, 1597–1607
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020 · 2020
Cited alongside, same era.
Rhythm is a Dancer: Music-Driven Motion Synthesis with Global Structure
Andreas Aristidou, Anastasios Yiannakidis, Kfir Aberman, Daniel Cohen-Or, Ariel Shamir, and Yiorgos Chrysanthou. 2021 · 2021
Cited alongside, same era.
CLIP2Video: Mastering Video-Text Retrieval via Image CLIP
Han Fang, Pengfei Xiong, Luhui Xu, and Yu Chen. 2021 · 2021
Cited alongside, same era.
Clipdraw: Exploring text-to-drawing synthesis through language-image encoders
Kevin Frans, LB Soros, and Olaf Witkowski. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision. In International Conference on Machine Learning . PMLR, 8748–8763
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
Zero-shot text-to-image generation. In International Conference on Machine Learning . PMLR, 8821–8831
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. 2021 · 2021
Later among the works it cites.
Clip-forge: Towards zero-shot text-to-shape generation
Aditya Sanghi, Hang Chu, Joseph G Lambourne, Ye Wang, Chin-Yi Cheng, and Marco Fumero. 2021 · 2021
Later among the works it cites.
CLIP-NeRF: Text-and-Image Driven Manipulation of Neural Radiance Fields
Can Wang, Menglei Chai, Mingming He, Dongdong Chen, and Jing Liao. 2021a · 2021
Later among the works it cites.
Multi-Person 3D Motion Prediction with Multi-Range Transformers
Jiashun Wang, Huazhe Xu, Medhini Narasimhan, and Xiaolong Wang. 2021b · 2021
Later among the works it cites.
Autoregressive Stylized Motion Synthesis with Generative Flow. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 13612–13621
Yu-Hui Wen, Zhipeng Yang, Hongbo Fu, Lin Gao, Yanan Sun, and Yong-Jin Liu. 2021 · 2021
Later among the works it cites.
MUGL: Large Scale Multi Person Conditional Action Generation with Locomotion. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision . 257–265
Shubh Maheshwari, Debtanu Gupta, and Ravi Kiran Sarvadevabhatla. 2022 · 2022
Closest in time.
CLIPasso: Semantically-Aware Object Sketching
Yael Vinker, Ehsan Pajouheshgar, Jessica Y Bo, Roman Christian Bachmann, Amit Haim Bermano, Daniel Cohen-Or, Amir Zamir, and Ariel Shamir. 2022 · 2022
Closest in time.
Action2motion: Conditioned generation of 3d human motions. In Proceedings of the 28th ACM International Conference on Multimedia . 2021–2029
Chuan Guo, Xinxin Zuo, Sen Wang, Shihao Zou, Qingyao Sun, Annan Deng, Minglun Gong, and Li Cheng. 2020 · 2029
Closest in time.
Styleclip: Text-driven manipulation of stylegan imagery. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 2085–2094
Or Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or, and Dani Lischinski. 2021 · 2094
Closest in time.