Fetching the paper…
Reading the bibliography…
Inspired by the recent success of LLMs, the field of human motion understanding has increasingly shifted toward developing large motion models.
Recurrent network models for human dynamics
Fragkiadaki, K., Levine, S., Felsen, P., and Malik, J · 2015
Earlier work this paper cites.
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A · 2015
Earlier work this paper cites.
Keep it smpl: Automatic estimation of 3d human pose and shape from a single image
Bogo, F., Kanazawa, A., Lassner, C., Gehler, P., Romero, J., and Black, M. J · 2016
Earlier work this paper cites.
The kit motion-language dataset
Plappert, M., Mandery, C., and Asfour, T · 2016
Earlier work this paper cites.
Learning human motion models for long-term predictions
Ghosh, P., Song, J., Aksan, E., and Hilliges, O · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Viola, F., Green, T., Back, T., Natsev, P., et al · 2017
Earlier work this paper cites.
Neural discrete representation learning
Van Den Oord, A., Vinyals, O., et al · 2017
Earlier work this paper cites.
Text2action: Generative adversarial synthesis from language to action
Ahn, H., Ha, T., Choi, Y., Yoo, H., and Oh, S · 2018
Earlier work this paper cites.
Language2pose: Natural language grounded pose forecasting
Ahuja, C. and Morency, L.-P · 2019
Earlier work this paper cites.
Amass: Archive of motion capture as surface shapes
Mahmood, N., Ghorbani, N., Troje, N. F., Pons-Moll, G., and Black, M. J · 2019
Earlier work this paper cites.
Expressive body capture: 3d hands, face, and body from a single image
Pavlakos, G., Choutas, V., Ghorbani, N., Bolkart, T., Osman, A. A., Tzionas, D., and Black, M. J · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
A stochastic conditioning scheme for diverse human motion prediction
Aliakbarian, S., Saleh, F. S., Salzmann, M., Petersson, L., and Gould, S · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Univl: A unified video and language pre-training model for multimodal understanding and generation
Luo, H., Ji, L., Shi, B., Huang, H., Duan, N., Li, T., Li, J., Bharti, T., and Zhou, M · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Grab: A dataset of whole-body human grasping of objects
Taheri, O., Ghorbani, N., Black, M. J., and Tzionas, D · 2020
Earlier work this paper cites.
Learning diverse stochastic human-action generators by learning smooth latent transitions
Wang, Z., Yu, P., Zhao, Y., Zhang, R., Zhou, Y., Yuan, J., and Chen, C · 2020
Earlier work this paper cites.
Haa500: Human-centric atomic action dataset with curated videos
Chung, J., Wuu, C.-h., Yang, H.-r., Tai, Y.-W., and Tang, C.-K · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Videogpt: Video generation using vq-vae and transformers
Yan, W., Zhang, Y., Abbeel, P., and Srinivas, A · 2021
Earlier work this paper cites.
Implicit neural representations for variable length human motion generation
Cervantes, P., Sekikawa, Y., Sato, I., and Shinoda, K · 2022
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2022
Cited alongside, same era.
Maskvit: Masked visual pre-training for video prediction
Gupta, A., Tian, S., Zhang, Y., Wu, J., Martín-Martín, R., and Fei-Fei, L · 2022
Cited alongside, same era.
Autoregressive image generation using residual quantization
Lee, D., Kim, C., Kim, S., Cho, M., and Han, W.-S · 2022
Cited alongside, same era.
Investigating pose representations and motion contexts modeling for 3d motion prediction
Liu, Z., Wu, S., Jin, S., Ji, S., Liu, Q., Lu, S., and Cheng, L · 2022
Cited alongside, same era.
Hierarchical generation of human-object interactions with diffusion probabilistic models
Pi, H., Peng, S., Yang, M., Zhou, X., and Bao, H · 2023
Later among the works it cites.
Learning 3d human pose estimation from dozens of datasets using a geometry-aware autoencoder to bridge between skeleton formats
Sárándi, I., Hermans, A., and Leibe, B · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
What is the best automated metric for text to motion generation?
Voas, J., Wang, Y., Huang, Q., and Mooney, R · 2023
Later among the works it cites.
mplug-owl: Modularization empowers large language models with multimodality
Ye, Q., Xu, H., Xu, G., Ye, J., Yan, M., Zhou, Y., Wang, J., Hu, A., Shi, P., Shi, Y., et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
Motionclip: Exposing human motion generation to clip space
Tevet, G., Gordon, B., Hertz, A., Bermano, A. H., and Cohen-Or, D · 2022
Cited alongside, same era.
Vitpose: Simple vision transformer baselines for human pose estimation
Xu, Y., Zhang, J., Zhang, Q., and Tao, D · 2022
Cited alongside, same era.
Locally hierarchical auto-regressive modeling for image generation
You, T., Kim, S., Kim, C., Lee, D., and Han, B · 2022
Cited alongside, same era.
Glamr: Global occlusion-aware human mesh recovery with dynamic cameras
Yuan, Y., Iqbal, U., Molchanov, P., Kitani, K., and Kautz, J · 2022
Cited alongside, same era.
Motiondiffuse: Text-driven human motion generation with diffusion model
Zhang, M., Cai, Z., Pan, L., Hong, F., Guo, X., Yang, L., and Liu, Z · 2022
Cited alongside, same era.
Executing your commands via motion diffusion in latent space
Chen, X., Jiang, B., Liu, W., Huang, Z., Fu, B., Chen, T., and Yu, G · 2023
Cited alongside, same era.
Later among the works it cites.
Language model beats diffusion–tokenizer is key to visual generation
Yu, L., Lezama, J., Gundavarapu, N. B., Versari, L., Sohn, K., Minnen, D., Cheng, Y., Gupta, A., Gu, X., Hauptmann, A. G., et al · 2023
Later among the works it cites.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Closest in time.
Videoorion: Tokenizing object dynamics in videos
Feng, Y., Li, Y., Zhang, W., Luo, H., Yue, Z., Zheng, S., and Lu, Z · 2024
Closest in time.
Momask: Generative masked modeling of 3d human motions
Guo, C., Mu, Y., Javed, M. G., Wang, S., and Cheng, L · 2024
Closest in time.
https://theorangeduck.com/page/animation-quality
Holden, D · 2024
Closest in time.
Motion-x: A large-scale 3d expressive whole-body human motion dataset
Lin, J., Zeng, A., Lu, S., Cai, Y., Zhang, R., Wang, H., and Zhang, L · 2024
Closest in time.
Quadrupedgpt: Towards a versatile quadruped agent in open-ended worlds
Mei, Y., Wang, Y., Zheng, S., and Jin, Q · 2024
Closest in time.
GPT-4o mini: advancing cost-efficient intelligence
OpenAI · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reid, M., Savinov, N., Teplyashin, D., Lepikhin, D., Lillicrap, T., Alayrac, J.-b., Soricut, R., Lazaridou, A., Firat, O., Schrittwieser, J., et al · 2024
Closest in time.
Wham: Reconstructing world-grounded humans with accurate 3d motion
Shin, S., Kim, J., Halilaj, E., and Black, M. J · 2024
Closest in time.
Motiongpt-2: A general-purpose motion-language model for motion generation and understanding
Wang, Y., Huang, D., Zhang, Y., Ouyang, W., Jiao, J., Feng, X., Zhou, Y., Wan, P., Tang, S., and Xu, D · 2024
Closest in time.
Motionllm: Multimodal motion-language learning with large language models
Wu, Q., Zhao, Y., Wang, Y., Tai, Y.-W., and Tang, C.-K · 2024
Closest in time.
Unicode: Learning a unified codebook for multimodal large language models
Zheng, S., Zhou, B., Feng, Y., Wang, Y., and Lu, Z · 2024
Closest in time.
Avatargpt: All-in-one framework for motion understanding planning generation and beyond
Zhou, Z., Wan, Y., and Wang, B · 2024
Closest in time.
Fg-t2m++: Llms-augmented fine-grained text driven human motion generation
Wang, Y., Li, M., Liu, J., Leng, Z., Li, F. W., Zhang, Z., and Liang, X · 2025
Closest in time.