Fetching the paper…
Reading the bibliography…
Generating human motion from text has been dominated by denoising motion models either through diffusion or generative masking process.
Plappert, M., Mandery, C., Asfour, T.: The KIT motion-language dataset. Big Data 4
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
Vaswani, A., Shazeer, N.M., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., Polosukhin, I.: Attention is all you need. In: Neural Information Processing Systems (2017), https://api.semanticscholar.org/CorpusID:13756489
2017
Earlier work this paper cites.
Ahuja, C., Morency, L.P.: Language2pose: Natural language grounded pose forecasting. In: 2019 International Conference on 3D Vision (3DV). pp. 719–728 (2019). https://doi.org/10.1109/3DV.2019.00084
2019
Earlier work this paper cites.
Devlin, J., Chang, M.W., Lee, K., Toutanova, K.: Bert: Pre-training of deep bidirectional transformers for language understanding. In: North American Chapter of the Association for Computational Linguistics (2019), https://api.semanticscholar.org/CorpusID:52967399
2019
Earlier work this paper cites.
Ghazvininejad, M., Levy, O., Liu, Y., Zettlemoyer, L.: Mask-predict: Parallel decoding of conditional masked language models. In: Conference on Empirical Methods in Natural Language Processing (2019), https://api.semanticscholar.org/CorpusID:202538740
2019
Earlier work this paper cites.
Mahmood, N., Ghorbani, N., Troje, N.F., Pons-Moll, G., Black, M.J.: Amass: Archive of motion capture as surface shapes. 2019 IEEE/CVF International Conference on Computer Vision (ICCV) pp. 5441–5450 (2019), https://api.semanticscholar.org/CorpusID:102351100
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
Guo, C., Zuo, X., Wang, S., Zou, S., Sun, Q., Deng, A., Gong, M., Cheng, L.: Action2motion: Conditioned generation of 3d human motions. Proceedings of the 28th ACM International Conference on Multimedia (2020), https://api.semanticscholar.org/CorpusID:220870974
2020
Earlier work this paper cites.
Ho, J., Jain, A., Abbeel, P.: Denoising diffusion probabilistic models. ArXiv abs/2006.11239
2020
Earlier work this paper cites.
Qian, L., Zhou, H., Bao, Y., Wang, M., Qiu, L., Zhang, W., Yu, Y., Li, L.: Glancing transformer for non-autoregressive neural machine translation. In: Annual Meeting of the Association for Computational Linguistics (2020), https://api.semanticscholar.org/CorpusID:221150562
2020
Earlier work this paper cites.
Song, J., Meng, C., Ermon, S.: Denoising diffusion implicit models. ArXiv abs/2010.02502
2020
Earlier work this paper cites.
Austin, J., Johnson, D.D., Ho, J., Tarlow, D., Van Den Berg, R.: Structured denoising diffusion models in discrete state-spaces. Advances in Neural Information Processing Systems 34
2021
Earlier work this paper cites.
Du, Z., Qian, Y., Liu, X., Ding, M., Qiu, J., Yang, Z., Tang, J.: Glm: General language model pretraining with autoregressive blank infilling. In: Annual Meeting of the Association for Computational Linguistics (2021), https://api.semanticscholar.org/CorpusID:247519241
2021
Earlier work this paper cites.
Ghosh, A., Cheema, N., Oguz, C., Theobalt, C., Slusallek, P.: Synthesis of compositional animations from textual descriptions. 2021 IEEE/CVF International Conference on Computer Vision (ICCV) pp. 1376–1386 (2021), https://api.semanticscholar.org/CorpusID:232404671
2021
Earlier work this paper cites.
Nichol, A., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., McGrew, B., Sutskever, I., Chen, M.: Glide: Towards photorealistic image generation and editing with text-guided diffusion models. In: International Conference on Machine Learning (2021), https://api.semanticscholar.org/CorpusID:245335086
2021
Earlier work this paper cites.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., Sutskever, I.: Learning transferable visual models from natural language supervision. In: International Conference on Machine Learning (2021), https://api.semanticscholar.org/CorpusID:231591445
2021
Earlier work this paper cites.
Zeghidour, N., Luebs, A., Omran, A., Skoglund, J., Tagliasacchi, M.: Soundstream: An end-to-end neural audio codec. IEEE/ACM Transactions on Audio, Speech, and Language Processing 30
2021
Cited alongside, same era.
Zhang, Z., Ma, J., Zhou, C., Men, R., Li, Z., Ding, M., Tang, J., Zhou, J., Yang, H.: M6-ufc: Unifying multi-modal controls for conditional image synthesis via non-autoregressive generative transformers (2021), https://api.semanticscholar.org/CorpusID:237204528
2021
Cited alongside, same era.
Zhang, Z., Ma, J., Zhou, C., Men, R., Li, Z., Ding, M., Tang, J., Zhou, J., Yang, H.: Ufc-bert: Unifying multi-modal controls for conditional image synthesis. In: Neural Information Processing Systems (2021), https://api.semanticscholar.org/CorpusID:235253928
2021
Cited alongside, same era.
Cmu graphics lab motion capture database, http://mocap.cs.cmu.edu/ , accessed: 2022-11-11
2022
Cited alongside, same era.
2022
Later among the works it cites.
2022
Later among the works it cites.
Yu, J., Xu, Y., Koh, J.Y., Luong, T., Baid, G., Wang, Z., Vasudevan, V., Ku, A., Yang, Y., Ayan, B.K., Hutchinson, B.C., Han, W., Parekh, Z., Li, X., Zhang, H., Baldridge, J., Wu, Y.: Scaling autoregressive models for content-rich text-to-image generation. Trans. Mach. Learn. Res. 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chang, H., Zhang, H., Jiang, L., Liu, C., Freeman, W.T.: Maskgit: Masked generative image transformer. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) pp. 11305–11315 (2022), https://api.semanticscholar.org/CorpusID:246680316
2022
Cited alongside, same era.
Chen, X., Jiang, B., Liu, W., Huang, Z., Fu, B., Chen, T., Yu, J., Yu, G.: Executing your commands via motion diffusion in latent space. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) pp. 18000–18010 (2022), https://api.semanticscholar.org/CorpusID:254408910
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
Guo, C., Zou, S., Zuo, X., Wang, S., Ji, W., Li, X., Cheng, L.: Generating diverse and natural 3d human motions from text. In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 5142–5151 (2022). https://doi.org/10.1109/CVPR52688.2022.00509
2022
Cited alongside, same era.
2022
Cited alongside, same era.
Ho, J., Salimans, T.: Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598 (2022)
2022
Cited alongside, same era.
Kim, J., Kim, J., Choi, S.: Flame: Free-form language-based motion synthesis & editing. In: AAAI Conference on Artificial Intelligence (2022), https://api.semanticscholar.org/CorpusID:251979380
2022
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
Guo, C., Mu, Y., Javed, M.G., Wang, S., Cheng, L.: Momask: Generative masked modeling of 3d human motions (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Yan, S., Liu, Y., Wang, H., Du, X., Liu, M., Liu, H.: Cross-modal retrieval for motion and text via doptriple loss (2023), https://api.semanticscholar.org/CorpusID:263610212
2023
Later among the works it cites.
Zhang, J., Zhang, Y., Cun, X., Huang, S., Zhang, Y., Zhao, H., Lu, H., Shen, X.: Generating human motion from textual descriptions with discrete representations. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) pp. 14730–14740 (2023), https://api.semanticscholar.org/CorpusID:255942203
2023
Later among the works it cites.
2023
Later among the works it cites.
Pinyoanuntapong, E., Wang, P., Lee, M., Chen, C.: Mmm: Generative masked motion model. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2024)
2024
Closest in time.