Fetching the paper…
Reading the bibliography…
Humanoid robots, with their human-like form, are uniquely suited for interacting in environments built for people.
Honda humanoid robots development
Hirose, M. and Ogawa, K · 1917
Earlier work this paper cites.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
Embodied Artificial Intelligence: Trends and Challenges , pp. 1–26
Pfeifer, R. and Iida, F · 2004
Earlier work this paper cites.
Denoising diffusion probabilistic models, 2020
Ho, J., Jain, A., and Abbeel, P · 2006
Earlier work this paper cites.
Image quality metrics: Psnr vs. ssim
Horé, A. and Ziou, D · 2010
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations, 2021
Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B · 2011
Earlier work this paper cites.
Generative adversarial networks, 2014
Goodfellow, I. J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics, 2015
Sohl-Dickstein, J., Weiss, E. A., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
Ha, D. and Schmidhuber, J · 2018
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium, 2018
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S · 2018
Earlier work this paper cites.
Neural discrete representation learning, 2018
van den Oord, A., Vinyals, O., and Kavukcuoglu, K · 2018
Earlier work this paper cites.
Humanoid Robotics: A Reference
Goswami, A. and Vadakkepat, P. (eds.) · 2019
Earlier work this paper cites.
Zero-shot text-to-image generation, 2021
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I · 2021
Earlier work this paper cites.
Human-humanoid interaction and cooperation: a review
Vianello, L., Penco, L., Gomes, W., You, Y., Anzalone, S. M., Maurice, P., Thomas, V., and Ivaldi, S · 2021
Earlier work this paper cites.
On the opportunities and risks of foundation models, 2022
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., Brynjolfsson, E., Buch, S., Card, D., Castellon, R., Chatterji, N., Chen, A., Creel, K., Davis, J. Q., Demszky, D., Donahue, C., Doumbouya, M., Durmus, E., Ermon, S., Etchemendy, J., Ethayarajh, K., Fei-Fei, L., Finn, C., Gale, T., Gillespie, L., Goel, K., Goodman, N., Grossman, S., Guha, N., Hashimoto, T., Henderson, P., Hewitt, J., Ho, D. E., Hong, J., Hsu, K., Huang, J., Icard, T., Jain, S., Jurafsky, D., Kalluri, P., Karamcheti, S., Keeling, G., Khani, F., Khattab, O., Koh, P. W., Krass, M., Krishna, R., Kuditipudi, R., Kumar, A., Ladhak, F., Lee, M., Lee, T., Leskovec, J., Levent, I., Li, X. L., Li, X., Ma, T., Malik, A., Manning, C. D., Mirchandani, S., Mitchell, E., Munyikwa, Z., Nair, S., Narayan, A., Narayanan, D., Newman, B., Nie, A., Niebles, J. C., Nilforoshan, H., Nyarko, J., Ogut, G., Orr, L., Papadimitriou, I., Park, J. S., Piech, C., Portelance, E., Potts, C., Raghunathan, A., Reich, R., Ren, H., Rong, F., Roohani, Y., Ruiz, C., Ryan, J., Ré, C., Sadigh, D., Sagawa, S., Santhanam, K., Shih, A., Srinivasan, K., Tamkin, A., Taori, R., Thomas, A. W., Tramèr, F., Wang, R. E., Wang, W., Wu, B., Wu, J., Wu, Y., Xie, S. M., Yasunaga, M., You, J., Zaharia, M., Zhang, M., Zhang, T., Zhang, X., Zhang, Y., Zheng, L., Zhou, K., and Liang, P · 2022
Earlier work this paper cites.
Conditional image generation by conditioning variational auto-encoders, 2022
Harvey, W., Naderiparizi, S., and Wood, F · 2022
Earlier work this paper cites.
Classifier-free diffusion guidance, 2022
Ho, J. and Salimans, T · 2022
Earlier work this paper cites.
Ho, J., Salimans, T., Gritsenko, A., Chan, W., Norouzi, M., and Fleet, D. J · 2022
Earlier work this paper cites.
Flow straight and fast: Learning to generate and transfer data with rectified flow, 2022
Liu, X., Gong, C., and Liu, Q · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models, 2022
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Earlier work this paper cites.
Stochastic interpolants: A unifying framework for flows and diffusions, 2023
Albergo, M. S., Boffi, N. M., and Vanden-Eijnden, E · 2023
Cited alongside, same era.
Du, Y., Yang, M., Florence, P., Xia, F., Wahid, A., Ichter, B., Sermanet, P., Yu, T., Abbeel, P., Tenenbaum, J. B., Kaelbling, L., Zeng, A., and Tompson, J · 2023
Cited alongside, same era.
Code as policies: Language model programs for embodied control, 2023
Liang, J., Huang, W., Xia, F., Xu, P., Hausman, K., Ichter, B., Florence, P., and Zeng, A · 2023
Cited alongside, same era.
Flow matching for generative modeling, 2023
Lipman, Y., Chen, R. T. Q., Ben-Hamu, H., Nickel, M., and Le, M · 2023
Cited alongside, same era.
Llm-enhanced scene graph learning for household rearrangement, 2024
Li, W., Yu, Z., She, Q., Yu, Z., Lan, Y., Zhu, C., Hu, R., and Xu, K · 2024
Later among the works it cites.
Liu, Y., Zhang, K., Li, Y., Yan, Z., Gao, C., Chen, R., Yuan, Z., Huang, Y., Sun, H., Gao, J., He, L., and Sun, L · 2024
Later among the works it cites.
Open-magvit2: An open-source project toward democratizing auto-regressive visual generation, 2024
Luo, Z., Shi, F., Ge, Y., Yang, Y., Wang, L., and Shan, Y · 2024
Later among the works it cites.
Movie gen: A cast of media foundation models, 2024
Polyak, A., Zohar, A., Brown, A., Tjandra, A., Sinha, A., Lee, A., Vyas, A., Shi, B., Ma, C.-Y., Chuang, C.-Y., Yan, D., Choudhary, D., Wang, D., Sethi, G., Pang, G., Ma, H., Misra, I., Hou, J., Wang, J., Jagadeesh, K., Li, K., Zhang, L., Singh, M., Williamson, M., Le, M., Yu, M., Singh, M. K., Zhang, P., Vajda, P., Duval, Q., Girdhar, R., Sumbaly, R., Rambhatla, S. S., Tsai, S., Azadi, S., Datta, S., Chen, S., Bell, S., Ramaswamy, S., Sheynin, S., Bhattacharya, S., Motwani, S., Xu, T., Li, T., Hou, T., Hsu, W.-N., Yin, X., Dai, X., Taigman, Y., Luo, Y., Liu, Y.-C., Wu, Y.-C., Zhao, Y., Kirstain, Y., He, Z., He, Z., Pumarola, A., Thabet, A., Sanakoyeu, A., Mallya, A., Guo, B., Araya, B., Kerr, B., Wood, C., Liu, C., Peng, C., Vengertsev, D., Schonfeld, E., Blanchard, E., Juefei-Xu, F., Nord, F., Liang, J., Hoffman, J., Kohler, J., Fire, K., Sivakumar, K., Chen, L., Yu, L., Gao, L., Georgopoulos, M., Moritz, R., Sampson, S. K., Li, S., Parmeggiani, S., Fine, S., Fowler, T., Petrovic, V., and Du, Y · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Peebles, W. and Xie, S · 2023
Cited alongside, same era.
Roformer: Enhanced transformer with rotary position embedding, 2023
Su, J., Lu, Y., Pan, S., Murtadha, A., Wen, B., and Liu, Y · 2023
Cited alongside, same era.
Attention is all you need, 2023
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2023
Cited alongside, same era.
Learning interactive real-world simulators
Yang, M., Du, Y., Ghasemipour, K., Tompson, J., Schuurmans, D., and Abbeel, P · 2023
Cited alongside, same era.
Magvit: Masked generative video transformer, 2023
Yu, L., Cheng, Y., Sohn, K., Lezama, J., Zhang, H., Chang, H., Hauptmann, A. G., Yang, M.-H., Hao, Y., Essa, I., and Jiang, L · 2023
Cited alongside, same era.
Copilot4d: Learning unsupervised world models for autonomous driving via discrete diffusion
Zhang, L., Xiong, Y., Yang, Z., Casas, S., Hu, R., and Urtasun, R · 2023
Cited alongside, same era.
Learning fine-grained bimanual manipulation with low-cost hardware, 2023
Zhao, T. Z., Kumar, V., Levine, S., and Finn, C · 2023
Cited alongside, same era.
1X World Model Challenge, June 2024
1X Technologies · 2024
Cited alongside, same era.
Later among the works it cites.
Eyes wide shut? exploring the visual shortcomings of multimodal llms, 2024
Tong, S., Liu, Z., Zhai, Y., Ma, Y., LeCun, Y., and Xie, S · 2024
Later among the works it cites.
Llm3:large language model-based task and motion planning with motion failure reasoning, 2024
Wang, S., Han, M., Jiao, Z., Zhang, Z., Wu, Y. N., Zhu, S.-C., and Liu, H · 2024
Later among the works it cites.
ivideogpt: Interactive videogpts are scalable world models
Wu, J., Yin, S., Feng, N., He, X., Li, D., Hao, J., and Long, M · 2024
Later among the works it cites.
Pandora: Towards general world model with natural language actions and video states, 2024
Xiang, J., Liu, G., Gu, Y., Gao, Q., Ning, Y., Zha, Y., Feng, Z., Tao, T., Hao, S., Shi, Y., Liu, Z., Xing, E. P., and Hu, Z · 2024
Later among the works it cites.
A survey on video diffusion models, 2024
Xing, Z., Feng, Q., Chen, H., Dai, Q., Hu, H., Xu, H., Wu, Z., and Jiang, Y.-G · 2024
Later among the works it cites.
Video as the new language for real-world decision making, 2024
Yang, S., Walker, J., Parker-Holder, J., Du, Y., Bruce, J., Barreto, A., Abbeel, P., and Schuurmans, D · 2024
Later among the works it cites.
Language model beats diffusion – tokenizer is key to visual generation, 2024
Yu, L., Lezama, J., Gundavarapu, N. B., Versari, L., Sohn, K., Minnen, D., Cheng, Y., Birodkar, V., Gupta, A., Gu, X., Hauptmann, A. G., Gong, B., Yang, M.-H., Essa, I., Ross, D. A., and Jiang, L · 2024
Later among the works it cites.
Vision-language models for vision tasks: A survey, 2024
Zhang, J., Huang, J., Jin, S., and Lu, S · 2024
Later among the works it cites.
Irasim: Learning interactive real-robot action simulators
Zhu, F., Wu, H., Guo, S., Liu, Y., Cheang, C., and Kong, T · 2024
Later among the works it cites.
Bar, A., Zhou, G., Tran, D., Darrell, T., and LeCun, Y · 2025
Closest in time.
Chen, C., Qian, R., Hu, W., Fu, T.-J., Tong, J., Wang, X., Li, L., Zhang, B., Schwing, A., Liu, W., and Yang, Y · 2025
Closest in time.
Auraflow: Generate high-fidelity 3d assets with diffusion models
fal.ai Blog · 2025
Closest in time.
Hunyuanvideo: A systematic framework for large video generative models, 2025
Kong, W., Tian, Q., Zhang, Z., Min, R., Dai, Z., Zhou, J., Xiong, J., Li, X., Wu, B., Zhang, J., Wu, K., Lin, Q., Yuan, J., Long, Y., Wang, A., Wang, A., Li, C., Huang, D., Yang, F., Tan, H., Wang, H., Song, J., Bai, J., Wu, J., Xue, J., Wang, J., Wang, K., Liu, M., Li, P., Li, S., Wang, W., Yu, W., Deng, X., Li, Y., Chen, Y., Cui, Y., Peng, Y., Yu, Z., He, Z., Xu, Z., Zhou, Z., Xu, Z., Tao, Y., Lu, Q., Liu, S., Zhou, D., Wang, H., Yang, Y., Wang, D., Liu, Y., Jiang, J., and Zhong, C · 2025
Closest in time.
Cosmos world foundation model platform for physical ai, 2025
NVIDIA, :, Agarwal, N., Ali, A., Bala, M., Balaji, Y., Barker, E., Cai, T., Chattopadhyay, P., Chen, Y., Cui, Y., Ding, Y., Dworakowski, D., Fan, J., Fenzi, M., Ferroni, F., Fidler, S., Fox, D., Ge, S., Ge, Y., Gu, J., Gururani, S., He, E., Huang, J., Huffman, J., Jannaty, P., Jin, J., Kim, S. W., Klár, G., Lam, G., Lan, S., Leal-Taixe, L., Li, A., Li, Z., Lin, C.-H., Lin, T.-Y., Ling, H., Liu, M.-Y., Liu, X., Luo, A., Ma, Q., Mao, H., Mo, K., Mousavian, A., Nah, S., Niverty, S., Page, D., Paschalidou, D., Patel, Z., Pavao, L., Ramezanali, M., Reda, F., Ren, X., Sabavat, V. R. N., Schmerling, E., Shi, S., Stefaniak, B., Tang, S., Tchapmi, L., Tredak, P., Tseng, W.-C., Varghese, J., Wang, H., Wang, H., Wang, H., Wang, T.-C., Wei, F., Wei, X., Wu, J. Z., Xu, J., Yang, W., Yen-Chen, L., Zeng, X., Zeng, Y., Zhang, J., Zhang, Q., Zhang, Y., Zhao, Q., and Zolkowski, A · 2025
Closest in time.
Cogvideox: Text-to-video diffusion models with an expert transformer, 2025
Yang, Z., Teng, J., Zheng, W., Ding, M., Huang, S., Xu, J., Yang, Y., Hong, W., Zhang, X., Feng, G., Yin, D., Zhang, Y., Wang, W., Cheng, Y., Xu, B., Gu, X., Dong, Y., and Tang, J · 2025
Closest in time.