Fetching the paper…
Reading the bibliography…
This paper introduces SceneCraft, a Large Language Model (LLM) Agent converting text descriptions into Blender-executable Python scripts which render complex scenes with up to a hundred 3D assets.
Wordseye: an automatic text-to-scene conversion system
Coyne, R. and Sproat, R · 2001
Earlier work this paper cites.
Real-time automatic 3d scene generation from natural language voice and text descriptions
Seversky, L. M. and Yin, L · 2006
Earlier work this paper cites.
Learning the visual interpretation of sentences
Zitnick, C. L., Parikh, D., and Vanderwende, L · 2013
Earlier work this paper cites.
Learning spatial knowledge for text to 3d scene generation
Chang, A. X., Savva, M., and Manning, C. D · 2014
Earlier work this paper cites.
Text to 3d scene generation with rich lexical grounding
Chang, A. X., Monroe, W., Savva, M., Potts, C., and Manning, C. D · 2015
Earlier work this paper cites.
Sceneseer: 3d scene design with natural language
Chang, A. X., Eric, M., Savva, M., and Manning, C. D · 2017
Earlier work this paper cites.
Language-driven synthesis of 3d scenes from scene databases
Ma, R., Patil, A. G., Fisher, M., Li, M., Pirk, S., Hua, B., Yeung, S., Tong, X., Guibas, L. J., and Zhang, H · 2018
Earlier work this paper cites.
FVD: A new metric for video generation
Unterthiner, T., van Steenkiste, S., Kurach, K., Marinier, R., Michalski, M., and Gelly, S · 2019
Earlier work this paper cites.
Unsupervised question decomposition for question answering
Perez, E., Lewis, P. S. H., Yih, W., Cho, K., and Kiela, D · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Earlier work this paper cites.
GODIVA: generating open-domain videos from natural descriptions
Wu, C., Huang, L., Zhang, Q., Li, B., Ji, L., Yang, F., Sapiro, G., and Duan, N · 2021
Earlier work this paper cites.
Dreamfusion: Text-to-3d using 2d diffusion
Poole, B., Jain, A., Barron, J. T., and Mildenhall, B · 2022
Earlier work this paper cites.
Vision-language models as a source of rewards
Baumli, K., Baveja, S., Behbahani, F. M. P., Chan, H., Comanici, G., Flennerhag, S., Gazeau, M., Holsheimer, K., Horgan, D., Laskin, M., Lyle, C., Masoom, H., McKinney, K., Mnih, V., Neitz, A., Pardo, F., Parker-Holder, J., Quan, J., Rocktäschel, T., Sahni, H., Schaul, T., Schroecker, Y., Spencer, S., Steigerwald, R., Wang, L., and Zhang, L · 2023
Cited alongside, same era.
RT-2: vision-language-action models transfer web knowledge to robotic control
Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Chen, X., Choromanski, K., Ding, T., Driess, D., Dubey, A., Finn, C., Florence, P., Fu, C., Arenas, M. G., Gopalakrishnan, K., Han, K., Hausman, K., Herzog, A., Hsu, J., Ichter, B., Irpan, A., Joshi, N. J., Julian, R., Kalashnikov, D., Kuang, Y., Leal, I., Lee, L., Lee, T. E., Levine, S., Lu, Y., Michalewski, H., Mordatch, I., Pertsch, K., Rao, K., Reymann, K., Ryoo, M. S., Salazar, G., Sanketi, P., Sermanet, P., Singh, J., Singh, A., Soricut, R., Tran, H. T., Vanhoucke, V., Vuong, Q., Wahid, A., Welker, S., Wohlhart, P., Wu, J., Xia, F., Xiao, T., Xu, P., Xu, S., Yu, T., and Zitkovich, B · 2023
Cited alongside, same era.
Universal self-consistency for large language model generation
Chen, X., Aksitov, R., Alon, U., Ren, J., Xiao, K., Yin, P., Prakash, S., Sutton, C., Wang, X., and Zhou, D · 2023
Cited alongside, same era.
Infinite photorealistic worlds using procedural generation
Raistrick, A., Lipson, L., Ma, Z., Mei, L., Wang, M., Zuo, Y., Kayan, K., Wen, H., Han, B., Wang, Y., Newell, A., Law, H., Goyal, A., Yang, K., and Deng, J · 2023
Later among the works it cites.
Vision-language models are zero-shot reward models for reinforcement learning
Rocamonde, J., Montesinos, V., Nava, E., Perez, E., and Lindner, D · 2023
Later among the works it cites.
Unsupervised traffic scene generation with synthetic 3d scene graphs
Savkin, A., Ellouze, R., Navab, N., and Tombari, F · 2023
Later among the works it cites.
Reflexion: an autonomous agent with dynamic memory and self-reflection
Shinn, N., Labash, B., and Gopinath, A · 2023
Later among the works it cites.
Roomdreamer: Text-driven 3d indoor scene synthesis with coherent geometry and texture
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deng, X., Gu, Y., Zheng, B., Chen, S., Stevens, S., Wang, B., Sun, H., and Su, Y · 2023
Cited alongside, same era.
AVIS: autonomous visual information seeking with large language models
Hu, Z., Iscen, A., Sun, C., Chang, K., Sun, Y., Ross, D. A., Schmid, C., and Fathi, A · 2023
Cited alongside, same era.
Videopoet: A large language model for zero-shot video generation
Kondratyuk, D., Yu, L., Gu, X., Lezama, J., Huang, J., Hornung, R., Adam, H., Akbari, H., Alon, Y., Birodkar, V., Cheng, Y., Chiu, M., Dillon, J., Essa, I., Gupta, A., Hahn, M., Hauth, A., Hendon, D., Martinez, A., Minnen, D., Ross, D. A., Schindler, G., Sirotenko, M., Sohn, K., Somandepalli, K., Wang, H., Yan, J., Yang, M., Yang, X., Seybold, B., and Jiang, L · 2023
Cited alongside, same era.
Magic3d: High-resolution text-to-3d content creation
Lin, C., Gao, J., Tang, L., Takikawa, T., Zeng, X., Huang, X., Kreis, K., Fidler, S., Liu, M., and Lin, T · 2023
Cited alongside, same era.
Agentbench: Evaluating llms as agents
Liu, X., Yu, H., Zhang, H., Xu, Y., Lei, X., Lai, H., Gu, Y., Ding, H., Men, K., Yang, K., Zhang, S., Deng, X., Zeng, A., Du, Z., Zhang, C., Shen, S., Zhang, T., Su, Y., Sun, H., Huang, M., Dong, Y., and Tang, J · 2023
Cited alongside, same era.
Gpt4motion: Scripting physical motions in text-to-video generation via blender-oriented gpt planning
Lv, J., Huang, Y., Yan, M., Huang, J., Liu, J., Liu, Y., Wen, Y., Chen, X., and Chen, S · 2023
Cited alongside, same era.
Eureka: Human-level reward design via coding large language models
Ma, Y. J., Liang, W., Wang, G., Huang, D., Bastani, O., Jayaraman, D., Zhu, Y., Fan, L., and Anandkumar, A · 2023
Cited alongside, same era.
Gpt-4v(ision) system card
OpenAI · 2023
Cited alongside, same era.
Advances in data-driven analysis and synthesis of 3d indoor scenes
Patil, A. G., Patil, S. G., Li, M., Fisher, M., Savva, M., and Zhang, H · 2023
Cited alongside, same era.
Song, L., Cao, L., Xu, H., Kang, K., Tang, F., Yuan, J., and Yang, Z · 2023
Later among the works it cites.
3d-gpt: Procedural 3d modeling with large language models
Sun, C., Han, J., Deng, W., Wang, X., Qin, Z., and Gould, S · 2023
Later among the works it cites.
Voyager: An open-ended embodied agent with large language models
Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., and Anandkumar, A · 2023
Later among the works it cites.
Lego-net: Learning regular rearrangements of objects in rooms
Wei, Q. A., Ding, S., Park, J. J., Sajnani, R., Poulenard, A., Sridhar, S., and Guibas, L. J · 2023
Later among the works it cites.
Least-to-most prompting enables complex reasoning in large language models
Zhou, D., Schärli, N., Hou, L., Wei, J., Scales, N., Wang, X., Schuurmans, D., Cui, C., Bousquet, O., Le, Q. V., and Chi, E. H · 2023
Later among the works it cites.
Mastering text-to-image diffusion: Recaptioning, planning, and generating with multimodal llms
Yang, L., Yu, Z., Meng, C., Xu, M., Ermon, S., and Cui, B · 2024
Closest in time.
Gpt-4v(ision) is a generalist web agent, if grounded
Zheng, B., Gou, B., Kil, J., Sun, H., and Su, Y · 2024
Closest in time.