Fetching the paper…
Reading the bibliography…
Diffusion-based image generation models such as DALL-E 3 and Stable Diffusion-XL demonstrate remarkable capabilities in generating images with realistic and unique compositions.
A Computational Approach to Edge Detection
Canny, J · 1986
Earlier work this paper cites.
Long Short-Term Memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
3D-R2N2: A Unified Approach for Single and Multi-view 3D Object Reconstruction
Choy, C., Xu, D., Gwak, J., Chen, K., and Savarese, S · 2016
Earlier work this paper cites.
Representation Learning for Grounded Spatial Reasoning
Janner, M., Narasimhan, K., and Barzilay, R · 2018
Earlier work this paper cites.
Can Language Models Encode Perceptual Structure Without Grounding? A Case Study in Color
Abdou, M., Kulmizev, A., Hershcovich, D., Frank, S., Pavlick, E., and Søgaard, A · 2021
Earlier work this paper cites.
SPARTQA: A Textual Question Answering Benchmark for Spatial Reasoning
Mirzaee, R., Rajaby Faghihi, H., Ning, Q., and Kordjamshidi, P · 2021
Earlier work this paper cites.
Training-Free Structured Diffusion Guidance for Compositional Text-to-Image Synthesis
Feng, W., He, X., Fu, T.-J., Jampani, V., Akula, A. R., Narayana, P., Basu, S., Wang, X. E., and Wang, W. Y · 2022
Earlier work this paper cites.
Inner Monologue: Embodied Reasoning through Planning with Language Models
Huang, W., Xia, F., Xiao, T., Chan, H., Liang, J., Florence, P., Zeng, A., Tompson, J., Mordatch, I., Chebotar, Y., Sermanet, P., Brown, N., Jackson, T., Luu, L., Levine, S., Hausman, K., and Ichter, B · 2022
Earlier work this paper cites.
Mapping Language Models to Grounded Conceptual Spaces
Patel, R. and Pavlick, E · 2022
Earlier work this paper cites.
DreamFusion: Text-to-3D using 2D Diffusion
Poole, B., Jain, A., Barron, J. T., and Mildenhall, B · 2022
Earlier work this paper cites.
Hierarchical Text-Conditional Image Generation with CLIP Latents, April 2022
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M · 2022
Earlier work this paper cites.
High-Resolution Image Synthesis With Latent Diffusion Models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Cited alongside, same era.
Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E., Ghasemipour, S. K. S., Gontijo-Lopes, R., Ayan, B. K., Salimans, T., Ho, J., Fleet, D. J., and Norouzi, M · 2022
Cited alongside, same era.
Scaling Autoregressive Models for Content-Rich Text-to-Image Generation
Yu, J., Xu, Y., Koh, J. Y., Luong, T., Baid, G., Wang, Z., Vasudevan, V., Ku, A., Yang, Y., Ayan, B. K., Hutchinson, B., Han, W., Parekh, Z., Li, X., Zhang, H., Baldridge, J., and Wu, Y · 2022
Cited alongside, same era.
Sparks of Artificial General Intelligence: Early experiments with GPT-4, April 2023
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., Nori, H., Palangi, H., Ribeiro, M. T., and Zhang, Y · 2023
Cited alongside, same era.
Attend-and-Excite: Attention-Based Semantic Guidance for Text-to-Image Diffusion Models
Chefer, H., Alaluf, Y., Vinker, Y., Wolf, L., and Cohen-Or, D · 2023
Generative Agents: Interactive Simulacra of Human Behavior
Park, J. S., O’Brien, J., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S · 2023
Later among the works it cites.
SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis, July 2023
Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Müller, J., Penna, J., and Rombach, R · 2023
Later among the works it cites.
Linguistic Binding in Diffusion Models: Enhancing Attribute Correspondence through Attention Map Alignment
Rassin, R., Hirsch, E., Glickman, D., Ravfogel, S., Goldberg, Y., and Chechik, G · 2023
Later among the works it cites.
Toolformer: Language Models Can Teach Themselves to Use Tools
Schick, T., Dwivedi-Yu, J., Dessi, R., Raileanu, R., Lomeli, M., Hambro, E., Zettlemoyer, L., Cancedda, N., and Scialom, T · 2023
Later among the works it cites.
HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face
Shen, Y., Song, K., Tan, X., Li, D., Lu, W., and Zhuang, Y · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
PAL: Program-aided Language Models
Gao, L., Madaan, A., Zhou, S., Alon, U., Liu, P., Yang, Y., Callan, J., and Neubig, G · 2023
Cited alongside, same era.
BlenderGPT
gd3kr · 2023
Cited alongside, same era.
OpenAGI: When LLM Meets Domain Experts
Ge, Y., Hua, W., Mei, K., Ji, J., Tan, J., Xu, S., Li, Z., and Zhang, Y · 2023
Cited alongside, same era.
Visual Programming: Compositional Visual Reasoning Without Training
Gupta, T. and Kembhavi, A · 2023
Cited alongside, same era.
Shap-E: Generating Conditional 3D Implicit Functions, May 2023
Jun, H. and Nichol, A · 2023
Cited alongside, same era.
GPT-4V(ision) technical work and authors, 2023
OpenAI · 2023
Cited alongside, same era.
Improving Image Generation with Better Captions
Betker, J., Goh, G., Jing, L., Brooks, T., Wang, J., Li, L., Ouyang, L., Zhuang, J., Lee, J., Guo, Y., Manassra, W., Dhariwal, P., Chu, C., Jiao, Y., and Ramesh, A
Cited in the paper.
Reflexion: Language agents with verbal reinforcement learning
Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K. R., and Yao, S · 2023
Later among the works it cites.
Voyager: An Open-Ended Embodied Agent with Large Language Models
Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., and Anandkumar, A · 2023
Later among the works it cites.
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Wu, C., Yin, S., Qi, W., Wang, X., Tang, Z., and Duan, N · 2023
Later among the works it cites.
Lumos: Learning agents with unified data, modular design, and open-source llms
Yin, D., Brahman, F., Ravichander, A., Chandu, K. R., Chang, K.-W., Choi, Y., and Lin, B. Y · 2023
Later among the works it cites.
Adding Conditional Control to Text-to-Image Diffusion Models
Zhang, L., Rao, A., and Agrawala, M · 2023
Later among the works it cites.
Mixtral of Experts, January 2024
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., de las Casas, D., Hanna, E. B., Bressand, F., Lengyel, G., Bour, G., Lample, G., Lavaud, L. R., Saulnier, L., Lachaux, M.-A., Stock, P., Subramanian, S., Yang, S., Antoniak, S., Scao, T. L., Gervet, T., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2024
Closest in time.