Fetching the paper…
Reading the bibliography…
While large vision-language models can generate motion graphics animations from text prompts, they regularly fail to include all spatio-temporal properties described in the prompt.
Semantics and Syntax of Motion
Leonard Talmy. 1975 · 1975
Earlier work this paper cites.
Maintaining knowledge about temporal intervals
James F. Allen. 1983 · 1983
Earlier work this paper cites.
How language structures space
Leonard Talmy. 1983 · 1983
Earlier work this paper cites.
VII.1 - DECOMPOSING A MATRIX INTO SIMPLE TRANSFORMATIONS
Spencer W. Thomas. 1991 · 1991
Earlier work this paper cites.
WordNet: A Lexical Database for English. In Human Language Technology: Proceedings of a Workshop held at Plainsboro, New Jersey, March 8-11, 1994
George A. Miller. 1994 · 1994
Earlier work this paper cites.
A model for reasoning about bidimensional temporal relations. In PRINCIPLES OF KNOWLEDGE REPRESENTATION AND REASONING-INTERNATIONAL CONFERENCE- . MORGAN KAUFMANN PUBLISHERS, 124–130
Philippe Balbiani, Jean-François Condotta, and L Farinas Del Cerro. 1998 · 1998
Earlier work this paper cites.
How space structures language
Barbara Tversky and Paul U Lee. 1998 · 1998
Earlier work this paper cites.
Animation control for real-time virtual humans
Norman I. Badler, Martha S. Palmer, and Rama Bindiganavale. 1999 · 1999
Earlier work this paper cites.
A suggestive interface for 3D drawing. In Proceedings of the 14th Annual ACM Symposium on User Interface Software and Technology (Orlando, Florida) (UIST ’01) . Association for Computing Machinery, New York, NY, USA, 173–181
Takeo Igarashi and John F. Hughes. 2001 · 2001
Earlier work this paper cites.
CarSim: an automatic 3D text-to-scene conversion system applied to road accident reports. In Demonstrations
Ola Åkerberg, Hans Svensson, Bastian Schulz, and Pierre Nugues. 2003 · 2003
Earlier work this paper cites.
An extensible SAT-solver. In International conference on theory and applications of satisfiability testing . Springer, 502–518
Niklas Eén and Niklas Sörensson. 2003 · 2003
Earlier work this paper cites.
Automatic conversion of natural language to 3D animation
Minhua Ma. 2006 · 2006
Earlier work this paper cites.
Weakly Supervised Learning of Semantic Parsers for Mapping Instructions to Actions
Yoav Artzi and Luke Zettlemoyer. 2013 · 2013
Earlier work this paper cites.
Spatial reasoning with rectangular cardinal relations: The convex tractable subalgebra
Isabel Navarrete, Antonio Morales, Guido Sciavicco, and M Antonia Cardenas-Viedma. 2013 · 2013
Earlier work this paper cites.
Inferring and Executing Programs for Visual Reasoning. In Proceedings of the IEEE International Conference on Computer Vision (ICCV)
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Judy Hoffman, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick. 2017 · 2017
Earlier work this paper cites.
Aishwarya Kamath and Rajarshi Das. 2018 · 2018
Earlier work this paper cites.
CARDINAL: Computer Assisted Authoring of Movie Scripts. In Proceedings of the 23rd International Conference on Intelligent User Interfaces (Tokyo, Japan) (IUI ’18) . Association for Computing Machinery, New York, NY, USA, 509–519
Marcel Marti, Jodok Vieli, Wojciech Witoń, Rushit Sanghrajka, Daniel Inversini, Diana Wotruba, Isabel Simo, Sasha Schriber, Mubbasir Kapadia, and Markus Gross. 2018 · 2018
Earlier work this paper cites.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Earlier work this paper cites.
Generating Animations from Screenplays. In Proceedings of the Eighth Joint Conference on Lexical and Computational Semantics (*SEM 2019) , Rada Mihalcea, Ekaterina Shutova, Lun-Wei Ku, Kilian Evang, and Soujanya Poria (Eds.). Association for Computational Linguistics, Minneapolis, Minnesota, 292–307
Yeyao Zhang, Eleftheria Tsipidi, Sasha Schriber, Mubbasir Kapadia, Markus Gross, and Ashutosh Modi. 2019 · 2019
Cited alongside, same era.
spaCy: Industrial-strength Natural Language Processing in Python
Matthew Honnibal, Ines Montani, Sofie Van Landeghem, and Adriane Boyd. 2020 · 2020
Cited alongside, same era.
CLIPScore: A Reference-free Evaluation Metric for Image Captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. 2021 · 2021
Cited alongside, same era.
Introduction to algorithms
Thomas H Cormen, Charles E Leiserson, Ronald L Rivest, and Clifford Stein. 2022 · 2022
Cited alongside, same era.
What’s Left? concept grounding with logic-enhanced foundation models. In Proceedings of the 37th International Conference on Neural Information Processing Systems (New Orleans, LA, USA) (NIPS ’23) . Curran Associates Inc., Red Hook, NY, USA, Article 1684, 17 pages
Joy Hsu, Jiayuan Mao, Joshua B. Tenenbaum, and Jiajun Wu. 2024 · 2024
Later among the works it cites.
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al · 2024
Later among the works it cites.
Text2CAD: Generating Sequential CAD Models from Beginner-to-Expert Level Text Prompts
Mohammad Sadil Khan, Sankalp Sinha, Talha Uddin Sheikh, Didier Stricker, Sk Aziz Ali, and Muhammad Zeshan Afzal. 2024 · 2024
Later among the works it cites.
pyrealb at the GEM‘24 Data-to-text Task: Symbolic English Text Generation from RDF Triples. In Proceedings of the 17th International Natural Language Generation Conference: Generation Challenges , Simon Mille and Miruna-Adriana Clinciu (Eds.). Association for Computational Linguistics, Tokyo, Japan, 54–58
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
F-coref: Fast, Accurate and Easy to Use Coreference Resolution. In Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing: System Demonstrations , Wray Buntine and Maria Liakata (Eds.). Association for Computational Linguistics, Taipei, Taiwan, 48–56
Shon Otmazgin, Arie Cattan, and Yoav Goldberg. 2022 · 2022
Cited alongside, same era.
Efficient Few-Shot Learning Without Prompts
Lewis Tunstall, Nils Reimers, Unso Eun Seo Jo, Luke Bates, Daniel Korat, Moshe Wasserblat, and Oren Pereg. 2022 · 2022
Cited alongside, same era.
A Review of Text-to-Animation Systems
Nacir Bouali and Violetta Cavalli-Sforza. 2023 · 2023
Cited alongside, same era.
Visual Programming: Compositional Visual Reasoning Without Training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 14953–14962
Tanmay Gupta and Aniruddha Kembhavi. 2023 · 2023
Cited alongside, same era.
TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) . 20406–20417
Yushi Hu, Benlin Liu, Jungo Kasai, Yizhong Wang, Mari Ostendorf, Ranjay Krishna, and Noah A. Smith. 2023 · 2023
Cited alongside, same era.
Holistic Evaluation of Text-To-Image Models
Tony Lee, Michihiro Yasunaga, Chenlin Meng, Yifan Mai, Joon Sung Park, Agrim Gupta, Yunzhi Zhang, Deepak Narayanan, Hannah Benita Teufel, Marco Bellagente, Minguk Kang, Taesung Park, Jure Leskovec, Jun-Yan Zhu, Li Fei-Fei, Jiajun Wu, Stefano Ermon, and Percy Liang. 2023 · 2023
Cited alongside, same era.
ViperGPT: Visual Inference via Python Execution for Reasoning
Dídac Surís, Sachit Menon, and Carl Vondrick. 2023 · 2023
Cited alongside, same era.
Programmatically Grounded, Compositionally Generalizable Robotic Manipulation. In The Eleventh International Conference on Learning Representations
Renhao Wang, Jiayuan Mao, Joy Hsu, Hang Zhao, Jiajun Wu, and Yang Gao. 2023 · 2023
Cited alongside, same era.
Guy Lapalme. 2024 · 2024
Later among the works it cites.
Design2code: How far are we from automating front-end engineering?
Chenglei Si, Yanzhe Zhang, Zhengyuan Yang, Ruibo Liu, and Diyi Yang. 2024 · 2024
Later among the works it cites.
Predicated Diffusion: Predicate Logic-Based Attention Guidance for Text-to-Image Diffusion Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 8651–8660
Kota Sueyoshi and Takashi Matsubara. 2024 · 2024
Later among the works it cites.
Jiao Sun, Yushi Hu, Deqing Fu, Royi Rassin, Sjoerd van Steenkiste, Dana Alon, Su Wang, Charles Herrmann, Ranjay Krishna, Da-Cheng Juan, and Cyrus Rashtchian. 2024 · 2024
Later among the works it cites.
Keyframer: Empowering Animation Design using Large Language Models
Tiffany Tseng, Ruijia Cheng, and Jeffrey Nichols. 2024 · 2024
Later among the works it cites.
Tooncrafter: Generative cartoon interpolation
Jinbo Xing, Hanyuan Liu, Menghan Xia, Yong Zhang, Xintao Wang, Ying Shan, and Tien-Tsin Wong. 2024b · 2024
Later among the works it cites.
Empowering LLMs to Understand and Generate Complex Vector Graphics
Ximing Xing, Juncheng Hu, Guotao Liang, Jing Zhang, Dong Xu, and Qian Yu. 2024a · 2024
Later among the works it cites.
The Scene Language: Representing Scenes with Programs, Words, and Embeddings
Yunzhi Zhang, Zizhang Li, Matt Zhou, Shangzhe Wu, and Jiajun Wu. 2024 · 2024
Later among the works it cites.
Greensock Animation Platform v3
Jack Doyle. 2025 · 2025
Closest in time.
anime.js ⋅ \cdot JavaScript animation engine
Julien Garnier. 2025 · 2025
Closest in time.
SceneCraft: an LLM agent for synthesizing 3D scenes as blender code. In Proceedings of the 41st International Conference on Machine Learning (Vienna, Austria) (ICML’24) . JMLR.org, Article 776, 31 pages
Ziniu Hu, Ahmet Iscen, Aashi Jain, Thomas Kipf, Yisong Yue, David A Ross, Cordelia Schmid, and Alireza Fathi. 2025 · 2025
Closest in time.
A Solver-Aided Hierarchical Language for LLM-Driven CAD Design
Benjamin T Jones, Felix Hähnlein, Zihan Zhang, Maaz Ahmad, Vladimir Kim, and Adriana Schulz. 2025 · 2025
Closest in time.
LogoMotion: Visually-Grounded Code Synthesis for Creating and Editing Animation. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25) . Association for Computing Machinery, New York, NY, USA, Article 157, 16 pages
Vivian Liu, Rubaiat Habib Kazi, Li-Yi Wei, Matthew Fisher, Timothy Langlois, Seth Walker, and Lydia Chilton. 2025 · 2025
Closest in time.
Using CSS animations - CSS: Cascading style sheets: MDN
MozDevNet. 2025 · 2025
Closest in time.
NeuralSVG: An Implicit Representation for Text-to-Vector Generation
Sagi Polaczek, Yuval Alaluf, Elad Richardson, Yael Vinker, and Daniel Cohen-Or. 2025 · 2025
Closest in time.