Fetching the paper…
Reading the bibliography…
Scientific diagrams are vital tools for communicating structured knowledge across disciplines.
Structure-mapping: A theoretical framework for analogy
Dedre Gentner. 1983 · 1983
Earlier work this paper cites.
Diagrams in the comprehension of scientific texts
Mary Hegarty, Patricia A Carpenter, and Marcel Adam Just. 1991 · 1991
Earlier work this paper cites.
Cognitive load theory
Jan L Plass, Roxana Moreno, and Roland Brünken. 2010 · 2010
Earlier work this paper cites.
The effects of visual complexity on cognitive load as influenced by field dependency and spatial ability
Christopher Gene Allen. 2011 · 2011
Earlier work this paper cites.
SVG essentials: Producing scalable vector graphics with XML
J David Eisenberg and Amelia Bellamy-Royds. 2014 · 2014
Earlier work this paper cites.
Best-worst scaling: theory and methods
Terry N Flynn and Anthony AJ Marley. 2014 · 2014
Earlier work this paper cites.
Spearman correlation coefficients, differences between
Leann Myers and Maria J Sirois. 2014 · 2014
Earlier work this paper cites.
Diagrams as tools for scientific reasoning
Adele Abrahamsen and William Bechtel. 2015 · 2015
Earlier work this paper cites.
Best-worst scaling: Theory, methods and applications
Jordan J Louviere, Terry N Flynn, and Anthony Alfred John Marley. 2015 · 2015
Earlier work this paper cites.
Fidelity vs. simplicity: a global approach to line drawing vectorization
Jean-Dominique Favreau, Florent Lafarge, and Adrien Bousseau. 2016 · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017 · 2017
Earlier work this paper cites.
Questionnaire design
Jon A Krosnick. 2017 · 2017
Earlier work this paper cites.
Pdf2latex: A deep learning system to convert mathematical documents from pdf to latex. In Proceedings of the ACM Symposium on Document Engineering 2020 . 1–10
Zelun Wang and Jyh-Charn Liu. 2020 · 2020
Earlier work this paper cites.
Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF international conference on computer vision . 9650–9660
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. 2021 · 2021
Earlier work this paper cites.
Clipscore: A reference-free evaluation metric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. 2021 · 2021
Earlier work this paper cites.
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2021 · 2021
Earlier work this paper cites.
Deepvecfont: synthesizing high-quality vector fonts via dual-modality learning
Yizhi Wang and Zhouhui Lian. 2021 · 2021
Earlier work this paper cites.
Improved Aesthetic Predictor
Christoph Schuhmann. 2022 · 2022
Earlier work this paper cites.
CLIP Models are Few-Shot Learners: Empirical Studies on VQA and Visual Entailment. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 6088–6100
Haoyu Song, Li Dong, Weinan Zhang, Ting Liu, and Furu Wei. 2022 · 2022
Earlier work this paper cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Cited alongside, same era.
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning. In Advances in Neural Information Processing Systems , A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Eds.), Vol. 36. Curran Associates, Inc., 49250–49267
Wenliang Dai, Junnan Li, DONGXU LI, Anthony Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale N Fung, and Steven Hoi. 2023 · 2023
Cited alongside, same era.
Code Llama: Open Foundation Models for Code
Wenhan Xiong Grattafiori, Alexandre Défossez, Jade Copet, Faisal Azhar, Hugo Touvron, Louis Martin, Nicolas Usunier, Thomas Scialom, and Gabriel Synnaeve. 2023 · 2023
Cited alongside, same era.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning . PMLR, 19730–19742
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023 · 2023
Cited alongside, same era.
Bridging the Gap: Leveraging Informal Software Architecture Artifacts for Structured Model Creation
Josh Kaplan and Luis Rabelo. 2024 · 2024
Later among the works it cites.
Kaixin Li, Yuchen Tian, Qisheng Hu, Ziyang Luo, Zhiyong Huang, and Jing Ma. 2024 · 2024
Later among the works it cites.
Hello GPT-4o
OpenAI. 2024 · 2024
Later among the works it cites.
Agent planning with world knowledge model
Shuofei Qiao, Runnan Fang, Ningyu Zhang, Yuqi Zhu, Xiang Chen, Shumin Deng, Yong Jiang, Pengjun Xie, Fei Huang, and Huajun Chen. 2024 · 2024
Later among the works it cites.
Image2struct: Benchmarking structure extraction for vision-language models
Josselin Roberts, Tony Lee, Chi Heem Wong, Michihiro Yasunaga, Yifan Mai, and Percy S Liang. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scigraphqa: A large-scale synthetic multi-turn question-answering dataset for scientific graphs
Shengzhi Li and Nima Tajbakhsh. 2023 · 2023
Cited alongside, same era.
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al · 2023
Cited alongside, same era.
Dinov2: Learning robust visual features without supervision
Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al · 2023
Cited alongside, same era.
Starvector: Generating scalable vector graphics code from images
Juan A Rodriguez, Shubham Agarwal, Issam H Laradji, Pau Rodriguez, David Vazquez, Christopher Pal, and Marco Pedersoli. 2023 · 2023
Cited alongside, same era.
Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 2609–2634
Lei Wang, Wanyu Xu, Yihuai Lan, Zhiqiang Hu, Yunshi Lan, Roy Ka-Wei Lee, and Ee-Peng Lim. 2023 · 2023
Cited alongside, same era.
Symbol-LLM: leverage language models for symbolic system in visual human activity reasoning
Xiaoqian Wu, Yong-Lu Li, Jianhua Sun, and Cewu Lu. 2023 · 2023
Cited alongside, same era.
Llama 3.2V: Vision Multimodal Large Model
Meta AI. 2024 · 2024
Cited alongside, same era.
Introducing Claude 3.5 Sonnet
Anthropic. 2024 · 2024
Cited alongside, same era.
Design2code: How far are we from automating front-end engineering?
Chenglei Si, Yanzhe Zhang, Zhengyuan Yang, Ruibo Liu, and Diyi Yang. 2024 · 2024
Later among the works it cites.
Chengyue Wu, Yixiao Ge, Qiushan Guo, Jiahao Wang, Zhixuan Liang, Zeyu Lu, Ying Shan, and Ping Luo. 2024a · 2024
Later among the works it cites.
Lotlip: Improving language-image pre-training for long text understanding
Wei Wu, Kecheng Zheng, Shuailei Ma, Fan Lu, Yuxin Guo, Yifei Zhang, Wei Chen, Qingpei Guo, Yujun Shen, and Zheng-Jun Zha. 2024b · 2024
Later among the works it cites.
WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling
Xudong Xie, Hao Yan, Liang Yin, Yang Liu, Jing Ding, Minghui Liao, Yuliang Liu, Wei Chen, and Xiang Bai. 2024 · 2024
Later among the works it cites.
Empowering LLMs to Understand and Generate Complex Vector Graphics
Ximing Xing, Juncheng Hu, Guotao Liang, Jing Zhang, Dong Xu, and Qian Yu. 2024 · 2024
Later among the works it cites.
Chartmimic: Evaluating lmm’s cross-modal reasoning capability via chart-to-code generation
Cheng Yang, Chufan Shi, Yaxin Liu, Bo Shui, Junjie Wang, Mohan Jing, Linran Xu, Xinyu Zhu, Siheng Li, Yuxiang Zhang, et al · 2024
Later among the works it cites.
ScImage: How good are multimodal large language models at scientific text-to-image generation?
Leixin Zhang, Steffen Eger, Yinjie Cheng, Weihe Zhai, Jonas Belouadi, Christoph Leiter, Simone Paolo Ponzetto, Fahimeh Moafian, and Zhixue Zhao. 2024 · 2024
Later among the works it cites.
Claude 3.7 Sonnet and Claude Code
Anthropic. 2025 · 2025
Closest in time.
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, et al · 2025
Closest in time.
Rongyao Fang, Chengqi Duan, Kun Wang, Linjiang Huang, Hao Li, Shilin Yan, Hao Tian, Xingyu Zeng, Rui Zhao, Jifeng Dai, et al · 2025
Closest in time.
LayerTracer: Cognitive-Aligned Layered SVG Synthesis via Diffusion Transformer
Yiren Song, Danze Chen, and Mike Zheng Shou. 2025 · 2025
Closest in time.
WebWalker: Benchmarking LLMs in Web Traversal
Jialong Wu, Wenbiao Yin, Yong Jiang, Zhenglin Wang, Zekun Xi, Runnan Fang, Linhai Zhang, Yulan He, Deyu Zhou, Pengjun Xie, et al · 2025
Closest in time.
Grok 3 Beta — The Age of Reasoning Agents
xAI. 2025 · 2025
Closest in time.
LongProc: Benchmarking Long-Context Language Models on Long Procedural Generation
Xi Ye, Fangcong Yin, Yinghui He, Joie Zhang, Howard Yen, Tianyu Gao, Greg Durrett, and Danqi Chen. 2025 · 2025
Closest in time.
ChartCoder: Advancing Multimodal Large Language Model for Chart-to-Code Generation
Xuanle Zhao, Xianzhen Luo, Qi Shi, Chi Chen, Shuo Wang, Wanxiang Che, Zhiyuan Liu, and Maosong Sun. 2025 · 2025
Closest in time.