Fetching the paper…
Reading the bibliography…
Automatically synthesizing figures from text captions is a compelling capability.
Automating the design of graphical presentations of relational information
Jock Mackinlay · 1986
Earlier work this paper cites.
Interactive graphic design using automatic presentation knowledge
Steven F. Roth, John Kolojejchick, Joe Mattis, and Jade Goldstein · 1994
Earlier work this paper cites.
Applied cryptography - protocols, algorithms, and source code in C, 2nd Edition
Bruce Schneier · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
A metric for distributions with applications to image databases
Y. Rubner, C. Tomasi, and L.J. Guibas · 1998
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Dimensionality reduction by learning an invariant mapping
R. Hadsell, S. Chopra, and Y. LeCun · 2006
Earlier work this paper cites.
Tucker’s congruence coefficient as a meaningful index of factor similarity
Urbano Lorenzo-Seva and Jos M. F. ten Berge · 2006
Earlier work this paper cites.
Zero-shot learning with semantic output codes
Mark Palatucci, Dean Pomerleau, Geoffrey E Hinton, and Tom M Mitchell · 2009
Earlier work this paper cites.
MaxDiff analysis: Simple counting, individual-level logit, and HB
Bryan K. Orme · 2009
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network, 2015
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Best–Worst Scaling: Theory, Methods and Applications
Jordan J. Louviere, Terry N. Flynn, and A. A. J. Marley · 2015
Earlier work this paper cites.
From word embeddings to document distances
Matt Kusner, Yu Sun, Nicholas Kolkin, and Kilian Weinberger · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Capturing reliable fine-grained sentiment associations by crowdsourcing and best–worst scaling
Svetlana Kiritchenko and Saif M. Mohammad · 2016
Earlier work this paper cites.
Neuro-symbolic program synthesis
Emilio Parisotto, Abdel rahman Mohamed, Rishabh Singh, Lihong Li, Dengyong Zhou, and Pushmeet Kohli · 2017
Earlier work this paper cites.
RobustFill: Neural program learning under noisy I/O
Jacob Devlin, Jonathan Uesato, Surya Bhupatiraju, Rishabh Singh, Abdel rahman Mohamed, and Pushmeet Kohli · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A. Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell · 2017
Earlier work this paper cites.
Best–worst scaling more reliable than rating scales: A case study on sentiment intensity annotation
Svetlana Kiritchenko and Saif Mohammad · 2017
Earlier work this paper cites.
Synthesizing programs for images using reinforced adversarial learning
Yaroslav Ganin, Tejas Kulkarni, Igor Babuschkin, S. M. Ali Eslami, and Oriol Vinyals · 2018
Earlier work this paper cites.
Learning to infer graphics programs from hand-drawn images
Kevin Ellis, Daniel Ritchie, Armando Solar-Lezama, and Josh Tenenbaum · 2018
Earlier work this paper cites.
CSGNet: Neural shape parser for constructive solid geometry
Gopal Sharma, Rishabh Goyal, Difan Liu, Evangelos Kalogerakis, and Subhransu Maji · 2018
Earlier work this paper cites.
Demystifying MMD GANs
Mikołaj Bińkowski, Dougal J. Sutherland, Michael Arbel, and Arthur Gretton · 2018
Earlier work this paper cites.
Write, execute, assess: Program synthesis with a REPL
Kevin Ellis, Maxwell Nye, Yewen Pu, Felix Sosa, Josh Tenenbaum, and Armando Solar-Lezama · 2019
Earlier work this paper cites.
Learning to infer and execute 3D shape programs
Yonglong Tian, Andrew Luo, Xingyuan Sun, Kevin Ellis, William T. Freeman, Joshua B. Tenenbaum, and Jiajun Wu · 2019
Earlier work this paper cites.
EED: Extended edit distance measure for machine translation
Peter Stanchev, Weiyue Wang, and Hermann Ney · 2019
Earlier work this paper cites.
MoverScore: Text generation evaluating with contextualized embeddings and earth mover distance
Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger · 2019
Earlier work this paper cites.
Generative adversarial networks
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2020
Earlier work this paper cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter, 2020
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 2020
Earlier work this paper cites.
S2ORC: The semantic scholar open research corpus
Kyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney, and Daniel Weld · 2020
Earlier work this paper cites.
On the limitations of cross-lingual encoders as exposed by reference-free machine translation evaluation
Wei Zhao, Goran Glavaš, Maxime Peyrard, Yang Gao, Robert West, and Steffen Eger · 2020
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Cited alongside, same era.
DreamCoder: bootstrapping inductive program synthesis with wake-sleep library learning
Kevin Ellis, Catherine Wong, Maxwell Nye, Mathias Sablé-Meyer, Lucas Morales, Luke Hewitt, Luc Cary, Armando Solar-Lezama, and Joshua B. Tenenbaum · 2021
Cited alongside, same era.
Synthesizing natural language to visualization (nl2vis) benchmarks from NL2SQL benchmarks
Yuyu Luo, Nan Tang, Guoliang Li, Chengliang Chai, Wenbo Li, and Xuedi Qin · 2021
Cited alongside, same era.
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever · 2021
Cited alongside, same era.
ImageBind: One embedding space to bind them all
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra · 2023
Later among the works it cites.
UScore: An effective approach to fully unsupervised evaluation metrics for machine translation
Jonas Belouadi and Steffen Eger · 2023
Later among the works it cites.
DiagrammerGPT: Generating open-domain, open-platform diagrams via LLM planning
Abhay Zala, Han Lin, Jaemin Cho, and Mohit Bansal · 2024
Later among the works it cites.
Re-thinking inverse graphics with large language models
Peter Kulits, Haiwen Feng, Weiyang Liu, Victoria Fernandez Abrevaya, and Michael J. Black · 2024
Later among the works it cites.
Is programming by example solved by LLMs?
Wen-Ding Li and Kevin Ellis · 2024
Later among the works it cites.
What matters when building vision-language models?
Hugo Laurençon, Leo Tronchon, Matthieu Cord, and Victor Sanh · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
CogView: Mastering text-to-image generation via transformers
Ming Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng, Chang Zhou, Da Yin, Junyang Lin, Xu Zou, Zhou Shao, Hongxia Yang, and Jie Tang · 2021
Cited alongside, same era.
Taming transformers for high-resolution image synthesis
Patrick Esser, Robin Rombach, and Bjorn Ommer · 2021
Cited alongside, same era.
Denoising diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon · 2021
Cited alongside, same era.
A comprehensive survey on transfer learning
Fuzhen Zhuang, Zhiyuan Qi, Keyu Duan, Dongbo Xi, Yongchun Zhu, Hengshu Zhu, Hui Xiong, and Qing He · 2021
Cited alongside, same era.
CLIPScore: A reference-free evaluation metric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi · 2021
Cited alongside, same era.
SentSim: Crosslingual semantic evaluation of machine translation
Yurun Song, Junchen Zhao, and Lucia Specia · 2021
Cited alongside, same era.
CLIPDraw: Exploring text-to-drawing synthesis through language-image encoders
Kevin Frans, Lisa B. Soros, and Olaf Witkowski · 2022
Cited alongside, same era.
Later among the works it cites.
Building and better understanding vision-language models: insights and future directions, 2024
Hugo Laurençon, Andrés Marafioti, Victor Sanh, and Léo Tronchon · 2024
Later among the works it cites.
Cambrian-1: A fully open, vision-centric exploration of multimodal LLMs
Shengbang Tong, Ellis L Brown II, Penghao Wu, Sanghyun Woo, Adithya Jairam Iyer, Sai Charitha Akula, Shusheng Yang, Jihan Yang, Manoj Middepogu, Ziteng Wang, Xichen Pan, Rob Fergus, Yann LeCun, and Saining Xie · 2024
Later among the works it cites.
Multimodal ArXiv: A dataset for improving scientific comprehension of large vision-language models
Lei Li, Yuqi Wang, Runxin Xu, Peiyi Wang, Xiachong Feng, Lingpeng Kong, and Qi Liu · 2024
Later among the works it cites.
A vision check-up for language models
Pratyusha Sharma, Tamar Rott Shaham, Manel Baradad, Stephanie Fu, Adrian Rodriguez-Munoz, Shivam Duggal, Phillip Isola, and Antonio Torralba · 2024
Later among the works it cites.
How Is ChatGPT’s Behavior Changing Over Time?
Lingjiao Chen, Matei Zaharia, and James Zou · 2024
Later among the works it cites.
Plots made quickly: An efficient approach for generating visualizations from natural language queries
Henrik Voigt, Kai Lawonn, and Sina Zarrieß · 2024
Later among the works it cites.
Automated data visualization from natural language via large language models: An exploratory study
Yang Wu, Yao Wan, Hongyu Zhang, Yulei Sui, Wucai Wei, Wei Zhao, Guandong Xu, and Hai Jin · 2024
Later among the works it cites.
A survey on multimodal large language models
Shukang Yin, Chaoyou Fu, Sirui Zhao, Ke Li, Xing Sun, Tong Xu, and Enhong Chen · 2024
Later among the works it cites.
CLIP-KD: An empirical study of CLIP model distillation
Chuanguang Yang, Zhulin An, Libo Huang, Junyu Bi, Xinqiang Yu, Han Yang, Boyu Diao, and Yongjun Xu · 2024
Later among the works it cites.
How far are we to GPT-4V? closing the gap to commercial multimodal models with open-source suites
Zhe Chen, Weiyun Wang, Hao Tian, Shenglong Ye, Zhangwei Gao, Erfei Cui, Wenwen Tong, Kongzhi Hu, Jiapeng Luo, Zheng Ma, Ji Ma, Jiaqi Wang, Xiaoyi Dong, Hang Yan, Hewei Guo, Conghui He, Botian Shi, Zhenjiang Jin, Chao Xu, Bin Wang, Xingjian Wei, Wei Li, Wenjian Zhang, Bo Zhang, Pinlong Cai, Licheng Wen, Xiangchao Yan, Min Dou, Lewei Lu, Xizhou Zhu, Tong Lu, Dahua Lin, Yu Qiao, Jifeng Dai, and Wenhai Wang · 2024
Later among the works it cites.
PaliGemma: A versatile 3B VLM for transfer, 2024
Lucas Beyer, Andreas Steiner, André Susano Pinto, Alexander Kolesnikov, Xiao Wang, Daniel Salz, Maxim Neumann, Ibrahim Alabdulmohsin, Michael Tschannen, Emanuele Bugliarello, Thomas Unterthiner, Daniel Keysers, Skanda Koppula, Fangyu Liu, Adam Grycner, Alexey Gritsenko, Neil Houlsby, Manoj Kumar, Keran Rong, Julian Eisenschlos, Rishabh Kabra, Matthias Bauer, Matko Bošnjak, Xi Chen, Matthias Minderer, Paul Voigtlaender, Ioana Bica, Ivana Balazevic, Joan Puigcerver, Pinelopi Papalampidi, Olivier Henaff, Xi Xiong, Radu Soricut, Jeremiah Harmsen, and Xiaohua Zhai · 2024
Later among the works it cites.
When does perceptual alignment benefit vision representations?
Shobhita Sundaram, Stephanie Fu, Lukas Muttenthaler, Netanel Yakir Tamir, Lucy Chai, Simon Kornblith, Trevor Darrell, and Phillip Isola · 2024
Later among the works it cites.
Qwen2.5-Coder technical report, 2024
Binyuan Hui, Jian Yang, Zeyu Cui, Jiaxi Yang, Dayiheng Liu, Lei Zhang, Tianyu Liu, Jiajun Zhang, Bowen Yu, Keming Lu, Kai Dang, Yang Fan, Yichang Zhang, An Yang, Rui Men, Fei Huang, Bo Zheng, Yibo Miao, Shanghaoran Quan, Yunlong Feng, Xingzhang Ren, Xuancheng Ren, Jingren Zhou, and Junyang Lin · 2024
Later among the works it cites.
Who evaluates the evaluations? objectively scoring text-to-image prompt coherence metrics with t2IScorescore (TS2)
Michael Saxon, Fatima Jahara, Mahsa Khoshnoodi, Yujie Lu, Aditya Sharma, and William Yang Wang · 2024
Later among the works it cites.
What if we recaption billions of web images with LLaMA-3?, 2024
Xianhang Li, Haoqin Tu, Mude Hui, Zeyu Wang, Bingchen Zhao, Junfei Xiao, Sucheng Ren, Jieru Mei, Qing Liu, Huangjie Zheng, Yuyin Zhou, and Cihang Xie · 2024
Later among the works it cites.
LanguageBind: Extending video-language pretraining to n-modality by language-based semantic alignment
Bin Zhu, Bin Lin, Munan Ning, Yang Yan, Jiaxi Cui, WANG HongFa, Yatian Pang, Wenhao Jiang, Junwu Zhang, Zongwei Li, Cai Wan Zhang, Zhifeng Li, Wei Liu, and Li Yuan · 2024
Later among the works it cites.
Learning to reason with LLMs, 2024
OpenAI · 2024
Later among the works it cites.
UltraEdit: Instruction-based fine-grained image editing at scale
Haozhe Zhao, Xiaojian Ma, Liang Chen, Shuzheng Si, Rujie Wu, Kaikai An, Peiyu Yu, Minjia Zhang, Qing Li, and Baobao Chang · 2024
Later among the works it cites.
ScImage: How good are multimodal large language models at scientific text-to-image generation?
Leixin Zhang, Yinjie Cheng, Weihe Zhai, Steffen Eger, Jonas Belouadi, Fahimeh Moafian, and Zhixue Zhao · 2025
Closest in time.
Diffusion on syntax trees for program synthesis
Shreyas Kapur, Erik Jenner, and Stuart Russell · 2025
Closest in time.
StarVector: Generating scalable vector graphics code from images and text
Juan A. Rodriguez, Abhay Puri, Shubham Agarwal, Issam H. Laradji, Pau Rodriguez, Sai Rajeswar, David Vazquez, Christopher Pal, and Marco Pedersoli · 2025
Closest in time.
MM1.5: Methods, analysis & insights from multimodal LLM fine-tuning
Haotian Zhang, Mingfei Gao, Zhe Gan, Philipp Dufter, Nina Wenzel, Forrest Huang, Dhruti Shah, Xianzhi Du, Bowen Zhang, Yanghao Li, Sam Dodge, Keen You, Zhen Yang, Aleksei Timofeev, Mingze Xu, Hong-You Chen, Jean-Philippe Fauconnier, Zhengfeng Lai, Haoxuan You, Zirui Wang, Afshin Dehghan, Peter Grasch, and Yinfei Yang · 2025
Closest in time.
Multi-LLM collaborative caption generation in scientific documents
Jaeyoung Kim, Jongho Lee, Hong-Jun Choi, Ting-Yao Hsu, Chieh-Yang Huang, Sungchul Kim, Ryan Rossi, Tong Yu, Clyde Lee Giles, Ting-Hao ‘Kenneth’ Huang, and Sungchul Choi · 2025
Closest in time.
NeuralSVG: An implicit representation for text-to-vector generation, 2025
Sagi Polaczek, Yuval Alaluf, Elad Richardson, Yael Vinker, and Daniel Cohen-Or · 2025
Closest in time.
BLIP3-KALE: Knowledge augmented large-scale dense captions
Anas Awadalla, Le Xue, Manli Shu, An Yan, Jun Wang, Senthil Purushwalkam, Sheng Shen, Hannah Lee, Oscar Lo, Jae Sung Park, Etash Kumar Guha, Silvio Savarese, Ludwig Schmidt, Yejin Choi, Caiming Xiong, and Ran Xu · 2025
Closest in time.