Fetching the paper…
Reading the bibliography…
Multimodal large language models (MLLMs) have made impressive progress in many applications in recent years.
A study of BFLOAT16 for deep learning training
Kalamkar, D.; Mudigere, D.; Mellempudi, N.; Das, D.; Banerjee, K.; Avancha, S.; Vooturi, D. T.; Jammalamadaka, N.; Huang, J.; Yuen, H.; et al. 2019 · 1905
Earlier work this paper cites.
Lewis, M.; Liu, Y.; Goyal, N.; Ghazvininejad, M.; Mohamed, A.; Levy, O.; Stoyanov, V.; and Zettlemoyer, L. 2019 · 1910
Earlier work this paper cites.
SMILES, a line notation and computerized interpreter for chemical structures
Anderson, E.; Veith, G. D.; and Weininger, D. 1987 · 1987
Earlier work this paper cites.
SMILES, a chemical language and information system. 1. Introduction to methodology and encoding rules
Weininger, D. 1988 · 1988
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y. 2004 · 2004
Earlier work this paper cites.
RDKit: Open-source cheminformatics
Landrum, G.; et al. 2006 · 2006
Earlier work this paper cites.
Natural language processing with Python: analyzing text with the natural language toolkit
Bird, S.; Klein, E.; and Loper, E. 2009 · 2009
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Cho, K.; Van Merriënboer, B.; Gulcehre, C.; Bahdanau, D.; Bougares, F.; Schwenk, H.; and Bengio, Y. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P.; and Ba, J. 2014 · 2014
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Sennrich, R.; Haddow, B.; and Birch, A. 2015 · 2015
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Szegedy, C.; Vanhoucke, V.; Ioffe, S.; Shlens, J.; and Wojna, Z. 2016 · 2016
Earlier work this paper cites.
Goh, G. B.; Siegel, C.; Vishnu, A.; Hodas, N. O.; and Baker, N. 2017 · 2017
Earlier work this paper cites.
Image-to-image translation with conditional adversarial networks
Isola, P.; Zhu, J.-Y.; Zhou, T.; and Efros, A. A. 2017 · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I.; and Hutter, F. 2017 · 2017
Earlier work this paper cites.
Neural discrete representation learning
Van Den Oord, A.; Vinyals, O.; et al. 2017 · 2017
Earlier work this paper cites.
Learning multimodal graph-to-graph translation for molecular optimization
Jin, W.; Yang, K.; Barzilay, R.; and Jaakkola, T. 2018 · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Zhang, R.; Isola, P.; Efros, A. A.; Shechtman, E.; and Wang, O. 2018 · 2018
Earlier work this paper cites.
CORE: Automatic Molecule Optimization using Copy and Refine Strategy
Fu, T.; Xiao, C.; and Sun, J. 2020 · 2020
Earlier work this paper cites.
DECIMER: towards deep learning for chemical image recognition
Rajan, K.; Zielesny, A.; and Steinbeck, C. 2020 · 2020
Cited alongside, same era.
Taming transformers for high-resolution image synthesis
Esser, P.; Rombach, R.; and Ommer, B. 2021 · 2021
Cited alongside, same era.
Therapeutics data Commons: machine learning datasets and tasks for therapeutics
Huang, K.; Fu, T.; Gao, W.; Zhao, Y.; Roohani, Y.; Leskovec, J.; Coley, C. W.; Xiao, C.; Sun, J.; and Zitnik, M. 2021 · 2021
Cited alongside, same era.
PubChem in 2021: new data content and improved web interfaces
Kim, S.; Chen, J.; Cheng, T.; Gindulyte, A.; He, J.; He, S.; Li, Q.; Shoemaker, B. A.; Thiessen, P. A.; Yu, B.; et al. 2021 · 2021
Cited alongside, same era.
Computer vision in chemistry: Automatic titration
Kosenkov, Y.; and Kosenkov, D. 2021 · 2021
Cited alongside, same era.
COT: an efficient Python tool for detecting marker genes among many subtypes
Visual instruction tuning
Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2023 · 2023
Later among the works it cites.
MolScribe: Robust Molecular Structure Recognition with Image-to-Graph Generation
Qian, Y.; Guo, J.; Tu, Z.; Li, Z.; Coley, C. W.; and Barzilay, R. 2023 · 2023
Later among the works it cites.
OpenEye Toolkits Documentation
Software, O. S. 2023 · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Team, G.; Anil, R.; Borgeaud, S.; Alayrac, J.-B.; Yu, J.; Soricut, R.; Schalkwyk, J.; Dai, A. M.; Hauth, A.; Millican, K.; et al. 2023 · 2023
Later among the works it cites.
Pytorch fsdp: experiences on scaling fully sharded data parallel
Zhao, Y.; Gu, A.; Varma, R.; Luo, L.; Huang, C.-C.; Xu, M.; Wright, L.; Shojanazeri, H.; Ott, M.; Shleifer, S.; et al. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lu, Y.; Wu, C.-T.; Parker, S. J.; Chen, L.; Saylor, G.; Van Eyk, J. E.; Herrington, D. M.; and Wang, Y. 2021 · 2021
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B.; Donahue, J.; Luc, P.; Miech, A.; Barr, I.; Hasson, Y.; Lenc, K.; Mensch, A.; Millican, K.; Reynolds, M.; et al. 2022 · 2022
Cited alongside, same era.
MolGenSurvey: A Systematic Survey in Machine Learning Models for Molecule Design
Du, Y.; Fu, T.; Sun, J.; and Liu, S. 2022 · 2022
Cited alongside, same era.
Translation between molecules and natural language
Edwards, C.; Lai, T.; Ros, K.; Honke, G.; Cho, K.; and Ji, H. 2022 · 2022
Cited alongside, same era.
Chemformer: a pre-trained transformer for computational chemistry
Irwin, R.; Dimitriadis, S.; He, J.; and Bjerrum, E. J. 2022 · 2022
Cited alongside, same era.
Openflamingo: An open-source framework for training large autoregressive vision-language models
Awadalla, A.; Gao, I.; Gardner, J.; Hessel, J.; Hanafy, Y.; Zhu, W.; Marathe, K.; Bitton, Y.; Gadre, S.; Sagawa, S.; et al. 2023 · 2023
Cited alongside, same era.
Qwen-VL: A Frontier Large Vision-Language Model with Versatile Abilities
Bai, J.; Bai, S.; Yang, S.; Wang, S.; Tan, Z.; Wang, P.; Chen, J.; Zhou, J.; and Zhou, J. 2023 · 2023
Cited alongside, same era.
Zhu, D.; Chen, J.; Shen, X.; Li, X.; and Elhoseiny, M. 2023 · 2023
Later among the works it cites.
Cai, Z.; Cao, M.; Chen, H.; Chen, K.; Chen, K.; Chen, X.; Chen, X.; Chen, Z.; Chen, Z.; Chu, P.; et al. 2024 · 2024
Later among the works it cites.
Liu, D.; Zhao, S.; Zhuo, L.; Lin, W.; Qiao, Y.; Li, H.; and Gao, P. 2024 · 2024
Later among the works it cites.
GPT-4V(ision) System Card
OpenAI. 2023 · 2024
Later among the works it cites.
GPT-4o: Our most advanced AI model
OpenAI. 2024 · 2024
Later among the works it cites.
Autoregressive model beats diffusion: Llama for scalable image generation
Sun, P.; Jiang, Y.; Chen, S.; Zhang, S.; Peng, B.; Luo, P.; and Yuan, Z. 2024 · 2024
Later among the works it cites.
Chameleon: Mixed-modal early-fusion foundation models
Team, C. 2024 · 2024
Later among the works it cites.
Emu3: Next-token prediction is all you need
Wang, X.; Zhang, X.; Luo, Z.; Sun, Q.; Cui, Y.; Wang, J.; Zhang, F.; Wang, Y.; Li, Z.; Yu, Q.; et al. 2024 · 2024
Later among the works it cites.
Show-o: One single transformer to unify multimodal understanding and generation
Xie, J.; Mao, W.; Bai, Z.; Zhang, D. J.; Wang, W.; Lin, K. Q.; Gu, Y.; Chen, Z.; Yang, Z.; and Shou, M. Z. 2024 · 2024
Later among the works it cites.
Transfusion: Predict the next token and diffuse images with one multi-modal model
Zhou, C.; Yu, L.; Babu, A.; Tirumala, K.; Yasunaga, M.; Shamis, L.; Kahn, J.; Ma, X.; Zettlemoyer, L.; and Levy, O. 2024 · 2024
Later among the works it cites.
Chemvlm: Exploring the power of multimodal large language models in chemistry area
Li, J.; Zhang, D.; Wang, X.; Hao, Z.; Lei, J.; Tan, Q.; Zhou, C.; Liu, W.; Yang, Y.; Xiong, X.; et al. 2025 · 2025
Closest in time.
Image-based generation for molecule design with SketchMol
Wang, Z.; Chen, Y.; Ma, P.; Yu, Z.; Wang, J.; Liu, Y.; Ye, X.; Sakurai, T.; and Zeng, X. 2025 · 2025
Closest in time.
Learning multi-view molecular representations with structured and unstructured knowledge
Luo, Y.; Yang, K.; Hong, M.; Liu, X. Y.; Nie, Z.; Zhou, H.; and Nie, Z. 2024 · 2093
Closest in time.