Fetching the paper…
Reading the bibliography…
While Large Language Models (LLMs) excel in world knowledge understanding, adapting them to specific subfields requires precise adjustments.
Vaswani, A., Shazeer, N., Parmar, et al.: Attention is all you need. NIPS 30
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
Lau, J.J., Gayen, et al.: A dataset of clinically generated visual questions and answers about radiology images. Scientific data 5
2018
Earlier work this paper cites.
Houlsby, N., Giurgiu, A., Jastrzebski, et al.: Parameter-efficient transfer learning for nlp. In: International Conference on Machine Learning. pp. 2790–2799. PMLR (2019)
2019
Earlier work this paper cites.
Brown, T., Mann, B., Ryder, et al.: Language models are few-shot learners. Advances in neural information processing systems 33
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
Liu, B., Zhan, L.M., Xu, et al.: Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering. In: 2021 ISBI. pp. 1650–1654 (2021)
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Radford, A., Kim, J.W., Hallacy, et al.: Learning transferable visual models from natural language supervision. In: International conference on machine learning. pp. 8748–8763. PMLR (2021)
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Li, P., Liu, G., He, et al.: Masked vision and language pre-training with unimodal and multimodal contrastive losses for medical visual question answering. In: MICCAI. pp. 374–383. Springer (2023)
2023
Later among the works it cites.
Liu, H., Li, C., Wu, et al.: Visual instruction tuning. arXiv preprint arXiv:2304.08485 (2023)
2023
Later among the works it cites.
Liu, X., Zheng, Y., Du, Z., Ding, et al.: Gpt understands, too. AI Open (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chen, Z., Du, Y., Hu, et al.: Multi-modal masked autoencoders for medical vision-and-language pre-training. In: MICCAI. pp. 679–689. Springer (2022)
2022
Cited alongside, same era.
Cong, F., Xu, S., et al.: Caption-aware medical vqa via semantic focusing and progressive cross-modality comprehension. In: ACM MM. pp. 3569–3577 (2022)
2022
Cited alongside, same era.
Li, J., Li, D., Xiong, et al.: Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In: ICCV. pp. 12888–12900 (2022)
2022
Cited alongside, same era.
Liu, H., Tam, D., Muqeeth, et al.: Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning. Advances in Neural Information Processing Systems 35
2022
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Later among the works it cites.
2024
Closest in time.
2024
Closest in time.