Fetching the paper…
Reading the bibliography…
Large Vision Language Models (LVLMs) have shown remarkable capabilities in multimodal tasks like visual question answering or image captioning.
A mathematical theory of communication
Shannon, C. E. (1948) · 1948
Earlier work this paper cites.
Beam search
Bisiani, R. (1987) · 1987
Earlier work this paper cites.
Bleu: A method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J. (2002) · 2002
Earlier work this paper cites.
Meta-classification: Combining multimodal classifiers
Lin, W.-H. and Hauptmann, A. (2003) · 2003
Earlier work this paper cites.
Least angle regression
Efron, B., Hastie, T., Johnstone, I., and Tibshirani, R. (2004) · 2004
Earlier work this paper cites.
The relationship between precision-recall and roc curves
Davis, J. and Goadrich, M. (2006) · 2006
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with high levels of correlation with human judgments
Lavie, A. and Agarwal, A. (2007) · 2007
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L. (2014) · 2014
Earlier work this paper cites.
Obtaining well calibrated probabilities using bayesian binning
Pakdaman Naeini, M., Cooper, G., and Hauskrecht, M. (2015) · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Vedantam, R., Zitnick, C. L., and Parikh, D. (2015) · 2015
Earlier work this paper cites.
Spice: Semantic propositional image caption evaluation
Anderson, P., Fernando, B., Johnson, M., and Gould, S. (2016) · 2016
Earlier work this paper cites.
A baseline for detecting misclassified and out-of-distribution examples in neural networks
Hendrycks, D. and Gimpel, K. (2017) · 2017
Earlier work this paper cites.
Neural baby talk
Lu, J., Yang, J., Batra, D., and Parikh, D. (2018) · 2018
Earlier work this paper cites.
Object hallucination in image captioning
Rohrbach, A., Hendricks, L. A., Burns, K., Darrell, T., and Saenko, K. (2018) · 2018
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
Tibshirani, R. (2018) · 2018
Earlier work this paper cites.
Confidence scoring using whitebox meta-models with linear classifier probes
Chen, T., Navratil, J., Iyengar, V., and Shanmugam, K. (2019) · 2019
Earlier work this paper cites.
Uncertainty measures and prediction quality rating for the semantic segmentation of nested multi resolution street scene images
Rottmann, M. and Schubert, M. (2019) · 2019
Earlier work this paper cites.
Towards better confidence estimation for neural models
Vasudevan, V. T., Sethy, A., and Ghias, A. R. (2019) · 2019
Earlier work this paper cites.
The curious case of neural text degeneration
Holtzman, A., Buys, J., Du Li, Forbes, M., and Choi, Y. (2020) · 2020
Cited alongside, same era.
Yodar: Uncertainty-based sensor fusion for vehicle detection with camera and radar sensors
Kowol, K., Rottmann, M., Bracke, S., and Gottschalk, H. (2020) · 2020
Cited alongside, same era.
Time-dynamic estimates of the reliability of deep semantic segmentation networks
Maag, K., Rottmann, M., and Gottschalk, H. (2020) · 2020
Cited alongside, same era.
Prediction error meta classification in semantic segmentation: Detection via aggregated dispersion measures of softmax probabilities
Rottmann, M., Colling, P., Paul Hack, T., Chan, R., Hüger, F., Schlicht, P., and Gottschalk, H. (2020) · 2020
Cited alongside, same era.
Bdd100k: A diverse driving dataset for heterogeneous multitask learning
Yu, F., Chen, H., Wang, X., Xian, W., Chen, Y., Liu, F., Madhavan, V., and Darrell, T. (2020) · 2020
Cited alongside, same era.
Evaluation and analysis of hallucination in large vision-language models
Wang, J., Zhou, Y., Xu, G., Shi, P., Zhao, C., Xu, H., Ye, Q., Yan, M., Zhang, J., Zhu, J., Sang, J., and Tang, H. (2023) · 2023
Later among the works it cites.
mplug-owl: Modularization empowers large language models with multimodality
Ye, Q., Xu, H., Xu, G., Ye, J., Yan, M., Zhou, Y., Wang, J., Hu, A., Shi, P., Shi, Y., Li, C., Xu, Y., Chen, H., Tian, J., Qi, Q., Zhang, J., and Huang, F. (2023) · 2023
Later among the works it cites.
Woodpecker: Hallucination correction for multimodal large language models
Yin, S., Fu, C., Zhao, S., Xu, T., Wang, H., Sui, D., Shen, Y., Li, K., Sun, X., and Chen, E. (2023) · 2023
Later among the works it cites.
Analyzing and mitigating object hallucination in large vision-language models
Zhou, Y., Cui, C., Yoon, J., Zhang, L., Deng, Z., Finn, C., Bansal, M., and Yao, H. (2023) · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Clipscore: A reference-free evaluation metric for image captioning
Hessel, J., Holtzman, A., Forbes, M., Le Bras, R., and Choi, Y. (2021) · 2021
Cited alongside, same era.
Improving video instance segmentation by light-weight temporal uncertainty estimates
Maag, K., Rottmann, M., Varghese, S., Hüger, F., Schlicht, P., and Gottschalk, H. (2021) · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I. (2021) · 2021
Cited alongside, same era.
Metadetect: Uncertainty quantification and prediction quality estimates for object detection
Schubert, M., Kahl, K., and Rottmann, M. (2021) · 2021
Cited alongside, same era.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Li, J., LI, D., Xiong, C., and Hoi, S. (2022) · 2022
Cited alongside, same era.
A token-level reference-free hallucination detection benchmark for free-form text generation
Liu, T., Zhang, Y., Brockett, C., Mao, Y., Sui, Z., Chen, W., and Dolan, B. (2022) · 2022
Cited alongside, same era.
Temporal performance prediction for deep convolutional long short-term memory networks
Fieback, L., Dash, B., Spiegelberg, J., and Gottschalk, H. (2023) · 2023
Cited alongside, same era.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M. (2023) · 2023
Later among the works it cites.
Unified hallucination detection for multimodal large language models
Chen, X., Wang, C., Xue, Y., Zhang, N., Yang, X., Li, Q., Shen, Y., Liang, L., Gu, J., and Chen, H. (2024) · 2024
Closest in time.
A survey for foundation models in autonomous driving
Gao, H., Li, Y., Long, K., Yang, M., and Shen, Y. (2024) · 2024
Closest in time.
Conformal alignment: Knowing when to trust foundation models with guarantees
Gui, Y., Jin, Y., and Ren, Z. (2024) · 2024
Closest in time.
Evaluating general vision-language models for clinical medicine
Jiang, Y., Omiye, J. A., Zakka, C., Moor, M., Gui, H., Alipour, S., Mousavi, S. S., Chen, J. H., Rajpurkar, P., and Daneshjou, R. (2024) · 2024
Closest in time.
Mmbench: Is your multi-modal model an all-around player?
Liu, Y., Duan, H., Zhang, Y., Li, B., Zhang, S., Zhao, W., Yuan, Y., Wang, J., He, C., Liu, Z., Chen, K., and Lin, D. (2025) · 2024
Closest in time.
Negative object presence evaluation (NOPE) to measure object hallucination in vision-language models
Lovenia, H., Dai, W., Cahyawijaya, S., Ji, Z., and Fung, P. (2024) · 2024
Closest in time.
Simple token-level confidence improves caption correctness
Petryk, S., Whitehead, S., Gonzalez, J., Darrell, T., Rohrbach, A., and Rohrbach, M. (2023) · 2024
Closest in time.
Drivevlm: The convergence of autonomous driving and large vision-language models
Tian, X., Gu, J., Li, B., Liu, Y., Hu, C., Wang, Y., Zhan, K., Jia, P., Lang, X., and Zhao, H. (2024) · 2024
Closest in time.
Mitigating fine-grained hallucination by fine-tuning large vision-language models with caption rewrites
Wang, L., He, J., Li, S., Liu, N., and Lim, E.-P. (2024) · 2024
Closest in time.
Logical closed loop: Uncovering object hallucinations in large vision-language models
Wu, J., Liu, Q., Wang, D., Zhang, J., Wu, S., Wang, L., and Tan, T. (2024) · 2024
Closest in time.
Xing, S., Zhao, F., Wu, Z., An, T., Chen, W., Li, C., Zhang, J., and Dai, X. (2024) · 2024
Closest in time.
MM-vet: Evaluating large multimodal models for integrated capabilities
Yu, W., Yang, Z., Li, L., Wang, J., Lin, K., Liu, Z., Wang, X., and Wang, L. (2024) · 2024
Closest in time.
Mitigating object hallucination in large vision-language models via classifier-free guidance
Zhao, L., Deng, Y., Zhang, W., and Gu, Q. (2024) · 2024
Closest in time.