Fetching the paper…
Reading the bibliography…
The advent of large vision-language models (LVLMs) represents a remarkable advance in the quest for artificial general intelligence.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., et al. (2020) · 1901
Earlier work this paper cites.
Frequency-tuned salient region detection
Achanta, R., Hemami, S., Estrada, F., Susstrunk, S. (2009) · 2009
Earlier work this paper cites.
Saliency filters: Contrast based filtering for salient region detection
Perazzi, F., Krähenbühl, P., Pritch, Y., Hornung, A. (2012) · 2012
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Bourdev, L., Girshick, R., Hays, J., et al. (2014) · 2014
Earlier work this paper cites.
Toward embedded detection of polyps in wce images for early diagnosis of colorectal cancer
Silva, J., Histace, A., Romain, O., Dray, X., Granado, B. (2014) · 2014
Earlier work this paper cites.
Automated polyp detection in colonoscopy videos using shape and context information
Tajbakhsh, N., Gurudu, S. R., Liang, J. (2015) · 2015
Earlier work this paper cites.
Skin lesion analysis toward melanoma detection: A challenge at the 2017 international symposium on biomedical imaging (isbi), hosted by the international skin imaging collaboration (isic)
Codella, N. C., Gutman, D., Celebi, M. E., Helba, B., Marchetti, M. A., Dusza, S. W., et al. (2018) · 2017
Earlier work this paper cites.
Learning to detect salient objects with image-level supervision
Wang, L., Lu, H., Wang, Y., Feng, M., Wang, D., Yin, B., et al. (2017) · 2017
Earlier work this paper cites.
Structure-measure: A new way to evaluate foreground maps
Fan, D.-P., Cheng, M.-M., Liu, Y., Li, T., Borji, A. (2017) · 2017
Earlier work this paper cites.
Salient objects in clutter: Bringing salient object detection to the foreground
Fan, D.-P., Cheng, M.-M., Liu, J.-J., Gao, S.-H., Hou, Q., Borji, A. (2018) · 2018
Earlier work this paper cites.
Deepside: A general deep framework for salient object detection
Fu, K., Zhao, Q., Gu, I. Y.-H., Yang, J. (2019) · 2019
Earlier work this paper cites.
Reconstruct and represent video contents for captioning via reinforcement learning
Zhang, W., Wang, B., Ma, L., Liu, W. (2019) · 2019
Earlier work this paper cites.
Jl-dcf: Joint learning and densely-cooperative fusion framework for rgb-d salient object detection
Fu, K., Fan, D.-P., Ji, G.-P., Zhao, Q. (2020) · 2020
Earlier work this paper cites.
Segmenting transparent objects in the wild
Xie, E., Wang, W., Wang, W., Ding, M., Shen, C., Luo, P. (2020) · 2020
Earlier work this paper cites.
An improved deep learning approach and its applications on colonic polyp images detection
Wang, W., Tian, J., Zhang, C., Luo, Y., Wang, X., Li, J. (2020) · 2020
Earlier work this paper cites.
Siamese network for rgb-d salient object detection and beyond
Fu, K., Fan, D.-P., Ji, G.-P., Zhao, Q., Shen, J., Zhu, C. (2021) · 2021
Earlier work this paper cites.
Concealed object detection
Fan, D.-P., Ji, G.-P., Cheng, M.-M., Shao, L. (2021) · 2021
Earlier work this paper cites.
The mvtec anomaly detection dataset: a comprehensive real-world dataset for unsupervised anomaly detection
Bergmann, P., Batzner, K., Fauser, M., Sattlegger, D., Steger, C. (2021) · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., et al. (2021) · 2021
Cited alongside, same era.
A comparative analysis of object detection metrics with a companion open-source toolkit
Padilla, R., Passos, W. L., Dias, T. L., Netto, S. L., Da Silva, E. A. (2021) · 2021
Cited alongside, same era.
Specificity-preserving rgb-d saliency detection
Zhou, T., Fu, H., Chen, G., Zhou, Y., Fan, D.-P., Shao, L. (2021) · 2021
Cited alongside, same era.
Rgb-d salient object detection via 3d convolutional neural networks
Chen, Q., Liu, Z., Zhang, Y., Fu, K., Zhao, Q., Du, H. (2021) · 2021
Cited alongside, same era.
Large ai models in health informatics: Applications, challenges, and the future
Qiu, J., Li, L., Sun, J., Peng, J., Shi, P., Zhang, R., et al. (2023) · 2023
Later among the works it cites.
Fu, C., Chen, P., Shen, Y., Qin, Y., Zhang, M., Lin, X., et al. (2023a) · 2023
Later among the works it cites.
3d visual saliency: an independent perceptual measure or a derivative of 2d image saliency?
Song, R., Zhang, W., Zhao, Y., Liu, Y., Rosin, P. L. (2023) · 2023
Later among the works it cites.
Vocabulary-free image classification
Conti, A., Fini, E., Mancini, M., Rota, P., Wang, Y., Ricci, E. (2023) · 2023
Later among the works it cites.
Lvlm-ehub: A comprehensive evaluation benchmark for large vision-language models
Xu, P., Shao, W., Zhang, K., Gao, P., Liu, S., Lei, M., et al. (2023) · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Depth quality-inspired feature manipulation for efficient rgb-d salient object detection
Zhang, W., Ji, G.-P., Wang, Z., Fu, K., Zhao, Q. (2021) · 2021
Cited alongside, same era.
Fast camouflaged object detection via edge-based reversible re-calibration network
Ji, G.-P., Zhu, L., Zhuge, M., Fu, K. (2022) · 2022
Cited alongside, same era.
Spot-the-difference self-supervised pre-training for anomaly detection and segmentation
Zou, Y., Jeong, J., Pemula, L., Zhang, D., Dabeer, O. (2022) · 2022
Cited alongside, same era.
Light field salient object detection: A review and benchmark
Fu, K., Jiang, Y., Ji, G.-P., Zhou, T., Zhao, Q., Fan, D.-P. (2022) · 2022
Cited alongside, same era.
Rgb-d salient object detection of using few-shot learning
He, J., Fu, K. (2022) · 2022
Cited alongside, same era.
Language models with image descriptors are strong few-shot video-language learners
Wang, Z., Li, M., Xu, R., Zhou, L., Lei, J., Lin, X., et al. (2022) · 2022
Cited alongside, same era.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., et al. (2023) · 2023
Cited alongside, same era.
Later among the works it cites.
Holistic analysis of hallucination in gpt-4v (ision): Bias and interference challenges
Cui, C., Zhou, Y., Yang, X., Wu, S., Zhang, L., Zou, J., et al. (2023) · 2023
Later among the works it cites.
Segment anything
Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., et al. (2023) · 2023
Later among the works it cites.
Analyzing and mitigating object hallucination in large vision-language models
Zhou, Y., Cui, C., Yoon, J., Zhang, L., Deng, Z., Finn, C., et al. (2023) · 2023
Later among the works it cites.
Exposing and mitigating spurious correlations for cross-modal retrieval
Kim, J. M., Koepke, A., Schmid, C., Akata, Z. (2023) · 2023
Later among the works it cites.
Improving cross-task generalization with step-by-step instructions
Wu, Y., Zhao, Y., Li, Z., Qin, B., Xiong, K. (2023) · 2023
Later among the works it cites.
Deepspeed-visualchat: Multi-round multi-image interleave chat via multi-modal causal attention
Yao, Z., Wu, X., Li, C., Zhang, M., Qi, H., Ruwase, O., et al. (2023) · 2023
Later among the works it cites.
How easy is it to fool your multimodal llms? an empirical analysis on deceptive prompts
Qian, Y., Zhang, H., Yang, Y., Gan, Z. (2024) · 2024
Closest in time.
Vigor: Improving visual grounding of large vision language models with fine-grained reward modeling
Yan, S., Bai, M., Chen, W., Zhou, X., Huang, Q., Li, L. E. (2024) · 2024
Closest in time.
Enhancing multimodal large language models with vision detection models: An empirical study
Jiao, Q., Chen, D., Huang, Y., Li, Y., Shen, Y. (2024) · 2024
Closest in time.
Zhong, L., Liao, X., Zhang, S., Zhang, X., Wang, G. (2024) · 2024
Closest in time.
Decoupling static and hierarchical motion perception for referring video segmentation
He, S., Ding, H. (2024) · 2024
Closest in time.