Fetching the paper…
Reading the bibliography…
We present a perception in reflection paradigm designed to transcend the limitations of current large vision-language models (LVLMs), which are expected yet often fail to achieve perfect perception initially.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Wordnet: a lexical database for english
Miller, G. A · 1995
Earlier work this paper cites.
Learning to rank for information retrieval
Liu, T.-Y. et al · 2009
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G., and Bagnell, D · 2011
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Chen, X., Fang, H., Lin, T.-Y., Vedantam, R., Gupta, S., Dollár, P., and Zitnick, C. L · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Ren, S., He, K., Girshick, R., and Sun, J · 2016
Earlier work this paper cites.
Mask r-cnn
He, K., Gkioxari, G., Dollár, P., and Girshick, R · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Nocaps: Novel object captioning at scale
Agrawal, H., Desai, K., Wang, Y., Chen, X., Jain, R., Johnson, M., Batra, D., Parikh, D., Lee, S., and Anderson, P · 2019
Earlier work this paper cites.
Neural text generation with unlikelihood training
Welleck, S., Kulikov, I., Roller, S., Dinan, E., Cho, K., and Weston, J · 2019
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Emerging properties in self-supervised vision transformers
Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., and Joulin, A · 2021
Earlier work this paper cites.
Deep learning methods for abstract visual reasoning: A survey on raven’s progressive matrices
Małkiński, M. and Mańdziuk, J · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
Reasoning about actions over visual and linguistic modalities: A survey
Sampat, S. K., Patel, M., Das, S., Yang, Y., and Baral, C · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Earlier work this paper cites.
React: Synergizing reasoning and acting in language models
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y · 2022
Earlier work this paper cites.
Qwen-vl: A frontier large vision-language model with versatile abilities
Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., and Zhou, J · 2023
Earlier work this paper cites.
Improving image generation with better captions
Betker, J., Goh, G., Jing, L., Brooks, T., Wang, J., Li, L., Ouyang, L., Zhuang, J., Lee, J., Guo, Y., et al · 2023
Earlier work this paper cites.
Shikra: Unleashing multimodal llm’s referential dialogue magic
Chen, K., Zhang, Z., Zeng, W., Zhang, R., Zhu, F., and Zhao, R · 2023
Cited alongside, same era.
Vision transformers need registers, 2023
Darcet, T., Oquab, M., Mairal, J., and Bojanowski, P · 2023
Cited alongside, same era.
Volcano: mitigating multimodal hallucination through self-feedback guided revision
Lee, S., Park, S. H., Jo, Y., and Seo, M · 2023
Cited alongside, same era.
Selfcheck: Using llms to zero-shot check their own step-by-step reasoning
Miao, N., Teh, Y. W., and Rainforth, T · 2023
Cited alongside, same era.
Inverse reinforcement learning without reinforcement learning
Swamy, G., Wu, D., Choudhury, S., Bagnell, D., and Wu, S · 2023
Cited alongside, same era.
Can feedback enhance semantic grounding in large vision-language models?
Liao, Y.-H., Mahmood, R., Fidler, S., and Acuna, D · 2024
Later among the works it cites.
Rule based rewards for language model safety
Mu, T., Helyar, A., Heidecke, J., Achiam, J., Vallone, A., Kivlichan, I., Lin, M., Beutel, A., Schulman, J., and Weng, L · 2024
Later among the works it cites.
Dreambench++: A human-aligned benchmark for personalized image generation
Peng, Y., Cui, Y., Tang, H., Qi, Z., Dong, R., Bai, J., Han, C., Ge, Z., Zhang, X., and Xia, S.-T · 2024
Later among the works it cites.
Recursive introspection: Teaching language model agents how to self-improve
Qu, Y., Zhang, T., Garg, N., and Kumar, A · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
mplug-owl: Modularization empowers large language models with multimodality
Ye, Q., Xu, H., Xu, G., Ye, J., Yan, M., Zhou, Y., Wang, J., Hu, A., Shi, P., Shi, Y., et al · 2023
Cited alongside, same era.
Merlin: Empowering multimodal llms with foresight minds
Yu, E., Zhao, L., Wei, Y., Yang, J., Wu, D., Kong, L., Wei, H., Wang, T., Ge, Z., Zhang, X., et al · 2023
Cited alongside, same era.
Chatspot: Bootstrapping multimodal llms via precise referring instruction tuning
Zhao, L., Yu, E., Ge, Z., Yang, J., Wei, H., Zhou, H., Sun, J., Peng, Y., Dong, R., Han, C., et al · 2023
Cited alongside, same era.
Minigpt-4: Enhancing vision-language understanding with advanced large language models, 2023
Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M · 2023
Cited alongside, same era.
Dualfocus: Integrating macro and micro perspectives in multi-modal large language models
Cao, Y., Zhang, P., Dong, X., Lin, D., and Wang, J · 2024
Cited alongside, same era.
Onechart: Purify the chart structural extraction via one auxiliary token
Chen, J., Kong, L., Wei, H., Liu, C., Ge, Z., Zhao, L., Sun, J., Han, C., and Zhang, X · 2024
Cited alongside, same era.
Instructblip: Towards general-purpose vision-language models with instruction tuning
Dai, W., Li, J., Li, D., Tiong, A. M. H., Zhao, J., Wang, W., Li, B., Fung, P. N., and Hoi, S · 2024
Cited alongside, same era.
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2024
Later among the works it cites.
Visual cot: Unleashing chain-of-thought reasoning in multi-modal language models
Shao, H., Qian, S., Xiao, H., Song, G., Zong, Z., Wang, L., Liu, Y., and Li, H · 2024
Later among the works it cites.
Reflexion: Language agents with verbal reinforcement learning
Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K., and Yao, S · 2024
Later among the works it cites.
Understanding the performance gap between online and offline alignment algorithms
Tang, Y., Guo, D. Z., Zheng, Z., Calandriello, D., Cao, Y., Tarassov, E., Munos, R., Pires, B. Á., Valko, M., Cheng, Y., et al · 2024
Later among the works it cites.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
Wang, P., Bai, S., Tan, S., Wang, S., Fan, Z., Bai, J., Chen, K., Liu, X., Wang, J., Ge, W., et al · 2024
Later among the works it cites.
V?: Guided visual search as a core mechanism in multimodal llms
Wu, P. and Xie, S · 2024
Later among the works it cites.
Large multimodal agents: A survey
Xie, J., Chen, Z., Zhang, R., Wan, X., and Li, G · 2024
Later among the works it cites.
Llava-critic: Learning to evaluate multimodal models
Xiong, T., Wang, X., Guo, D., Ye, Q., Fan, H., Gu, Q., Huang, H., and Li, C · 2024
Later among the works it cites.
Restful-llama: Connecting user queries to restful apis
Xu, H., Zhao, R., Wang, J., and Chen, H · 2024
Later among the works it cites.
Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback
Yu, T., Yao, Y., Zhang, H., He, T., Han, Y., Cui, G., Hu, J., Liu, Z., Zheng, H.-T., Sun, M., et al · 2024
Later among the works it cites.
Llava-next: A strong zero-shot video understanding model, April 2024
Zhang, Y., Li, B., Liu, h., Lee, Y. j., Gui, L., Fu, D., Feng, J., Liu, Z., and Li, C · 2024
Later among the works it cites.
Self-supervised visual preference alignment
Zhu, K., Zhao, L., Ge, Z., and Zhang, X · 2024
Later among the works it cites.
Sharegpt4v: Improving large multi-modal models with better captions
Chen, L., Li, J., Dong, X., Zhang, P., He, C., Wang, J., Zhao, F., and Lin, D · 2025
Closest in time.
Unhackable temporal rewarding for scalable video mllms
Yu, E., Lin, K., Zhao, L., Wei, Y., Zhu, Z., Wei, H., Sun, J., Ge, Z., Zhang, X., Wang, J., et al · 2025
Closest in time.
Perpo: Perceptual preference optimization via discriminative rewarding
Zhu, Z., Zhao, L., Lin, K., Yang, J., Yu, E., Liu, C., Wei, H., Sun, J., Ge, Z., and Zhang, X · 2025
Closest in time.