Fetching the paper…
Reading the bibliography…
This paper presents Perceptual Preference Optimization (PerPO), a perception alignment method aimed at addressing the visual discrimination challenges in generative pre-trained multimodal large language models (MLLMs).
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W. J · 2002
Earlier work this paper cites.
Empirical risk minimization for support vector classifiers
Pérez-Cruz, F., Navia-Vázquez, Á., Figueiras-Vidal, A. R., and Artés-Rodríguez, A · 2003
Earlier work this paper cites.
Discriminative methods for multi-labeled classification
Godbole, S. and Sarawagi, S · 2004
Earlier work this paper cites.
On a method of empirical risk minimization
Golubev, G. K · 2004
Earlier work this paper cites.
Learning to rank using gradient descent
Burges, C., Shaked, T., Renshaw, E., Lazier, A., Deeds, M., Hamilton, N., and Hullender, G · 2005
Earlier work this paper cites.
Coarse-to-fine n-best parsing and maxent discriminative reranking
Charniak, E. and Johnson, M · 2005
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev, B · 2005
Earlier work this paper cites.
Learning to rank with nonsmooth cost functions
Burges, C. J. C., Ragno, R., and Le, Q. V · 2006
Earlier work this paper cites.
Time course of visual perception: coarse-to-fine processing and beyond
Hegdé, J · 2008
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
Mao, J., Huang, J., Toshev, A., Camburu, O., and Murphy, K · 2016
Earlier work this paper cites.
Modeling context in referring expressions
Yu, L., Poirson, P., Yang, S., Berg, A. C., and Berg, T. L · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Making the V in VQA matter: Elevating the role of image understanding in visual question answering
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D · 2017
Earlier work this paper cites.
Focal loss for dense object detection
Lin, T · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A · 2018
Earlier work this paper cites.
Planning and decision-making for autonomous vehicles
Schwarting, W., Alonso-Mora, J., and Rus, D · 2018
Earlier work this paper cites.
Learning discriminative model prediction for tracking
Bhat, G., Danelljan, M., Gool, L. V., and Timofte, R · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B · 2020
Earlier work this paper cites.
Computer vision for autonomous vehicles: Problems, datasets and state of the art
Janai, J., Güney, F., Behl, A., and Geiger, A · 2020
Earlier work this paper cites.
Learning to summarize from human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D. M., and Christiano, P · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Cited alongside, same era.
Webgpt: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Cited alongside, same era.
Decision-making and planning method for autonomous vehicles based on motivation and risk assessment
Wang, Y., Wang, C., Zhao, W., and Xu, C · 2021
Cited alongside, same era.
Deformable DETR: deformable transformers for end-to-end object detection
Feedback-generation for programming exercises with GPT-4
Azaiz, I., Kiesler, N., and Strickroth, S · 2024
Later among the works it cites.
A general theoretical paradigm to understand learning from human preferences
Azar, M. G., Guo, Z. D., Piot, B., Munos, R., Rowland, M., Valko, M., and Calandriello, D · 2024
Later among the works it cites.
Dreamllm: Synergistic multimodal comprehension and creation
Dong, R., Han, C., Peng, Y., Qi, Z., Ge, Z., Yang, J., Zhao, L., Sun, J., Zhou, H., Wei, H., Kong, X., Zhang, X., Ma, K., and Yi, L · 2024
Later among the works it cites.
Kto: Model alignment as prospect theoretic optimization
Ethayarajh, K., Xu, W., Muennighoff, N., Jurafsky, D., and Kiela, D · 2024
Later among the works it cites.
Deepseek-coder: When the large language model meets programming - the rise of code intelligence
Guo, D., Zhu, Q., Yang, D., Xie, Z., Dong, K., Zhang, W., Chen, G., Bi, X., Wu, Y., Li, Y. K., Luo, F., Xiong, Y., and Liang, W · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhu, X., Su, W., Lu, L., Li, B., Wang, X., and Dai, J · 2021
Cited alongside, same era.
Constitutional ai: Harmlessness from ai feedback
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., et al · 2022
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P. F., Leike, J., and Lowe, R · 2022
Cited alongside, same era.
Defining and characterizing reward hacking
Skalse, J., Howe, N. H. R., Krasheninnikov, D., and Krueger, D · 2022
Cited alongside, same era.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., et al · 2023
Cited alongside, same era.
Mathematical capabilities of chatgpt
Frieder, S., Pinchetti, L., Chevalier, A., Griffiths, R., Salvatori, T., Lukasiewicz, T., Petersen, P., and Berner, J · 2023
Cited alongside, same era.
Later among the works it cites.
Orpo: Monolithic preference optimization without reference model, 2024
Hong, J., Lee, N., and Thorne, J · 2024
Later among the works it cites.
sdpo: Don’t use your data all at once
Kim, D., Kim, Y., Song, W., Kim, H., Kim, Y., Kim, S., and Park, C · 2024
Later among the works it cites.
Llava-next: Improved reasoning, ocr, and world knowledge, January 2024b
Liu, H., Li, C., Li, Y., Li, B., Zhang, Y., Shen, S., and Lee, Y. J · 2024
Later among the works it cites.
Simpo: Simple preference optimization with a reference-free reward
Meng, Y., Xia, M., and Chen, D · 2024
Later among the works it cites.
Hello gpt-4o
OpenAI · 2024
Later among the works it cites.
Smaug: Fixing failure modes of preference optimisation with dpo-positive
Pal, A., Karkhanis, D., Dooley, S., Roberts, M., Naidu, S., and White, C · 2024
Later among the works it cites.
Rio: A benchmark for reasoning intention-oriented objects in open environments
Qu, M., Wu, Y., Liu, W., Liang, X., Song, J., Zhao, Y., and Wei, Y · 2024
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2024
Later among the works it cites.
The effect of sampling temperature on problem solving in large language models
Renze, M. and Guven, E · 2024
Later among the works it cites.
Math-llava: Bootstrapping mathematical reasoning for multimodal large language models
Shi, W., Hu, Z., Bin, Y., Liu, J., Yang, Y., Ng, S., Bing, L., and Lee, R. K · 2024
Later among the works it cites.
Preference ranking optimization for human alignment
Song, F., Yu, B., Li, M., Yu, H., Huang, F., Li, Y., and Wang, H · 2024
Later among the works it cites.
On the road with gpt-4v (ision): Explorations of utilizing visual-language model as autonomous driving agent
Wen, L., Yang, X., Fu, D., Wang, X., Cai, P., Li, X., Tao, M., Li, Y., Linran, X., Shang, D., et al · 2024
Later among the works it cites.
Knowledge-based and generative-ai-driven pedagogical conversational agents: A comparative study of grice’s cooperative principles and trust
Wölfel, M., Shirzad, M. B., Reich, A., and Anderer, K · 2024
Later among the works it cites.
A survey on multilingual large language models: Corpora, alignment, and bias
Xu, Y., Hu, L., Zhao, J., Qiu, Z., Ye, Y., and Gu, H · 2024
Later among the works it cites.
Yang, A., Yang, B., Hui, B., Zheng, B., Yu, B., Zhou, C., Li, C., Li, C., Liu, D., Huang, F., Dong, G., Wei, H., Lin, H., Tang, J., Wang, J., Yang, J., Tu, J., Zhang, J., Ma, J., Yang, J., Xu, J., Zhou, J., Bai, J., He, J., Lin, J., Dang, K., Lu, K., Chen, K., Yang, K., Li, M., Xue, M., Ni, N., Zhang, P., Wang, P., Peng, R., Men, R., Gao, R., Lin, R., Wang, S., Bai, S., Tan, S., Zhu, T., Li, T., Liu, T., Ge, W., Deng, X., Zhou, X., Ren, X., Zhang, X., Wei, X., Ren, X., Liu, X., Fan, Y., Yao, Y., Zhang, Y., Wan, Y., Chu, Y., Liu, Y., Cui, Z., Zhang, Z., Guo, Z., and Fan, Z · 2024
Later among the works it cites.
MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI
Yue, X., Ni, Y., Zheng, T., Zhang, K., Liu, R., Zhang, G., Stevens, S., Jiang, D., Ren, W., Sun, Y., Wei, C., Yu, B., Yuan, R., Sun, R., Yin, M., Zheng, B., Yang, Z., Liu, Y., Huang, W., Sun, H., Su, Y., and Chen, W · 2024
Later among the works it cites.
Self-supervised visual preference alignment
Zhu, K., Zhao, L., Ge, Z., and Zhang, X · 2024
Later among the works it cites.