Fetching the paper…
Reading the bibliography…
Iterative self-improvement, a concept extending beyond personal growth, has found powerful applications in machine learning, particularly in transforming weak models into strong ones.
The architecture of complexity
Simon, H. A. 1962 · 1962
Earlier work this paper cites.
Contextual Prerequisites for Understanding: Some Investigations of Comprehension and Recall
Bransford, J. D.; and Johnson, M. K. 1972 · 1972
Earlier work this paper cites.
Schema Theory: An Introduction
Anderson, R. C. 1984 · 1984
Earlier work this paper cites.
The role of knowledge in discourse comprehension: A construction-integration model
Kintsch, W. 1988 · 1988
Earlier work this paper cites.
The strength of weak learnability
Schapire, R. E. 1990 · 1990
Earlier work this paper cites.
Mindset: The new psychology of success
Dweck, C. S. 2006 · 2006
Earlier work this paper cites.
Video Question Answering via Gradually Refined Attention over Appearance and Motion
Xu, D.; Zhao, Z.; Xiao, J.; Wu, F.; Zhang, H.; He, X.; and Zhuang, Y. 2017 · 2017
Earlier work this paper cites.
Learning Transferable Visual Models From Natural Language Supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021 · 2021
Earlier work this paper cites.
LoRA: Low-Rank Adaptation of Large Language Models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2022 · 2022
Earlier work this paper cites.
Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision
Burns, C.; Izmailov, P.; Kirchner, J. H.; Baker, B.; Gao, L.; Aschenbrenner, L.; Chen, Y.; Ecoffet, A.; Joglekar, M.; Leike, J.; Sutskever, I.; and Wu, J. 2023 · 2023
Earlier work this paper cites.
Jin, P.; Takanobu, R.; Zhang, C.; Cao, X.; and Yuan, L. 2023 · 2023
Cited alongside, same era.
RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
Lee, H.; Phatale, S.; Mansoor, H.; Mesnard, T.; Ferret, J.; Lu, K.; Bishop, C.; Hall, E.; Carbune, V.; Rastogi, A.; and Prakash, S. 2023 · 2023
Cited alongside, same era.
LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Li, Y.; Wang, C.; and Jia, J. 2023 · 2023
Cited alongside, same era.
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Lin, B.; Zhu, B.; Ye, Y.; Ning, M.; Jin, P.; and Yuan, L. 2023 · 2023
Cited alongside, same era.
Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback
Ahn, D.; Choi, Y.; Yu, Y.; Kang, D.; and Choi, J. 2024 · 2024
Closest in time.
Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
Chen, Z.; Deng, Y.; Yuan, H.; Ji, K.; and Gu, Q. 2024 · 2024
Closest in time.
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Li, K.; Wang, Y.; He, Y.; Li, Y.; Wang, Y.; Liu, Y.; Wang, Z.; Xu, J.; Chen, G.; Luo, P.; Wang, L.; and Qiao, Y. 2024 · 2024
Closest in time.
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Maaz, M.; Rasheed, H.; Khan, S.; and Khan, F. S. 2024 · 2024
Closest in time.
Iterative Reasoning Preference Optimization
Pang, R. Y.; Yuan, W.; Cho, K.; He, H.; Sukhbaatar, S.; and Weston, J. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Liu, R.; Li, C.; Tang, H.; Ge, Y.; Shan, Y.; and Li, G. 2023 · 2023
Cited alongside, same era.
Self-Refine: Iterative Refinement with Self-Feedback
Madaan, A.; Tandon, N.; Gupta, P.; Hallinan, S.; Gao, L.; Wiegreffe, S.; Alon, U.; Dziri, N.; Prabhumoye, S.; Yang, Y.; Gupta, S.; Majumder, B. P.; Hermann, K.; Welleck, S.; Yazdanbakhsh, A.; and Clark, P. 2023 · 2023
Cited alongside, same era.
A Long Way to Go: Investigating Length Correlations in RLHF
Prasann Singhal, J. X., Tanya Goyal; and Durrett, G. 2023 · 2023
Cited alongside, same era.
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Rafailov, R.; Sharma, A.; Mitchell, E.; Manning, C. D.; Ermon, S.; and Finn, C. 2023 · 2023
Cited alongside, same era.
Verbosity bias in preference labeling by large language models
Saito, K.; Wachi, A.; Wataoka, K.; and Akimoto, Y. 2023 · 2023
Cited alongside, same era.
LLaMA: Open and Efficient Foundation Language Models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; Rodriguez, A.; Joulin, A.; Grave, E.; and Lample, G. 2023a
Cited in the paper.
Llama 2: Open Foundation and Fine-Tuned Chat Models
Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; Bikel, D.; Blecher, L.; Ferrer, C. C.; Chen, M.; Cucurull, G.; Esiobu, D.; Fernandes, J.; Fu, J.; Fu, W.; Fuller, B.; Gao, C.; Goswami, V.; Goyal, N.; Hartshorn, A.; Hosseini, S.; Hou, R.; Inan, H.; Kardas, M.; Kerkez, V.; Khabsa, M.; Kloumann, I.; Korenev, A.; Koura, P. S.; Lachaux, M.-A.; Lavril, T.; Lee, J.; Liskovich, D.; Lu, Y.; Mao, Y.; Martinet, X.; Mihaylov, T.; Mishra, P.; Molybog, I.; Nie, Y.; Poulton, A.; Reizenstein, J.; Rungta, R.; Saladi, K.; Schelten, A.; Silva, R.; Smith, E. M.; Subramanian, R.; Tan, X. E.; Tang, B.; Taylor, R.; Williams, A.; Kuan, J. X.; Xu, P.; Yan, Z.; Zarov, I.; Zhang, Y.; Fan, A.; Kambadur, M.; Narang, S.; Rodriguez, A.; Stojnic, R.; Edunov, S.; and Scialom, T. 2023b
Cited in the paper.
Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward
Zhang, R.; Gui, L.; Sun, Z.; Feng, Y.; Xu, K.; Zhang, Y.; Fu, D.; Li, C.; Hauptmann, A.; Bisk, Y.; and Yang, Y. 2024a
Cited in the paper.
Park, R.; Rafailov, R.; Ermon, S.; and Finn, C. 2024 · 2024
Closest in time.
PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Xu, L.; Zhao, Y.; Zhou, D.; Lin, Z.; Ng, S. K.; and Feng, J. 2024 · 2024
Closest in time.
Self-Rewarding Language Models
Yuan, W.; Pang, R. Y.; Cho, K.; Li, X.; Sukhbaatar, S.; Xu, J.; and Weston, J. 2024 · 2024
Closest in time.
Calibrated Self-Rewarding Vision Language Models
Zhou, Y.; Fan, Z.; Cheng, D.; Yang, S.; Chen, Z.; Cui, C.; Wang, X.; Li, Y.; Zhang, L.; and Yao, H. 2024 · 2024
Closest in time.