Fetching the paper…
Reading the bibliography…
While recent advancements in reasoning optimization have significantly enhanced the capabilities of large language models (LLMs), existing efforts to improve reasoning have been limited to solving mathematical problems and focusing on visual graphical inputs, neglecting broader applications in general video understanding.This paper proposes video-SALMONN-o1, the first open-source reasoning-enhanced audio-visual LLM designed for general video understanding tasks.
Librispeech: An ASR corpus based on public domain audio books
Panayotov, V., Chen, G., Povey, D., and Khudanpur, S · 2015
Earlier work this paper cites.
Gaussian error linear units (GELUs)
Hendrycks, D. and Gimpel, K · 2016
Earlier work this paper cites.
Audio visual scene-aware dialog
Alamri, H., Cartillier, V., Das, A., Wang, J., Cherian, A., Essa, I., Batra, D., Marks, T. K., Hori, C., and Anderson, P · 2019
Earlier work this paper cites.
AudioCaps: Generating captions for audios in the wild
Kim, C. D., Kim, B., Lee, H., and Kim, G · 2019
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Earlier work this paper cites.
Value: A multi-task benchmark for video-and-language understanding evaluation
Li, L., Lei, J., Gan, Z., Yu, L., Chen, Y.-C., Pillai, R., Cheng, Y., Zhou, L., Wang, X. E., Wang, W. Y., et al · 2021
Earlier work this paper cites.
Next-QA: Next phase of question-answering to explaining temporal actions
Xiao, J., Shang, X., Yao, A., and Chua, T.-S · 2021
Earlier work this paper cites.
Pano-AVQA: Grounded audio-visual question answering on 360deg videos
Yun, H., Yu, Y., Yang, W., Lee, K., and Kim, G · 2021
Earlier work this paper cites.
LoRA: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2022
Earlier work this paper cites.
Learning to answer questions in dynamic audio-visual scenarios
Li, G., Wei, Y., Tian, Y., Xu, C., Wen, J.-R., and Hu, D · 2022
Earlier work this paper cites.
Solving math word problems with process- and outcome-based feedback
Uesato, J., Kushman, N., Kumar, R., Song, F., Siegel, N., Wang, L., Creswell, A., Irving, G., and Higgins, I · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q., and Zhou, D · 2022
Earlier work this paper cites.
Reasoning with language model is planning with world model
Hao, S., Gu, Y., Ma, H., Hong, J. J., Wang, Z., Wang, D. Z., and Hu, Z · 2023
Earlier work this paper cites.
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K · 2023
Earlier work this paper cites.
Egoschema: A diagnostic benchmark for very long-form video language understanding
Mangalam, K., Akshulakov, R., and Malik, J · 2023
Earlier work this paper cites.
Video-bench: A comprehensive benchmark and toolkit for evaluating video-based large language models
Ning, M., Zhu, B., Xie, Y., Lin, B., Cui, J., Yuan, L., Chen, D., and Yuan, L · 2023
Cited alongside, same era.
Robust Speech Recognition via Large-scale Weak Supervision
Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I · 2023
Cited alongside, same era.
Tree of Thoughts: Deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., and Narasimhan, K · 2023
Cited alongside, same era.
Sigmoid loss for language image pre-training
Zhai, X., Mustafa, B., Kolesnikov, A., and Beyer, L · 2023
Cited alongside, same era.
M 3 AV: A multimodal, multigenre, and multipurpose audio-visual academic lecture dataset
Chen, Z., Liu, H., Yu, W., Sun, G., Liu, H., Wu, J., Zhang, C., Wang, Y., and Wang, Y · 2024
Improve mathematical reasoning in language models by automated process supervision
Luo, L., Liu, Y., Liu, R., Phatale, S., Guo, M., Lara, H., Li, Y., Shu, L., Zhu, Y., Meng, L., Sun, J., and Rastogi, A · 2024
Later among the works it cites.
Learning to reason with large language models, 2024
OpenAI · 2024
Later among the works it cites.
OpenAI Team · 2024
Later among the works it cites.
Direct Preference Optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2024
Later among the works it cites.
Scaling LLM test-time compute optimally can be more effective than scaling model parameters
Snell, C., Lee, J., Xu, K., and Kumar, A · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
VideoLLaMA 2: Advancing spatial-temporal modeling and audio understanding in Video-LLMs
Cheng, Z., Leng, S., Zhang, H., Xin, Y., Li, X., Chen, G., Zhu, Y., Zhang, W., Luo, Z., Zhao, D., and Bing, L · 2024
Cited alongside, same era.
Deepseek-r1-lite-preview is now live: unleashing supercharged reasoning power, 2024
DeepSeek Team · 2024
Cited alongside, same era.
CoT-ST: Enhancing LLM-based speech translation with multimodal chain-of-thought
Du, Y., Ma, Z., Yang, Y., Deng, K., Chen, X., Yang, B., Xiang, Y., Liu, M., and Qin, B · 2024
Cited alongside, same era.
MMBench-Video: A long-form multi-shot benchmark for holistic video understanding
Fang, X., Mao, K., Duan, H., Zhao, X., Li, Y., Lin, D., and Chen, K · 2024
Cited alongside, same era.
Alphazero-like tree-search can guide large language model decoding and training
Feng, X., Wan, Z., Wen, M., Wen, Y., Zhang, W., and Wang, J · 2024
Cited alongside, same era.
Video-MME: The first-ever comprehensive evaluation benchmark of multi-modal LLMs in video analysis
Fu, C., Dai, Y., Luo, Y., Li, L., Ren, S., Zhang, R., Wang, Z., Zhou, C., Shen, Y., Zhang, M., et al · 2024
Cited alongside, same era.
Think before you speak: Training language models with pause tokens
Goyal, S., Ji, Z., Rawat, A. S., Menon, A. K., Kumar, S., and Nagarajan, V · 2024
Cited alongside, same era.
Gemini: A family of highly capable multimodal models
Team, G., Anil, R., Borgeaud, S., Alayrac, J.-B., Yu, J., et al · 2024
Later among the works it cites.
FunQA: Towards surprising video comprehension
Xie, B., Zhang, S., Zhou, Z., Li, B., Zhang, Y., Hessel, J., Yang, J., and Liu, Z · 2024
Later among the works it cites.
LLaVA-CoT: Let vision language models reason step-by-step
Xu, G., Jin, P., Li, H., Song, Y., Sun, L., and Yuan, L · 2024
Later among the works it cites.
Qwen2.5-Math technical report: Toward mathematical expert model via self-improvement
Yang, A., Zhang, B., Hui, B., Gao, B., Yu, B., Li, C., Liu, D., Tu, J., Zhou, J., Lin, J., Lu, K., Xue, M., Lin, R., Liu, T., Ren, X., and Zhang, Z · 2024
Later among the works it cites.
InternLM-Math: Open math large language models toward verifiable reasoning
Ying, H., Zhang, S., Li, L., Zhou, Z., Shao, Y., Fei, Z., Ma, Y., Hong, J., Liu, K., Wang, Z., Wang, Y., Wu, Z., Li, S., Zhou, F., Liu, H., Zhang, S., Zhang, W., Yan, H., Qiu, X., Wang, J., Chen, K., and Lin, D · 2024
Later among the works it cites.
OVM, outcome-supervised value models for planning in mathematical reasoning
Yu, F., Gao, A., and Wang, B · 2024
Later among the works it cites.
Advancing LLM reasoning generalists with preference trees
Yuan, L., Cui, G., Wang, H., Ding, N., Wang, X., Deng, J., Shan, B., Chen, H., Xie, R., Lin, Y., Liu, Z., Zhou, B., Peng, H., Liu, Z., and Sun, M · 2024
Later among the works it cites.
Marco-o1: Towards open reasoning models for open-ended solutions
Zhao, Y., Yin, H., Zeng, B., Wang, H., Shi, T., Lyu, C., Wang, L., Luo, W., and Zhang, K · 2024
Later among the works it cites.
Virgo: A preliminary exploration on reproducing o1-like mllm
Du, Y., Liu, Z., Li, Y., Zhao, W. X., Huo, Y., Wang, B., Chen, W., Liu, Z., Wang, Z., and Wen, J.-R · 2025
Closest in time.