Fetching the paper…
Reading the bibliography…
Recently, reinforcement learning (RL) has been shown to greatly enhance the reasoning capabilities of large language models (LLMs), and RL-based approaches have been progressively applied to visual multimodal tasks.
2017
Earlier work this paper cites.
2021
Earlier work this paper cites.
P. Yang, X. Wang, X. Duan, H. Chen, R. Hou, C. Jin, and W. Zhu, “AVQA: A dataset for audio-visual question answering on videos,” in Proceedings of the 30th ACM International Conference on Multimedia , ser. MM ’22, 2022, p. 3480–3491
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. K. Li, Y. Wu, and D. Guo, “Deepseekmath: Pushing the limits of mathematical reasoning in open language models,” 2402.03300 , 2024
2024
Earlier work this paper cites.
J. Kim, J. Jung, M. Jeon, S. H. Woo, and J. Lee, “Expanding on EnCLAP with auxiliary retrieval model for automated audio captioning,” DCASE2024 Challenge, Tech. Rep. 108, May 2024
2024
Earlier work this paper cites.
2024
Cited alongside, same era.
W. Chen, X. Li, Z. Ma, Y. Liang, A. Jiang, Z. Zheng, Y. Qian, P. Fan, W.-Q. Zhang, C. Lu, J. Liu, and X. Chen, “SJTU-THU Automated Audio Captioning System for DCASE 2024,” DCASE2024 Challenge, Tech. Rep. 2, May 2024
2024
Cited alongside, same era.
J. Liu, G. Li, C. Liu, J. Zhang, H. Dinkel, Y. Wang, Z. Yan, Y. Wang, and B. Wang, “Leveraging CED encoder and large language models for automated audio captioning,” DCASE2024 Challenge, Tech. Rep. 42, May 2024
2024
Cited alongside, same era.
J. Liu, G. Li, J. Zhang, H. Dinkel, Y. Wang, Z. Yan, Y. Wang, and B. Wang, “Enhancing automated audio captioning via large language models with optimized audio encoding,” in Interspeech 2024 , 2024, pp. 1135–1139
2024
Cited alongside, same era.
2025
Closest in time.
2025
Closest in time.
S. Ghosh, Z. Kong, S. Kumar, S. Sakshi, J. Kim, W. Ping, R. Valle, D. Manocha, and B. Catanzaro, “Audio Flamingo 2: An audio-language model with long-audio understanding and expert reasoning abilities,” 2025
2025
Closest in time.
2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
OpenAI, “Learning to reason with llms,” 2024. [Online]. Available: https://openai.com/index/learning-to-reason-with-llms/
2024
Cited alongside, same era.
2025
Cited alongside, same era.
L. Chen, L. Li, H. Zhao, Y. Song, and Vinci, “R1-V: Reinforcing super generalization ability in vision-language models with less than $3,” 2025, accessed: 2025-02-02. [Online]. Available: https://github.com/Deep-Agent/R1-V
2025
Cited alongside, same era.
2025
Closest in time.
W. Chen, Z. Ma, X. Li, X. Xu, Y. Liang, Z. Zheng, K. Yu, and X. Chen, “Slam-aac: Enhancing audio captioning with paraphrasing augmentation and clap-refine through llms,” in ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2025, pp. 1–5
2025
Closest in time.