Fetching the paper…
Reading the bibliography…
We present a comprehensive framework for enhancing Retrieval-Augmented Generation (RAG) systems through dynamic retrieval strategies and reinforcement fine-tuning.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Yang, Z., Qi, P., Zhang, S., Bengio, Y., Cohen, W. W., Salakhutdinov, R., and Manning, C. D · 2018
Earlier work this paper cites.
Pubmedqa: A dataset for biomedical research question answering
Jin, Q., Dhingra, B., Liu, Z., Cohen, W. W., and Lu, X · 2019
Earlier work this paper cites.
Leveraging passage retrieval with generative models for open domain question answering
Izacard, G. and Grave, E · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., et al · 2020
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T., Fu, D., Ermon, S., Rudra, A., and Ré, C · 2022
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdhery, A., and Zhou, D · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Earlier work this paper cites.
Accelerating large language model decoding with speculative sampling
Chen, C., Borgeaud, S., Irving, G., Lespiau, J.-B., Sifre, L., and Jumper, J · 2023
Earlier work this paper cites.
Flashattention-2: Faster attention with better parallelism and work partitioning
Dao, T · 2023
Earlier work this paper cites.
Alphazero-like tree-search can guide large language model decoding and training
Feng, X., Wan, Z., Wen, M., McAleer, S. M., Wen, Y., Zhang, W., and Wang, J · 2023
Earlier work this paper cites.
Fast inference from transformers via speculative decoding
Leviathan, Y., Kalman, M., and Matias, Y · 2023
Earlier work this paper cites.
Liu, X., Hu, L., Bailis, P., Cheung, A., Deng, Z., Stoica, I., and Zhang, H · 2023
Earlier work this paper cites.
Don’t do rag: When cache-augmented generation is all you need for knowledge tasks
Chan, B. J., Chen, C.-T., Cheng, J.-H., and Huang, H.-H · 2024
Earlier work this paper cites.
A simple and provable scaling law for the test-time compute of large language models
Chen, Y., Pan, X., Li, Y., Ding, B., and Zhou, J · 2024
Earlier work this paper cites.
Inference-aware fine-tuning for best-of-n sampling in large language models
Chow, Y., Tennenholtz, G., Gur, I., Zhuang, V., Dai, B., Thiagarajan, S., Boutilier, C., Agarwal, R., Kumar, A., and Faust, A · 2024
Earlier work this paper cites.
Finch: Prompt-guided key-value cache compression for large language models
Corallo, G. and Papotti, P · 2024
Earlier work this paper cites.
Entropy guided extrapolative decoding to improve factuality in large language models
Das, S., Jin, L., Song, L., Mi, H., Peng, B., and Yu, D · 2024
Earlier work this paper cites.
A simple and effective l _ 2 l\_2 norm-based strategy for kv cache compression
Devoto, A., Zhao, Y., Scardapane, S., and Minervini, P · 2024
Cited alongside, same era.
Ada-kv: Optimizing kv cache eviction by adaptive budget allocation for efficient llm inference
Feng, Y., Lv, J., Cao, Y., Xie, X., and Zhou, S. K · 2024
Cited alongside, same era.
Break the sequential dependency of llm inference using lookahead decoding
Fu, Y., Bailis, P., Stoica, I., and Zhang, H · 2024
Cited alongside, same era.
Interpretable contrastive monte carlo tree search reasoning
Gao, Z., Niu, B., He, X., Xu, H., Liu, H., Liu, A., Hu, X., and Wen, L · 2024
Cited alongside, same era.
Technical report: Enhancing llm reasoning with reward-guided tree search
Marco-o1: Towards open reasoning models for open-ended solutions
Zhao, Y., Yin, H., Zeng, B., Wang, H., Shi, T., Lyu, C., Wang, L., Luo, W., and Zhang, K · 2024
Later among the works it cites.
Scaling up test-time compute with latent reasoning: A recurrent depth approach
Geiping, J., McLeish, S., Jain, N., Kirchenbauer, J., Singh, S., Bartoldson, B. R., Kailkhura, B., Bhatele, A., and Goldstein, T · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al · 2025
Closest in time.
Kvquant: Towards 10 million context length llm inference with kv cache quantization
Hooper, C., Kim, S., Mohammadzadeh, H., Mahoney, M. W., Shao, S., Keutzer, K., and Gholami, A · 2025
Closest in time.
Test-time computing: from system-1 thinking to system-2 thinking
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jiang, J., Chen, Z., Min, Y., Chen, J., Cheng, X., Wang, J., Tang, Y., Sun, H., Deng, J., Zhao, W. X., et al · 2024
Cited alongside, same era.
Liu, A., Feng, B., Xue, B., Wang, B., Wu, B., Lu, C., Zhao, C., Deng, C., Zhang, C., Ruan, C., et al · 2024
Cited alongside, same era.
Gorilla: Large language model connected with massive apis
Patil, S. G., Zhang, T., Wang, X., and Gonzalez, J. E · 2024
Cited alongside, same era.
Mutual reasoning makes smaller llms stronger problem-solvers
Qi, Z., Ma, M., Xu, J., Zhang, L. L., Yang, F., and Yang, M · 2024
Cited alongside, same era.
Memorag: Moving towards next-gen rag via memory-inspired knowledge discovery
Qian, H., Zhang, P., Liu, Z., Mao, K., and Dou, Z · 2024
Cited alongside, same era.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Bi, X., Zhang, H., Zhang, M., Li, Y., Wu, Y., et al · 2024
Cited alongside, same era.
Su, W., Tang, Y., Ai, Q., Wu, Z., and Liu, Y · 2024
Cited alongside, same era.
Dawn-icl: Strategic planning of problem-solving trajectories for zero-shot in-context learning
Tang, X., Wang, X., Zhao, W. X., and Wen, J.-R · 2024
Cited alongside, same era.
Ji, Y., Li, J., Ye, H., Wu, K., Xu, J., Mo, L., and Zhang, M · 2025
Closest in time.
Snapkv: Llm knows what you are looking for before generation
Li, Y., Huang, Y., Yang, B., Venkitesh, B., Locatelli, A., Ye, H., Cai, T., Lewis, P., and Chen, D · 2025
Closest in time.
Qlass: Boosting language agent inference via q-guided stepwise search
Lin, Z., Tang, Y., Yao, X., Yin, D., Hu, Z., Sun, Y., and Chang, K.-W · 2025
Closest in time.
Can 1b llm surpass 405b llm? rethinking compute-optimal test-time scaling
Liu, R., Gao, J., Zhao, J., Zhang, K., Li, X., Qi, B., Ouyang, W., and Zhou, B · 2025
Closest in time.
Muennighoff, N., Yang, Z., Shi, W., Li, X. L., Fei-Fei, L., Hajishirzi, H., Zettlemoyer, L., Liang, P., Candès, E., and Hashimoto, T · 2025
Closest in time.
Entropy adaptive decoding: Dynamic model switching for efficient inference
Simonds, T · 2025
Closest in time.
Parametric retrieval augmented generation
Su, W., Tang, Y., Ai, Q., Yan, J., Wang, C., Wang, H., Ye, Z., Zhou, Y., and Liu, Y · 2025
Closest in time.
Chain-of-retrieval augmented generation
Wang, L., Chen, H., Yang, N., Huang, X., Dou, Z., and Wei, F · 2025
Closest in time.
Boosting multimodal reasoning with mcts-automated structured thinking
Wu, J., Feng, M., Zhang, S., Jin, R., Che, F., Wen, Z., and Tao, J · 2025
Closest in time.
Kvlink: Accelerating large language models via efficient kv cache reuse
Yang, J., Hou, B., Wei, W., Bao, Y., and Chang, S · 2025
Closest in time.
Monte carlo tree diffusion for system 2 planning
Yoon, J., Cho, H., Baek, D., Bengio, Y., and Ahn, S · 2025
Closest in time.
Generating symbolic world models via test-time scaling of large language models
Yu, Z., Yuan, Y., Xiao, T. Z., Xia, F. F., Fu, J., Zhang, G., Lin, G., and Liu, W · 2025
Closest in time.
Zeng, Z., Cheng, Q., Yin, Z., Zhou, Y., and Qiu, X · 2025
Closest in time.