Fetching the paper…
Reading the bibliography…
Diagnosis-Related Group (DRG) codes are essential for hospital reimbursement and operations but require labor-intensive assignment.
After the revolution: Drgs at age 30
K. Quinn · 2014
Earlier work this paper cites.
Icd-10-cm/pcs ms-drg v34. 0 definitions manual
CMS · 2016
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
J. Rasley, S. Rajbhandari, O. Ruwase, and Y. He · 2020
Earlier work this paper cites.
Trl: Transformer reinforcement learning
L. von Werra, Y. Belkada, L. Tunstall, E. Beeching, T. Thrush, N. Lambert, S. Huang, K. Rasul, and Q. Gallouédec · 2020
Earlier work this paper cites.
Early prediction of diagnostic-related groups and estimation of hospital cost by processing clinical notes
J. Liu, D. Capurro, A. Nguyen, and K. Verspoor · 2021
Earlier work this paper cites.
Automated clinical coding: what, why, and where we are?
H. Dong, M. Falis, W. Whiteley, B. Alex, J. Matterson, S. Ji, J. Chen, and H. Wu · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Earlier work this paper cites.
Drgcoder: Explainable clinical coding for the early prediction of diagnostic-related groups
D. Hajialigol, D. Kaknes, T. Barbour, D. Yao, C. North, J. Sun, D. Liem, and X. Wang · 2023
Earlier work this paper cites.
Mimic-iv, a freely accessible electronic health record dataset
A. E. Johnson, L. Bulgarelli, L. Shen, A. Gayles, A. Shammout, S. Horng, T. J. Pollard, S. Hao, B. Moody, B. Gow, et al · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica · 2023
Earlier work this paper cites.
Surpassing gpt-4 medical coding with a two-stage approach
Z. Yang, S. S. Batra, J. Stremmel, and E. Halperin · 2023
Earlier work this paper cites.
Huatuogpt-o1, towards medical complex reasoning with llms
J. Chen, Z. Cai, K. Ji, X. Wang, W. Liu, R. Wang, J. Hou, and B. Wang · 2024
Earlier work this paper cites.
A. Jaech, A. Kalai, A. Lerer, A. Richardson, A. El-Kishky, A. Low, A. Helyar, A. Madry, A. Beutel, A. Carney, et al · 2024
Earlier work this paper cites.
Exploring llm multi-agents for icd coding
R. Li, X. Wang, and H. Yu · 2024
Earlier work this paper cites.
Rule based rewards for language model safety
T. Mu, A. Helyar, J. Heidecke, J. Achiam, A. Vallone, I. Kivlichan, M. Lin, A. Beutel, J. Schulman, and L. Weng · 2024
Cited alongside, same era.
Open r1 update 3: Steady progress and a new technical report, 2024
G. Penedo, L. Tunstall, A. Lozhkov, H. Kydlicek, E. Beeching, L. B. Allal, Q. Gallouédec, L. von Werra, A. P. Lajarín, and N. Habib · 2024
Cited alongside, same era.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. Li, Y. Wu, et al · 2024
Cited alongside, same era.
Large language models are poor medical coders—benchmarking of medical code querying
A. Soroush, B. S. Glicksberg, E. Zimlichman, Y. Barash, R. Freeman, A. W. Charney, G. N. Nadkarni, and E. Klang · 2024
Cited alongside, same era.
Dr-llava: Visual instruction tuning with symbolic clinical grounding
S. Sun, A. Schubert, G. M. Goldgof, Z. Sun, T. Hartvigsen, A. J. Butte, and A. Alaa · 2024
Y. Ji, S. Zhao, X. Tian, H. Wang, S. Chen, Y. Peng, H. Zhao, and X. Li · 2025
Closest in time.
Med-r1: Reinforcement learning for generalizable medical reasoning in vision-language models
Y. Lai, J. Zhong, M. Li, S. Zhao, and X. Yang · 2025
Closest in time.
W. Lan, W. Wang, C. Ji, G. Yang, Y. Zhang, X. Liu, S. Wu, and G. Wang · 2025
Closest in time.
Cppo: Accelerating the training of group relative policy optimization-based reasoning models
Z. Lin, M. Lin, Y. Xie, and R. Ji · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Drg-llama: tuning llama model to predict diagnosis-related group for hospitalized patients
H. Wang, C. Gao, C. Dantona, B. Hull, and J. Sun · 2024
Cited alongside, same era.
A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, et al · 2024
Cited alongside, same era.
Online difficulty filtering for reasoning oriented reinforcement learning
S. Bae, J. Hong, M. Y. Lee, H. Kim, J. Nam, and D. Kwak · 2025
Cited alongside, same era.
Reasoning models don’t always say what they think
Y. Chen, J. Benton, A. Radhakrishnan, J. Uesato, C. Denison, J. Schulman, A. Somani, P. Hase, M. Wagner, F. Roger, et al · 2025
Cited alongside, same era.
Gpg: A simple and strong reinforcement learning baseline for model reasoning
X. Chu, H. Huang, X. Zhang, F. Wei, and Y. Wang · 2025
Cited alongside, same era.
Open r1: A fully open reproduction of deepseek-r1
H. Face · 2025
Cited alongside, same era.
Can large language models replace coding specialists? evaluating gpt performance in medical coding tasks
Y. Feng · 2025
Cited alongside, same era.
Z. Liu, C. Chen, W. Li, P. Qi, T. Pang, C. Du, W. S. Lee, and M. Lin · 2025
Closest in time.
Learning what reinforcement learning can’t: Interleaved online fine-tuning for hardest questions
L. Ma, H. Liang, M. Qiang, L. Tang, X. Ma, Z. H. Wong, J. Niu, C. Shen, R. He, B. Cui, et al · 2025
Closest in time.
Reasoning models can be effective without thinking
W. Ma, J. He, C. Snell, T. Griggs, S. Min, and M. Zaharia · 2025
Closest in time.
Kimi k1. 5: Scaling reinforcement learning with llms
K. Team, A. Du, B. Gao, B. Xing, C. Jiang, C. Chen, C. Li, C. Xiao, C. Du, C. Liao, et al · 2025
Closest in time.
Logic-rl: Unleashing llm reasoning with rule-based reinforcement learning
T. Xie, Z. Gao, Q. Ren, H. Luo, Y. Hong, B. Dai, J. Zhou, K. Qiu, Z. Wu, and C. Luo · 2025
Closest in time.
A minimalist approach to llm reasoning: from rejection sampling to reinforce
W. Xiong, J. Yao, Y. Xu, B. Pang, L. Wang, D. Sahoo, J. Li, N. Jiang, T. Zhang, C. Xiong, et al · 2025
Closest in time.
Learning to reason under off-policy guidance
J. Yan, Y. Li, Z. Hu, Z. Wang, G. Cui, X. Qu, Y. Cheng, and Y. Zhang · 2025
Closest in time.
Dapo: An open-source llm reinforcement learning system at scale
Q. Yu, Z. Zhang, R. Zhu, Y. Yuan, X. Zuo, Y. Yue, T. Fan, G. Liu, L. Liu, X. Liu, et al · 2025
Closest in time.
Does reinforcement learning really incentivize reasoning capacity in llms beyond the base model?
Y. Yue, Z. Chen, R. Lu, A. Zhao, Z. Wang, S. Song, and G. Huang · 2025
Closest in time.
Simplerl-zoo: Investigating and taming zero reinforcement learning for open base models in the wild
W. Zeng, Y. Huang, Q. Liu, W. Liu, K. He, Z. Ma, and J. He · 2025
Closest in time.
Srpo: A cross-domain implementation of large-scale reinforcement learning on llm
X. Zhang, J. Wang, Z. Cheng, W. Zhuang, Z. Lin, M. Zhang, S. Wang, Y. Cui, C. Wang, J. Peng, et al · 2025
Closest in time.