Fetching the paper…
Reading the bibliography…
A brief, fluent, and relevant summary can be helpful during program comprehension; however, such a summary does require significant human effort to produce.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Verification of forecasts expressed in terms of probability
Glenn W Brier. 1950 · 1950
Earlier work this paper cites.
On the Combination of Forecast Probabilities for Consecutive Precipitation Periods. In Weather and Forecasting , Vol. 5. 640––650
Daniel S. Wilks. 1990 · 1990
Earlier work this paper cites.
Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods
John Platt et al · 1999
Earlier work this paper cites.
Obtaining calibrated probability estimates from decision trees and naive bayesian classifiers. In Icml , Vol. 1. 609–616
Bianca Zadrozny and Charles Elkan. 2001 · 2001
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics . 311–318
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Transforming classifier scores into accurate multiclass probability estimates. In Proceedings of the eighth ACM SIGKDD international conference on Knowledge discovery and data mining . 694–699
Bianca Zadrozny and Charles Elkan. 2002 · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries. In Text summarization branches out . 74–81
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization . 65–72
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Do Code and Comments Co-Evolve? On the Relation between Source Code and Comment Changes. In 14th Working Conference on Reverse Engineering (WCRE 2007) . 70–79
Beat Fluri, Michael Wursch, and Harald C. Gall. 2007 · 2007
Earlier work this paper cites.
The probabilistic relevance framework: BM25 and beyond
Stephen Robertson, Hugo Zaragoza, et al · 2009
Earlier work this paper cites.
Towards automatically generating summary comments for java methods. In Proceedings of the 25th IEEE/ACM international conference on Automated software engineering . 43–52
Giriprasad Sridhara, Emily Hill, Divya Muppaneni, Lori Pollock, and K Vijay-Shanker. 2010 · 2010
Earlier work this paper cites.
Nonparametric Statistical Methods (3rd ed.)
Myles Hollander, Douglas A. Wolfe, and Eric Chicken. 2013 · 2013
Earlier work this paper cites.
Obtaining well calibrated probabilities using bayesian binning. In Proceedings of the AAAI conference on artificial intelligence , Vol. 29
Mahdi Pakdaman Naeini, Gregory Cooper, and Milos Hauskrecht. 2015 · 2015
Earlier work this paper cites.
An Introduction to Statistical Methods and Data Analysis
R.L. Ott and M.T. Longnecker. 2015 · 2015
Earlier work this paper cites.
On calibration of modern neural networks. In International conference on machine learning . PMLR, 1321–1330
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger. 2017 · 2017
Earlier work this paper cites.
Analyzing Uncertainty in Neural Machine Translation
Myle Ott, Michael Auli, David Grangier, and Marc’Aurelio Ranzato. 2018 · 2018
Earlier work this paper cites.
Automatic code summarization: A systematic literature review
Yuxiang Zhu and Minxue Pan. 2019 · 2019
Earlier work this paper cites.
CodeBERT: A Pre-Trained Model for Programming and Natural Languages. In Findings of the Association for Computational Linguistics: EMNLP 2020 . 1536–1547
Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, et al · 2020
Cited alongside, same era.
A human study of comprehension and code summarization. In Proceedings of the 28th International Conference on Program Comprehension . 2–13
Sean Stapleton, Yashmeet Gambhir, Alexander LeClair, Zachary Eberhart, Westley Weimer, Kevin Leach, and Yu Huang. 2020 · 2020
Cited alongside, same era.
On the Inference Calibration of Neural Machine Translation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). Association for Computational Linguistics, Online, 3070–3079
Shuo Wang, Zhaopeng Tu, Shuming Shi, and Yang Liu. 2020 · 2020
Cited alongside, same era.
Unified pre-training for program understanding and generation
Wasi Uddin Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang. 2021 · 2021
Retrieval-based prompt selection for code-related few-shot learning. In 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE) . IEEE, 2450–2462
Noor Nashid, Mifta Sintaha, and Ali Mesbah. 2023 · 2023
Later among the works it cites.
Code llama: Open foundation models for code
Baptiste Roziere, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, et al · 2023
Later among the works it cites.
scipy.stats.false_discovery_control
Pauli Virtanen, Ralf Gommers, Travis E. Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, Stéfan J. van der Walt, Matthew Brett, Joshua Wilson, K. Jarrod Millman, Nikolay Mayorov, Andrew R. J. Nelson, Eric Jones, Robert Kern, Eric Larson, C J Carey, İlhan Polat, Yu Feng, Eric W. Moore, Jake VanderPlas, Denis Laxalde, Josef Perktold, Robert Cimrman, Ian Henriksen, E. A. Quintero, Charles R. Harris, Anne M. Archibald, Antônio H. Ribeiro, Fabian Pedregosa, Paul van Mulbregt, and SciPy 1.0 Contributors. 2023 · 2023
Later among the works it cites.
Dylan Zhang, Xuchao Zhang, Chetan Bansal, Pedro Las-Casas, Rodrigo Fonseca, and Saravan Rajmohan. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Codexglue: A machine learning benchmark dataset for code understanding and generation
Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin Clement, Dawn Drain, Daxin Jiang, Duyu Tang, et al · 2021
Cited alongside, same era.
Uncertainty Estimation in Autoregressive Structured Prediction. In International Conference on Learning Representations
Andrey Malinin and Mark Gales. 2021 · 2021
Cited alongside, same era.
Reassessing automatic evaluation metrics for code summarization tasks. In Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering . 1105–1116
Devjeet Roy, Sarah Fakhoury, and Venera Arnaoudova. 2021 · 2021
Cited alongside, same era.
Yue Wang, Weishi Wang, Shafiq Joty, and Steven CH Hoi. 2021 · 2021
Cited alongside, same era.
Semantic similarity metrics for evaluating source code summarization. In Proceedings of the 30th IEEE/ACM International Conference on Program Comprehension . 36–47
Sakib Haque, Zachary Eberhart, Aakash Bansal, and Collin McMillan. 2022 · 2022
Cited alongside, same era.
Practitioners’ expectations on automated code comment generation. In Proceedings of the 44th International Conference on Software Engineering . 1693–1705
Xing Hu, Xin Xia, David Lo, Zhiyuan Wan, Qiuyuan Chen, and Thomas Zimmermann. 2022 · 2022
Cited alongside, same era.
On the evaluation of neural code summarization. In Proceedings of the 44th international conference on software engineering . 1597–1608
Ensheng Shi, Yanlin Wang, Lun Du, Junjie Chen, Shi Han, Hongyu Zhang, Dongmei Zhang, and Hongbin Sun. 2022 · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Cited alongside, same era.
Later among the works it cites.
Wetter Und Klima - Deutscher Wetterdienst - Our Services - Skill Measure: Brier Skill Score
2024 · 2024
Closest in time.
Automatic semantic augmentation of language model prompts (for code summarization). In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering . 1–13
Toufique Ahmed, Kunal Suresh Pai, Premkumar Devanbu, and Earl Barr. 2024 · 2024
Closest in time.
Large language models are few-shot summarizers: Multi-intent comment generation via in-context learning. In Proceedings of the 46th IEEE/ACM International Conference on Software Engineering . 1–13
Mingyang Geng, Shangwen Wang, Dezun Dong, Haotian Wang, Ge Li, Zhi Jin, Xiaoguang Mao, and Xiangke Liao. 2024 · 2024
Closest in time.
DeepSeek-Coder: When the Large Language Model Meets Programming–The Rise of Code Intelligence
Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Y Wu, YK Li, et al · 2024
Closest in time.
Evaluating Code Summarization Techniques: A New Metric and an Empirical Characterization
Antonio Mastropaolo, Matteo Ciniselli, Massimiliano Di Penta, and Gabriele Bavota. 2024 · 2024
Closest in time.
Context-aware Code Summary Generation
Chia-Yi Su, Aakash Bansal, Yu Huang, Toby Jia-Jun Li, and Collin McMillan. 2024 · 2024
Closest in time.
Distilled GPT for source code summarization
Chia-Yi Su and Collin McMillan. 2024 · 2024
Closest in time.
Source Code Summarization in the Era of Large Language Models
Weisong Sun, Yun Miao, Yuekang Li, Hongyu Zhang, Chunrong Fang, Yi Liu, Gelei Deng, Yang Liu, and Zhenyu Chen. 2024 · 2024
Closest in time.
How Good Are FiveThirtyEight Forecasts?
Jay Boice Wezerek, Gus. 2023 · 2024
Closest in time.
On Calibration of Pre-trained Code Models. In 2024 IEEE/ACM 46th International Conference on Software Engineering (ICSE) . IEEE Computer Society, 861–861
Zhenhao Zhou, Chaofeng Sha, and Xin Peng. 2024 · 2024
Closest in time.
Resource-Efficient & Effective Code Summarization
Saima Afrin, Joseph Call, Khai-Nguyen Nguyen, Oscar Chaparro, and Antonio Mastropaolo. 2025 · 2025
Closest in time.
Look Before You Leap: An Exploratory Study of Uncertainty Analysis for Large Language Models
Yuheng Huang, Jiayang Song, Zhijie Wang, Shengming Zhao, Huaming Chen, Felix Juefei-Xu, and Lei Ma. 2025 · 2025
Closest in time.
Calibration and correctness of language models for code. In Proceedings, ICSE 2025
Claudio Spiess, David Gros, Kunal Suresh Pai, Michael Pradel, Md Rafiqul Islam Rabin, Amin Alipour, Susmit Jha, Prem Devanbu, and Toufique Ahmed. 2025 · 2025
Closest in time.