Fetching the paper…
Reading the bibliography…
Although Foundation Models (FMs), such as GPT-4, are increasingly used in domains like finance and software engineering, reliance on textual interfaces limits these models' real-world interaction.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
The cathedral and the bazaar
Eric Raymond. 1999 · 1999
Earlier work this paper cites.
A systematic review of statistical power in software engineering experiments
Tore Dybå, Vigdis By Kampenes, and Dag IK Sjøberg. 2006 · 2006
Earlier work this paper cites.
Pixy: A static analysis tool for detecting web application vulnerabilities. In 2006 IEEE Symposium on Security and Privacy (S&P’06) . IEEE, 6–pp
Nenad Jovanovic, Christopher Kruegel, and Engin Kirda. 2006 · 2006
Earlier work this paper cites.
Evaluating static analysis defect warnings on production software. In Proceedings of the 7th ACM SIGPLAN-SIGSOFT workshop on Program analysis for software tools and engineering . 1–8
Nathaniel Ayewah, William Pugh, J David Morgenthaler, John Penix, and YuQian Zhou. 2007 · 2007
Earlier work this paper cites.
Vulnerability type distributions in CVE
Steve Christey and Robert A Martin. 2007 · 2007
Earlier work this paper cites.
Determinism and evolution. In Proceedings of the 2008 international working conference on Mining software repositories . 1–10
Israel Herraiz, Jesus M Gonzalez-Barahona, and Gregorio Robles. 2008 · 2008
Earlier work this paper cites.
Time for some a priori thinking about post hoc testing
Graeme D Ruxton and Guy Beauchamp. 2008 · 2008
Earlier work this paper cites.
An empirical study on the maintenance of source code clones
Suresh Thummalapenta, Luigi Cerulo, Lerina Aversano, and Massimiliano Di Penta. 2010 · 2010
Earlier work this paper cites.
Cliff’s Delta Calculator: A non-parametric effect size program for two groups of observations
Guillermo Macbeth, Eugenia Razumiejczyk, and Rubén Daniel Ledesma. 2011 · 2011
Earlier work this paper cites.
An integrated security governance framework for effective PCI DSS implementation
Mathew Nicho, Hussein Fakhry, and Charles Haiber. 2011 · 2011
Earlier work this paper cites.
Factors affecting the success of Open Source Software
Vishal Midha and Prashant Palvia. 2012 · 2012
Earlier work this paper cites.
Do code smells reflect important maintainability aspects?. In 2012 28th IEEE international conference on software maintenance (ICSM) . IEEE, 306–315
Aiko Yamashita and Leon Moonen. 2012 · 2012
Earlier work this paper cites.
SonarQube in action
G Ann Campbell and Patroklos P Papapetrou. 2013 · 2013
Earlier work this paper cites.
Empirical studies on the functional complexity of software in large-scale software systems
Yingxu Wang and Vincent Chiew. 2013 · 2013
Earlier work this paper cites.
Code smells as system-level indicators of maintainability: An empirical study
Aiko Yamashita and Steve Counsell. 2013 · 2013
Earlier work this paper cites.
Why do automated builds break? an empirical study. In 2014 IEEE international conference on software maintenance and evolution . IEEE, 41–50
Noureddine Kerzazi, Foutse Khomh, and Bram Adams. 2014 · 2014
Earlier work this paper cites.
A large scale study of programming languages and code quality in github. In Proceedings of the 22nd ACM SIGSOFT international symposium on foundations of software engineering . 155–165
Baishakhi Ray, Daryl Posnett, Vladimir Filkov, and Premkumar Devanbu. 2014 · 2014
Earlier work this paper cites.
The vision of software clone management: Past, present, and future (keynote paper). In 2014 Software Evolution Week-IEEE Conference on Software Maintenance, Reengineering, and Reverse Engineering (CSMR-WCRE) . IEEE, 18–33
Chanchal K Roy, Minhaz F Zibran, and Rainer Koschke. 2014 · 2014
Earlier work this paper cites.
Predicting vulnerable components: Software metrics vs text mining. In 2014 IEEE 25th international symposium on software reliability engineering . IEEE, 23–33
James Walden, Jeff Stuckman, and Riccardo Scandariato. 2014 · 2014
Earlier work this paper cites.
Fleiss’ kappa statistic without paradoxes
Rosa Falotico and Piero Quatto. 2015 · 2015
Earlier work this paper cites.
Predicting software future sustainability: A longitudinal perspective
Amir Hossein Ghapanchi. 2015 · 2015
Earlier work this paper cites.
An analysis of HTML and CSS syntax errors in a web development course
Thomas H Park, Brian Dorn, and Andrea Forte. 2015 · 2015
Earlier work this paper cites.
When and why your code starts to smell bad. In 2015 IEEE/ACM 37th IEEE International Conference on Software Engineering , Vol. 1. IEEE, 403–414
Michele Tufano, Fabio Palomba, Gabriele Bavota, Rocco Oliveto, Massimiliano Di Penta, Andrea De Lucia, and Denys Poshyvanyk. 2015 · 2015
Earlier work this paper cites.
Understanding the factors that impact the popularity of GitHub repositories. In 2016 IEEE international conference on software maintenance and evolution (ICSME) . IEEE, 334–344
Hudson Borges, Andre Hora, and Marco Tulio Valente. 2016 · 2016
Earlier work this paper cites.
Detecting code smells in python programs. In 2016 international conference on Software Analysis, Testing and Evolution (SATE) . IEEE, 18–23
Zhifei Chen, Lin Chen, Wanwangying Ma, and Baowen Xu. 2016 · 2016
Earlier work this paper cites.
Usage, costs, and benefits of continuous integration in open-source projects. In Proceedings of the 31st IEEE/ACM international conference on automated software engineering . 426–437
Michael Hilton, Timothy Tunnell, Kai Huang, Darko Marinov, and Danny Dig. 2016 · 2016
Earlier work this paper cites.
Kruskal–Wallis H-test for oneway analysis of variance (ANOVA) by ranks
Thomas W MacFarland, Jan M Yates, Thomas W MacFarland, and Jan M Yates. 2016 · 2016
Earlier work this paper cites.
Empirical evaluation of the impact of object-oriented code refactoring on quality attributes: A systematic literature review
Jehad Al Dallal and Anas Abdin. 2017 · 2017
Earlier work this paper cites.
Evaluating code complexity triggers, use of complexity measures and the influence of code complexity on maintenance time
Vard Antinyan, Miroslaw Staron, and Anna Sandberg. 2017 · 2017
Earlier work this paper cites.
How do apps evolve in their permission requests? a preliminary study. In 2017 IEEE/ACM 14th International Conference on Mining Software Repositories (MSR) . IEEE, 37–41
Paolo Calciati and Alessandra Gorla. 2017 · 2017
Earlier work this paper cites.
Gradual typing with union and intersection types
Giuseppe Castagna and Victor Lanvin. 2017 · 2017
Earlier work this paper cites.
Why modern open source projects fail. In Proceedings of the 2017 11th Joint meeting on foundations of software engineering . 186–196
Jailton Coelho and Marco Tulio Valente. 2017 · 2017
Earlier work this paper cites.
Software complexity analysis using halstead metrics. In 2017 international conference on trends in electronics and informatics (ICEI) . IEEE, 1109–1113
T Hariprasad, G Vidhyagaran, K Seenu, and Chandrasegar Thirumalai. 2017 · 2017
Earlier work this paper cites.
Curating github for engineered software projects
Nuthan Munaiah, Steven Kroh, Craig Cabrey, and Meiyappan Nagappan. 2017 · 2017
Earlier work this paper cites.
Detecting argument selection defects
Andrew Rice, Edward Aftandilian, Ciera Jaspan, Emily Johnston, Michael Pradel, and Yulissa Arroyo-Paredes. 2017 · 2017
Earlier work this paper cites.
An empirical study of code smells in javascript projects. In 2017 IEEE 24th international conference on software analysis, evolution and reengineering (SANER) . IEEE, 294–305
Amir Saboury, Pooya Musavi, Foutse Khomh, and Giulio Antoniol. 2017 · 2017
Earlier work this paper cites.
(No) influence of continuous integration on the commit activity in GitHub projects. In Proceedings of the 4th ACM SIGSOFT International Workshop on Software Analytics . 1–7
Sebastian Baltes, Jascha Knack, Daniel Anastasiou, Ralf Tymann, and Stephan Diehl. 2018 · 2018
Earlier work this paper cites.
How do defects hurt qualities? an empirical study on characterizing a software maintainability ontology in open source software. In 2018 IEEE International Conference on Software Quality, Reliability and Security (QRS) . IEEE, 226–237
Celia Chen, Shi Lin, Michael Shoga, Qing Wang, and Barry Boehm. 2018 · 2018
Earlier work this paper cites.
On the impact of security vulnerabilities in the npm package dependency network. In Proceedings of the 15th international conference on mining software repositories . 181–191
Alexandre Decan, Tom Mens, and Eleni Constantinou. 2018 · 2018
Earlier work this paper cites.
Refactoring: improving the design of existing code
Martin Fowler. 2018 · 2018
Earlier work this paper cites.
Causes, impacts, and detection approaches of code smell: a survey. In Proceedings of the 2018 ACM Southeast Conference . 1–8
Md Shariful Haque, Jeff Carver, and Travis Atkison. 2018 · 2018
Earlier work this paper cites.
On the diffuseness and the impact on maintainability of code smells: a large scale empirical investigation. In Proceedings of the 40th International Conference on Software Engineering . 482–482
Fabio Palomba, Gabriele Bavota, Massimiliano Di Penta, Fausto Fasano, Rocco Oliveto, and Andrea De Lucia. 2018 · 2018
Earlier work this paper cites.
Vulnerable open source dependencies: Counting those that matter. In Proceedings of the 12th ACM/IEEE international symposium on empirical software engineering and measurement . 1–10
Ivan Pashchenko, Henrik Plate, Serena Elisa Ponta, Antonino Sabetta, and Fabio Massacci. 2018 · 2018
Earlier work this paper cites.
Evaluation of software reusability based on coupling and cohesion
G Priyalakshmi and R Latha. 2018 · 2018
Earlier work this paper cites.
An empirical analysis of vulnerabilities in python packages for web applications. In 2018 9th International Workshop on Empirical Software Engineering in Practice (IWESEP) . IEEE, 25–30
Jukka Ruohonen. 2018 · 2018
Earlier work this paper cites.
Identifying various code-smells and refactoring opportunities in object-oriented software system: a systematic literature review
Randeep Singh and Ashok Kumar. 2018 · 2018
Earlier work this paper cites.
A systematic literature review: Refactoring for disclosing code smells in object oriented software
Satwinder Singh and Sharanpreet Kaur. 2018 · 2018
Earlier work this paper cites.
How developers diagnose potential security vulnerabilities with a static analysis tool
Justin Smith, Brittany Johnson, Emerson Murphy-Hill, Bill Chu, and Heather Richter Lipford. 2018 · 2018
Earlier work this paper cites.
Ecosystem-level determinants of sustained activity in open-source projects: A case study of the PyPI ecosystem. In Proceedings of the 2018 26th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering . 644–655
Marat Valiev, Bogdan Vasilescu, and James Herbsleb. 2018 · 2018
Earlier work this paper cites.
Why are android apps removed from google play? a large-scale empirical study. In Proceedings of the 15th International Conference on Mining Software Repositories . 231–242
Haoyu Wang, Hao Li, Li Li, Yao Guo, and Guoai Xu. 2018 · 2018
Earlier work this paper cites.
An empirical analysis of technical lag in npm package dependencies. In International conference on software reuse . Springer, 95–110
Ahmed Zerouali, Eleni Constantinou, Tom Mens, Gregorio Robles, and Jesús González-Barahona. 2018 · 2018
Earlier work this paper cites.
On the abandonment and survival of open source projects: An empirical investigation. In 2019 ACM/IEEE International Symposium on Empirical Software Engineering and Measurement (ESEM) . IEEE, 1–12
Guilherme Avelino, Eleni Constantinou, Marco Tulio Valente, and Alexander Serebrenik. 2019 · 2019
Earlier work this paper cites.
A large scale study of long-time contributor prediction for github projects
Lingfeng Bao, Xin Xia, David Lo, and Gail C Murphy. 2019 · 2019
Earlier work this paper cites.
An empirical analysis of the python package index (pypi)
Ethan Bommarito and Michael Bommarito. 2019 · 2019
Cited alongside, same era.
The seven sins: Security smells in infrastructure as code scripts. In 2019 IEEE/ACM 41st International Conference on Software Engineering (ICSE) . IEEE, 164–175
Akond Rahman, Chris Parnin, and Laurie Williams. 2019 · 2019
Cited alongside, same era.
LoopFix: An approach to automatic repair of buggy loops
Weichao Wang, Zhaopeng Meng, Zan Wang, Shuang Liu, and Jianye Hao. 2019 · 2019
Cited alongside, same era.
Small world with high risks: A study of security threats in the npm ecosystem. In 28th USENIX security symposium (USENIX Security 19) . 995–1010
Markus Zimmermann, Cristian-Alexandru Staicu, Cam Tenny, and Michael Pradel. 2019 · 2019
Cited alongside, same era.
Branch use in practice: A large-scale empirical study of 2,923 projects on github. In 2019 IEEE 19th International Conference on Software Quality, Reliability and Security (QRS) . IEEE, 306–317
How to refactor this code? An exploratory study on developer-ChatGPT refactoring conversations. In Proceedings of the 21st International Conference on Mining Software Repositories . 202–206
Eman Abdullah AlOmar, Anushkrishna Venkatakrishnan, Mohamed Wiem Mkaouer, Christian Newman, and Ali Ouni. 2024 · 2024
Later among the works it cites.
How do machine learning projects use continuous integration practices? an empirical study on GitHub actions. In Proceedings of the 21st International Conference on Mining Software Repositories . 665–676
João Helis Bernardo, Daniel Alencar Da Costa, Sérgio Queiroz de Medeiros, and Uirá Kulesza. 2024 · 2024
Later among the works it cites.
Potential of large language models in health care: Delphi study
Kerstin Denecke, Richard May, LLMHealthGroup, and Octavio Rivera Romero. 2024 · 2024
Later among the works it cites.
Ahmed E Hassan, Gustavo A Oliva, Dayi Lin, Boyuan Chen, Zhen Ming, et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Weiqin Zou, Weiqiang Zhang, Xin Xia, Reid Holmes, and Zhenyu Chen. 2019 · 2019
Cited alongside, same era.
Empirical study of the relationship between design patterns and code smells
Mahmoud Alfadel, Khalid Aljasser, and Mohammad Alshayeb. 2020 · 2020
Cited alongside, same era.
Buildfast: History-aware build outcome prediction for fast feedback and reduced cost in continuous integration. In Proceedings of the 35th IEEE/ACM International Conference on Automated Software Engineering . 42–53
Bihuan Chen, Linlin Chen, Chen Zhang, and Xin Peng. 2020 · 2020
Cited alongside, same era.
On relating technical, social factors, and the introduction of bugs. In 2020 IEEE 27th International Conference on Software Analysis, Evolution and Reengineering (SANER) . IEEE, 378–388
Filipe Falcão, Caio Barbosa, Baldoino Fonseca, Alessandro Garcia, Márcio Ribeiro, and Rohit Gheyi. 2020 · 2020
Cited alongside, same era.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020 · 2020
Cited alongside, same era.
Code smells and refactoring: A tertiary systematic review of challenges and observations
Guilherme Lacerda, Fabio Petrillo, Marcelo Pimenta, and Yann Gaël Guéhéneuc. 2020 · 2020
Cited alongside, same era.
Are sonarqube rules inducing bugs?. In 2020 IEEE 27th international conference on software analysis, evolution and reengineering (SANER) . IEEE, 501–511
Valentina Lenarduzzi, Francesco Lomio, Heikki Huttunen, and Davide Taibi. 2020 · 2020
Cited alongside, same era.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al · 2020
Cited alongside, same era.
Cost of a Data Breach Report 2024
IBM. 2025 · 2024
Later among the works it cites.
A large-scale study of ml-related python projects. In Proceedings of the 39th ACM/SIGAPP Symposium on Applied Computing . 1272–1281
Samuel Idowu, Yorick Sens, Thorsten Berger, Jacob Krüger, and Michael Vierhauser. 2024 · 2024
Later among the works it cites.
Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement
Jaehun Jung, Faeze Brahman, and Yejin Choi. 2024 · 2024
Later among the works it cites.
Comparative analysis of real issues in open-source machine learning projects
Tuan Dung Lai, Anj Simmons, Scott Barnett, Jean-Guy Schneider, and Rajesh Vasa. 2024 · 2024
Later among the works it cites.
Hao Li, Cor-Paul Bezemer, and Ahmed E Hassan. 2024 · 2024
Later among the works it cites.
Jiahuei Lin, Dayi Lin, Sky Zhang, and Ahmed E Hassan. 2024 · 2024
Later among the works it cites.
OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, et al · 2024
Later among the works it cites.
Judging the judges: A systematic study of position bias in llm-as-a-judge
Lin Shi, Chiyu Ma, Wenhua Liang, Xingjian Diao, Weicheng Ma, and Soroush Vosoughi. 2024 · 2024
Later among the works it cites.
A systematic literature review on automated software vulnerability detection using machine learning
Nima Shiri Harzevili, Alvine Boaye Belle, Junjie Wang, Song Wang, Zhen Ming Jiang, and Nachiappan Nagappan. 2024 · 2024
Later among the works it cites.
Exception Miner: Multi-language Static Analysis Tool to Identify Exception Handling Anti-Patterns. In Simpósio Brasileiro de Engenharia de Software (SBES) . SBC, 741–747
Jairo Souza, Tales Alves, Robson Oliveira, Leopoldo Teixeira, and Baldoino Fonseca. 2024 · 2024
Later among the works it cites.
Significant productivity gains through programming with large language models
Thomas Weber, Maximilian Brandmaier, Albrecht Schmidt, and Sven Mayer. 2024 · 2024
Later among the works it cites.
iSMELL: Assembling LLMs with Expert Toolsets for Code Smell Detection and Refactoring. In Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering . 1345–1357
Di Wu, Fangwen Mu, Lin Shi, Zhaoqiang Guo, Kui Liu, Weiguang Zhuang, Yuqi Zhong, and Li Zhang. 2024 · 2024
Later among the works it cites.
Machine learning-based methods for code smell detection: a survey
Pravin Singh Yadav, Rajwant Singh Rao, Alok Mishra, and Manjari Gupta. 2024 · 2024
Later among the works it cites.
Artificial Intelligence and Dynamic Analysis-Based Web Application Vulnerability Scanner
Mehmet Ali Yalçinkaya and Ecir Uğur Küçüksille. 2024 · 2024
Later among the works it cites.
Can LLMs Replace Manual Annotation of Software Engineering Artifacts?. In IEEE/ACM International Conference on Mining Software Repositories
Toufique Ahmed, Premkumar Devanbu, Christoph Treude, and Michael Pradel. 2025 · 2025
Closest in time.
App Review Guidelines
Apple. 2025 · 2025
Closest in time.
AutoGen: A framework for building AI agents and applications
Microsoft Autogen. 2025 · 2025
Closest in time.
Community Health Analytics in Open Source Software: Topic - All Metrics
CHAOSS Project. 2025a · 2025
Closest in time.
Practitioner Guide: Responsiveness
CHAOSS Project. 2025b · 2025
Closest in time.
Cloudflare Agents Docs: Model Context Protocol (MCP)
Cloudflare. 2025 · 2025
Closest in time.
CrewAI: The leading multi-agent platform
CrewAI. 2025 · 2025
Closest in time.
An extensible multilanguage static code analyzer
Andrey Loskutov Keith Lea David Hovemeyer, Bill Pugh. 2025 · 2025
Closest in time.
Dify: Build Production Ready Agentic Solution
Dify. 2025 · 2025
Closest in time.
How Do Generative Models Draw a Software Engineer? A Case Study on Stable Diffusion Bias
Tosin Fadahunsi, Giordano d’Aloisio, Antinisca Di Marco, and Federica Sarro. 2025 · 2025
Closest in time.
State of Software 2025: A Global Report on the Hidden Costs and Risks of Software
Software Improvement Group. 2025 · 2025
Closest in time.
The replication package of our study on MCP Servers
Mohammed Mehedi Hasan, Hao Li, Emad Fallahzadeh, Gopi Krishnan Rajbahadur, Bram Adams, and A. E Hassan. 2025 · 2025
Closest in time.
Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions
Xinyi Hou, Yanjie Zhao, Shenao Wang, and Haoyu Wang. 2025 · 2025
Closest in time.
Security checklist
Alphabet Inc. 2025 · 2025
Closest in time.
MCP Guardian: A Security-First Layer for Safeguarding MCP-Based AI System
Sonu Kumar, Anubhav Girdhar, Ritesh Patil, and Divyansh Tripathi. 2025 · 2025
Closest in time.
Introducing MCP-Scan: Protecting MCP with Invariant
Invariant Lab. 2025 · 2025
Closest in time.
LangChain: composable framework to build with LLMs
LangChain. 2025 · 2025
Closest in time.
Bridging the language gap: an empirical study of bindings for open source machine learning libraries across software package ecosystems
Hao Li and Cor-Paul Bezemer. 2025 · 2025
Closest in time.
LlamaIndex: LlamaIndex is the leading framework for building LLM-powered agents over your data with LLMs and workflows
LlamaIndex. 2025 · 2025
Closest in time.
GenAI Toolbox: An introduction to MCP Toolbox
Google LLC. 2025 · 2025
Closest in time.
Function Calling: Llama 3.1 models now officially supports function calling
Meta. 2025 · 2025
Closest in time.
Common Weakness Enumeration: : A community developed list of SW & HW weakness that can become vulnerabilities
Mitre. 2025 · 2025
Closest in time.
Enterprise-Grade Security for the Model Context Protocol (MCP): Frameworks and Mitigation Strategies
Vineeth Sai Narajala and Idan Habler. 2025 · 2025
Closest in time.
An extensible multilanguage static code analyzer
PMD. 2025 · 2025
Closest in time.
DevSecOps practices and tools
Luís Prates and Rúben Pereira. 2025 · 2025
Closest in time.
Model Context Protocol servers
Model Context Protocol. 2025 · 2025
Closest in time.
MCP Safety Audit: LLMs with the Model Context Protocol Allow Major Security Exploits
Brandon Radosevich and John Halloran. 2025 · 2025
Closest in time.
Opportunities and security risks of technical leverage: A replication study on the NPM ecosystem
Haya Samaana, Diego Elias Costa, Ahmad Abdellatif, and Emad Shihab. 2025 · 2025
Closest in time.
Finetuning large language models for vulnerability detection
Aleksei Shestov, Rodion Levichev, Ravil Mussabayev, Evgeny Maslov, Pavel Zadorozhny, Anton Cheshkov, Rustam Mussabayev, Alymzhan Toleu, Gulmira Tolegen, and Alexander Krassovitskiy. 2025 · 2025
Closest in time.
Ten most expensive bugs in history (part 2)
sixsentix. 2025 · 2025
Closest in time.
Smithery: Your Agent’s Gateway to the World
smithery. 2025 · 2025
Closest in time.
The Stripe Model Context Protocol server allows you to integrate with Stripe APIs through function calling
Stripe. 2025 · 2025
Closest in time.
ProfBench: Multi-Domain Rubrics requiring Professional Knowledge to Answer and Judge
Zhilin Wang, Jaehun Jung, Ximing Lu, Shizhe Diao, Ellie Evans, Jiaqi Zeng, Pavlo Molchanov, Yejin Choi, Jan Kautz, and Yi Dong. 2025 · 2025
Closest in time.
PersonaEval: Are LLM Evaluators Human Enough to Judge Role-Play?
Lingfeng Zhou, Jialing Zhang, Jin Gao, Mohan Jiang, and Dequan Wang. 2025 · 2025
Closest in time.
A comparative study of static code analysis tools for vulnerability detection in c/c++ and java source code
Arvinder Kaur and Ruchikaa Nayyar. 2020 · 2029
Closest in time.