Fetching the paper…
Reading the bibliography…
Software issue localization, the task of identifying the precise code locations (files, classes, or functions) relevant to a natural language issue description (e.g., bug report, feature request), is a critical yet time-consuming aspect of software development.
Okapi at trec-3
Stephen E. Robertson, Steve Walker, Susan Jones, Micheline Hancock-Beaulieu, and Mike Gatford · 1994
Earlier work this paper cites.
An enhanced approach for software bug localization using map reduce technique based apriori (mrtba) algorithm
A Adhiselvam, E Kirubakaran, and R Sukumar · 2015
Earlier work this paper cites.
Spectrum-based software fault localization: A survey of techniques, advances, and challenges
Higor A de Souza, Marcos L Chaim, and Fabio Kon · 2016
Earlier work this paper cites.
Chapter three - fault localization using hybrid static/dynamic analysis
E. Elsaka · 2016
Earlier work this paper cites.
A survey on software fault localization
W Eric Wong, Ruizhi Gao, Yihao Li, Rui Abreu, and Franz Wotawa · 2016
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Earlier work this paper cites.
FEVER: a large-scale dataset for fact extraction and VERification
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal · 2018
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D Manning · 2018
Earlier work this paper cites.
Codesearchnet challenge: Evaluating the state of semantic code search
Hamel Husain, Ho-Hsiang Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt · 2019
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al · 2019
Earlier work this paper cites.
Sentence-BERT: Sentence embeddings using Siamese BERT-networks
Nils Reimers and Iryna Gurevych · 2019
Earlier work this paper cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He · 2020
Earlier work this paper cites.
Locating faults with program slicing: an empirical analysis
Ezekiel Soremekun, Lukas Kirschner, Marcel Böhme, and Andreas Zeller · 2021
Earlier work this paper cites.
Codet5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation
Wang Yue, Wang Weishi, Shafiq Joty, and Steven C.H. Hoi · 2021
Earlier work this paper cites.
The stack: 3 tb of permissively licensed source code
Denis Kocetkov, Raymond Li, LI Jia, Chenghao Mou, Yacine Jernite, Margaret Mitchell, Carlos Muñoz Ferrandis, Sean Hughes, Thomas Wolf, Dzmitry Bahdanau, et al · 2022
Earlier work this paper cites.
Matryoshka representation learning
Aditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford, Aditya Sinha, Vivek Ramanujan, William Howard-Snyder, Kaifeng Chen, Sham Kakade, Prateek Jain, and Ali Farhadi · 2022
Earlier work this paper cites.
Code summarization: Do transformers really understand code?
Ankita Nandkishor Sontakke, Manasi Patwardhan, Lovekesh Vig, Raveendra Kumar Medicherla, Ravindra Naik, and Gautam Shroff · 2022
Earlier work this paper cites.
Jina embeddings: A novel set of high-performance sentence embedding models
Michael Günther, Georgios Mastrapas, Bo Wang, Han Xiao, and Jonathan Geuter · 2023
Earlier work this paper cites.
Jina embeddings: A novel set of high-performance sentence embedding models, 2023
Michael Günther, Louis Milliken, Jonathan Geuter, Georgios Mastrapas, Bo Wang, and Han Xiao · 2023
Cited alongside, same era.
Neftune: Noisy embeddings improve instruction finetuning, 2023
Neel Jain, Ping yeh Chiang, Yuxin Wen, John Kirchenbauer, Hong-Min Chu, Gowthami Somepalli, Brian R. Bartoldson, Bhavya Kailkhura, Avi Schwarzschild, Aniruddha Saha, Micah Goldblum, Jonas Geiping, and Tom Goldstein · 2023
Cited alongside, same era.
Swe-bench: Can language models resolve real-world github issues?
Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan · 2023
Cited alongside, same era.
Towards general text embeddings with multi-stage contrastive learning
Zehan Li, Xin Zhang, Yanzhao Zhang, Dingkun Long, Pengjun Xie, and Meishan Zhang · 2023
Cited alongside, same era.
Agentfl: Scaling llm-based fault localization to project-level context
Yihao Qin, Shangwen Wang, Yiling Lou, Jinhao Dong, Kaixin Wang, Xiaoling Li, and Xiaoguang Mao · 2024
Later among the works it cites.
First: Faster improved listwise reranking with single token decoding
Revanth Gangi Reddy, JaeHyeok Doo, Yifei Xu, Md Arafat Sultan, Deevya Swain, Avirup Sil, and Heng Ji · 2024
Later among the works it cites.
Cornstack: High-quality contrastive data for better code ranking
Tarun Suresh, Revanth Gangi Reddy, Yifei Xu, Zach Nussbaum, Andriy Mulyar, Brandon Duderstadt, and Heng Ji · 2024
Later among the works it cites.
Arctic-embed 2.0: Multilingual retrieval without compromise
Puxuan Yu, Luke Merrick, Gaurav Nuti, and Daniel Campos · 2024
Later among the works it cites.
CODE REPRESENTATION LEARNING AT SCALE
Dejiao Zhang, Wasi Uddin Ahmad, Ming Tan, Hantian Ding, Ramesh Nallapati, Dan Roth, Xiaofei Ma, and Bing Xiang · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
GitHub Copilot—Your AI pair programmer, 2023
Microsoft · 2023
Cited alongside, same era.
Mteb: Massive text embedding benchmark
Niklas Muennighoff, Nouamane Tazi, Loic Magne, and Nils Reimers · 2023
Cited alongside, same era.
Rap-gen: Retrieval-augmented patch generation with codet5 for automatic program repair
Weishi Wang, Yue Wang, Shafiq Joty, and Steven C.H. Hoi · 2023
Cited alongside, same era.
React: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao · 2023
Cited alongside, same era.
Understanding the use of spectrum-based fault localization
Higor Amario de Souza, Marcelo de Souza Lauretto, Fabio Kon, and Marcos Lordello Chaim · 2024
Cited alongside, same era.
Are large language models a threat to programming platforms? an exploratory study
Md Mustakim Billah, Palash Ranjan Roy, Zadia Codabux, and Banani Roy · 2024
Cited alongside, same era.
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al · 2024
Cited alongside, same era.
Qwen2.5-coder technical report
Binyuan Hui, Jian Yang, Zeyu Cui, Jiaxi Yang, Dayiheng Liu, Lei Zhang, Tianyu Liu, Jiajun Zhang, Bowen Yu, Kai Dang, An Yang, Rui Men, Fei Huang, Shanghaoran Quan, Xingzhang Ren, Xuancheng Ren, Jingren Zhou, and Junyang Lin · 2024
Cited alongside, same era.
Later among the works it cites.
Moatless tools, 2024
Albert Örwall · 2024
Later among the works it cites.
Claude: Conversational ai by anthropic, 2023
Anthropic · 2025
Closest in time.
Locagent: Graph-guided llm agents for code localization
Zhaoling Chen, Xiangru Tang, Gangda Deng, Fang Wu, Jialong Wu, Zhiwei Jiang, Viktor Prasanna, Arman Cohan, and Xingyao Wang · 2025
Closest in time.
Devin: The First AI Software Engineer
Cognition AI · 2025
Closest in time.
Cursor: The AI Code Editor
Cursor · 2025
Closest in time.
How aider scored sota 26.3% on swe bench lite — aider, 2024
Paul Gauthier · 2025
Closest in time.
Gemini embedding: Generalizable embeddings from gemini
Jinhyuk Lee, Feiyang Chen, Sahil Dua, Daniel Cer, Madhuri Shanbhogue, Iftekhar Naim, Gustavo Hernández Ábrego, Zhe Li, Kaifeng Chen, Henrique Schechter Vera, et al · 2025
Closest in time.
Openhands: An open platform for AI software developers as generalist agents
Xingyao Wang, Boxuan Li, Yufan Song, Frank F. Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, Hoang H. Tran, Fuqiang Li, Ren Ma, Mingzhang Zheng, Bill Qian, Yanjun Shao, Niklas Muennighoff, Yizhe Zhang, Binyuan Hui, Junyang Lin, Robert Brennan, Hao Peng, Heng Ji, and Graham Neubig · 2025
Closest in time.
Windsurf Editor: The AI‑Native IDE
Windsurf · 2025
Closest in time.
Swe-fixer: Training open-source llms for effective and efficient github issue resolution
Chengxing Xie, Bowen Li, Chang Gao, He Du, Wai Lam, Difan Zou, and Kai Chen · 2025
Closest in time.
Swe-smith: Scaling data for software engineering agents, 2025
John Yang, Kilian Lieret, Carlos E. Jimenez, Alexander Wettig, Kabir Khandpur, Yanzhe Zhang, Binyuan Hui, Ofir Press, Ludwig Schmidt, and Diyi Yang · 2025
Closest in time.
Orcaloca: An llm agent framework for software issue localization
Zhongming Yu, Hejia Zhang, Yujie Zhao, Hanxian Huang, Matrix Yao, Ke Ding, and Jishen Zhao · 2025
Closest in time.
Qwen3 embedding: Advancing text embedding and reranking through foundation models, 2025
Yanzhao Zhang, Mingxin Li, Dingkun Long, Xin Zhang, Huan Lin, Baosong Yang, Pengjun Xie, An Yang, Dayiheng Liu, Junyang Lin, Fei Huang, and Jingren Zhou · 2025
Closest in time.