Fetching the paper…
Reading the bibliography…
Language model (LM) agents are increasingly being used to automate complicated tasks in digital environments.
Human-computer interaction: psychology as a science of design
J. M. Carroll · 1997
Earlier work this paper cites.
About face 3: the essentials of interaction design
A. Cooper, R. Reimann, and D. Cronin · 2007
Earlier work this paper cites.
Defects4J: A Database of existing faults to enable controlled testing studies for Java programs
R. Just, D. Jalali, and M. D. Ernst · 2014
Earlier work this paper cites.
Entropy guided spectrum based bug localization using statistical language model
S. Chakraborty, Y. Li, M. Irvine, R. Saha, and B. Ray · 2018
Earlier work this paper cites.
How often do single-statement bugs occur? the manysstubs4j dataset
R.-M. Karampatsis and C. Sutton · 2019
Earlier work this paper cites.
Understanding human intelligence through human limitations
T. L. Griffiths · 2020
Earlier work this paper cites.
Program synthesis with large language models, 2021
J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le, and C. Sutton · 2021
Earlier work this paper cites.
Evaluating large language models trained on code, 2021
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. de Oliveira Pinto, and J. K. et. al · 2021
Earlier work this paper cites.
Measuring coding challenge competence with apps, 2021
D. Hendrycks, S. Basart, S. Kadavath, M. Mazeika, A. Arora, E. Guo, C. Burns, S. Puranik, H. He, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Codexglue: A machine learning benchmark dataset for code understanding and generation, 2021
S. Lu, D. Guo, S. Ren, J. Huang, A. Svyatkovskiy, A. Blanco, C. Clement, D. Drain, D. Jiang, D. Tang, G. Li, L. Zhou, L. Shou, L. Zhou, M. Tufano, M. Gong, M. Zhou, N. Duan, N. Sundaresan, S. K. Deng, S. Fu, and S. Liu · 2021
Earlier work this paper cites.
Multipl-e: A scalable and extensible approach to benchmarking neural code generation, 2022
F. Cassano, J. Gouwar, D. Nguyen, S. Nguyen, L. Phipps-Costin, D. Pinckney, M.-H. Yee, Y. Zi, C. J. Anderson, M. Q. Feldman, A. Guha, M. Greenberg, and A. Jangda · 2022
Earlier work this paper cites.
Ds-1000: A natural and reliable benchmark for data science code generation, 2022
Y. Lai, C. Li, Y. Wang, T. Zhang, R. Zhong, L. Zettlemoyer, S. W. tau Yih, D. Fried, S. Wang, and T. Yu · 2022
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback, 2022
R. Nakano, J. Hilton, S. Balaji, J. Wu, L. Ouyang, C. Kim, C. Hesse, S. Jain, V. Kosaraju, W. Saunders, X. Jiang, K. Cobbe, T. Eloundou, G. Krueger, K. Button, M. Knight, B. Chess, and J. Schulman · 2022
Earlier work this paper cites.
Lamda: Language models for dialog applications, 2022
R. Thoppilan, D. D. Freitas, J. Hall, N. Shazeer, A. Kulshreshtha, H.-T. Cheng, A. Jin, T. Bos, L. Baker, Y. Du, Y. Li, H. Lee, H. S. Zheng, A. Ghafouri, M. Menegali, Y. Huang, M. Krikun, D. Lepikhin, J. Qin, D. Chen, Y. Xu, Z. Chen, A. Roberts, M. Bosma, V. Zhao, Y. Zhou, C.-C. Chang, I. Krivokon, W. Rusch, M. Pickett, P. Srinivasan, L. Man, K. Meier-Hellstern, M. R. Morris, T. Doshi, R. D. Santos, T. Duke, J. Soraker, B. Zevenbergen, V. Prabhakaran, M. Diaz, B. Hutchinson, K. Olson, A. Molina, E. Hoffman-John, J. Lee, L. Aroyo, R. Rajakumar, A. Butryna, M. Lamm, V. Kuzmina, J. Fenton, A. Cohen, R. Bernstein, R. Kurzweil, B. Aguera-Arcas, C. Cui, M. Croak, E. Chi, and Q. Le · 2022
Earlier work this paper cites.
Less training, more repairing please: revisiting automated program repair via zero-shot learning
C. S. Xia and L. Zhang · 2022
Earlier work this paper cites.
Natural language to code generation in interactive data science notebooks, 2022
P. Yin, W.-D. Li, K. Xiao, A. Rao, Y. Wen, K. Shi, J. Howland, P. Bailey, M. Catasta, H. Michalewski, A. Polozov, and C. Sutton · 2022
Earlier work this paper cites.
Parsel: Algorithmic reasoning with language models by composing decompositions, 2022
E. Zelikman, Q. Huang, G. Poesia, N. D. Goodman, and N. Haber · 2022
Earlier work this paper cites.
Crosscodeeval: A diverse and multilingual benchmark for cross-file code completion
Y. Ding, Z. Wang, W. U. Ahmad, H. Ding, M. Tan, N. Jain, M. K. Ramanathan, R. Nallapati, P. Bhatia, D. Roth, and B. Xiang · 2023
Earlier work this paper cites.
Classeval: A manually-crafted benchmark for evaluating llms on class-level code generation, 2023
X. Du, M. Liu, K. Wang, H. Wang, J. Liu, Y. Chen, J. Feng, C. Sha, X. Peng, and Y. Lou · 2023
Cited alongside, same era.
Automated repair of programs from large language models, 2023
Z. Fan, X. Gao, M. Mirchev, A. Roychoudhury, and S. H. Tan · 2023
Cited alongside, same era.
Metagpt: Meta programming for a multi-agent collaborative framework, 2023
S. Hong, M. Zhuge, J. Chen, X. Zheng, Y. Cheng, C. Zhang, J. Wang, Z. Wang, S. K. S. Yau, Z. Lin, L. Zhou, C. Ran, L. Xiao, C. Wu, and J. Schmidhuber · 2023
Cited alongside, same era.
Large language models are few-shot testers: Exploring llm-based general bug reproduction, 2023
S. Kang, J. Yoon, and S. Yoo · 2023
Cited alongside, same era.
R. T. McCoy, S. Yao, D. Friedman, M. Hardy, and T. L. Griffiths · 2023
Cited alongside, same era.
Mlagentbench: Evaluating language agents on machine learning experimentation, 2024
Q. Huang, J. Vora, P. Liang, and J. Leskovec · 2024
Closest in time.
Livecodebench: Holistic and contamination free evaluation of large language models for code, 2024
N. Jain, K. Han, A. Gu, W.-D. Li, F. Yan, T. Zhang, S. Wang, A. Solar-Lezama, K. Sen, and I. Stoica · 2024
Closest in time.
SWE-bench: Can language models resolve real-world github issues?
C. E. Jimenez, J. Yang, A. Wettig, S. Yao, K. Pei, O. Press, and K. R. Narasimhan · 2024
Closest in time.
Visualwebarena: Evaluating multimodal agents on realistic visual web tasks, 2024
J. Y. Koh, R. Lo, L. Jang, V. Duvvur, M. C. Lim, P.-Y. Huang, G. Neubig, S. Zhou, R. Salakhutdinov, and D. Fried · 2024
Closest in time.
Octopack: Instruction tuning code large language models
N. Muennighoff, Q. Liu, A. R. Zebaze, Q. Zheng, B. Hui, T. Y. Zhuo, S. Singh, X. Tang, L. V. Werra, and S. Longpre · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Measuring and narrowing the compositionality gap in language models
O. Press, M. Zhang, S. Min, L. Schmidt, N. Smith, and M. Lewis · 2023
Cited alongside, same era.
Reflexion: Language agents with verbal reinforcement learning, 2023
N. Shinn, F. Cassano, E. Berman, A. Gopinath, K. Narasimhan, and S. Yao · 2023
Cited alongside, same era.
An analysis of the automatic bug fixing performance of chatgpt, 2023
D. Sobania, M. Briesch, C. Hanna, and J. Petke · 2023
Cited alongside, same era.
Hierarchical prompting assists large language model on web navigation, 2023
A. Sridhar, R. Lo, F. F. Xu, H. Zhu, and S. Zhou · 2023
Cited alongside, same era.
Cognitive architectures for language agents, 2023
T. Sumers, S. Yao, K. Narasimhan, and T. L. Griffiths · 2023
Cited alongside, same era.
The rise and potential of large language model based agents: A survey, 2023
Z. Xi, W. Chen, X. Guo, W. He, Y. Ding, B. Hong, M. Zhang, J. Wang, S. Jin, E. Zhou, R. Zheng, X. Fan, X. Wang, L. Xiong, Y. Zhou, W. Wang, C. Jiang, Y. Zou, X. Liu, Z. Yin, S. Dou, R. Weng, W. Cheng, Q. Zhang, W. Qin, Y. Zheng, X. Qiu, X. Huang, and T. Gui · 2023
Cited alongside, same era.
Universal fuzzing via large language models
C. S. Xia, M. Paltenghi, J. L. Tian, M. Pradel, and L. Zhang · 2023
Cited alongside, same era.
C. Packer, S. Wooders, K. Lin, V. Fang, S. G. Patil, I. Stoica, and J. E. Gonzalez · 2024
Closest in time.
An empirical evaluation of llms for solving offensive security challenges, 2024
M. Shao, B. Chen, S. Jancheska, B. Dolan-Gavitt, S. Garg, R. Karri, and M. Shafique · 2024
Closest in time.
Ehragent: Code empowers large language models for few-shot complex tabular reasoning on electronic health records, 2024
W. Shi, R. Xu, Y. Zhuang, Y. Yu, J. Zhang, H. Wu, Y. Zhu, J. Ho, C. Yang, and M. D. Wang · 2024
Closest in time.
Medagents: Large language models as collaborators for zero-shot medical reasoning, 2024
X. Tang, A. Zou, Z. Zhang, Z. Li, Y. Zhao, X. Zhang, A. Cohan, and M. Gerstein · 2024
Closest in time.
An in-context learning agent for formal theorem-proving, 2024
A. Thakur, G. Tsoukalas, Y. Wen, J. Xin, and S. Chaudhuri · 2024
Closest in time.
Automating the enterprise with foundation models, 2024
M. Wornow, A. Narayan, K. Opsahl-Ong, Q. McIntyre, N. H. Shah, and C. Re · 2024
Closest in time.
Os-copilot: Towards generalist computer agents with self-improvement, 2024
Z. Wu, C. Han, Z. Ding, Z. Weng, Z. Liu, S. Yao, T. Yu, and L. Kong · 2024
Closest in time.
Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments, 2024
T. Xie, D. Zhang, J. Chen, X. Li, S. Zhao, R. Cao, T. J. Hua, Z. Cheng, D. Shin, F. Lei, Y. Liu, Y. Xu, S. Zhou, S. Savarese, C. Xiong, V. Zhong, and T. Yu · 2024
Closest in time.
Large language models for test-free fault localization
A. Z. H. Yang, C. Le Goues, R. Martins, and V. Hellendoorn · 2024
Closest in time.
Self-taught optimizer (stop): Recursively self-improving code generation, 2024
E. Zelikman, E. Lorch, L. Mackey, and A. T. Kalai · 2024
Closest in time.
Training language model agents without modifying language models, 2024
S. Zhang, J. Zhang, J. Liu, L. Song, C. Wang, R. Krishna, and Q. Wu · 2024
Closest in time.
A survey on large language model based autonomous agents
L. Wang, C. Ma, X. Feng, Z. Zhang, H. Yang, J. Zhang, Z. Chen, J. Tang, X. Chen, Y. Lin, W. X. Zhao, Z. Wei, and J. Wen · 2095
Closest in time.