Fetching the paper…
Reading the bibliography…
The rise of Agent AI and Large Language Model-powered Multi-Agent Systems (LLM-MAS) has underscored the need for responsible and dependable system operation.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 1909
Earlier work this paper cites.
A pragmatic bdi architecture
Fischer, K., Müller, J. P., and Pischel, M · 1995
Earlier work this paper cites.
Comparison of probability and possibility for design against catastrophic failure under uncertainty
Nikolaidis, E., Chen, S., Cudney, H., Haftka, R. T., and Rosca, R · 2004
Earlier work this paper cites.
Conflicting agents: conflict management in multi-agent systems , volume 1
Tessier, C., Chaudron, L., and Müller, H.-J · 2005
Earlier work this paper cites.
Firewall configuration: An application of multiagent metalevel argumentation
Applebaum, A., Li, Z., Levitt, K., Parsons, S., Rowe, J., and Sklar, E. I · 2016
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Irving, G., Christiano, P., and Amodei, D · 2018
Earlier work this paper cites.
Preference learning along multiple criteria: A game-theoretic perspective
Bhatia, K., Pananjady, A., Bartlett, P., Dragan, A., and Wainwright, M. J · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., et al · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F · 2020
Earlier work this paper cites.
What do you mean? a review on recovery strategies to overcome conversational breakdowns of conversational agents
Benner, D., Elshan, E., Schöbel, S., and Janson, A · 2021
Earlier work this paper cites.
Conformal bayesian computation
Fong, E. and Holmes, C. C · 2021
Earlier work this paper cites.
A survey of deep meta-learning
Huisman, M., Van Rijn, J. N., and Plaat, A · 2021
Earlier work this paper cites.
Retrieval augmentation reduces hallucination in conversation
Shuster, K., Poff, S., Chen, M., Kiela, D., and Weston, J · 2021
Earlier work this paper cites.
Neural-symbolic integration: A compositional perspective
Tsamoura, E., Hospedales, T., and Michael, L · 2021
Earlier work this paper cites.
Information theoretic approach to detect collusion in multi-agent games
Bonjour, T., Aggarwal, V., and Bhargava, B · 2022
Earlier work this paper cites.
Improving alignment of dialogue agents via targeted human judgements, 2022
Glaese, A., McAleese, N., Trębacz, M., Aslanides, J., Firoiu, V., Ewalds, T., Rauh, M., Weidinger, L., Chadwick, M., Thacker, P., Campbell-Gillingham, L., Uesato, J., Huang, P.-S., Comanescu, R., Yang, F., See, A., Dathathri, S., Greig, R., Chen, C., Fritz, D., Elias, J. S., Green, R., Mokrá, S., Fernando, N., Wu, B., Foley, R., Young, S., Gabriel, I., Isaac, W., Mellor, J., Hassabis, D., Kavukcuoglu, K., Hendricks, L. A., and Irving, G · 2022
Earlier work this paper cites.
Researching alignment research: Unsupervised analysis, 2022
Kirchner, J. H., Smith, L., Thibodeau, J., McDonell, K., and Reynolds, L · 2022
Earlier work this paper cites.
A path towards autonomous machine intelligence version 0.9. 2, 2022-06-27
LeCun, Y · 2022
Earlier work this paper cites.
Workflow provenance in the lifecycle of scientific machine learning
Souza, R., Azevedo, L. G., Lourenço, V., Soares, E., Thiago, R., Brandão, R., Civitarese, D., Vital Brazil, E., Moreno, M., Valduriez, P., et al · 2022
Earlier work this paper cites.
Threats to training: A survey of poisoning attacks and defenses on machine learning systems
Wang, Z., Ma, J., Wang, X., Hu, J., Qin, Z., and Ren, K · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Earlier work this paper cites.
Establishing data provenance for responsible artificial intelligence systems
Werder, K., Ramesh, B., and Zhang, R · 2022
Earlier work this paper cites.
Dependency tracking for risk mitigation in machine learning (ml) systems
Xu, X., Wang, C., Wang, Z., Lu, Q., and Zhu, L · 2022
Earlier work this paper cites.
Multi-agent planning based on causal graphs: From theory to experiment
Zeng, L., Chen, G., and Yang, B · 2022
Earlier work this paper cites.
What, indeed, is an achievable provable guarantee for learning-enabled safety-critical systems
Bensalem, S., Cheng, C.-H., Huang, W., Huang, X., Wu, C., and Zhao, X · 2023
Earlier work this paper cites.
Sparks of artificial general intelligence: Early experiments with gpt-4, 2023
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., Nori, H., Palangi, H., Ribeiro, M. T., and Zhang, Y · 2023
Earlier work this paper cites.
Enhancing chat language models by scaling high-quality instructional conversations
Ding, N., Chen, Y., Xu, B., Qin, Y., Hu, S., Liu, Z., Sun, M., and Zhou, B · 2023
Earlier work this paper cites.
RAFT: Reward ranked finetuning for generative foundation model alignment
Dong, H., Xiong, W., Goyal, D., Zhang, Y., Chow, W., Pan, R., Diao, S., Zhang, J., SHUM, K., and Zhang, T · 2023
Earlier work this paper cites.
Bridging the gap: A survey on integrating (human) feedback for natural language generation
Fernandes, P., Madaan, A., Liu, E., Farinhas, A., Martins, P. H., Bertsch, A., de Souza, J. G. C., Zhou, S., Wu, T., Neubig, G., and Martins, A. F. T · 2023
Earlier work this paper cites.
Händler, T · 2023
Earlier work this paper cites.
The safety filter: A unified view of safety-critical control in autonomous systems
Hsu, K.-C., Hu, H., and Fisac, J. F · 2023
Earlier work this paper cites.
Language models, agent models, and world models: The law for machine reasoning and planning, 2023
Hu, Z. and Shu, T · 2023
Earlier work this paper cites.
Large language models can self-improve
Huang, J., Gu, S., Hou, L., Wu, Y., Wang, X., Yu, H., and Han, J · 2023
Earlier work this paper cites.
The consensus game: Language model generation via equilibrium search
Jacob, A. P., Shen, Y., Farina, G., and Andreas, J · 2023
Earlier work this paper cites.
Survey of hallucination in natural language generation
Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., and Fung, P · 2023
Earlier work this paper cites.
The past, present and better future of feedback learning in large language models for subjective human preferences and values
Kirk, H. R., Bean, A. M., Vidgen, B., Röttger, P., and Hale, S. A · 2023
Earlier work this paper cites.
Rlaif vs. rlhf: Scaling reinforcement learning from human feedback with ai feedback
Lee, H., Phatale, S., Mansoor, H., Lu, K., Mesnard, T., Bishop, C., Carbune, V., and Rastogi, A · 2023
Earlier work this paper cites.
Reimagining multi-criterion decision making by data-driven methods based on machine learning: A literature review
Liao, H., He, Y., Wu, X., Wu, Z., and Bausys, R · 2023
Earlier work this paper cites.
Pan, L., Saxon, M., Xu, W., Nathani, D., Wang, X., and Wang, W. Y · 2023
Earlier work this paper cites.
Instruction tuning with gpt-4, 2023
Peng, B., Li, C., He, P., Galley, M., and Gao, J · 2023
Cited alongside, same era.
Phelps, S. and Ranson, R · 2023
Cited alongside, same era.
Robots that ask for help: Uncertainty alignment for large language model planners
Ren, A., Dixit, A., Bodrova, A., Singh, S., Tu, S., Brown, N., Xu, P., Takayama, L., Xia, F., Varley, J., Xu, Z., Sadigh, D., Zeng, A., and Majumdar, A · 2023
Cited alongside, same era.
" good enough" agents: Investigating reliability imperfections in human-ai interactions across parallel task domains
Rodriguez, S. S., Karahalios, K., Lane, H. C., Chin, J., and Schaffer, J · 2023
Cited alongside, same era.
MoralDial: A framework to train and evaluate moral dialogue systems via moral discussions
Towards uncertainty-aware language agent
Han, J., Buntine, W., and Shareghi, E · 2024
Later among the works it cites.
Adaptive guardrails for large language models via trust modeling and in-context learning
Hu, J., Dong, Y., and Huang, X · 2024
Later among the works it cites.
A survey of safety and trustworthiness of large language models through the lens of verification and validation
Huang, X., Ruan, W., Huang, W., Jin, G., Dong, Y., Wu, C., Bensalem, S., Mu, R., Qi, Y., Zhao, X., et al · 2024
Later among the works it cites.
Flooding spread of manipulated knowledge in llm-based multi-agent communities
Ju, T., Wang, Y., Ma, X., Cheng, P., Zhao, H., Wang, Y., Liu, L., Xie, J., Zhang, Z., and Liu, G · 2024
Later among the works it cites.
Mediq: Question-asking llms for adaptive and reliable clinical reasoning
Li, S. S., Balachandran, V., Feng, S., Ilgen, J., Pierson, E., Koh, P. W., and Tsvetkov, Y · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sun, H., Zhang, Z., Mi, F., Wang, Y., Liu, W., Cui, J., Wang, B., Liu, Q., and Huang, M · 2023
Cited alongside, same era.
Self-criticism: Aligning large language models with their understanding of helpfulness, honesty, and harmlessness
Tan, X., Shi, S., Qiu, X., Qu, C., Qi, Z., Xu, Y., and Qi, Y · 2023
Cited alongside, same era.
Alpaca: A strong, replicable instruction-following model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Cited alongside, same era.
Creating large language model applications utilizing langchain: A primer on developing llm apps fast
Topsakal, O. and Akinci, T. C · 2023
Cited alongside, same era.
Zephyr: Direct distillation of lm alignment, 2023
Tunstall, L., Beeching, E., Lambert, N., Rajani, N., Rasul, K., Belkada, Y., Huang, S., von Werra, L., Fourrier, C., Habib, N., Sarrazin, N., Sanseviero, O., Rush, A. M., and Wolf, T · 2023
Cited alongside, same era.
Secure consensus of multi-agent systems under denial-of-service attacks
Wen, G., Wang, P., Lv, Y., Chen, G., and Zhou, J · 2023
Cited alongside, same era.
Towards reasoning in large language models via multi-agent peer review collaboration, 2023
Xu, Z., Shi, S., Hu, B., Yu, J., Li, D., Zhang, M., and Wu, Y · 2023
Cited alongside, same era.
Yang, Y., Li, H., Wang, Y., and Wang, Y · 2023
Cited alongside, same era.
Later among the works it cites.
Encouraging divergent thinking in large language models through multi-agent debate
Liang, T., He, Z., Jiao, W., Wang, X., Wang, Y., Wang, R., Yang, Y., Shi, S., and Tu, Z · 2024
Later among the works it cites.
Strategic collusion of LLM agents: Market division in multi-commodity competitions
Lin, R. Y., Ojha, S., Cai, K., and Chen, M · 2024
Later among the works it cites.
Secret collusion among AI agents: Multi-agent deception via steganography
Motwani, S. R., Baranchuk, M., Strohmeier, M., Bolina, V., Torr, P., Hammond, L., and de Witt, C. S · 2024
Later among the works it cites.
Agentcoord: Visually exploring coordination strategy for llm-based multi-agent collaboration
Pan, B., Lu, J., Wang, K., Zheng, L., Wen, Z., Feng, Y., Zhu, M., and Chen, W · 2024
Later among the works it cites.
In-context unlearning: Language models as few-shot unlearners
Pawelczyk, M., Neel, S., and Lakkaraju, H · 2024
Later among the works it cites.
Jailbreaking and mitigation of vulnerabilities in large language models
Peng, B., Bi, Z., Niu, Q., Liu, M., Feng, P., Wang, T., Yan, L. K., Wen, Y., Zhang, Y., and Yin, C. H · 2024
Later among the works it cites.
Hwmp-based secure communication of multi-agent systems
Ren, S., Liu, J., Ge, S. S., and Li, D · 2024
Later among the works it cites.
Large language model uncertainty measurement and calibration for medical diagnosis and treatment
Savage, T., Wang, J., Gallo, R., Boukil, A., Patel, V., Ahmad Safavi-Naini, S. A., Soroush, A., and Chen, J. H · 2024
Later among the works it cites.
In-context learning agents are asymmetric belief updaters
Schubert, J. A., Jagadish, A. K., Binz, M., and Schulz, E · 2024
Later among the works it cites.
Shorinwa, O., Mei, Z., Lidard, J., Ren, A. Z., and Majumdar, A · 2024
Later among the works it cites.
Moral alignment for llm agents
Tennant, E., Hailes, S., and Musolesi, M · 2024
Later among the works it cites.
“tipping the balance”: Human intervention in large language model multi-agent debate
Triem, H. and Ding, Y · 2024
Later among the works it cites.
Tsai, Y.-H. H., Talbott, W., and Zhang, J · 2024
Later among the works it cites.
Improving decision-making in open-world agents with conformal prediction and monty hall
Vishwakarma, H., Mishler, A., Cook, T., Dalmasso, N., Raman, N., and Ganesh, S · 2024
Later among the works it cites.
Wei, H., He, S., Xia, T., Wong, A., Lin, J., and Han, M · 2024
Later among the works it cites.
Shall we team up: Exploring spontaneous cooperation of competing llm agents
Wu, Z., Peng, R., Zheng, S., Liu, Q., Han, X., Kwon, B., Onizuka, M., Tang, S., and Xiao, C · 2024
Later among the works it cites.
MAgIC: Investigation of large language model powered multi-agent in cognition, adaptability, rationality and collaboration
Xu, L., Hu, Z., Zhou, D., Ren, H., Dong, Z., Keutzer, K., Ng, S.-K., and Feng, J · 2024
Later among the works it cites.
The earth is flat because…: Investigating LLMs’ belief towards misinformation via persuasive conversation
Xu, R., Lin, B., Yang, S., Zhang, T., Shi, W., Zhang, T., Fang, Z., Xu, W., and Qiu, H · 2024
Later among the works it cites.
Unsupervised information refinement training of large language models for retrieval-augmented generation
Xu, S., Pang, L., Yu, M., Meng, F., Shen, H., Cheng, X., and Zhou, J · 2024
Later among the works it cites.
Confidence calibration and rationalization for LLMs via multi-agent deliberation
Yang, R., Rajagopal, D., Hayati, S. A., Hu, B., and Kang, D · 2024
Later among the works it cites.
LLM4drive: A survey of large language models for autonomous driving
Yang, Z., Jia, X., Li, H., and Yan, J · 2024
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T., Cao, Y., and Narasimhan, K · 2024
Later among the works it cites.
Yoffe, L., Amayuelas, A., and Wang, W. Y · 2024
Later among the works it cites.
Advancing llm reasoning generalists with preference trees
Yuan, L., Cui, G., Wang, H., Ding, N., Wang, X., Deng, J., Shan, B., Chen, H., Xie, R., Lin, Y., et al · 2024
Later among the works it cites.
Uncertainty is fragile: Manipulating uncertainty in large language models
Zeng, Q., Jin, M., Yu, Q., Wang, Z., Hua, W., Zhou, Z., Sun, G., Meng, Y., Ma, S., Wang, Q., Juefei-Xu, F., Ding, K., Yang, F., Tang, R., and Zhang, Y · 2024
Later among the works it cites.
Exploring collaboration mechanisms for LLM agents: A social psychology view
Zhang, J., Xu, X., Zhang, N., Liu, R., Hooi, B., and Deng, S · 2024
Later among the works it cites.
An electoral approach to diversify LLM-based multi-agent collective decision-making
Zhao, X., Wang, K., and Peng, W · 2024
Later among the works it cites.
Evaluating uncertainty-based failure detection for closed-loop llm planners
Zheng, Z., Feng, Q., Li, H., Knoll, A., and Feng, J · 2024
Later among the works it cites.
Semantic information extraction and multi-agent communication optimization based on generative pre-trained transformer
Zhou, L., Deng, X., Wang, Z., Zhang, X., Dong, Y., Hu, X., Ning, Z., and Wei, J · 2024
Later among the works it cites.
Multi-agent consensus seeking via large language models, 2025
Chen, H., Ji, W., Xu, L., and Zhao, S · 2025
Closest in time.
Security and privacy challenges of large language models: A survey
Das, B. C., Amini, M. H., and Wu, Y · 2025
Closest in time.
Llm-based multi-agent systems for software engineering: Literature review, vision and the road ahead
He, J., Treude, C., and Lo, D · 2025
Closest in time.
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2025
Closest in time.
Zeeshan, T., Kumar, A., Pirttikangas, S., and Tarkoma, S · 2025
Closest in time.