Fetching the paper…
Reading the bibliography…
There is an increasing interest in using language models (LMs) for automated decision-making, with multiple countries actively testing LMs to aid in military crisis decision-making.
A new measure of rank correlation
Maurice G Kendall · 1938
Earlier work this paper cites.
War games and national security with a grain of SALT
Garry D Brewer and Bruce G Blair · 1979
Earlier work this paper cites.
False Warnings of Soviet Missile Attacks Put U.S. Forces on Alert in 1979-1980, 2020
National Security Archive · 1979
Earlier work this paper cites.
Proud prophet - 83, 1983
National Defense University · 1983
Earlier work this paper cites.
The complete wargames handbook
James F Dunnigan · 1992
Earlier work this paper cites.
This Week in EUCOM History: January 23-29, 1995, 2012
EUCOM History Office · 1995
Earlier work this paper cites.
Wargames handbook: How to play and design commercial and professional wargames
James F Dunnigan · 2000
Earlier work this paper cites.
False alarm, nuclear danger
Geoffrey Forden, Pavel Podvig, and Theodore A Postol · 2000
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
MC02 Final Report, 2002
United States Joint Forces Command · 2002
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
Automation bias in intelligent time critical decision support systems
Mary L Cummings · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman · 2018
Earlier work this paper cites.
International Humanitarian Law and the Challenges of Contemporary Armed Conflicts
International Committee of the Red Cross · 2019
Earlier work this paper cites.
MoverScore: Text generation evaluating with contextualized embeddings and earth mover distance
Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger · 2019
Earlier work this paper cites.
BERTScore: Evaluating Text Generation with BERT
Tianyi Zhang*, Varsha Kishore*, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi · 2020
Earlier work this paper cites.
Moral Choices Without Moral Language: 1950s Political-Military Wargaming at the RAND Corporation (Fall 2021)
John R Emery · 2021
Earlier work this paper cites.
”A Fine-Grained Analysis of BERTScore”
Michael Hanna and Ondřej Bojar · 2021
Earlier work this paper cites.
DeBERTa: Decoding-enhanced BERT with Disentangled Attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen · 2021
Earlier work this paper cites.
Human-level play in the game of Diplomacy by combining language models with strategic reasoning
FAIR, Anton Bakhtin, Noam Brown, Emily Dinan, Gabriele Farina, Colin Flaherty, Daniel Fried, Andrew Goff, Jonathan Gray, Hengyuan Hu, et al · 2022
Earlier work this paper cites.
TruthfulQA: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans · 2022
Earlier work this paper cites.
Dangerous straits: Wargaming a future conflict over Taiwan
Stacie Pettyjohn, Becca Wasser, and Chris Dougherty · 2022
Earlier work this paper cites.
Never Give Artificial Intelligence the Nuclear Codes
Ross Andersen · 2023
Cited alongside, same era.
The First Battle of the Next War: Wargaming a Chinese Invasion of Taiwan
Mark F Cancian, Matthew Cancian, and Eric Heginbotham · 2023
Cited alongside, same era.
Palantir demos how AI can be used in the military, 2023
Ryan Daws · 2023
Cited alongside, same era.
Strategic reasoning with language models
Kanishk Gandhi, Dorsa Sadigh, and Noah D Goodman · 2023
Cited alongside, same era.
Reducing the Risks of Artificial Intelligence for Military Decision Advantage
Wyatt Hoffman and Heeu Millie Kim Kim · 2023
Cited alongside, same era.
War and peace (waragent): Large language model-based multi-agent simulation of world wars
On Large Language Models in National Security Applications
William N Caballero and Phillip R Jenkins · 2024
Closest in time.
Battlefield information and tactics engine (BITE): a multimodal large language model approach for battlespace management
Brian J Connolly · 2024
Closest in time.
Pentagon explores military uses of large language models, 2024
Eva Dou, Nitasha Tiku, and Gerrit De Vynck · 2024
Closest in time.
The Evolution of LLMs in Healthcare, 2024
Brian Eastwood · 2024
Closest in time.
Detecting hallucinations in large language models using semantic entropy
Sebastian Farquhar, Jannik Kossen, Lorenz Kuhn, and Yarin Gal · 2024
Closest in time.
Risks from Language Models for Automated Mental Healthcare: Ethics and Structure for Implementation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wenyue Hua, Lizhou Fan, Lingyao Li, Kai Mei, Jianchao Ji, Yingqiang Ge, Libby Hemphill, and Yongfeng Zhang · 2023
Cited alongside, same era.
How Large-Language Models Can Revolutionize Military Planning, April 2023
Benjamin Jensen and Dan Tadross · 2023
Cited alongside, same era.
Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation
Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar · 2023
Cited alongside, same era.
G-eval: Nlg evaluation using gpt-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu · 2023
Cited alongside, same era.
SelfCheckGPT: Zero-resource black-box hallucination detection for generative large language models
Potsawee Manakul, Adian Liusie, and Mark Gales · 2023
Cited alongside, same era.
The US Military Is Taking Generative AI Out for a Spin, 2023
Katrina Manson · 2023
Cited alongside, same era.
Block Nuclear Launch by Autonomous AI Act
Ed Markey · 2023
Cited alongside, same era.
Declan Grabb, Max Lamparth, and Nina Vasan · 2024
Closest in time.
Hadean builds large language model for British Army virtual training space, February 2024
John Hill · 2024
Closest in time.
Open-Ended Wargames with Large Language Models
Daniel P Hogan and Andrea Brennen · 2024
Closest in time.
Max Lamparth, Anthony Corso, Jacob Ganz, Oriana Skylar Mastro, Jacquelyn Schneider, and Harold Trinkunas · 2024
Closest in time.
Strategic behavior of large language models and the role of game structure versus contextual framing
Nunzio Lorè and Babak Heydari · 2024
Closest in time.
The Impact of Large Language Models in Finance: Towards Trustworthy Adoption
Carsten Maple, Alpay Sabuncuoglu, Lukasz Szpruch, Andrew Elliott, and Tony Zemaitis Gesine Reinert · 2024
Closest in time.
China have built an AI army general using LLMs like ChatGPT, 2024
Christopher McFadden · 2024
Closest in time.
Are large language models consistent over value-laden questions?
Jared Moore, Tanvi Deshpande, and Diyi Yang · 2024
Closest in time.
Models, 2024
OpenAI · 2024
Closest in time.
Large language models sensitivity to the order of options in multiple-choice questions
Pouya Pezeshkpour and Estevam Hruschka · 2024
Closest in time.
Open problems in technical ai governance
Anka Reuel, Ben Bucknall, Stephen Casper, Tim Fist, Lisa Soder, Onni Aarne, Lewis Hammond, Lujain Ibrahim, Alan Chan, Peter Wills, et al · 2024
Closest in time.
Escalation risks from language models in military and diplomatic decision-making
Juan-Pablo Rivera, Gabriel Mukobi, Anka Reuel, Max Lamparth, Chandler Smith, and Jacquelyn Schneider · 2024
Closest in time.
Evaluating Consistency and Reasoning Capabilities of Large Language Models
Yash Saxena, Sarthak Chopra, and Arunendra Mani Tripathi · 2024
Closest in time.
Scale AI Partners with DoD’s Chief Digital and Artificial Intelligence Office (CDAO) to Test and Evaluate LLMs, 2024
Scale · 2024
Closest in time.
Evaluating the moral beliefs encoded in llms
Nino Scherrer, Claudia Shi, Amir Feder, and David Blei · 2024
Closest in time.
The Most Useful Military Applications of AI in 2024 and Beyond, 2024
Sentinent Digital · 2024
Closest in time.
Ai-powered autonomous weapons risk geopolitical instability and threaten ai research
Riley Simmons-Edler, Ryan Badman, Shayne Longpre, and Kanaka Rajan · 2024
Closest in time.
LLM as a Mastermind: A Survey of Strategic Reasoning with Large Language Models
Yadong Zhang, Shaoguang Mao, Tao Ge, Xun Wang, Adrian de Wynter, Yan Xia, Wenshan Wu, Ting Song, Man Lan, and Furu Wei · 2024
Closest in time.