Fetching the paper…
Reading the bibliography…
As large-scale language models increasingly impact safety-critical domains, ensuring their reliable adherence to well-defined principles remains a fundamental challenge.
“The Basic AI Drives”
Stephen. Omohundro · 2008
Earlier work this paper cites.
“Superintelligence: Paths, Dangers, Strategies”
Nick Bostrom · 2014
Earlier work this paper cites.
“Deep Reinforcement Learning from Human Preferences”
Paul Christiano et al · 2017
Earlier work this paper cites.
“Human compatible: Artificial intelligence and the problem of control”
Stuart Russell · 2019
Earlier work this paper cites.
“A General Language Assistant as a Laboratory for Alignment”
Amanda Askell et al · 2021
Earlier work this paper cites.
“BBQ: A Hand-Built Bias Benchmark for Question Answering”
Alicia Parrish et al · 2021
Earlier work this paper cites.
“Training language models to follow instructions with human feedback”
Long Ouyang et al · 2022
Earlier work this paper cites.
“Solving math word problems with process-and outcome-based feedback”
Jonathan Uesato et al · 2022
Earlier work this paper cites.
“Learning by Distilling Context”
Charlie Snell, Dan Klein and Ruiqi Zhong · 2022
Earlier work this paper cites.
“Constitutional AI: Harmlessness from AI Feedback”
Yuntao Bai et al · 2022
Earlier work this paper cites.
Josh Achiam et al · 2023
Earlier work this paper cites.
“Universal and transferable adversarial attacks on aligned language models”
Andy Zou et al · 2023
Earlier work this paper cites.
“Self-Refine: Iterative Refinement with Self-Feedback”
Aman Madaan et al · 2023
Earlier work this paper cites.
Liangming Pan et al · 2023
Cited alongside, same era.
“Generating sequences by learning to self-correct”
Sean Welleck et al · 2023
Cited alongside, same era.
“Large Language Model Programs”
Imanol Schlag et al · 2023
Cited alongside, same era.
“Learning to Reason with LLMs”, 2024
OpenAI · 2024
Cited alongside, same era.
Abhimanyu Dubey et al · 2024
Cited alongside, same era.
“Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context”
“Jailbreaking Black Box Large Language Models in Twenty Queries”
Patrick Chao et al · 2024
Closest in time.
“JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models”
Patrick Chao et al · 2024
Closest in time.
“Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents”
Priyanshu Kumar et al · 2024
Closest in time.
“O1 System Card”, 2024
OpenAI · 2024
Closest in time.
“GPT-4o System Card”, 2024
OpenAI · 2024
Closest in time.
“Model Card Addendum: Claude 3.5 Haiku and Upgraded Claude 3.5 Sonnet”, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Machel Reid et al · 2024
Cited alongside, same era.
“Jailbroken: How does llm safety training fail?”
Alexander Wei, Nika Haghtalab and Jacob Steinhardt · 2024
Cited alongside, same era.
“Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks”
Maksym Andriushchenko, Francesco Croce and Nicolas Flammarion · 2024
Cited alongside, same era.
“A StrongREJECT for Empty Jailbreaks”
Alexandra Souly et al · 2024
Cited alongside, same era.
“XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models”
Paul Röttger et al · 2024
Cited alongside, same era.
“Introducing the Model Spec”, 2024
OpenAI · 2024
Cited alongside, same era.
“WildChat: 1M ChatGPT Interaction Logs in the Wild”
Wenting Zhao et al · 2024
Cited alongside, same era.
Anthropic · 2024
Closest in time.
“Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context”
Google Gemini · 2024
Closest in time.
“Measuring short-form factuality in large language models”
Jason Wei et al · 2024
Closest in time.
“Direct Preference Optimization: Your Language Model is Secretly a Reward Model”
Rafael Rafailov et al · 2024
Closest in time.
“Backtracking improves generation safety”
Yiming Zhang et al · 2024
Closest in time.
“Uncovering Deceptive Tendencies in Language Models: A Simulated Company AI Assistant”
Olli Järviniemi and Evan Hubinger · 2024
Closest in time.
“Deception abilities emerged in large language models”
Thilo Hagendorff · 2024
Closest in time.