Fetching the paper…
Reading the bibliography…
One way to increase confidence in the outputs of Large Language Models (LLMs) is to support them with reasoning that is clear and easy to check -- a property we call legibility.
Rank analysis of incomplete block designs: I. the method of paired comparisons
R. A. Bradley and M. E. Terry · 1952
Earlier work this paper cites.
Trading group theory for randomness
L. Babai · 1985
Earlier work this paper cites.
Computationally sound proofs
S. Micali · 2000
Earlier work this paper cites.
Evasion attacks against machine learning at test time
B. Biggio, I. Corona, D. Maiorca, B. Nelson, N. Šrndić, P. Laskov, G. Giacinto, and F. Roli · 2013
Earlier work this paper cites.
Legibility and predictability of robot motion
A. D. Dragan, K. C. Lee, and S. S. Srinivasa · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow, and R. Fergus · 2013
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
SymPy: symbolic computing in Python
A. Meurer, C. P. Smith, M. Paprocki, O. Čertík, S. B. Kirpichev, M. Rocklin, A. Kumar, S. Ivanov, J. K. Moore, S. Singh, T. Rathnayake, S. Vig, B. E. Granger, R. P. Muller, F. Bonazzi, H. Gupta, S. Vats, F. Johansson, F. Pedregosa, M. J. Curry, A. R. Terrel, v. Roučka, A. Saboo, I. Fernando, S. Kulal, R. Cimrman, and A. Scopatz · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Supervising strong learners by amplifying weak experts
P. Christiano, B. Shlegeris, and D. Amodei · 2018
Earlier work this paper cites.
Adversarial examples that fool both computer vision and time-limited humans
G. Elsayed, S. Shankar, B. Cheung, N. Papernot, A. Kurakin, I. Goodfellow, and J. Sohl-Dickstein · 2018
Earlier work this paper cites.
G. Irving, P. Christiano, and D. Amodei · 2018
Earlier work this paper cites.
Scalable agent alignment via reward modeling: a research direction
J. Leike, D. Krueger, T. Everitt, M. Martic, V. Maini, and S. Legg · 2018
Earlier work this paper cites.
On evaluating adversarial robustness
N. Carlini, A. Athalye, N. Papernot, W. Brendel, J. Rauber, D. Tsipras, I. Goodfellow, A. Madry, and A. Kurakin · 2019
Earlier work this paper cites.
The knowledge complexity of interactive proof-systems
S. Goldwasser, S. Micali, and C. Rackoff · 2019
Earlier work this paper cites.
Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead
C. Rudin · 2019
Earlier work this paper cites.
Robustbench: a standardized adversarial robustness benchmark
F. Croce, M. Andriushchenko, V. Sehwag, E. Debenedetti, N. Flammarion, M. Chiang, P. Mittal, and M. Hein · 2020
Earlier work this paper cites.
AI safety via market making
E. Hubinger · 2020
Earlier work this paper cites.
Evaluating code readability and legibility: An examination of human-centric studies
D. Oliveira, R. Bruno, F. Madeiral, and F. Castor · 2020
Earlier work this paper cites.
Learning to give checkable answers with prover-verifier games
C. Anil, G. Zhang, Y. Wu, and R. Grosse · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al · 2021
Earlier work this paper cites.
Interactive proofs for verifying machine learning
S. Goldwasser, G. N. Rothblum, J. Shafer, and A. Yehudayoff · 2021
Cited alongside, same era.
Recursively summarizing books with human feedback
J. Wu, L. Ouyang, D. M. Ziegler, N. Stiennon, R. Lowe, J. Leike, and P. Christiano · 2021
Cited alongside, same era.
Constitutional AI: Harmlessness from AI feedback
Y. Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McKinnon, et al · 2022
Cited alongside, same era.
Measuring progress on scalable oversight for large language models
S. R. Bowman, J. Hyun, E. Perez, E. Chen, C. Pettit, S. Heiner, K. Lukošiūtė, A. Askell, A. Jones, A. Chen, et al · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Progress measures for grokking via mechanistic interpretability
N. Nanda, L. Chan, T. Lieberum, J. Smith, and J. Steinhardt · 2023
Later among the works it cites.
Anthropic Fall 2023 Debate Progress Update
A. Radhakrishnan · 2023
Later among the works it cites.
Question decomposition improves the faithfulness of model-generated reasoning
A. Radhakrishnan, K. Nguyen, A. Chen, C. Chen, C. Denison, D. Hernandez, E. Durmus, E. Hubinger, J. Kernion, K. Lukošiūtė, et al · 2023
Later among the works it cites.
Scalable and transferable black-box jailbreaks for language models via persona modulation
R. Shah, S. Pour, A. Tagade, S. Casper, J. Rando, et al · 2023
Later among the works it cites.
Self-Consistency Improves Chain of Thought Reasoning in Language Models
X. Wang, J. Wei, D. Schuurmans, Q. V. Le, E. H. Chi, S. Narang, A. Chowdhery, and D. Zhou · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
QuALITY: Question Answering with Long Input Texts, Yes!
R. Y. Pang, A. Parrish, N. Joshi, N. Nangia, J. Phang, A. Chen, V. Padmakumar, J. Ma, J. Thompson, H. He, and S. Bowman · 2022
Cited alongside, same era.
Two-Turn Debate Doesn’t Help Humans Answer Hard Reading Comprehension Questions
A. Parrish, H. Trivedi, N. Nangia, V. Padmakumar, J. Phang, A. S. Saimbhi, and S. R. Bowman · 2022
Cited alongside, same era.
Self-critiquing models for assisting human evaluators
W. Saunders, C. Yeh, J. Wu, S. Bills, L. Ouyang, J. Ward, and J. Leike · 2022
Cited alongside, same era.
Defining and characterizing reward gaming
J. Skalse, N. Howe, D. Krasheninnikov, and D. Krueger · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Cited alongside, same era.
Adversarial training for high-stakes reliability
D. Ziegler, S. Nix, L. Chan, T. Bauman, P. Schmidt-Nielsen, T. Lin, A. Scherlis, N. Nabeshima, B. Weinstein-Raun, D. de Haas, et al · 2022
Cited alongside, same era.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Cited alongside, same era.
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
A. Zou, Z. Wang, J. Z. Kolter, and M. Fredrikson · 2023
Later among the works it cites.
Models that prove their own correctness
N. Amit, S. Goldwasser, O. Paradise, and G. Rothblum · 2024
Closest in time.
Are aligned neural networks adversarially aligned?
N. Carlini, M. Nasr, C. A. Choquette-Choo, M. Jagielski, I. Gao, P. W. W. Koh, D. Ippolito, F. Tramer, and L. Schmidt · 2024
Closest in time.
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
C. Denison, M. MacDiarmid, F. Barez, D. Duvenaud, S. Kravec, S. Marks, N. Schiefer, R. Soklaski, A. Tamkin, J. Kaplan, et al · 2024
Closest in time.
Teaching large language models to reason with reinforcement learning
A. Havrilla, Y. Du, S. C. Raparthy, C. Nalmpantis, J. Dwivedi-Yu, M. Zhuravinskyi, E. Hambro, S. Sukhbaatar, and R. Raileanu · 2024
Closest in time.
Query-Based Adversarial Prompt Generation
J. Hayase, E. Borevkovic, N. Carlini, F. Tramèr, and M. Nasr · 2024
Closest in time.
Sleeper agents: Training deceptive LLMs that persist through safety training
E. Hubinger, C. Denison, J. Mu, M. Lambert, M. Tong, M. MacDiarmid, T. Lanham, D. M. Ziegler, T. Maxwell, N. Cheng, et al · 2024
Closest in time.
Debating with More Persuasive LLMs Leads to More Truthful Answers
A. Khan, J. Hughes, D. Valentine, L. Ruis, K. Sachan, A. Radhakrishnan, E. Grefenstette, S. R. Bowman, T. Rocktäschel, and E. Perez · 2024
Closest in time.
Distinguishing three alignment taxes, 2022
J. Leike · 2024
Closest in time.
Let’s Verify Step by Step
H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe · 2024
Closest in time.
LLM Critics Help Catch LLM Bugs
N. McAleese, Rai, J. F. C. Uribe, E. Nitishinskaya, M. Trąbacz, and J. Leike · 2024
Closest in time.
Interpretability guarantees with merlin-arthur classifiers
S. Wäldchen, K. Sharma, B. Turan, M. Zimmer, and S. Pokutta · 2024
Closest in time.
Learning Task Decomposition to Assist Humans in Competitive Programming
J. Wen, R. Zhong, P. Ke, Z. Shao, H. Wang, and M. Huang · 2024
Closest in time.
Bitonic sorter — Wikipedia, the Free Encyclopedia, 2023
Wikipedia contributors · 2024
Closest in time.
Explainability for large language models: A survey
H. Zhao, H. Chen, F. Yang, N. Liu, H. Deng, H. Cai, S. Wang, D. Yin, and M. Du · 2024
Closest in time.