Fetching the paper…
Reading the bibliography…
One way to address safety risks from large language models (LLMs) is to censor dangerous knowledge from their training data.
Noise-tolerant learning, the parity problem, and the statistical query model
A. Blum, A. Kalai, and H. Wasserman · 2003
Earlier work this paper cites.
Fast learning requires good memory: A time-space lower bound for parity learning
R. Raz · 2018
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel, et al · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2021
Earlier work this paper cites.
Locating and editing factual associations in gpt
K. Meng, D. Bau, A. Andonian, and Y. Belinkov · 2022
Earlier work this paper cites.
A property induction framework for neural language models
K. Misra, J. T. Rayz, and A. Ettinger · 2022
Earlier work this paper cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
A. Srivastava, A. Rastogi, A. Rao, A. A. M. Shoeb, A. Abid, A. Fisch, A. R. Brown, A. Santoro, A. Gupta, A. Garriga-Alonso, et al · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Earlier work this paper cites.
Managing ai risks in an era of rapid progress
Y. Bengio, G. Hinton, A. Yao, D. Song, P. Abbeel, Y. N. Harari, Y.-Q. Zhang, L. Xue, S. Shalev-Shwartz, G. Hadfield, et al · 2023
Earlier work this paper cites.
Implicit meta-learning may lead language models to trust more reliable sources
D. Krasheninnikov, E. Krasheninnikov, B. Mlodozeniec, T. Maharaj, and D. Krueger · 2023
Cited alongside, same era.
Tell, don’t show: Declarative facts influence how llms generalize
A. Meinke and O. Evans · 2023
Cited alongside, same era.
Can lms learn new entities from descriptions? challenges in propagating injected knowledge
Y. Onoe, M. J. Zhang, S. Padmanabhan, G. Durrett, and E. Choi · 2023
Cited alongside, same era.
Mquake: Assessing knowledge editing in language models via multi-hop questions
Z. Zhong, Z. Wu, C. D. Manning, C. Potts, and D. Chen · 2023
Cited alongside, same era.
What algorithms can transformers learn? a study in length generalization
The case for ensuring that powerful AIs are controlled
R. Greenblatt and B. Shlegeris · 2024
Closest in time.
Sleeper agents: Training deceptive llms that persist through safety training
E. Hubinger, C. Denison, J. Mu, M. Lambert, M. Tong, M. MacDiarmid, T. Lanham, D. M. Ziegler, T. Maxwell, N. Cheng, et al · 2024
Closest in time.
The Operational Risks of AI in Large-Scale Biological Attacks: Results of a Red-Team Study
C. A. Mouton, C. Lucas, and E. Guest · 2024
Closest in time.
Openai documentation
OpenAI · 2024
Closest in time.
Openai api, 2024
OpenAI · 2024
Closest in time.
Testing the general deductive reasoning capacity of large language models using ood examples
A. Saparov, R. Y. Pang, V. Padmakumar, N. Joshi, M. Kazemi, N. Kim, and H. He · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Zhou, A. Bradley, E. Littwin, N. Razin, O. Saremi, J. Susskind, S. Bengio, and P. Nakkiran · 2023
Cited alongside, same era.
Llama 3 model card, 2024
AI@Meta · 2024
Cited alongside, same era.
Evaluating the ripple effects of knowledge editing in language models
R. Cohen, E. Biran, O. Yoran, A. Globerson, and M. Geva · 2024
Cited alongside, same era.
Without specific countermeasures, the easiest path to transformative ai likely leads to ai takeover
A. Cotra · 2024
Cited alongside, same era.
Geonames database, 2024
GeoNames · 2024
Cited alongside, same era.
Physics of language models: Part 3.2, knowledge manipulation
Z. Allen-Zhu and Y. Li
Cited in the paper.
Physics of language models: Part 3.1, knowledge storage and extraction
Z. Allen-Zhu and Y. Li
Cited in the paper.
Path independent equilibrium models can better exploit test-time computation
C. Anil, A. Pokle, K. Liang, J. Treutlein, Y. Wu, S. Bai, J. Z. Kolter, and R. B. Grosse
Cited in the paper.
Closest in time.
Do large language models latently perform multi-hop reasoning?
S. Yang, E. Gribovskaya, N. Kassner, M. Geva, and S. Riedel · 2024
Closest in time.
Llamafactory: Unified efficient fine-tuning of 100+ language models
Y. Zheng, R. Zhang, J. Zhang, Y. Ye, Z. Luo, and Y. Ma · 2024
Closest in time.