Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) present a dual-use dilemma: they enable beneficial applications while harboring potential for harm, particularly through conversational interactions.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020) · 1901
Earlier work this paper cites.
A mathematical theory of communication
Shannon, C. E. (1948) · 1948
Earlier work this paper cites.
Three approaches to the quantitative definition of information
Kolmogorov, A. N. (1965) · 1965
Earlier work this paper cites.
Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning
Liu, H., Tam, D., Muqeeth, M., Mohta, J., Huang, T., Bansal, M., and Raffel, C. A. (2022) · 1965
Earlier work this paper cites.
On the length of programs for computing finite binary sequences
Chaitin, G. J. (1966) · 1966
Earlier work this paper cites.
Universal sequential search problems
Levin, L. A. (1973) · 1973
Earlier work this paper cites.
Conceptualizing conversational complexity
Daly, J. A., Bell, R. A., Glenn, P. J., and Lawrence, S. (1985) · 1985
Earlier work this paper cites.
Highly optimized tolerance: A mechanism for power laws in designed systems
Carlson, J. M. and Doyle, J. (1999) · 1999
Earlier work this paper cites.
Navigation in a small world
Kleinberg, J. M. (2000) · 2000
Earlier work this paper cites.
Information theory, inference and learning algorithms
MacKay, D. J. (2003) · 2003
Earlier work this paper cites.
Common pitfalls using the normalized compression distance: What to watch out for in a compressor
Alfonseca, M., Cebrián, M., and Ortega, A. (2005) · 2005
Earlier work this paper cites.
The google similarity distance
Cilibrasi, R. L. and Vitanyi, P. M. (2007) · 2007
Earlier work this paper cites.
Emergence of zipf’s law in the evolution of communication
Corominas-Murtra, B., Fortuny, J., and Solé, R. V. (2011) · 2011
Earlier work this paper cites.
The structure and dynamics of networks
Newman, M., Barabási, A.-L., and Watts, D. J. (2011) · 2011
Earlier work this paper cites.
Lossless data compression with neural networks
Bellard, F. (2019) · 2019
Earlier work this paper cites.
An introduction to Kolmogorov complexity and its applications, 4th Edition
Li, M., Vitányi, P., et al. (2019) · 2019
Earlier work this paper cites.
Machine behaviour
Rahwan, I., Cebrian, M., Obradovich, N., Bongard, J., Bonnefon, J.-F., Breazeal, C., Crandall, J. W., Christakis, N. A., Couzin, I. D., Jackson, M. O., et al. (2019) · 2019
Earlier work this paper cites.
A Review of Methods for Estimating Algorithmic Complexity: Options, Challenges, and New Directions
Zenil, H. (2020) · 2020
Earlier work this paper cites.
On the safety of conversational models: Taxonomy, dataset, and benchmark
Sun, H., Xu, G., Deng, J., Cheng, J., Zheng, C., Zhou, H., Peng, N., Zhu, X., and Huang, M. (2021) · 2021
Cited alongside, same era.
Ethical and social risks of harm from language models
Weidinger, L., Mellor, J., Rauh, M., Griffin, C., Uesato, J., Huang, P.-S., Cheng, M., Glaese, M., Balle, B., Kasirzadeh, A., et al. (2021) · 2021
Cited alongside, same era.
Shelley: A crowd-sourced collaborative horror writer
Yanardag, P., Cebrian, M., and Rahwan, I. (2021) · 2021
Cited alongside, same era.
Talking with strangers is surprisingly informative
Atir, S., Wald, K. A., and Epley, N. (2022) · 2022
Cited alongside, same era.
Constitutional ai: Harmlessness from ai feedback
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., et al. (2022) · 2022
Llm censorship: A machine learning challenge or a computer security problem?
Glukhov, D., Shumailov, I., Gal, Y., Papernot, N., and Papyan, V. (2023) · 2023
Later among the works it cites.
Generative AI red teaming challenge: Transparency report
Humane Intelligence (2024) · 2023
Later among the works it cites.
Evaluating language-model agents on realistic autonomous tasks
Kinniment, M., Sato, L. J. K., Du, H., Goodrich, B., Hasin, M., Chan, L., Miles, L. H., Lin, T. R., Wijk, H., Burget, J., et al. (2023) · 2023
Later among the works it cites.
Jailbreaking chatgpt via prompt engineering: An empirical study
Liu, Y., Deng, G., Xu, Z., Li, Y., Zheng, Y., Zhang, Y., Zhao, L., Zhang, T., and Liu, Y. (2023) · 2023
Later among the works it cites.
A Conversation With Bing’s Chatbot Left Me Deeply Unsettled
Roose, K. (2023) · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Emergent language-based coordination in deep multi-agent systems
Baroni, M., Dessì, R., and Lazaridou, A. (2022) · 2022
Cited alongside, same era.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Ganguli, D., Lovitt, L., Kernion, J., Askell, A., Bai, Y., Kadavath, S., Mann, B., Perez, E., Schiefer, N., Ndousse, K., et al. (2022) · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. (2022) · 2022
Cited alongside, same era.
Red teaming language models with language models
Perez, E., Huang, S., Song, F., Cai, T., Ring, R., Aslanides, J., Glaese, A., McAleese, N., and Irving, G. (2022) · 2022
Cited alongside, same era.
Large pre-trained language models contain human-like biases of what is right and wrong to do
Schramowski, P., Turan, C., Andersen, N., Rothkopf, C. A., and Kersting, K. (2022) · 2022
Cited alongside, same era.
Methods and Applications of Algorithmic Complexity: Beyond Statistical Lossless Compression
Zenil, H., Toscano, F. S., and Gauvrit, N. (2022) · 2022
Cited alongside, same era.
Hh-rlhf: Codebase for rlhf (reinforcement learning from human feedback)
Anthropic (2023) · 2023
Cited alongside, same era.
Role play with large language models
Shanahan, M., McDonell, K., and Reynolds, L. (2023) · 2023
Later among the works it cites.
Towards Understanding Sycophancy in Language Models
Sharma, M., Tong, M., Korbak, T., Duvenaud, D., Askell, A., Bowman, S. R., Cheng, N., Durmus, E., Hatfield-Dodds, Z., Johnston, S. R., Kravec, S., Maxwell, T., McCandlish, S., Ndousse, K., Rausch, O., Schiefer, N., Yan, D., Zhang, M., and Perez, E. (2023) · 2023
Later among the works it cites.
Red teaming language model detectors with language models
Shi, Z., Wang, Y., Yin, F., Chen, X., Chang, K.-W., and Hsieh, C.-J. (2023) · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. (2023) · 2023
Later among the works it cites.
On the roles of function and selection in evolving systems
Wong, M. L., Cleland, C. E., Arend Jr, D., Bartlett, S., Cleaves, H. J., Demarest, H., Prabhu, A., Lunine, J. I., and Hazen, R. M. (2023) · 2023
Later among the works it cites.
Instructions as backdoors: Backdoor vulnerabilities of instruction tuning for large language models
Xu, J., Ma, M. D., Wang, F., Xiao, C., and Chen, M. (2023) · 2023
Later among the works it cites.
Complementary explanations for effective in-context learning
Ye, X., Iyer, S., Celikyilmaz, A., Stoyanov, V., Durrett, G., and Pasunuru, R. (2023) · 2023
Later among the works it cites.
Many-shot jailbreaking
Anil, C., Durmus, E., Sharma, M., Benton, J., Kundu, S., Batson, J., Rimsky, N., Tong, M., Mu, J., Ford, D., et al. (2024) · 2024
Closest in time.
From ”um” to ”yeah”: Producing, predicting, and regulating information flow in human conversation
Bergey, C. A. and DeDeo, S. (2024) · 2024
Closest in time.
A complexity-based theory of compositionality
Elmoznino, E., Jiralerspong, T., Bengio, Y., and Lajoie, G. (2024) · 2024
Closest in time.
Evaluating frontier models for dangerous capabilities
Phuong, M., Aitchison, M., Catt, E., Cogan, S., Kaskasoli, A., Krakovna, V., Lindner, D., Rahtz, M., Assael, Y., Hodkinson, S., et al. (2024) · 2024
Closest in time.
Great, now write an article about that: The crescendo multi-turn llm jailbreak attack
Russinovich, M., Salem, A., and Eldan, R. (2024) · 2024
Closest in time.