Fetching the paper…
Reading the bibliography…
Large Language Model (LLMs) such as ChatGPT that exhibit generative AI capabilities are facing accelerated adoption and innovation.
Transformer-xl: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q.V., Salakhutdinov, R., 2019 · 1901
Earlier work this paper cites.
Parameter-efficient transfer learning for nlp
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., de Laroussilhe, Q., Gesmundo, A., Attariyan, M., Gelly, S., 2019 · 1902
Earlier work this paper cites.
Predicting the type and target of offensive posts in social media
Zampieri, M., Malmasi, S., Nakov, P., Rosenthal, S., Farra, N., Kumar, R., 2019 · 1902
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Zhang, T., Kishore, V., Wu, F., Weinberger, K.Q., Artzi, Y., 2019 · 1904
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., Stoyanov, V., 2019 · 1907
Earlier work this paper cites.
Release strategies and the social impacts of language models
Solaiman, I., Brundage, M., Clark, J., Askell, A., Herbert-Voss, A., Wu, J., Radford, A., Krueger, G., Kim, J.W., Kreps, S., et al., 2019 · 1908
Earlier work this paper cites.
Universal adversarial triggers for attacking and analyzing nlp
Wallace, E., Feng, S., Kandpal, N., Gardner, M., Singh, S., 2021a · 1908
Earlier work this paper cites.
Emergent tool use from multi-agent autocurricula
Baker, B., Kanitscheider, I., Markov, T., Wu, Y., Powell, G., McGrew, B., Mordatch, I., 2020 · 1909
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., Soricut, R., 2019 · 1909
Earlier work this paper cites.
Fine-tuning language models from human preferences
Ziegler, D.M., Stiennon, N., Wu, J., Brown, T.B., Radford, A., Amodei, D., Christiano, P., Irving, G., 2019 · 1909
Earlier work this paper cites.
Fine-tuning language models from human preferences
Ziegler, D.M., Stiennon, N., Wu, J., Brown, T.B., Radford, A., Amodei, D., Christiano, P., Irving, G., 2020 · 1909
Earlier work this paper cites.
Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., Zettlemoyer, L., 2019 · 1910
Earlier work this paper cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Peng, X.B., Kumar, A., Zhang, G., Levine, S., 2019 · 1910
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., Wolf, T., 2019 · 1910
Earlier work this paper cites.
A divergence minimization perspective on imitation learning methods
Ghasemipour, S.K.S., Zemel, R., Gu, S., 2019 · 1911
Earlier work this paper cites.
Negated and misprimed probes for pretrained language models: Birds can talk, but cannot fly
Kassner, N., Schütze, H., 2019 · 1911
Earlier work this paper cites.
Blockwise self-attention for long document understanding
Qiu, J., Ma, H., Levy, O., Yih, S.W.t., Wang, S., Tang, J., 2019 · 1911
Earlier work this paper cites.
Towards robust toxic content classification
Kurita, K., Belova, A., Anastasopoulos, A., 2019 · 1912
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R.A., Terry, M.E., 1952 · 1952
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation, in: Solla, S., Leen, T., Müller, K. (Eds.), Advances in Neural Information Processing Systems, MIT Press
Sutton, R.S., McAllester, D., Singh, S., Mansour, Y., 1999 · 1999
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T.B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., Amodei, D., 2020 · 2001
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning, in: Proceedings of the Nineteenth International Conference on Machine Learning, Morgan Kaufmann Publishers Inc., San Francisco, CA, USA. p. 267–274
Kakade, S., Langford, J., 2002 · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation, in: Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pp. 311–318
Papineni, K., Roukos, S., Ward, T., Zhu, W.J., 2002 · 2002
Earlier work this paper cites.
Electra: Pre-training text encoders as discriminators rather than generators
Clark, K., Luong, M.T., Le, Q.V., Manning, C.D., 2020 · 2003
Earlier work this paper cites.
Bias and fairness in natural language processing, in: Baldwin, T., Carpuat, M. (Eds.), Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP): Tutorial Abstracts, Association for Computational Linguistics, Hong Kong, China
Chang, K.W., Prabhakaran, V., Ordonez, V., 2019 · 2004
Earlier work this paper cites.
The state and fate of linguistic diversity and inclusion in the nlp world
Joshi, P., Santy, S., Budhiraja, A., Bali, K., Choudhury, M., 2020 · 2004
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries, in: Text summarization branches out, pp. 74–81
Lin, C.Y., 2004 · 2004
Earlier work this paper cites.
Bleurt: Learning robust metrics for text generation
Sellam, T., Das, D., Parikh, A.P., 2020 · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments, in: Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization, pp. 65–72
Banerjee, S., Lavie, A., 2005 · 2005
Earlier work this paper cites.
Language models are few-shot learners
Brown, T.B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D.M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., Amodei, D., 2020 · 2005
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., tau Yih, W., Rocktäschel, T., Riedel, S., Kiela, D., 2021 · 2005
Earlier work this paper cites.
Model compression, in: Proceedings of the 12th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Association for Computing Machinery, New York, NY, USA. p. 535–541
Buciluundefined, C., Caruana, R., Niculescu-Mizil, A., 2006 · 2006
Earlier work this paper cites.
Deberta: Decoding-enhanced bert with disentangled attention
He, P., Liu, X., Gao, J., Chen, W., 2020 · 2006
Earlier work this paper cites.
Awac: Accelerating online reinforcement learning with offline datasets
Nair, A., Gupta, A., Dalal, M., Levine, S., 2021 · 2006
Earlier work this paper cites.
Linformer: Self-attention with linear complexity
Wang, S., Li, B.Z., Khabsa, M., Fang, H., Ma, H., 2020 · 2006
Earlier work this paper cites.
Aligning ai with shared human values
Hendrycks, D., Burns, C., Basart, S., Critch, A., Li, J., Song, D., Steinhardt, J., 2023 · 2008
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Gehman, S., Gururangan, S., Sap, M., Choi, Y., Smith, N.A., 2020 · 2009
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., Steinhardt, J., 2021 · 2009
Earlier work this paper cites.
Learning to summarize from human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D.M., Lowe, R., Voss, C., Radford, A., Amodei, D., Christiano, P., 2022 · 2009
Earlier work this paper cites.
Crows-pairs: A challenge dataset for measuring social biases in masked language models
Nangia, N., Vania, C., Bhalerao, R., Bowman, S.R., 2020 · 2010
Earlier work this paper cites.
Concealed data poisoning attacks on nlp models
Wallace, E., Zhao, T.Z., Feng, S., Singh, S., 2021b · 2010
Earlier work this paper cites.
Extracting training data from large language models
Carlini, N., Tramer, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, U., Oprea, A., Raffel, C., 2021 · 2012
Earlier work this paper cites.
Risk-sensitive reinforcement learning
Shen, Y., Tobia, M.J., Sommer, T., Obermayer, K., 2014 · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., Dean, J., 2015 · 2015
Earlier work this paper cites.
Concrete problems in ai safety
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., Mané, D., 2016 · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Wu, Y., Schuster, M., Chen, Z., Le, Q.V., Norouzi, M., Macherey, W., Krikun, M., Cao, Y., Gao, Q., Macherey, K., et al., 2016 · 2016
Earlier work this paper cites.
Constrained policy optimization
Achiam, J., Held, D., Tamar, A., Abbeel, P., 2017 · 2017
Earlier work this paper cites.
Rasa: Open source language understanding and dialogue management
Bocklisch, T., Faulkner, J., Pawlowski, N., Nichol, A., 2017 · 2017
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
Kirkpatrick, J., Pascanu, R., Rabinowitz, N., Veness, J., Desjardins, G., Rusu, A.A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., et al., 2017 · 2017
Earlier work this paper cites.
Media manipulation and disinformation online
Marwick, A.E., Lewis, R., 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I., 2017 · 2017
Earlier work this paper cites.
Patient and consumer safety risks when using conversational assistants for medical information: an observational study of siri, alexa, and google assistant
Bickmore, T.W., Trinh, H., Olafsson, S., O’Leary, T.K., Asadi, R., Rickles, N.M., Cruz, R., 2018 · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.W., Lee, K., Toutanova, K., 2018 · 2018
Earlier work this paper cites.
Universal language model fine-tuning for text classification, in: Gurevych, I., Miyao, Y. (Eds.), Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Association for Computational Linguistics, Melbourne, Australia. pp. 328–339
Howard, J., Ruder, S., 2018 · 2018
Earlier work this paper cites.
Detecting offensive content in open-domain conversations using two stage semi-supervision
Khatri, C., Hedayatnia, B., Goel, R., Venkatesh, A., Gabriel, R., Mandal, A., 2018 · 2018
Earlier work this paper cites.
Adversarial learning of task-oriented neural dialog models
Liu, B., Lane, I., 2018 · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., 2018 · 2018
Earlier work this paper cites.
Self-attention with relative position representations
Shaw, P., Uszkoreit, J., Vaswani, A., 2018 · 2018
Earlier work this paper cites.
Ruse: Regressor using sentence embeddings for automatic machine translation evaluation, in: Proceedings of the Third Conference on Machine Translation: Shared Task Papers, pp. 751–758
Shimanaka, H., Kajiwara, T., Komachi, M., 2018 · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al., 2019 · 2019
Earlier work this paper cites.
Defending against neural fake news
Zellers, R., Holtzman, A., Rashkin, H., Bisk, Y., Farhadi, A., Roesner, F., Choi, Y., 2019 · 2019
Earlier work this paper cites.
ETC: Encoding long and structured inputs in transformers, in: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Association for Computational Linguistics, Online. pp. 268–284
Ainslie, J., Ontanon, S., Alberti, C., Cvicek, V., Fisher, Z., Pham, P., Ravula, A., Sanghai, S., Wang, Q., Yang, L., 2020 · 2020
Earlier work this paper cites.
Climbing towards nlu: On meaning, form, and understanding in the age of data, in: Proceedings of the 58th annual meeting of the association for computational linguistics, pp. 5185–5198
Bender, E.M., Koller, A., 2020 · 2020
Earlier work this paper cites.
Gpt-3: Its nature, scope, limits, and consequences
Floridi, L., Chiriatti, M., 2020 · 2020
Earlier work this paper cites.
Artificial intelligence, values, and alignment
Gabriel, I., 2020 · 2020
Earlier work this paper cites.
All the news that’s fit to fabricate: Ai-generated text as a tool of media misinformation
Kreps, S., McCain, R.M., Brundage, M., 2022 · 2020
Earlier work this paper cites.
Misinformation in action: Fake news exposure is linked to lower trust in media, higher trust in government when your side is in power
Ognyanova, K., Lazer, D., Robertson, R.E., Wilson, C., 2020 · 2020
Earlier work this paper cites.
Privacy risks of general-purpose language models, in: 2020 IEEE Symposium on Security and Privacy (SP), IEEE. pp. 1314–1331
Pan, X., Zhang, M., Ji, S., Yang, M., 2020 · 2020
Earlier work this paper cites.
Toxicity detection: Does context really matter?
Pavlopoulos, J., Sorensen, J., Dixon, L., Thain, N., Androutsopoulos, I., 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., Liu, P.J., 2020 · 2020
Earlier work this paper cites.
Information leakage in embedding models, in: Proceedings of the 2020 ACM SIGSAC conference on computer and communications security, pp. 377–390
Song, C., Raghunathan, A., 2020 · 2020
Earlier work this paper cites.
Big bird: Transformers for longer sequences
Zaheer, M., Guruganesh, G., Dubey, K.A., Ainslie, J., Alberti, C., Ontanon, S., Pham, P., Ravula, A., Wang, Q., Yang, L., et al., 2020 · 2020
Earlier work this paper cites.
Analyzing information leakage of updates to natural language models, in: Proceedings of the 2020 ACM SIGSAC conference on computer and communications security, pp. 363–375
Zanella-Béguelin, S., Wutschitz, L., Tople, S., Rühle, V., Paverd, A., Ohrimenko, O., Köpf, B., Brockschmidt, M., 2020 · 2020
Earlier work this paper cites.
A general language assistant as a laboratory for alignment
Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., DasSarma, N., Elhage, N., Hatfield-Dodds, Z., Hernandez, D., Kernion, J., Ndousse, K., Olsson, C., Amodei, D., Brown, T., Clark, J., McCandlish, S., Olah, C., Kaplan, J., 2021 · 2021
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?, in: Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pp. 610–623
Bender, E.M., Gebru, T., McMillan-Major, A., Shmitchell, S., 2021 · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al., 2021 · 2021
Earlier work this paper cites.
Quantifying social biases in nlp: A generalization and empirical comparison of extrinsic fairness metrics
Czarnowska, P., Vyas, Y., Shah, K., 2021 · 2021
Earlier work this paper cites.
Lack of transparency and potential bias in artificial intelligence data sets and algorithms: a scoping review
Daneshjou, R., Smith, M.P., Sun, M.D., Rotemberg, V., Zou, J., 2021 · 2021
Earlier work this paper cites.
A survey on bias in deep nlp
Garrido-Muñoz, I., Montejo-Ráez, A., Mart’ınez-Santiago, F., Ureña-López, L.A., 2021 · 2021
Earlier work this paper cites.
Kenton, Z., Everitt, T., Weidinger, L., Gabriel, I., Mikulik, V., Irving, G., 2021 · 2021
Earlier work this paper cites.
Bias out-of-the-box: An empirical analysis of intersectional occupational biases in popular generative language models, in: Ranzato, M., Beygelzimer, A., Dauphin, Y., Liang, P., Vaughan, J.W. (Eds.), Advances in Neural Information Processing Systems, Curran Associates, Inc.. pp. 2611–2624
Kirk, H.R., Jun, Y., Volpin, F., Iqbal, H., Benussi, E., Dreyer, F., Shtedritski, A., Asano, Y., 2021 · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
Lester, B., Al-Rfou, R., Constant, N., 2021 · 2021
Earlier work this paper cites.
Prefix-tuning: Optimizing continuous prompts for generation
Li, X.L., Liang, P., 2021 · 2021
Cited alongside, same era.
Truthfulqa: Measuring how models mimic human falsehoods
Lin, S., Hilton, J., Evans, O., 2021 · 2021
Cited alongside, same era.
What makes good in-context examples for gpt- 3 3 ?
Liu, J., Shen, D., Zhang, Y., Dolan, B., Carin, L., Chen, W., 2021 · 2021
Cited alongside, same era.
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Lu, Y., Bartolo, M., Moore, A., Riedel, S., Stenetorp, P., 2021 · 2021
Cited alongside, same era.
A survey on bias and fairness in machine learning
Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., Galstyan, A., 2021 · 2021
The unlocking spell on base llms: Rethinking alignment via in-context learning
Lin, B.Y., Ravichander, A., Lu, X., Dziri, N., Sclar, M., Chandu, K., Bhagavatula, C., Choi, Y., 2023 · 2023
Later among the works it cites.
Evaluating statistical language models as pragmatic reasoners
Lipkin, B., Wong, L., Grand, G., Tenenbaum, J.B., 2023 · 2023
Later among the works it cites.
Trustworthy llms: a survey and guideline for evaluating large language models’ alignment
Liu, Y., Yao, Y., Ton, J.F., Zhang, X., Cheng, R.G.H., Klochkov, Y., Taufiq, M.F., Li, H., 2023 · 2023
Later among the works it cites.
Using in-context learning to improve dialogue safety
Meade, N., Gella, S., Hazarika, D., Gupta, P., Jin, D., Reddy, S., Liu, Y., Hakkani-Tür, D., 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Metaicl: Learning to learn in context
Min, S., Lewis, M., Zettlemoyer, L., Hajishirzi, H., 2021 · 2021
Cited alongside, same era.
Cross-task generalization via natural language crowdsourcing instructions
Mishra, S., Khashabi, D., Baral, C., Hajishirzi, H., 2021 · 2021
Cited alongside, same era.
Probing toxic content in large pre-trained language models, in: Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 4262–4274
Ousidhoum, N., Zhao, X., Fang, T., Song, Y., Yeung, D.Y., 2021 · 2021
Cited alongside, same era.
Learning to retrieve prompts for in-context learning
Rubin, O., Herzig, J., Berant, J., 2021 · 2021
Cited alongside, same era.
Multitask prompted training enables zero-shot task generalization
Sanh, V., Webson, A., Raffel, C., Bach, S.H., Sutawika, L., Alyafeai, Z., Chaffin, A., Stiegler, A., Scao, T.L., Raja, A., et al., 2021 · 2021
Cited alongside, same era.
Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in nlp
Schick, T., Udupa, S., Schütze, H., 2021 · 2021
Cited alongside, same era.
Does knowledge distillation really work?
Stanton, S., Izmailov, P., Kirichenko, P., Alemi, A.A., Wilson, A.G., 2021 · 2021
Cited alongside, same era.
Morris, M.R., Sohl-dickstein, J., Fiedel, N., Warkentin, T., Dafoe, A., Faust, A., Farabet, C., Legg, S., 2023 · 2023
Later among the works it cites.
Misinformation monitor
NewsGuard, 2023 · 2023
Later among the works it cites.
The alignment problem from a deep learning perspective
Ngo, R., Chan, L., Mindermann, S., 2023 · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C.D., Finn, C., 2023 · 2023
Later among the works it cites.
Nemo guardrails: A toolkit for controllable and safe llm applications with programmable rails
Rebedea, T., Dinu, R., Sreedhar, M., Parisien, C., Cohen, J., 2023 · 2023
Later among the works it cites.
Are emergent abilities of large language models a mirage?
Schaeffer, R., Miranda, B., Koyejo, S., 2023 · 2023
Later among the works it cites.
Role-play with large language models
Shanahan, M., McDonell, K., Reynolds, L., 2023 · 2023
Later among the works it cites.
Towards understanding sycophancy in language models
Sharma, M., Tong, M., Korbak, T., Duvenaud, D., Askell, A., Bowman, S.R., Cheng, N., Durmus, E., Hatfield-Dodds, Z., Johnston, S.R., Kravec, S., Maxwell, T., McCandlish, S., Ndousse, K., Rausch, O., Schiefer, N., Yan, D., Zhang, M., Perez, E., 2023 · 2023
Later among the works it cites.
Sociotechnical harms of algorithmic systems: Scoping a taxonomy for harm reduction
Shelby, R., Rismani, S., Henne, K., Moon, A., Rostamzadeh, N., Nicholas, P., Yilla, N., Gallegos, J., Smart, A., Garcia, E., Virk, G., 2023 · 2023
Later among the works it cites.
Model evaluation for extreme risks
Shevlane, T., Farquhar, S., Garfinkel, B., Phuong, M., Whittlestone, J., Leung, J., Kokotajlo, D., Marchal, N., Anderljung, M., Kolt, N., Ho, L., Siddarth, D., Avin, S., Hawkins, W., Kim, B., Gabriel, I., Bolina, V., Clark, J., Bengio, Y., Christiano, P., Dafoe, A., 2023 · 2023
Later among the works it cites.
Distilling reasoning capabilities into smaller language models
Shridhar, K., Stolfo, A., Sachan, M., 2023 · 2023
Later among the works it cites.
Exploiting large language models (llms) through deception techniques and persuasion principles
Singh, S., Abri, F., Namin, A.S., 2023 · 2023
Later among the works it cites.
Evaluating the social impact of generative ai systems in systems and society
Solaiman, I., Talat, Z., Agnew, W., Ahmad, L., Baker, D., Blodgett, S.L., au2, H.D.I., Dodge, J., Evans, E., Hooker, S., Jernite, Y., Luccioni, A.S., Lusoli, A., Mitchell, M., Newman, J., Png, M.T., Strait, A., Vassilev, A., 2023 · 2023
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A.A.M., Abid, A., Fisch, A., Brown, A.R., Santoro, A., Gupta, A., Garriga-Alonso, A., Kluska, A., Lewkowycz, A., Agarwal, A., Power, A., Ray, A., Warstadt, A., Kocurek, A.W., Safaya, A., Tazarv, A., Xiang, A., Parrish, A., Nie, A., Hussain, A., Askell, A., Dsouza, A., Slone, A., Rahane, A., Iyer, A.S., Andreassen, A., Madotto, A., Santilli, A., Stuhlmüller, A., Dai, A., La, A., Lampinen, A., Zou, A., Jiang, A., Chen, A., Vuong, A., Gupta, A., Gottardi, A., Norelli, A., Venkatesh, A., Gholamidavoodi, A., Tabassum, A., Menezes, A., Kirubarajan, A., Mullokandov, A., Sabharwal, A., Herrick, A., Efrat, A., Erdem, A., Karakaş, A., Roberts, B.R., Loe, B.S., Zoph, B., Bojanowski, B., Özyurt, B., Hedayatnia, B., Neyshabur, B., Inden, B., Stein, B., Ekmekci, B., Lin, B.Y., Howald, B., Orinion, B., Diao, C., Dour, C., Stinson, C., Argueta, C., Ramírez, C.F., Singh, C., Rathkopf, C., Meng, C., Baral, C., Wu, C., Callison-Burch, C., Waites, C., Voigt, C., Manning, C.D., Potts, C., Ramirez, C., Rivera, C.E., Siro, C., Raffel, C., Ashcraft, C., Garbacea, C., Sileo, D., Garrette, D., Hendrycks, D., Kilman, D., Roth, D., Freeman, D., Khashabi, D., Levy, D., González, D.M., Perszyk, D., Hernandez, D., Chen, D., Ippolito, D., Gilboa, D., Dohan, D., Drakard, D., Jurgens, D., Datta, D., Ganguli, D., Emelin, D., Kleyko, D., Yuret, D., Chen, D., Tam, D., Hupkes, D., Misra, D., Buzan, D., Mollo, D.C., Yang, D., Lee, D.H., Schrader, D., Shutova, E., Cubuk, E.D., Segal, E., Hagerman, E., Barnes, E., Donoway, E., Pavlick, E., Rodola, E., Lam, E., Chu, E., Tang, E., Erdem, E., Chang, E., Chi, E.A., Dyer, E., Jerzak, E., Kim, E., Manyasi, E.E., Zheltonozhskii, E., Xia, F., Siar, F., Martínez-Plumed, F., Happé, F., Chollet, F., Rong, F., Mishra, G., Winata, G.I., de Melo, G., Kruszewski, G., Parascandolo, G., Mariani, G., Wang, G., Jaimovitch-López, G., Betz, G., Gur-Ari, G., Galijasevic, H., Kim, H., Rashkin, H., Hajishirzi, H., Mehta, H., Bogar, H., Shevlin, H., Schütze, H., Yakura, H., Zhang, H., Wong, H.M., Ng, I., Noble, I., Jumelet, J., Geissinger, J., Kernion, J., Hilton, J., Lee, J., Fisac, J.F., Simon, J.B., Koppel, J., Zheng, J., Zou, J., Kocoń, J., Thompson, J., Wingfield, J., Kaplan, J., Radom, J., Sohl-Dickstein, J., Phang, J., Wei, J., Yosinski, J., Novikova, J., Bosscher, J., Marsh, J., Kim, J., Taal, J., Engel, J., Alabi, J., Xu, J., Song, J., Tang, J., Waweru, J., Burden, J., Miller, J., Balis, J.U., Batchelder, J., Berant, J., Frohberg, J., Rozen, J., Hernandez-Orallo, J., Boudeman, J., Guerr, J., Jones, J., Tenenbaum, J.B., Rule, J.S., Chua, J., Kanclerz, K., Livescu, K., Krauth, K., Gopalakrishnan, K., Ignatyeva, K., Markert, K., Dhole, K.D., Gimpel, K., Omondi, K., Mathewson, K., Chiafullo, K., Shkaruta, K., Shridhar, K., McDonell, K., Richardson, K., Reynolds, L., Gao, L., Zhang, L., Dugan, L., Qin, L., Contreras-Ochando, L., Morency, L.P., Moschella, L., Lam, L., Noble, L., Schmidt, L., He, L., Colón, L.O., Metz, L., Şenel, L.K., Bosma, M., Sap, M., ter Hoeve, M., Farooqi, M., Faruqui, M., Mazeika, M., Baturan, M., Marelli, M., Maru, M., Quintana, M.J.R., Tolkiehn, M., Giulianelli, M., Lewis, M., Potthast, M., Leavitt, M.L., Hagen, M., Schubert, M., Baitemirova, M.O., Arnaud, M., McElrath, M., Yee, M.A., Cohen, M., Gu, M., Ivanitskiy, M., Starritt, M., Strube, M., Swędrowski, M., Bevilacqua, M., Yasunaga, M., Kale, M., Cain, M., Xu, M., Suzgun, M., Walker, M., Tiwari, M., Bansal, M., Aminnaseri, M., Geva, M., Gheini, M., T, M.V., Peng, N., Chi, N.A., Lee, N., Krakover, N.G.A., Cameron, N., Roberts, N., Doiron, N., Martinez, N., Nangia, N., Deckers, N., Muennighoff, N., Keskar, N.S., Iyer, N.S., Constant, N., Fiedel, N., Wen, N., Zhang, O., Agha, O., Elbaghdadi, O., Levy, O., Evans, O., Casares, P.A.M., Doshi, P., Fung, P., Liang, P.P., Vicol, P., Alipoormolabashi, P., Liao, P., Liang, P., Chang, P., Eckersley, P., Htut, P.M., Hwang, P., Miłkowski, P., Patil, P., Pezeshkpour, P., Oli, P., Mei, Q., Lyu, Q., Chen, Q., Banjade, R., Rudolph, R.E., Gabriel, R., Habacker, R., Risco, R., Millière, R., Garg, R., Barnes, R., Saurous, R.A., Arakawa, R., Raymaekers, R., Frank, R., Sikand, R., Novak, R., Sitelew, R., LeBras, R., Liu, R., Jacobs, R., Zhang, R., Salakhutdinov, R., Chi, R., Lee, R., Stovall, R., Teehan, R., Yang, R., Singh, S., Mohammad, S.M., Anand, S., Dillavou, S., Shleifer, S., Wiseman, S., Gruetter, S., Bowman, S.R., Schoenholz, S.S., Han, S., Kwatra, S., Rous, S.A., Ghazarian, S., Ghosh, S., Casey, S., Bischoff, S., Gehrmann, S., Schuster, S., Sadeghi, S., Hamdan, S., Zhou, S., Srivastava, S., Shi, S., Singh, S., Asaadi, S., Gu, S.S., Pachchigar, S., Toshniwal, S., Upadhyay, S., Shyamolima, Debnath, Shakeri, S., Thormeyer, S., Melzi, S., Reddy, S., Makini, S.P., Lee, S.H., Torene, S., Hatwar, S., Dehaene, S., Divic, S., Ermon, S., Biderman, S., Lin, S., Prasad, S., Piantadosi, S.T., Shieber, S.M., Misherghi, S., Kiritchenko, S., Mishra, S., Linzen, T., Schuster, T., Li, T., Yu, T., Ali, T., Hashimoto, T., Wu, T.L., Desbordes, T., Rothschild, T., Phan, T., Wang, T., Nkinyili, T., Schick, T., Kornev, T., Tunduny, T., Gerstenberg, T., Chang, T., Neeraj, T., Khot, T., Shultz, T., Shaham, U., Misra, V., Demberg, V., Nyamai, V., Raunak, V., Ramasesh, V., Prabhu, V.U., Padmakumar, V., Srikumar, V., Fedus, W., Saunders, W., Zhang, W., Vossen, W., Ren, X., Tong, X., Zhao, X., Wu, X., Shen, X., Yaghoobzadeh, Y., Lakretz, Y., Song, Y., Bahri, Y., Choi, Y., Yang, Y., Hao, Y., Chen, Y., Belinkov, Y., Hou, Y., Hou, Y., Bai, Y., Seid, Z., Zhao, Z., Wang, Z., Wang, Z.J., Wang, Z., Wu, Z., 2023 · 2023
Later among the works it cites.
Principle-driven self-alignment of language models from scratch with minimal human supervision
Sun, Z., Shen, Y., Zhou, Q., Zhang, H., Chen, Z., Cox, D., Yang, Y., Gan, C., 2023 · 2023
Later among the works it cites.
Ul2: Unifying language learning paradigms
Tay, Y., Dehghani, M., Tran, V.Q., Garcia, X., Wei, J., Wang, X., Chung, H.W., Shakeri, S., Bahri, D., Schuster, T., Zheng, H.S., Zhou, D., Houlsby, N., Metzler, D., 2023 · 2023
Later among the works it cites.
Chatglm-6b
THUDM, 2023 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al., 2023 · 2023
Later among the works it cites.
Zephyr: Direct distillation of lm alignment
Tunstall, L., Beeching, E., Lambert, N., Rajani, N., Rasul, K., Belkada, Y., Huang, S., von Werra, L., Fourrier, C., Habib, N., Sarrazin, N., Sanseviero, O., Rush, A.M., Wolf, T., 2023 · 2023
Later among the works it cites.
Emergent analogical reasoning in large language models
Webb, T., Holyoak, K.J., Lu, H., 2023 · 2023
Later among the works it cites.
Ensuring safe, secure, and trustworthy ai
WhiteHouse, · 2023
Later among the works it cites.
Harnessing the power of llms in practice: A survey on chatgpt and beyond
Yang, J., Jin, H., Tang, R., Han, X., Feng, Q., Jiang, H., Yin, B., Hu, X., 2023 · 2023
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T.L., Cao, Y., Narasimhan, K., 2023 · 2023
Later among the works it cites.
Counterfactual memorization in neural language models
Zhang, C., Ippolito, D., Lee, K., Jagielski, M., Tramèr, F., Carlini, N., 2023 · 2023
Later among the works it cites.
A survey of large language models
Zhao, W.X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., Du, Y., Yang, C., Chen, Y., Chen, Z., Jiang, J., Ren, R., Li, Y., Tang, X., Liu, Z., Liu, P., Nie, J.Y., Wen, J.R., 2023 · 2023
Later among the works it cites.
Large language models are human-level prompt engineers
Zhou, Y., Muresanu, A.I., Han, Z., Paster, K., Pitis, S., Chan, H., Ba, J., 2023 · 2023
Later among the works it cites.
Fine-tuning language models with advantage-induced policy alignment
Zhu, B., Sharma, H., Frujeri, F.V., Dong, S., Zhu, C., Jordan, M.I., Jiao, J., 2023 · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Zou, A., Wang, Z., Carlini, N., Nasr, M., Kolter, J.Z., Fredrikson, M., 2023 · 2023
Later among the works it cites.
Securing large language models: Threats, vulnerabilities and responsible practices
Abdali, S., Anarfi, R., Barberan, C., He, J., 2024 · 2024
Closest in time.
Investigating the prompt leakage effect and black-box defenses for multi-turn llm interactions
Agarwal, D., Fabbri, A.R., Laban, P., Risher, B., Joty, S., Xiong, C., Wu, C.S., 2024 · 2024
Closest in time.
Defending against social engineering attacks in the age of llms
Ai, L., Kumarage, T., Bhattacharjee, A., Liu, Z., Hui, Z., Davinroy, M., Cook, J., Cassani, L., Trapeznikov, K., Kirchner, M., Basharat, A., Hoogs, A., Garland, J., Liu, H., Hirschberg, J., 2024 · 2024
Closest in time.
Jailbreaking leading safety-aligned llms with simple adaptive attacks
Andriushchenko, M., Croce, F., Flammarion, N., 2024 · 2024
Closest in time.
Many-shot jailbreaking
Anthropic, 2024 · 2024
Closest in time.
Foundational challenges in assuring alignment and safety of large language models
Anwar, U., Saparov, A., Rando, J., Paleka, D., Turpin, M., Hase, P., Lubana, E.S., Jenner, E., Casper, S., Sourbut, O., Edelman, B.L., Zhang, Z., Günther, M., Korinek, A., Hernandez-Orallo, J., Hammond, L., Bigelow, E., Pan, A., Langosco, L., Korbak, T., Zhang, H., Zhong, R., h’Eigeartaigh, S.O., Recchia, G., Corsi, G., Chan, A., Anderljung, M., Edwards, L., Bengio, Y., Chen, D., Albanie, S., Maharaj, T., Foerster, J., Tramer, F., He, H., Kasirzadeh, A., Choi, Y., Krueger, D., 2024 · 2024
Closest in time.
An introduction to vision-language modeling
Bordes, F., Pang, R.Y., Ajay, A., Li, A.C., Bardes, A., Petryk, S., Mañas, O., Lin, Z., Mahmoud, A., Jayaraman, B., Ibrahim, M., Hall, M., Xiong, Y., Lebensold, J., Ross, C., Jayakumar, S., Guo, C., Bouchacourt, D., Al-Tahan, H., Padthe, K., Sharma, V., Xu, H., Tan, X.E., Richards, M., Lavoie, S., Astolfi, P., Hemmat, R.A., Chen, J., Tirumala, K., Assouel, R., Moayeri, M., Talattof, A., Chaudhuri, K., Liu, Z., Chen, X., Garrido, Q., Ullrich, K., Agrawal, A., Saenko, K., Celikyilmaz, A., Chandra, V., 2024 · 2024
Closest in time.
Stealing part of a production language model
Carlini, N., Paleka, D., Dvijotham, K.D., Steinke, T., Hayase, J., Cooper, A.F., Lee, K., Jagielski, M., Nasr, M., Conmy, A., Wallace, E., Rolnick, D., Tramèr, F., 2024 · 2024
Closest in time.
Scaling synthetic data creation with 1,000,000,000 personas
Chan, X., Wang, X., Yu, D., Mi, H., Yu, D., 2024 · 2024
Closest in time.
From persona to personalization: A survey on role-playing language agents
Chen, J., Wang, X., Xu, R., Yuan, S., Zhang, Y., Shi, W., Xie, J., Li, S., Yang, R., Zhu, T., Chen, A., Li, N., Chen, L., Hu, C., Wu, S., Ren, S., Fu, Z., Xiao, Y., 2024 · 2024
Closest in time.
Towards guaranteed safe ai: A framework for ensuring robust and reliable ai systems
Dalrymple, D., Skalse, J., Bengio, Y., Russell, S., Tegmark, M., Seshia, S., Omohundro, S., Szegedy, C., Goldhaber, B., Ammann, N., Abate, A., Halpern, J., Barrett, C., Zhao, D., Zhi-Xuan, T., Wing, J., Tenenbaum, J., 2024 · 2024
Closest in time.
Security and privacy challenges of large language models: A survey
Das, B.C., Amini, M.H., Wu, Y., 2024 · 2024
Closest in time.
Sycophancy to subterfuge: Investigating reward-tampering in large language models
Denison, C., MacDiarmid, M., Barez, F., Duvenaud, D., Kravec, S., Marks, S., Schiefer, N., Soklaski, R., Tamkin, A., Kaplan, J., Shlegeris, B., Bowman, S.R., Perez, E., Hubinger, E., 2024 · 2024
Closest in time.
garak: A framework for security probing large language models
Derczynski, L., Galinkin, E., Martin, J., Majumdar, S., Inie, N., 2024 · 2024
Closest in time.
Large language models (llms): Deployment, tokenomics and sustainability
Dong, H., Xie, S., 2024 · 2024
Closest in time.
Building guardrails for large language models
Dong, Y., Mu, R., Jin, G., Qi, Y., Hu, J., Zhao, X., Meng, J., Ruan, W., Huang, X., 2024 · 2024
Closest in time.
Kto: Model alignment as prospect theoretic optimization
Ethayarajh, K., Xu, W., Muennighoff, N., Jurafsky, D., Kiela, D., 2024 · 2024
Closest in time.
The ethics of advanced ai assistanning ts
Gabriel, I., Manzini, A., Keeling, G., Hendricks, L.A., Rieser, V., Iqbal, H., Tomašev, N., Ktena, I., Kenton, Z., Rodriguez, M., et al., 2024 · 2024
Closest in time.
Bias and fairness in large language models: A survey
Gallegos, I.O., Rossi, R.A., Barrow, J., Tanjim, M.M., Kim, S., Dernoncourt, F., Yu, T., Zhang, R., Ahmed, N.K., 2024 · 2024
Closest in time.
Adaptive horizon actor-critic for policy learning in contact-rich differentiable simulation
Georgiev, I., Srinivasan, K., Xu, J., Heiden, E., Garg, A., 2024 · 2024
Closest in time.
Designing for human-agent alignment: Understanding what humans want from their agents, in: Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, ACM
Goyal, N., Chang, M., Terry, M., 2024 · 2024
Closest in time.
Minillm: Knowledge distillation of large language models
Gu, Y., Dong, L., Wei, F., Huang, M., 2024 · 2024
Closest in time.
Connecting large language models with evolutionary algorithms yields powerful prompt optimizers
Guo, Q., Wang, R., Guo, J., Li, B., Song, K., Tan, X., Liu, G., Bian, J., Yang, Y., 2024 · 2024
Closest in time.
Risk and response in large language models: Evaluating key threat categories
Harandizadeh, B., Salinas, A., Morstatter, F., 2024 · 2024
Closest in time.
Contrastive preference learning: Learning from human feedback without rl
Hejna, J., Rafailov, R., Sikchi, H., Finn, C., Niekum, S., Knox, W.B., Sadigh, D., 2024 · 2024
Closest in time.
Why so gullible? enhancing the robustness of retrieval-augmented models against counterfactual noise
Hong, G., Kim, J., Kang, J., Myaeng, S.H., Whang, J.J., 2024 · 2024
Closest in time.
Ai alignment: A comprehensive survey
Ji, J., Qiu, T., Chen, B., Zhang, B., Lou, H., Wang, K., Duan, Y., He, Z., Zhou, J., Zhang, Z., Zeng, F., Ng, K.Y., Dai, J., Pan, X., O’Gara, A., Lei, Y., Xu, H., Tse, B., Fu, J., McAleer, S., Yang, Y., Wang, Y., Zhu, S.C., Guo, Y., Gao, W., 2024 · 2024
Closest in time.
A survey on human preference learning for large language models
Jiang, R., Chen, K., Bai, X., He, Z., Li, J., Yang, M., Zhao, T., Nie, L., Zhang, M., 2024 · 2024
Closest in time.
Injecting undetectable backdoors in deep learning and language models
Kalavasis, A., Karbasi, A., Oikonomou, A., Sotiraki, K., Velegkas, G., Zampetakis, M., 2024 · 2024
Closest in time.
Understanding catastrophic forgetting in language models via implicit inference
Kotha, S., Springer, J.M., Raghunathan, A., 2024 · 2024
Closest in time.
Rewardbench: Evaluating reward models for language modeling
Lambert, N., Pyatkin, V., Morrison, J., Miranda, L., Lin, B.Y., Chandu, K., Dziri, N., Kumar, S., Zick, T., Choi, Y., Smith, N.A., Hajishirzi, H., 2024 · 2024
Closest in time.
Lu, K., Yu, B., Zhou, C., Zhou, J., 2024 · 2024
Closest in time.
Beyond accuracy: Evaluating the reasoning behavior of large language models – a survey
Mondorf, P., Plank, B., 2024 · 2024
Closest in time.
Nezhurina, M., Cipolina-Kun, L., Cherti, M., Jitsev, J., 2024 · 2024
Closest in time.
Reward hacking behavior can generalize across tasks — ai alignment forum
Nishimura-Gasparian, K., Dunn, I., Sleight, H., Turpin, M., Hubinger, E., Denison, C., Perez, E., 2024 · 2024
Closest in time.
OpenAI, Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F.L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., Avila, R., Babuschkin, I., Balaji, S., Balcom, V., Baltescu, P., Bao, H., Bavarian, M., Belgum, J., Bello, I., Berdine, J., Bernadett-Shapiro, G., Berner, C., Bogdonoff, L., Boiko, O., Boyd, M., Brakman, A.L., Brockman, G., Brooks, T., Brundage, M., Button, K., Cai, T., Campbell, R., Cann, A., Carey, B., Carlson, C., Carmichael, R., Chan, B., Chang, C., Chantzis, F., Chen, D., Chen, S., Chen, R., Chen, J., Chen, M., Chess, B., Cho, C., Chu, C., Chung, H.W., Cummings, D., Currier, J., Dai, Y., Decareaux, C., Degry, T., Deutsch, N., Deville, D., Dhar, A., Dohan, D., Dowling, S., Dunning, S., Ecoffet, A., Eleti, A., Eloundou, T., Farhi, D., Fedus, L., Felix, N., Fishman, S.P., Forte, J., Fulford, I., Gao, L., Georges, E., Gibson, C., Goel, V., Gogineni, T., Goh, G., Gontijo-Lopes, R., Gordon, J., Grafstein, M., Gray, S., Greene, R., Gross, J., Gu, S.S., Guo, Y., Hallacy, C., Han, J., Harris, J., He, Y., Heaton, M., Heidecke, J., Hesse, C., Hickey, A., Hickey, W., Hoeschele, P., Houghton, B., Hsu, K., Hu, S., Hu, X., Huizinga, J., Jain, S., Jain, S., Jang, J., Jiang, A., Jiang, R., Jin, H., Jin, D., Jomoto, S., Jonn, B., Jun, H., Kaftan, T., Łukasz Kaiser, Kamali, A., Kanitscheider, I., Keskar, N.S., Khan, T., Kilpatrick, L., Kim, J.W., Kim, C., Kim, Y., Kirchner, J.H., Kiros, J., Knight, M., Kokotajlo, D., Łukasz Kondraciuk, Kondrich, A., Konstantinidis, A., Kosic, K., Krueger, G., Kuo, V., Lampe, M., Lan, I., Lee, T., Leike, J., Leung, J., Levy, D., Li, C.M., Lim, R., Lin, M., Lin, S., Litwin, M., Lopez, T., Lowe, R., Lue, P., Makanju, A., Malfacini, K., Manning, S., Markov, T., Markovski, Y., Martin, B., Mayer, K., Mayne, A., McGrew, B., McKinney, S.M., McLeavey, C., McMillan, P., McNeil, J., Medina, D., Mehta, A., Menick, J., Metz, L., Mishchenko, A., Mishkin, P., Monaco, V., Morikawa, E., Mossing, D., Mu, T., Murati, M., Murk, O., Mély, D., Nair, A., Nakano, R., Nayak, R., Neelakantan, A., Ngo, R., Noh, H., Ouyang, L., O’Keefe, C., Pachocki, J., Paino, A., Palermo, J., Pantuliano, A., Parascandolo, G., Parish, J., Parparita, E., Passos, A., Pavlov, M., Peng, A., Perelman, A., de Avila Belbute Peres, F., Petrov, M., de Oliveira Pinto, H.P., Michael, Pokorny, Pokrass, M., Pong, V.H., Powell, T., Power, A., Power, B., Proehl, E., Puri, R., Radford, A., Rae, J., Ramesh, A., Raymond, C., Real, F., Rimbach, K., Ross, C., Rotsted, B., Roussez, H., Ryder, N., Saltarelli, M., Sanders, T., Santurkar, S., Sastry, G., Schmidt, H., Schnurr, D., Schulman, J., Selsam, D., Sheppard, K., Sherbakov, T., Shieh, J., Shoker, S., Shyam, P., Sidor, S., Sigler, E., Simens, M., Sitkin, J., Slama, K., Sohl, I., Sokolowsky, B., Song, Y., Staudacher, N., Such, F.P., Summers, N., Sutskever, I., Tang, J., Tezak, N., Thompson, M.B., Tillet, P., Tootoonchian, A., Tseng, E., Tuggle, P., Turley, N., Tworek, J., Uribe, J.F.C., Vallone, A., Vijayvergiya, A., Voss, C., Wainwright, C., Wang, J.J., Wang, A., Wang, B., Ward, J., Wei, J., Weinmann, C., Welihinda, A., Welinder, P., Weng, J., Weng, L., Wiethoff, M., Willner, D., Winter, C., Wolrich, S., Wong, H., Workman, L., Wu, S., Wu, J., Wu, M., Xiao, K., Xu, T., Yoo, S., Yu, K., Yuan, Q., Zaremba, W., Zellers, R., Zhang, C., Zhang, M., Zhao, S., Zheng, T., Zhuang, J., Zhuk, W., Zoph, B., 2024 · 2024
Closest in time.
Is value learning really the main bottleneck in offline rl?
Park, S., Frans, K., Levine, S., Kumar, A., 2024 · 2024
Closest in time.
Advprompter: Fast adaptive adversarial prompting for llms
Paulus, A., Zharmagambetov, A., Guo, C., Amos, B., Tian, Y., 2024 · 2024
Closest in time.
Safety alignment should be made more than just a few tokens deep
Qi, X., Panda, A., Lyu, K., Ma, X., Roy, S., Beirami, A., Mittal, P., Henderson, P., 2024 · 2024
Closest in time.
Jailbreakeval: An integrated toolkit for evaluating jailbreak attempts against large language models
Ran, D., Liu, J., Gong, Y., Zheng, J., He, X., Cong, T., Wang, A., 2024 · 2024
Closest in time.
Safe and responsible large language model development
Raza, S., Bamgbose, O., Ghuge, S., Reji, D.J., 2024 · 2024
Closest in time.
Schiller, C.A., 2024 · 2024
Closest in time.
Shen, H., Knearem, T., Ghosh, R., Alkiek, K., Krishna, K., Liu, Y., Ma, Z., Petridis, S., Peng, Y.H., Qiwei, L., Rakshit, S., Si, C., Xie, Y., Bigham, J.P., Bentley, F., Chai, J., Lipton, Z., Mei, Q., Mihalcea, R., Terry, M., Yang, D., Morris, M.R., Resnick, P., Jurgens, D., 2024 · 2024
Closest in time.
Understanding preference fine-tuning through the lens of coverage
Song, Y., Swamy, G., Singh, A., Bagnell, J.A., Sun, W., 2024 · 2024
Closest in time.
Preference fine-tuning of llms should leverage suboptimal, on-policy data
Tajwar, F., Singh, A., Sharma, A., Rafailov, R., Schneider, J., Xie, T., Ermon, S., Finn, C., Kumar, A., 2024 · 2024
Closest in time.
Simple synthetic data reduces sycophancy in large language models
Wei, J., Huang, D., Lu, Y., Zhou, D., Le, Q.V., 2024 · 2024
Closest in time.
Mirai: Evaluating llm agents for event forecasting
Ye, C., Hu, Z., Deng, Y., Huang, Z., Ma, M.D., Zhu, Y., Wang, W., 2024 · 2024
Closest in time.
Improved few-shot jailbreaking can circumvent aligned language models and their defenses
Zheng, X., Pang, T., Du, C., Liu, Q., Jiang, J., Lin, M., 2024 · 2024
Closest in time.
Lima: Less is more for alignment
Zhou, C., Liu, P., Xu, P., Iyer, S., Sun, J., Mao, Y., Ma, X., Efrat, A., Yu, P., Yu, L., et al., 2024 · 2024
Closest in time.
Piccolo: Exposing complex backdoors in nlp transformer models, in: 2022 IEEE Symposium on Security and Privacy (SP), IEEE. pp. 2025–2042
Liu, Y., Shen, G., Tao, G., An, S., Ma, S., Zhang, X., 2022 · 2042
Closest in time.