Fetching the paper…
Reading the bibliography…
Prompt engineering and finetuning aim to maximize language model performance on a given metric (like toxicity reduction).
Probability plotting methods for the analysis of data
Ramanathan Gnanadesikan and Martin B Wilk · 1968
Earlier work this paper cites.
The irrelevance of turing machines to artificial intelligence
Aaron Sloman · 2002
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean · 2013
Earlier work this paper cites.
Linguistic regularities in continuous space word representations
Tomáš Mikolov, Wen-tau Yih, and Geoffrey Zweig · 2013
Earlier work this paper cites.
nltk.tokenize.punkt module
Jan Strunk · 2013
Earlier work this paper cites.
Autoencoding beyond pixels using a learned similarity metric, 2016
Anders Boesen Lindbo Larsen, Søren Kaae Sønderby, Hugo Larochelle, and Ole Winther · 2016
Earlier work this paper cites.
Sequence level training with recurrent neural networks, 2016
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Sampling generative networks, 2016
Tom White · 2016
Earlier work this paper cites.
Trainable greedy decoding for neural machine translation
Jiatao Gu, Kyunghyun Cho, and Victor O.K. Li · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes, 2018
Guillaume Alain and Yoshua Bengio · 2018
Earlier work this paper cites.
Delete, retrieve, generate: A simple approach to sentiment and style transfer, 2018
Juncen Li, Robin Jia, He He, and Percy Liang · 2018
Earlier work this paper cites.
Openwebtext
Joshua Peterson, Stephan Meylan, and David Bourgin · 2018
Earlier work this paper cites.
Bias correction of learned generative models using likelihood-free importance weighting, 2019
Aditya Grover, Jiaming Song, Alekh Agarwal, Kenneth Tran, Ashish Kapoor, Eric Horvitz, and Stefano Ermon · 2019
Earlier work this paper cites.
Are sixteen heads really better than one?
Paul Michel, Omer Levy, and Graham Neubig · 2019
Earlier work this paper cites.
Language models as knowledge bases?
F. Petroni, T. Rocktäschel, A. H. Miller, P. Lewis, A. Bakhtin, Y. Wu, and S. Riedel · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences, 2019
Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 2019
Earlier work this paper cites.
Language models are few-shot learners, 2020
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Plug and play language models: A simple approach to controlled text generation, 2020
Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, and Rosanne Liu · 2020
Cited alongside, same era.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith · 2020
Cited alongside, same era.
AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts
Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh · 2020
Cited alongside, same era.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, et al · 2021
Cited alongside, same era.
The power of scale for parameter-efficient prompt tuning, 2021
Brian Lester, Rami Al-Rfou, and Noah Constant · 2021
Cited alongside, same era.
Does localization inform editing? surprising differences in causality-based localization vs. knowledge editing in language models, 2023
Peter Hase, Mohit Bansal, Been Kim, and Asma Ghandeharioun · 2023
Closest in time.
Inspecting and editing knowledge representations in language models, 2023
Evan Hernandez, Belinda Z. Li, and Jacob Andreas · 2023
Closest in time.
Editing models with task arithmetic, 2023
Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi · 2023
Closest in time.
Language models and cognitive automation for economic research
Anton Korinek · 2023
Closest in time.
In-context Vectors: Making in context learning more effective and controllable through latent space steering, 2023
Sheng Liu, Lei Xing, and James Zou · 2023
Closest in time.
Locating and editing factual associations in GPT, 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Prefix-Tuning: Optimizing continuous prompts for generation, 2021
Xiang Lisa Li and Percy Liang · 2021
Cited alongside, same era.
GPT-J-6B: 6B jax-based transformer
Ben Wang and Aran Komatsuzaki · 2021
Cited alongside, same era.
FUDGE: Controlled text generation with future discriminators
Kevin Yang and Dan Klein · 2021
Cited alongside, same era.
TransformerLens: A library for mechanistic interpretability of generative language models
Joseph Bloom and Neel Nanda · 2022
Cited alongside, same era.
Discovering latent knowledge in language models without supervision, 2022
Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt · 2022
Cited alongside, same era.
Toy models of superposition, 2022
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah · 2022
Cited alongside, same era.
Controllable text generation via probability density estimation in the latent space
Yuxuan Gu, Xiaocheng Feng, Sicheng Ma, Lingyuan Zhang, Heng Gong, Weihong Zhong, and Bing Qin · 2022
Cited alongside, same era.
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2023
Closest in time.
Understanding and controlling a maze-solving policy network, 2023
Ulisse Mini, Peli Grietzer, Mrinank Sharma, Austin Meek, Monte MacDiarmid, and Alexander Matt Turner · 2023
Closest in time.
Relative representations enable zero-shot latent space communication, 2023
Luca Moschella, Valentino Maiorca, Marco Fumero, Antonio Norelli, Francesco Locatello, and Emanuele Rodolà · 2023
Closest in time.
Actually, othello-gpt has a linear emergent world representation
Neel Nanda · 2023
Closest in time.
Distributed representations: Composition & superposition
Christopher Olah · 2023
Closest in time.
The linear representation hypothesis and the geometry of large language models
Kiho Park, Yo Joong Choe, and Victor Veitch · 2023
Closest in time.
PREADD: prefix-adaptive decoding for controlled text generation
Jonathan Pei, Kevin Yang, and Dan Klein · 2023
Closest in time.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson · 2023
Closest in time.
LLaMA: Open and efficient foundation language models, 2023
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Closest in time.
Eva-KELLM: A new benchmark for evaluating knowledge editing of LLMs, 2023
Suhang Wu, Minlong Peng, Yue Chen, Jinsong Su, and Mingming Sun · 2023
Closest in time.
Air-Decoding: Attribute distribution reconstruction for decoding-time controllable text generation
Tianqi Zhong, Quan Wang, Jingxuan Han, Yongdong Zhang, and Zhendong Mao · 2023
Closest in time.
Representation engineering: A top-down approach to ai transparency, 2023
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J. Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, J. Zico Kolter, and Dan Hendrycks · 2023
Closest in time.
Keeping llms aligned after fine-tuning: The crucial role of prompt templates, 2024
Kaifeng Lyu, Haoyu Zhao, Xinran Gu, Dingli Yu, Anirudh Goyal, and Sanjeev Arora · 2024
Closest in time.
Meta Llama 3
Meta · 2024
Closest in time.
Prompt engineering in consistency and reliability with the evidence-based guideline for llms
Li Wang, Xi Chen, XiangWen Deng, Hao Wen, MingKe You, WeiZhi Liu, Qi Li, and Jian Li · 2024
Closest in time.
A comprehensive study of knowledge editing for large language models, 2024
Ningyu Zhang, Yunzhi Yao, Bozhong Tian, Peng Wang, Shumin Deng, Mengru Wang, Zekun Xi, Shengyu Mao, Jintian Zhang, Yuansheng Ni, Siyuan Cheng, Ziwen Xu, Xin Xu, Jia-Chen Gu, Yong Jiang, Pengjun Xie, Fei Huang, Lei Liang, Zhiqiang Zhang, Xiaowei Zhu, Jun Zhou, and Huajun Chen · 2024
Closest in time.