Fetching the paper…
Reading the bibliography…
Reinforcement learning from human feedback (RLHF) is widely used to train large language models (LLMs).
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
V. Sanh, L. Debut, J. Chaumond, and T. Wolf · 1910
Earlier work this paper cites.
Sparse coding with an overcomplete basis set: A strategy employed by v1?
B. A. Olshausen and D. J. Field · 1997
Earlier work this paper cites.
Quantifying differences in reward functions
A. Gleave, M. Dennis, S. Legg, S. Russell, and J. Leike · 2006
Earlier work this paper cites.
The sparsity and bias of the lasso selection in high dimensional linear regression
C.-H. Zhang and J. Huang · 2008
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Learning word vectors for sentiment analysis
A. L. Maas, R. E. Daly, P. T. Pham, D. Huang, A. Y. Ng, and C. Potts · 2011
Earlier work this paper cites.
Understanding learned reward functions
E. J. Michaud, A. Gleave, and S. Russell · 2012
Earlier work this paper cites.
Do recommender systems manipulate consumer preferences? A study of anchoring effects
G. Adomavicius, J. C. Bockstedt, S. P. Curley, and J. Zhang · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Vader: A parsimonious rule-based model for sentiment analysis of social media text
C. Hutto and E. Gilbert · 2014
Earlier work this paper cites.
Visualizing and understanding recurrent networks
A. Karpathy, J. Johnson, and L. Fei-Fei · 2015
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
G. Alain and Y. Bengio · 2016
Earlier work this paper cites.
Network dissection: Quantifying interpretability of deep visual representations
D. Bau, B. Zhou, A. Khosla, A. Oliva, and A. Torralba · 2017
Earlier work this paper cites.
Feature visualization
C. Olah, A. Mordvintsev, and L. Schubert · 2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
What failure looks like
P. Christiano · 2019
Cited alongside, same era.
Probing neural network comprehension of natural language arguments
T. Niven and H.-Y. Kao · 2019
Cited alongside, same era.
spacy: Industrial-strength natural language processing in python, 2020
M. Honnibal, I. Montani, S. Van Landeghem, and A. Boyd · 2020
Cited alongside, same era.
Gpt-neo: Large scale autoregressive language modeling with mesh-tensorflow
S. Black, L. Gao, P. Wang, C. Leahy, and S. Biderman · 2021
Language models can explain neurons in language models
S. Bills, N. Cammarata, D. Mossing, H. Tillman, L. Gao, G. Goh, I. Sutskever, J. Leike, J. Wu, and W. Saunders · 2023
Closest in time.
Towards monosemanticity: Decomposing language models with dictionary learning
T. Bricken, A. Templeton, J. Batson, B. Chen, A. Jermyn, T. Conerly, N. Turner, C. Anil, C. Denison, A. Askell, R. Lasenby, Y. Wu, S. Kravec, N. Schiefer, T. Maxwell, N. Joseph, Z. Hatfield-Dodds, A. Tamkin, K. Nguyen, B. McLean, J. E. Burke, T. Hume, S. Carter, T. Henighan, and C. Olah · 2023
Closest in time.
Neuron to graph: Interpreting language model neurons at scale
A. Foote, N. Nanda, E. Kran, I. Konstas, S. Cohen, and F. Barez · 2023
Closest in time.
Finding neurons in a haystack: Case studies with sparse probing
W. Gurnee, N. Nanda, M. Pauly, K. Harvey, D. Troitskii, and D. Bertsimas · 2023
Closest in time.
Toxic-dpo dataset v0.2, 2023
Unalignment · 2023
Closest in time.
distilbert-imdb, 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Transformer visualization via dictionary learning: contextualized embedding as a linear superposition of transformer factors
Z. Yun, Y. Chen, B. Olshausen, and Y. LeCun · 2021
Cited alongside, same era.
N. Elhage, T. Hume, C. Olsson, N. Schiefer, T. Henighan, S. Kravec, Z. Hatfield-Dodds, R. Lasenby, D. Drain, C. Chen, R. Grosse, S. McCandlish, J. Kaplan, D. Amodei, M. Wattenberg, and C. Olah · 2022
Cited alongside, same era.
Preprocessing reward functions for interpretability
E. Jenner and A. Gleave · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. F. Christiano, J. Leike, and R. Lowe · 2022
Cited alongside, same era.
Taking features out of superposition with sparse autoencoders
L. Sharkey, D. Braun, and B. Millidge · 2022
Cited alongside, same era.
Anthropic hh-rlhf dataset, 2023
Anthropic · 2023
Cited alongside, same era.
L. von Werra · 2023
Closest in time.
TRL: Transformer Reinforcement Learning, 2023
L. von Werra, Y. Belkada, L. Tunstall, E. Beeching, T. Thrush, and N. Lambert · 2023
Closest in time.
Fundamental limitations of alignment in large language models
Y. Wolf, N. Wies, O. Avnery, Y. Levine, and A. Shashua · 2023
Closest in time.
Sparse autoencoders find composed features in small toy models
E. Anders, C. Neo, J. Hoelscher-Obermaier, and J. Howard · 2024
Closest in time.
Sparse autoencoders find highly interpretable features in language models
H. Cunningham, A. Ewart, L. Riggs, R. Huben, and L. Sharkey · 2024
Closest in time.
https://www.lesswrong.com/posts/bducmgmjjnctc7jkc/research-report-sparse-autoencoders-find-only-9-180-board
R. Huben · 2024
Closest in time.
Llama 3: Open and efficient foundation language models
Meta · 2024
Closest in time.
Do sparse autoencoders find "true features"?
D. Till · 2024
Closest in time.