Fetching the paper…
Reading the bibliography…
Feature visualization, also known as "dreaming", offers insights into vision models by optimizing the inputs to maximize a neuron's activation or other internal component.
“AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts”, 2020
Taylor Shin et al · 2010
Earlier work this paper cites.
“Intriguing properties of neural networks”, 2014
Christian Szegedy et al · 2014
Earlier work this paper cites.
“Inceptionism: Going Deeper into Neural Networks”, 2015
Alexander Mordvintsev, Christopher Olah and Mike Tyka · 2015
Earlier work this paper cites.
“Understanding Neural Networks Through Deep Visualization”, 2015
Jason Yosinski et al · 2015
Earlier work this paper cites.
“Categorical Reparameterization with Gumbel-Softmax”, 2017
Eric Jang, Shixiang Gu and Ben Poole · 2017
Earlier work this paper cites.
“Feature Visualization” https://distill.pub/2017/feature-visualization
Chris Olah, Alexander Mordvintsev and Ludwig Schubert · 2017
Earlier work this paper cites.
“What does BERT dream of?”, 2018
Alex Bäuerle and James Wexler · 2018
Earlier work this paper cites.
“HotFlip: White-Box Adversarial Examples for Text Classification”, 2018
Javid Ebrahimi, Anyi Rao, Daniel Lowd and Dejing Dou · 2018
Earlier work this paper cites.
“The Building Blocks of Interpretability” https://distill.pub/2018/building-blocks
Chris Olah et al · 2018
Earlier work this paper cites.
“Interpretable Textual Neuron Representations for NLP”
Nina Poerner, Benjamin Roth and Hinrich Schütze · 2018
Cited alongside, same era.
“Thread: Circuits” https://distill.pub/2020/circuits
Nick Cammarata et al · 2020
Cited alongside, same era.
“An Interpretability Illusion for BERT”, 2021
Tolga Bolukbasi et al · 2021
Cited alongside, same era.
“A Mathematical Framework for Transformer Circuits” https://transformer-circuits.pub/2021/framework/index.html
Nelson Elhage et al · 2021
Cited alongside, same era.
“Toy Models of Superposition” https://transformer-circuits.pub/2022/toy_model/index.html
Nelson Elhage et al · 2022
Cited alongside, same era.
“Gradient-Based Constrained Sampling from Language Models”, 2022
“Towards Monosemanticity: Decomposing Language Models With Dictionary Learning” https://transformer-circuits.pub/2023/monosemantic-features/index.html
Trenton Bricken et al · 2023
Later among the works it cites.
“Sparse Autoencoders Find Highly Interpretable Features in Language Models”, 2023
Hoagy Cunningham et al · 2023
Later among the works it cites.
In 88 Federal Register 75191 , 2023
“E.O. 14110 of Oct 30, 2023: Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence” · 2023
Later among the works it cites.
“Finding Neurons in a Haystack: Case Studies with Sparse Probing”, 2023
Wes Gurnee et al · 2023
Later among the works it cites.
“Automatically Auditing Large Language Models via Discrete Optimization”, 2023
Erik Jones, Anca Dragan, Aditi Raghunathan and Jacob Steinhardt · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sachin Kumar, Biswajit Paria and Yulia Tsvetkov · 2022
Cited alongside, same era.
Weijia Shi et al · 2022
Cited alongside, same era.
“Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling”, 2023
Stella Biderman et al · 2023
Cited alongside, same era.
“Language models can explain neurons in language models”, https://openaipublic.blob.core.windows.net/neuron-explainer/paper/index.html , 2023
Steven Bills et al · 2023
Cited alongside, same era.
Later among the works it cites.
“Hard Prompts Made Easy: Gradient-Based Discrete Optimization for Prompt Tuning and Discovery”, 2023
Yuxin Wen et al · 2023
Later among the works it cites.
“Bridge the Gap Between CV and NLP! A Gradient-based Textual Adversarial Attack Framework”, 2023
Lifan Yuan, Yichi Zhang, Yangyi Chen and Wei Wei · 2023
Later among the works it cites.
“AutoDAN: Automatic and Interpretable Adversarial Attacks on Large Language Models”
Sicheng Zhu et al · 2023
Later among the works it cites.
“Universal and Transferable Adversarial Attacks on Aligned Language Models”, 2023
Andy Zou, Zifan Wang, J. Kolter and Matt Fredrikson · 2023
Later among the works it cites.