Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) tend to prioritize adherence to user prompts over providing veracious responses, leading to the sycophancy issue.
Distributed representations
Hinton, G. E., McClelland, J. L., and Rumelhart, D. E · 1986
Earlier work this paper cites.
Causal diagrams for empirical research
Pearl, J · 1995
Earlier work this paper cites.
The do-calculus revisited
Pearl, J · 2012
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Mikolov, T., Sutskever, I., Chen, K., et al · 2013
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
Kirkpatrick, J., Pascanu, R., Rabinowitz, N. C., et al · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P., Leike, J., Brown, T. B., et al · 2017
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Joshi, M., Choi, E., Weld, D. S., et al · 2017
Earlier work this paper cites.
Program induction by rationale generation: Learning to solve and explain algebraic word problems
Ling, W., Yogatama, D., Dyer, C., et al · 2017
Earlier work this paper cites.
Learning to generate reviews and discovering sentiment
Radford, A., Józefowicz, R., and Sutskever, I · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N. M., Parmar, N., et al · 2017
Earlier work this paper cites.
Attention is not explanation
Jain, S. and Wallace, B. C · 2019
Earlier work this paper cites.
Optimal design of the resonant tank of the soft-switching solid-state transformer
Mauger, M., Kandula, P., and Divan, D · 2019
Earlier work this paper cites.
The bottom-up evolution of representations in the transformer: A study with machine translation and language modeling objectives
Voita, E., Sennrich, R., and Titov, I · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., et al · 2020
Earlier work this paper cites.
Transformer feed-forward layers are key-value memories
Geva, M., Schuster, R., Berant, J., et al · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., et al · 2020
Earlier work this paper cites.
Zoom in: An introduction to circuits
Olah, C., Cammarata, N., Schubert, L., et al · 2020
Earlier work this paper cites.
Investigating gender bias in language models using causal mediation analysis
Vig, J., Gehrmann, S., Belinkov, Y., et al · 2020
Earlier work this paper cites.
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., et al · 2020
Cited alongside, same era.
Why ai alignment could be hard with modern deep learning
Cotra, A · 2021
Cited alongside, same era.
A mathematical framework for transformer circuits
Elhage, N., Nanda, N., Olsson, C., et al · 2021
Cited alongside, same era.
Measuring mathematical problem solving with the math dataset
Hendrycks, D., Burns, C., Kadavath, S., et al · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J., Wallis, P., Allen-Zhu, Z., et al · 2021
Cited alongside, same era.
Truthfulqa: Measuring how models mimic human falsehoods
Lin, S. C., Hilton, J., and Evans, O · 2021
Cited alongside, same era.
Sparse autoencoders find highly interpretable features in language models
Cunningham, H., Ewart, A., Riggs, L., et al · 2023
Later among the works it cites.
Finding neurons in a haystack: Case studies with sparse probing
Gurnee, W., Nanda, N., Pauly, M., et al · 2023
Later among the works it cites.
Hanna, M., Liu, O., and Variengien, A · 2023
Later among the works it cites.
Jiang, A. Q., Sablayrolles, A., Mensch, A., et al · 2023
Later among the works it cites.
Inference-time intervention: Eliciting truthful answers from a language model
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Calibrate before use: Improving few-shot performance of language models
Zhao, T., Wallace, E., Feng, S., et al · 2021
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y., Jones, A., Ndousse, K., et al · 2022
Cited alongside, same era.
Delta tuning: A comprehensive study of parameter efficient methods for pre-trained language models
Ding, N., Qin, Y., Yang, G., et al · 2022
Cited alongside, same era.
Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space
Geva, M., Caciularu, A., Wang, K., et al · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., et al · 2022
Cited alongside, same era.
Three types of incremental learning
Van de Ven, G. M., Tuytelaars, T., and Tolias, A. S · 2022
Cited alongside, same era.
Li, K., Patel, O., Vi’egas, F., et al · 2023
Later among the works it cites.
Lieberum, T., Rahtz, M., Kram’ar, J., et al · 2023
Later among the works it cites.
Gpt-4 technical report
OpenAI · 2023
Later among the works it cites.
Question decomposition improves the faithfulness of model-generated reasoning
Radhakrishnan, A., Nguyen, K., Chen, A., et al · 2023
Later among the works it cites.
Steering llama 2 via contrastive activation addition
Rimsky, N., Gabrieli, N., Schulz, J., et al · 2023
Later among the works it cites.
Towards understanding sycophancy in language models
Sharma, M., Tong, M., Korbak, T., et al · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., et al · 2023
Later among the works it cites.
Label words are anchors: An information flow perspective for understanding in-context learning
Wang, L., Li, L., Dai, D., et al · 2023
Later among the works it cites.
Simple synthetic data reduces sycophancy in large language models
Wei, J. W., Huang, D., Lu, Y., et al · 2023
Later among the works it cites.
Language models are super mario: Absorbing abilities from homologous models as a free lunch
Yu, L., Bowen, Y., Yu, H., et al · 2023
Later among the works it cites.
How well do large language models perform in arithmetic tasks?
Yuan, Z., Yuan, H., Tan, C., et al · 2023
Later among the works it cites.
Explainability for large language models: A survey
Zhao, H., Chen, H., Yang, F., et al · 2023
Later among the works it cites.
Representation engineering: A top-down approach to ai transparency
Zou, A., Phan, L., Chen, S., et al · 2023
Later among the works it cites.