Fetching the paper…
Reading the bibliography…
We introduce Inference-Time Intervention (ITI), a technique designed to enhance the "truthfulness" of large language models (LLMs).
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020) · 1901
Earlier work this paper cites.
Bert rediscovers the classical nlp pipeline
Tenney, I., Das, D., and Pavlick, E. (2019) · 1905
Earlier work this paper cites.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G. (2019) · 1909
Earlier work this paper cites.
Plug and play language models: A simple approach to controlled text generation
Dathathri, S., Madotto, A., Lan, J., Hung, J., Frank, E., Molino, P., Yosinski, J., and Liu, R. (2019) · 1912
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. (2020) · 2009
Earlier work this paper cites.
Gedi: Generative discriminator guided sequence generation
Krause, B., Gotmare, A. D., McCann, B., Keskar, N. S., Joty, S., Socher, R., and Rajani, N. F. (2020) · 2009
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C. (2015) · 2015
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Alain, G. and Bengio, Y. (2016) · 2016
Earlier work this paper cites.
Probing classifiers: Promises, shortcomings, and advances
Belinkov, Y. (2016) · 2016
Earlier work this paper cites.
Arbitrary style transfer in real-time with adaptive instance normalization
Huang, X. and Belongie, S. (2017) · 2017
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Joshi, M., Choi, E., Weld, D. S., and Zettlemoyer, L. (2017) · 2017
Earlier work this paper cites.
Learning to generate reviews and discovering sentiment
Radford, A., Jozefowicz, R., and Sutskever, I. (2017) · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017) · 2017
Earlier work this paper cites.
Parameter-efficient transfer learning for nlp
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S. (2019) · 2019
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Devlin, J., Lee, K., et al. (2019) · 2019
Earlier work this paper cites.
A general language assistant as a laboratory for alignment
Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., DasSarma, N., et al. (2021) · 2021
Earlier work this paper cites.
A mathematical framework for transformer circuits
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., et al. (2021) · 2021
Earlier work this paper cites.
A framework for few-shot language model evaluation
Gao, L., Tow, J., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., McDonell, K., Muennighoff, N., Phang, J., Reynolds, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A. (2021) · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. (2021) · 2021
Cited alongside, same era.
Truthfulqa: Measuring how models mimic human falsehoods
Lin, S., Hilton, J., and Evans, O. (2021) · 2021
Cited alongside, same era.
Webgpt: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al. (2021) · 2021
Cited alongside, same era.
True few-shot learning with language models
Perez, E., Kiela, D., and Cho, K. (2021) · 2021
Cited alongside, same era.
Scaling language models: Methods, analysis & insights from training gopher
Mechanistic interpretability, variables, and the importance of interpretable bases
Olah, C. (2022) · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. (2022) · 2022
Later among the works it cites.
Discovering language model behaviors with model-written evaluations
Perez, E., Ringer, S., Lukošiūtė, K., Nguyen, K., Chen, E., Heiner, S., Pettit, C., Olsson, C., Kundu, S., Kadavath, S., et al. (2022) · 2022
Later among the works it cites.
Self-critiquing models for assisting human evaluators
Saunders, W., Yeh, C., Wu, J., Bills, S., Ouyang, L., Ward, J., and Leike, J. (2022) · 2022
Later among the works it cites.
Extracting latent steering vectors from pretrained language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rae, J. W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, F., Aslanides, J., Henderson, S., Ring, R., Young, S., et al. (2021) · 2021
Cited alongside, same era.
Retrieval augmentation reduces hallucination in conversation
Shuster, K., Poff, S., Chen, M., Kiela, D., and Weston, J. (2021) · 2021
Cited alongside, same era.
Language models are open knowledge graphs
Wang, C., Liu, X., and Song, D. (2021) · 2021
Cited alongside, same era.
Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models
Zaken, E. B., Ravfogel, S., and Goldberg, Y. (2021) · 2021
Cited alongside, same era.
Discovering latent knowledge in language models without supervision
Burns, C., Ye, H., Klein, D., and Steinhardt, J. (2022) · 2022
Cited alongside, same era.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., et al. (2022) · 2022
Cited alongside, same era.
How do new models from openai, deepmind and anthropic perform on truthfulqa
Evans, O., Lin, S., and Hilton, J. (2022) · 2022
Cited alongside, same era.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Ganguli, D., Lovitt, L., Kernion, J., Askell, A., Bai, Y., Kadavath, S., Mann, B., Perez, E., Schiefer, N., Ndousse, K., et al. (2022) · 2022
Cited alongside, same era.
Subramani, N., Suresh, N., and Peters, M. E. (2022) · 2022
Later among the works it cites.
Self-instruct: Aligning language model with self generated instructions
Wang, Y., Kordi, Y., Mishra, S., Liu, A., Smith, N. A., Khashabi, D., and Hajishirzi, H. (2022) · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Chi, E., Le, Q., and Zhou, D. (2022) · 2022
Later among the works it cites.
Robustness of edited neural networks
Brown, D., Godfrey, C., Nizinski, C., Tu, J., and Kvinge, H. (2023) · 2023
Closest in time.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., et al. (2023) · 2023
Closest in time.
Hase, P., Bansal, M., Kim, B., and Ghandeharioun, A. (2023) · 2023
Closest in time.
Measuring and manipulating knowledge representations in language models
Hernandez, E., Li, B. Z., and Andreas, J. (2023) · 2023
Closest in time.
Do large language models learn world models or just surface statistics?
Li, K. (2023) · 2023
Closest in time.
Emergent world representations: Exploring a sequence model trained on a synthetic task
Li, K., Hopkins, A. K., Bau, D., Viégas, F., Pfister, H., and Wattenberg, M. (2023) · 2023
Closest in time.
Editing implicit assumptions in text-to-image diffusion models
Orgad, H., Kawar, B., and Belinkov, Y. (2023) · 2023
Closest in time.
What discovering latent knowledge did and did not find
Roger, F. (2023) · 2023
Closest in time.
Alpaca: A strong, replicable instruction-following model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B. (2023) · 2023
Closest in time.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al. (2023) · 2023
Closest in time.
Steering gpt-2-xl by adding an activation vector
Turner, A., M, M., Udell, D., Theirgart, L., and Mini, U. (2023) · 2023
Closest in time.