Fetching the paper…
Reading the bibliography…
Language models influence the external world: they query APIs that read and write to web pages, generate content that shapes human behavior, and run system commands as autonomous agents.
An economic interpretation of optimal control theory
Dorfman, R · 1969
Earlier work this paper cites.
Econometric policy evaluation: A critique
Lucas Jr, R. E · 1976
Earlier work this paper cites.
Learning from Delayed Rewards
Watkins, C. J. C. H · 1989
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Varieties of confirmation bias
Klayman, J · 1995
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, A. Y., Harada, D., and Russell, S · 1999
Earlier work this paper cites.
Language (technology) is power: A critical survey of" bias" in nlp
Blodgett, S. L., Barocas, S., Daumé III, H., and Wallach, H · 2005
Earlier work this paper cites.
Language (technology) is power: A critical survey of" bias" in nlp
Blodgett, S. L., Barocas, S., Daumé III, H., and Wallach, H · 2005
Earlier work this paper cites.
Nash equilibria of static prediction games
Brückner, M. and Scheffer, T · 2009
Earlier work this paper cites.
Counterfactual reasoning and learning systems: The example of computational advertising
Bottou, L., Peters, J., Quiñonero-Candela, J., Charles, D. X., Chickering, D. M., Portugaly, E., Ray, D., Simard, P., and Snelson, E · 2013
Earlier work this paper cites.
Feedback control theory
Doyle, J. C., Francis, B. A., and Tannenbaum, A. R · 2013
Earlier work this paper cites.
Faulty reward functions in the wild, Dec 2016
Clark, J. and Amodei, D · 2016
Earlier work this paper cites.
Strategic classification
Hardt, M., Megiddo, N., Papadimitriou, C., and Wootters, M · 2016
Earlier work this paper cites.
Engineering a safer world: Systems thinking applied to safety
Leveson, N. G · 2016
Earlier work this paper cites.
Control principles of complex systems
Liu, Y.-Y. and Barabási, A.-L · 2016
Earlier work this paper cites.
Deconvolving feedback loops in recommender systems
Sinha, A., Gleich, D. F., and Ramani, K · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Active preference-based learning of reward functions
Sadigh, D., Dragan, A. D., Sastry, S., and Seshia, S. A · 2017
Earlier work this paper cites.
A game-theoretic approach to recommendation systems with strategic content providers
Ben-Porat, O. and Tennenholtz, M · 2018
Earlier work this paper cites.
How algorithmic confounding in recommendation systems increases homogeneity and decreases utility
Chaney, A. J., Stewart, B. M., and Engelhardt, B. E · 2018
Earlier work this paper cites.
Fairness without demographics in repeated loss minimization
Hashimoto, T., Srivastava, M., Namkoong, H., and Liang, P · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Earlier work this paper cites.
Extracting training data from large language models. arxiv
Carlini, N., Tramer, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, Ú., et al · 2020
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Gehman, S., Gururangan, S., Sap, M., Choi, Y., and Smith, N. A · 2020
Earlier work this paper cites.
Detoxify
Hanu, L. and Unitary team · 2020
Earlier work this paper cites.
Hidden incentives for auto-induced distributional shift
Krueger, D., Maharaj, T., and Leike, J · 2020
Earlier work this paper cites.
Feedback loop and bias amplification in recommender systems
Mansoury, M., Abdollahpouri, H., Pechenizkiy, M., Mobasher, B., and Burke, R · 2020
Earlier work this paper cites.
Performative prediction
Perdomo, J., Zrnic, T., Mendler-Dünner, C., and Hardt, M · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F · 2020
Earlier work this paper cites.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al · 2021
Earlier work this paper cites.
Unsolved problems in ml safety
Hendrycks, D., Carlini, N., Schulman, J., and Steinhardt, J · 2021
Earlier work this paper cites.
Spotify and the democratisation of music
Hodgson, T · 2021
Cited alongside, same era.
Truthfulqa: Measuring how models mimic human falsehoods
Lin, S., Hilton, J., and Evans, O · 2021
Cited alongside, same era.
Ethical and social risks of harm from language models
Weidinger, L., Mellor, J., Rauh, M., Griffin, C., Uesato, J., Huang, P.-S., Cheng, M., Glaese, M., Balle, B., Kasirzadeh, A., et al · 2021
Cited alongside, same era.
Constitutional ai: Harmlessness from ai feedback
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., et al · 2022
Cited alongside, same era.
Estimating and penalizing induced preference shifts in recommender systems
Carroll, M. D., Dragan, A. D., Russell, S., and Hadfield-Menell, D · 2022
Alpacaeval: An automatic evaluator of instruction-following models
Li, X., Zhang, T., Dubois, Y., Taori, R., Gulrajani, I., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Later among the works it cites.
My a.i. lover
Liang, C · 2023
Later among the works it cites.
Self-refine: Iterative refinement with self-feedback
Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., et al · 2023
Later among the works it cites.
choix, 2023
Maystre, L · 2023
Later among the works it cites.
Augmented language models: a survey
Mialon, G., Dessì, R., Lomeli, M., Nalmpantis, C., Pasunuru, R., Raileanu, R., Rozière, B., Schick, T., Dwivedi-Yu, J., Celikyilmaz, A., et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Preference dynamics under personalized recommendations
Dean, S. and Morgenstern, J · 2022
Cited alongside, same era.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Ganguli, D., Lovitt, L., Kernion, J., Askell, A., Bai, Y., Kadavath, S., Mann, B., Perez, E., Schiefer, N., Ndousse, K., et al · 2022
Cited alongside, same era.
Supply-side equilibria in recommender systems
Jagadeesan, M., Garg, N., and Steinhardt, J · 2022
Cited alongside, same era.
Breaking feedback loops in recommender systems with causal inference
Krauth, K., Wang, Y., and Jordan, M. I · 2022
Cited alongside, same era.
A new generation of perspective api: Efficient multilingual character-level transformers
Lees, A., Tran, V. Q., Tay, Y., Sorensen, J., Gupta, J., Metzler, D., and Vasserman, L · 2022
Cited alongside, same era.
Holistic evaluation of language models
Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., et al · 2022
Cited alongside, same era.
The effects of reward misspecification: Mapping and mitigating misaligned models
Pan, A., Bhatia, K., and Steinhardt, J · 2022
Cited alongside, same era.
Mu, N., Chen, S., Wang, Z., Chen, S., Karamardian, D., Aljeraisy, L., Hendrycks, D., and Wagner, D · 2023
Later among the works it cites.
Newsguard
NewsGuard · 2023
Later among the works it cites.
Pan, A., Shern, C. J., Zou, A., Li, N., Basart, S., Woodside, T., Ng, J., Zhang, H., Emmons, S., and Hendrycks, D · 2023
Later among the works it cites.
Leveraging implicit feedback from deployment data in dialogue
Pang, R. Y., Roller, S., Cho, K., He, H., and Weston, J · 2023
Later among the works it cites.
The new ai-powered bing is threatening users. that’s no laughing matter
Perrigo, B · 2023
Later among the works it cites.
Auto-gpt, 2023
Richards, T. B · 2023
Later among the works it cites.
Google’s bard just got more powerful. it’s still erratic
Roose, K · 2023
Later among the works it cites.
A conversation with bing’s chatbot left me deeply unsettled
Roose, K · 2023
Later among the works it cites.
Identifying the risks of lm agents with an lm-emulated sandbox
Ruan, Y., Dong, H., Wang, A., Pitis, S., Zhou, Y., Ba, J., Dubois, Y., Maddison, C. J., and Hashimoto, T · 2023
Later among the works it cites.
Scheurer, J., Balesni, M., and Hobbhahn, M · 2023
Later among the works it cites.
Reflexion: language agents with verbal reinforcement learning
Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K. R., and Yao, S · 2023
Later among the works it cites.
The curse of recursion: Training on generated data makes models forget
Shumailov, I., Shumaylov, Z., Zhao, Y., Gal, Y., Papernot, N., and Anderson, R · 2023
Later among the works it cites.
Emergent deception and emergent optimization, 2023
Steinhardt, J · 2023
Later among the works it cites.
Microsoft has been secretly testing its bing chatbot ‘sydney’ for years
Warren, T · 2023
Later among the works it cites.
Large language models as optimizers
Yang, C., Wang, X., Lu, Y., Liu, H., Le, Q. V., Zhou, D., and Chen, X · 2023
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., and Narasimhan, K · 2023
Later among the works it cites.
Self-taught optimizer (stop): Recursively self-improving code generation
Zelikman, E., Lorch, E., Mackey, L., and Kalai, A. T · 2023
Later among the works it cites.
Siren’s song in the ai ocean: a survey on hallucination in large language models
Zhang, Y., Li, Y., Cui, L., Cai, D., Liu, L., Fu, T., Huang, X., Zhao, E., Zhang, Y., Chen, Y., et al · 2023
Later among the works it cites.
Introducing the next generation of Claude
Anthropic · 2024
Closest in time.
Ai alignment with changing and influenceable reward functions
Carroll, M. D., Foote, D., Siththaranjan, A., Russell, S., and Dragan, A. D · 2024
Closest in time.
Sleeper agents: Training deceptive llms that persist through safety training, 2024
Hubinger, E., Denison, C., Mu, J., Lambert, M., Tong, M., MacDiarmid, M., Lanham, T., Ziegler, D. M., Maxwell, T., Cheng, N., Jermyn, A., Askell, A., Radhakrishnan, A., Anil, C., Duvenaud, D., Ganguli, D., Barez, F., Clark, J., Ndousse, K., Sachan, K., Sellitto, M., Sharma, M., DasSarma, N., Grosse, R., Kravec, S., Bai, Y., Witten, Z., Favaro, M., Brauner, J., Karnofsky, H., Christiano, P., Bowman, S. R., Graham, L., Kaplan, J., Mindermann, S., Greenblatt, R., Shlegeris, B., Schiefer, N., and Perez, E · 2024
Closest in time.
Measuring and controlling persona drift in language model dialogs
Li, K., Liu, T., Bashkansky, N., Bau, D., Viégas, F., Pfister, H., and Wattenberg, M · 2024
Closest in time.
Escalation risks from language models in military and diplomatic decision-making
Rivera, J.-P., Mukobi, G., Reuel, A., Lamparth, M., Smith, C., and Schneider, J · 2024
Closest in time.
Meet my a.i. friends
Roose, K · 2024
Closest in time.
Perils of self-feedback: Self-bias amplifies in large language models
Xu, W., Zhu, G., Zhao, X., Pan, L., Li, L., and Wang, W. Y · 2024
Closest in time.