Fetching the paper…
Reading the bibliography…
Manipulation is a common concern in many domains, such as social media, advertising, and chatbots.
Degenerate Feedback Loops in Recommender Systems
Ray Jiang, Silvia Chiappa, Tor Lattimore, András György, and Pushmeet Kohli. 2019 · 1902
Earlier work this paper cites.
Tom Everitt, Marcus Hutter, Ramana Kumar, and Victoria Krakovna. 2021b · 1908
Earlier work this paper cites.
Correcting for Social Desirability Response Sets in Opinion-Attitude Survey Research
David Horton Smith. 1967 · 1967
Earlier work this paper cites.
Problems of Monetary Management: the UK Experience in Papers in Monetary Economics
Charles Goodhart. 1975 · 1975
Earlier work this paper cites.
Knee-Deep in the Big Muddy: a Study of Escalating Commitment to a Chosen Course of Action
Barry M. Staw. 1976 · 1976
Earlier work this paper cites.
Intention, plans, and practical reason
Michael Bratman. 1987 · 1987
Earlier work this paper cites.
Decision-Theoretic Foundations for Causal Reasoning
D. Heckerman and R. Shachter. 1995 · 1995
Earlier work this paper cites.
Captology: the Study of Computers as Persuasive Technologies. In CHI 98 Conference Summary on Human Factors in Computing Systems . 385
Brian J Fogg. 1998 · 1998
Earlier work this paper cites.
The Incentives that Shape Behaviour
Ryan Carey, Eric Langlois, Tom Everitt, and Shane Legg. 2020 · 2001
Earlier work this paper cites.
Who’s Pulling Your Strings?: How to Break the Cycle of Manipulation and Regain Control of Your Life: How to Break the Cycle of Manipulation and Regain Control of Your Life
Harriet Braiker. 2003 · 2003
Earlier work this paper cites.
Persuasive Technology
Brian J Fogg. 2003 · 2003
Earlier work this paper cites.
Do Defaults Save Lives?
Eric J. Johnson and Daniel Goldstein. 2003 · 2003
Earlier work this paper cites.
Lying and Falsely Implicating
Jörg Meibauer. 2005 · 2004
Earlier work this paper cites.
Feedback Loop and Bias Amplification in Recommender Systems
Masoud Mansoury, Himan Abdollahpouri, Mykola Pechenizkiy, Bamshad Mobasher, and Robin Burke. 2020 · 2007
Earlier work this paper cites.
Preference change: approaches from philosophy, economics and psychology
Till Grüne-Yanoff and Sven Ove Hansson (Eds.). 2009 · 2009
Earlier work this paper cites.
Redefining Market Manipulation in Australia: The Role of an Implied Intent Element
(Robin) Hui Huang. 2009 · 2009
Earlier work this paper cites.
Nudge: Improving Decisions about Health, Wealth and Happiness (revised edition, new international edition ed.)
Richard H. Thaler and Cass R. Sunstein. 2009 · 2009
Earlier work this paper cites.
Lying and Deception: Theory and Practice
Thomas L. Carson. 2010 · 2010
Earlier work this paper cites.
Seeking Better Health Care Outcomes: the Ethics of Using the "Nudge"
J. S. Blumenthal-Barby and Hadley Burroughs. 2012 · 2011
Earlier work this paper cites.
Bayesian Persuasion
Emir Kamenica and Matthew Gentzkow. 2011 · 2011
Earlier work this paper cites.
What Makes Online Content Viral?
Jonah Berger and Katherine L. Milkman. 2012 · 2012
Earlier work this paper cites.
Open Problems in Cooperative AI
Allan Dafoe, Edward Hughes, Yoram Bachrach, Tantum Collins, Kevin R. McKee, Joel Z. Leibo, Kate Larson, and Thore Graepel. 2020 · 2012
Earlier work this paper cites.
Personalized Persuasion: Tailoring Persuasive Appeals to Recipients’ Personality Traits
Jacob B. Hirsh, Sonia K. Kang, and Galen V. Bodenhausen. 2012 · 2012
Earlier work this paper cites.
Do Recommender Systems Manipulate Consumer Preferences? A Study of Anchoring Effects
Gediminas Adomavicius, Jesse C. Bockstedt, Shawn P. Curley, and Jingjing Zhang. 2013 · 2013
Earlier work this paper cites.
Antidisruptive Practices Authority Interpretative Guidance and Policy Statement
CFTC. 2013 · 2013
Earlier work this paper cites.
Combining Multiple Influence Strategies to Increase Consumer Compliance
Maurits Kaptein and Steven Duplinsky. 2013 · 2013
Earlier work this paper cites.
The Mens Rea and Moral Status of Manipulation
Marcia Baron. 2014 · 2014
Earlier work this paper cites.
Digital Market Manipulation
M. Ryan Calo. 2014 · 2014
Earlier work this paper cites.
Implementing the "Wisdom of the Crowd"
Ilan Kremer, Yishay Mansour, and Motty Perry. 2014 · 2014
Earlier work this paper cites.
Survey Research in HCI
Hendrik Müller, Aaron Sedley, and Elizabeth Ferrall-Nunge. 2014 · 2014
Earlier work this paper cites.
Transformative experience (1st ed ed.)
L. A. Paul. 2014 · 2014
Earlier work this paper cites.
Auditing Algorithms: Research Methods for Detecting Discrimination on Internet Platforms
Christian Sandvig, Kevin Hamilton, Karrie Karahalios, and Cedric Langbort. 2014 · 2014
Earlier work this paper cites.
Coercion, Manipulation, Exploitation
Allen W. Wood. 2014 · 2014
Earlier work this paper cites.
Motivated Value Selection for Artificial Agents
Stuart Armstrong. 2015 · 2015
Earlier work this paper cites.
The Mysterious Ethics of High-Frequency Trading
Ricky Cooper, Michael Davis, and Ben Van Vliet. 2016 · 2015
Earlier work this paper cites.
The Search Engine Manipulation Effect (SEME) and its Possible Impact on the Outcomes of Elections
Robert Epstein and Ronald E. Robertson. 2015 · 2015
Earlier work this paper cites.
Do Automated Trading Systems Dream of Manipulating the Price of Futures contracts? Policing Markets for Improper Trading Practices by Algorithmic Robots
Gregory Scopino. 2015 · 2015
Earlier work this paper cites.
Concrete Problems in AI Safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané. 2016 · 2016
Earlier work this paper cites.
FCA Handbook: MAR 1 Market Abuse
Financial Conduct AuthorityA. 2016 · 2016
Earlier work this paper cites.
Strategic Classification. In Proceedings of the 2016 ACM Conference on Innovations in Theoretical Computer Science (ITCS ’16) . Association for Computing Machinery, New York, NY, USA, 111–122
Moritz Hardt, Nimrod Megiddo, Christos Papadimitriou, and Mary Wootters. 2016 · 2016
Earlier work this paper cites.
The Definition of Lying and Deception
James Edwin Mahon. 2016 · 2016
Earlier work this paper cites.
‘Hypernudge’: Big Data as a mode of regulation by design
Karen Yeung. 2017 · 2016
Earlier work this paper cites.
Effects of Online Recommendations on Consumers’ Willingness to Pay
Gediminas Adomavicius, Jesse C. Bockstedt, Shawn P. Curley, and Jingjing Zhang. 2018 · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017 · 2017
Earlier work this paper cites.
The New Market Manipulation
Tom C. W. Lin. 2017 · 2017
Earlier work this paper cites.
The Artificial Intelligence Black Box and the Failure of Intent and Causation
Yavar Bathaee. 2018 · 2018
Earlier work this paper cites.
A Game-Theoretic Approach to Recommendation Systems with Strategic Content Providers. In Advances in Neural Information Processing Systems , S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett (Eds.), Vol. 31. Curran Associates, Inc
Omer Ben-Porat and Moshe Tennenholtz. 2018 · 2018
Earlier work this paper cites.
Deceptive Design - User Interfaces Crafted to Trick You
Harry Brignull. 2018 · 2018
Earlier work this paper cites.
Aspiration: The Agency of Becoming
Agnes Callard. 2018 · 2018
Earlier work this paper cites.
How Algorithmic Confounding in Recommendation Systems Increases Homogeneity and Decreases Utility
Allison J. B. Chaney, Brandon M. Stewart, and Barbara E. Engelhardt. 2018 · 2018
Earlier work this paper cites.
In the Shades of the Uncanny Valley: An Experimental Study of Human–Chatbot Interaction
Leon Ciechanowski, Aleksandra Przegalinska, Mikolaj Magnuski, and Peter Gloor. 2019 · 2018
Earlier work this paper cites.
The Dark (Patterns) Side of UX Design. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems . ACM, Montreal QC Canada, 1–14
Colin M. Gray, Yubo Kou, Bryan Battles, Joseph Hoggatt, and Austin L. Toombs. 2018 · 2018
Earlier work this paper cites.
Towards Formal Definitions of Blameworthiness, Intention, and Moral responsibility. In Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence and Thirtieth Innovative Applications of Artificial Intelligence Conference and Eighth AAAI Symposium on Educational Advances in Artificial Intelligence (AAAI’18/IAAI’18/EAAI’18) . AAAI Press, New Orleans, Louisiana, USA, 1853–1860
Joseph Y. Halpern and Max Kleiman-Weiner. 2018 · 2018
Earlier work this paper cites.
Causal Structure Learning
Christina Heinze-Deml, Marloes H. Maathuis, and Nicolai Meinshausen. 2018 · 2018
Earlier work this paper cites.
Coercion and Deception in Persuasive Technologies. In 20th International Trust Workshop (co-located with AAMAS/IJCAI/ECAI/ICML 2018), Stockholm, Sweden, 14 July, 2018 . CEUR-WS, 38–49
Timotheus Kampik, Juan Carlos Nieves, and Helena Lindgren. 2018 · 2018
Earlier work this paper cites.
On Analyzing User Preference Dynamics with Temporal Social Networks
Fabíola S. F. Pereira, João Gama, Sandra de Amo, and Gina M. B. Oliveira. 2018 · 2018
Earlier work this paper cites.
The Good, the Bad and the Bait: Detecting and Characterizing Clickbait on YouTube. In 2018 IEEE Security and Privacy Workshops (SPW) . 63–69
Savvas Zannettou, Sotirios Chatzis, Kostantinos Papadamou, and Michael Sirivianos. 2018 · 2018
Earlier work this paper cites.
Online Political Microtargeting: Promises and Threats for Democracy
Frederik Zuiderveen Borgesius, Judith Moeller, Sanne Kruikemeier, Ronan Ó Fathaigh, Kristina Irion, Tom Dobber, Balázs Bodó, and Claes H. de Vreese. 2018 · 2018
Earlier work this paper cites.
"Reinforcement Learning for Recommender Systems: A Case Study on Youtube," by Minmin Chen
Association for Computing Machinery (ACM). 2019 · 2019
Earlier work this paper cites.
Where’s the Reward?
Shayan Doroudi, Vincent Aleven, and Emma Brunskill. 2019 · 2019
Earlier work this paper cites.
Horizon: Facebook’s Open Source Applied Reinforcement Learning Platform
Jason Gauci, Edoardo Conti, Yitao Liang, Kittipat Virochsiri, Yuchen He, Zachary Kaden, Vivek Narayanan, Xiaohui Ye, Zhengxing Chen, and Scott Fujimoto. 2019 · 2019
Earlier work this paper cites.
Social media addiction: Its impact, mediation, and intervention
Yubo Hou, Dan Xiong, Tonglin Jiang, Lily Song, and Qi Wang. 2019 · 2019
Cited alongside, same era.
The Disparate Effects of Strategic Manipulation. In Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT* ’19) . Association for Computing Machinery, New York, NY, USA, 259–268
Lily Hu, Nicole Immorlica, and Jennifer Wortman Vaughan. 2019 · 2019
Cited alongside, same era.
Human-level Performance in 3D Multiplayer Games with Population-Based Reinforcement Learning
Max Jaderberg, Wojciech M. Czarnecki, Iain Dunning, Luke Marris, Guy Lever, Antonio Garcia Castañeda, Charles Beattie, Neil C. Rabinowitz, Ari S. Morcos, Avraham Ruderman, Nicolas Sonnerat, Tim Green, Louise Deason, Joel Z. Leibo, David Silver, Demis Hassabis, Koray Kavukcuoglu, and Thore Graepel. 2019 · 2019
Cited alongside, same era.
How Do Classifiers Induce Agents to Invest Effort Strategically?. In Proceedings of the 2019 ACM Conference on Economics and Computation (EC ’19) . Association for Computing Machinery, New York, NY, USA, 825–844
Jon Kleinberg and Manish Raghavan. 2019 · 2019
Cited alongside, same era.
The Problem of Behaviour and Preference Manipulation in AI Systems. In The AAAI-22 Workshop on Artificial Intelligence Safety (SafeAI 2022)
Hal Ashton and Matija Franklin. 2022a · 2022
Later among the works it cites.
AI-driven Market Manipulation and Limits of the EU Law Enforcement Regime to Credible Deterrence
Alessio Azzutti. 2022 · 2022
Later among the works it cites.
On the Opportunities and Risks of Foundation Models
Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, Erik Brynjolfsson, Shyamal Buch, Dallas Card, Rodrigo Castellon, Niladri Chatterji, Annie Chen, Kathleen Creel, Jared Quincy Davis, Dora Demszky, Chris Donahue, Moussa Doumbouya, Esin Durmus, Stefano Ermon, John Etchemendy, Kawin Ethayarajh, Li Fei-Fei, Chelsea Finn, Trevor Gale, Lauren Gillespie, Karan Goel, Noah Goodman, Shelby Grossman, Neel Guha, Tatsunori Hashimoto, Peter Henderson, John Hewitt, Daniel E. Ho, Jenny Hong, Kyle Hsu, Jing Huang, Thomas Icard, Saahil Jain, Dan Jurafsky, Pratyusha Kalluri, Siddharth Karamcheti, Geoff Keeling, Fereshte Khani, Omar Khattab, Pang Wei Koh, Mark Krass, Ranjay Krishna, Rohith Kuditipudi, Ananya Kumar, Faisal Ladhak, Mina Lee, Tony Lee, Jure Leskovec, Isabelle Levent, Xiang Lisa Li, Xuechen Li, Tengyu Ma, Ali Malik, Christopher D. Manning, Suvir Mirchandani, Eric Mitchell, Zanele Munyikwa, Suraj Nair, Avanika Narayan, Deepak Narayanan, Ben Newman, Allen Nie, Juan Carlos Niebles, Hamed Nilforoshan, Julian Nyarko, Giray Ogut, Laurel Orr, Isabel Papadimitriou, Joon Sung Park, Chris Piech, Eva Portelance, Christopher Potts, Aditi Raghunathan, Rob Reich, Hongyu Ren, Frieda Rong, Yusuf Roohani, Camilo Ruiz, Jack Ryan, Christopher Ré, Dorsa Sadigh, Shiori Sagawa, Keshav Santhanam, Andy Shih, Krishnan Srinivasan, Alex Tamkin, Rohan Taori, Armin W. Thomas, Florian Tramèr, Rose E. Wang, William Wang, Bohan Wu, Jiajun Wu, Yuhuai Wu, Sang Michael Xie, Michihiro Yasunaga, Jiaxuan You, Matei Zaharia, Michael Zhang, Tianyi Zhang, Xikun Zhang, Yuhui Zhang, Lucia Zheng, Kaitlyn Zhou, and Percy Liang. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Victoria Krakovna, Laurent Orseau, Ramana Kumar, Miljan Martic, and Shane Legg. 2019 · 2019
Cited alongside, same era.
Categorizing Variants of Goodhart’s Law
David Manheim and Scott Garrabrant. 2019 · 2019
Cited alongside, same era.
The Social Cost of Strategic Classification. In Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT* ’19) . Association for Computing Machinery, New York, NY, USA, 230–239
Smitha Milli, John Miller, Anca D. Dragan, and Moritz Hardt. 2019 · 2019
Cited alongside, same era.
Choosing for Changing Selves (1 ed.)
Richard Pettigrew. 2019 · 2019
Cited alongside, same era.
Actionable Auditing. In Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society . ACM
Inioluwa Deborah Raji and Joy Buolamwini. 2019 · 2019
Cited alongside, same era.
Invisible Influence: Artificial Intelligence and the Ethics of Adaptive Choice Architectures. In Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society (AIES ’19) . Association for Computing Machinery, New York, NY, USA, 403–408
Daniel Susser. 2019 · 2019
Cited alongside, same era.
Online Manipulation: Hidden Influences in a Digital World
Daniel Susser, Beate Roessler, and Helen Nissenbaum. 2019a · 2019
Cited alongside, same era.
Technology, Autonomy, and Manipulation
Daniel Susser, Beate Roessler, and Helen Nissenbaum. 2019b · 2019
Cited alongside, same era.
Later among the works it cites.
Estimating and Penalizing Induced Preference Shifts in Recommender Systems
Micah Carroll, Anca Dragan, Stuart Russell, and Dylan Hadfield-Menell. 2022 · 2022
Later among the works it cites.
Towards Psychologically-Grounded Dynamic Preference Models. In Proceedings of the 16th ACM Conference on Recommender Systems . 35–48
Mihaela Curmei, Andreas A. Haupt, Benjamin Recht, and Dylan Hadfield-Menell. 2022 · 2022
Later among the works it cites.
Path-Specific Objectives for Safer Agent Incentives
Sebastian Farquhar, Ryan Carey, and Tom Everitt. 2022 · 2022
Later among the works it cites.
Matija Franklin, Hal Ashton, Rebecca Gorman, and Stuart Armstrong. 2022 · 2022
Later among the works it cites.
Predictability and Surprise in Large Generative Models. In 2022 ACM Conference on Fairness, Accountability, and Transparency . ACM
Deep Ganguli, Danny Hernandez, Liane Lovitt, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova Dassarma, Dawn Drain, Nelson Elhage, Sheer El Showk, Stanislav Fort, Zac Hatfield-Dodds, Tom Henighan, Scott Johnston, Andy Jones, Nicholas Joseph, Jackson Kernian, Shauna Kravec, Ben Mann, Neel Nanda, Kamal Ndousse, Catherine Olsson, Daniela Amodei, Tom Brown, Jared Kaplan, Sam McCandlish, Christopher Olah, Dario Amodei, and Jack Clark. 2022 · 2022
Later among the works it cites.
An Axiomatic Approach to Revising Preferences
Adrian Haret and Johannes Peter Wallner. 2022 · 2022
Later among the works it cites.
Simulators
Janus. 2022 · 2022
Later among the works it cites.
Survey of Hallucination in Natural Language Generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Yejin Bang, Andrea Madotto, and Pascale Fung. 2022 · 2022
Later among the works it cites.
Language Models (Mostly) Know What They Know
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, Scott Johnston, Sheer El-Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yuntao Bai, Sam Bowman, Stanislav Fort, Deep Ganguli, Danny Hernandez, Josh Jacobson, Jackson Kernion, Shauna Kravec, Liane Lovitt, Kamal Ndousse, Catherine Olsson, Sam Ringer, Dario Amodei, Tom Brown, Jack Clark, Nicholas Joseph, Ben Mann, Sam McCandlish, Chris Olah, and Jared Kaplan. 2022 · 2022
Later among the works it cites.
Causal Machine Learning: A Survey and Open Problems
Jean Kaddour, Aengus Lynch, Qi Liu, Matt J. Kusner, and Ricardo Silva. 2022 · 2022
Later among the works it cites.
Zachary Kenton, Ramana Kumar, Sebastian Farquhar, Jonathan Richens, Matt MacDermott, and Tom Everitt. 2022 · 2022
Later among the works it cites.
Targeting Recommendation Algorithms to Ideal Preferences Makes Users Better Off
Poruz Khambatta, Shwetha Mariadassou, Joshua Morris, and S Christian Wheeler. 2022 · 2022
Later among the works it cites.
Goal Misgeneralization in Deep Reinforcement Learning. In Proceedings of the 39th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 162) , Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvari, Gang Niu, and Sivan Sabato (Eds.). PMLR, 12004–12019
Lauro Langosco Di Langosco, Jack Koch, Lee D Sharkey, Jacob Pfau, and David Krueger. 2022 · 2022
Later among the works it cites.
Human-Level Play in the Game of Diplomacy by Combining Language Models with Strategic Reasoning
Meta Fundamental AI Research Diplomacy Team (FAIR), Anton Bakhtin, Noam Brown, Emily Dinan, Gabriele Farina, Colin Flaherty, Daniel Fried, Andrew Goff, Jonathan Gray, Hengyuan Hu, Athul Paul Jacob, Mojtaba Komeili, Karthik Konath, Minae Kwon, Adam Lerer, Mike Lewis, Alexander H. Miller, Sasha Mitts, Adithya Renduchintala, Stephen Roller, Dirk Rowe, Weiyan Shi, Joe Spisak, Alexander Wei, David Wu, Hugh Zhang, and Markus Zijlstra. 2022 · 2022
Later among the works it cites.
Finding the ‘Nudge’ in Hypernudge
Stuart Mills. 2022 · 2022
Later among the works it cites.
Can We Design Artificial Persons without Being Manipulative?
Maciej Musiał. 2022 · 2022
Later among the works it cites.
The Ethics of Manipulation
Robert Noggle. 2022 · 2022
Later among the works it cites.
In-context Learning and Induction Heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. 2022 · 2022
Later among the works it cites.
Choosing for Changing Selves
L. A. Paul. 2022 · 2022
Later among the works it cites.
Modelling Persuasion through Misuse of Rhetorical Appeals
Amalie Brogaard Pauli, Leon Derczynski, and Ira Assent. 2022 · 2022
Later among the works it cites.
Nudging for Changing Selves
Richard Pettigrew. 2022 · 2022
Later among the works it cites.
Human Autonomy in the Age of Artificial Intelligence
Carina Prunkl. 2022 · 2022
Later among the works it cites.
Outsider Oversight: Designing a Third Party Audit Ecosystem for AI Governance. In Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society . ACM
Inioluwa Deborah Raji, Peggy Xu, Colleen Honigsberg, and Daniel Ho. 2022 · 2022
Later among the works it cites.
Counterfactual Harm. In Advances in Neural Information Processing Systems , Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho (Eds.)
Jonathan Richens, Rory Beard, and Daniel H. Thompson. 2022 · 2022
Later among the works it cites.
Goal Misgeneralization: Why Correct Specifications Aren’t Enough For Correct Goals
Rohin Shah, Vikrant Varma, Ramana Kumar, Mary Phuong, Victoria Krakovna, Jonathan Uesato, and Zac Kenton. 2022 · 2022
Later among the works it cites.
What’s In a Name?
caroline sinders. 2022 · 2022
Later among the works it cites.
Defining and Characterizing Reward Hacking
Joar Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, and David Krueger. 2022 · 2022
Later among the works it cites.
The Galactica AI Model was Trained on Scientific Knowledge – but it Spat Out Alarmingly Plausible Nonsense
Aaron J. Snoswell and Jean Burgess. 2022 · 2022
Later among the works it cites.
How Platform Recommenders Work
Luke Thorburn. 2022 · 2022
Later among the works it cites.
What Will “Amplification” Mean in Court?
Luke Thorburn, Jonathan Stray, and Priyanjana Bengani. 2022 · 2022
Later among the works it cites.
On Agent Incentives to Manipulate Human Feedback in Multi-Agent Reward Learning Scenarios
Francis Rhys Ward. 2022 · 2022
Later among the works it cites.
A Causal Perspective on AI Deception in Games
Francis Rhys Ward, Francesca Toni, and Francesco Belardinelli. 2022 · 2022
Later among the works it cites.
Emergent Abilities of Large Language Models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. 2022 · 2022
Later among the works it cites.
Taxonomy of Risks posed by Language Models. In 2022 ACM Conference on Fairness, Accountability, and Transparency . ACM
Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa Kasirzadeh, Courtney Biles, Sasha Brown, Zac Kenton, Will Hawkins, Tom Stepleton, Abeba Birhane, Lisa Anne Hendricks, Laura Rimell, William Isaac, Julia Haas, Sean Legassick, Geoffrey Irving, and Iason Gabriel. 2022 · 2022
Later among the works it cites.
Understanding or Manipulation: Rethinking Online Performance Gains of Modern Recommender Systems
Zhengbang Zhu, Rongjun Qin, Junjie Huang, Xinyi Dai, Yang Yu, Yong Yu, and Weinan Zhang. 2022 · 2022
Later among the works it cites.
Artificial Intelligence Can Persuade Humans
Hui Bai. 2023 · 2023
Closest in time.
Discovering Latent Knowledge in Language Models Without Supervision. In The Eleventh International Conference on Learning Representations
Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt. 2023 · 2023
Closest in time.
Reinforcing User Retention in a Billion Scale Short Video Recommender System
Qingpeng Cai, Shuchang Liu, Xueliang Wang, Tianyou Zuo, Wentao Xie, Bin Yang, Dong Zheng, Peng Jiang, and Kun Gai. 2023 · 2023
Closest in time.
Harms from Increasingly Agentic Algorithmic Systems
Alan Chan, Rebecca Salganik, Alva Markelius, Chris Pang, Nitarshan Rajkumar, Dmitrii Krasheninnikov, Lauro Langosco, Zhonghao He, Yawen Duan, Micah Carroll, Michelle Lin, Alex Mayhew, Katherine Collins, Maryam Molamohammadi, John Burden, Wanru Zhao, Shalaleh Rismani, Konstantinos Voudouris, Umang Bhatt, Adrian Weller, David Krueger, and Tegan Maharaj. 2023 · 2023
Closest in time.
Susceptibility to Influence of Large Language Models
Lewis D. Griffin, Bennett Kleinberg, Maximilian Mozes, Kimberly T. Mai, Maria Vau, Matthew Caldwell, and Augustine Marvor-Parker. 2023 · 2023
Closest in time.
Learning to Influence Human Behavior with Offline Reinforcement Learning
Joey Hong, Anca Dragan, and Sergey Levine. 2023 · 2023
Closest in time.
Co-Writing with Opinionated Language Models Affects Users’ Views
Maurice Jakesch, Advait Bhat, Daniel Buschek, Lior Zalmanson, and Mor Naaman. 2023 · 2023
Closest in time.
Theory of Mind May Have Spontaneously Emerged in Large Language Models
Michal Kosinski. 2023 · 2023
Closest in time.
Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task. In The Eleventh International Conference on Learning Representations
Kenneth Li, Aspen K. Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. 2023 · 2023
Closest in time.
Dissociating language and thought in large language models: a cognitive perspective
Kyle Mahowald, Anna A. Ivanova, Idan A. Blank, Nancy Kanwisher, Joshua B. Tenenbaum, and Evelina Fedorenko. 2023 · 2023
Closest in time.
Engagement, User Satisfaction, and the Amplification of Divisive Content on Social Media
Smitha Milli, Micah Carroll, Yike Wang, Sashrika Pandey, Sebastian Zhao, and Anca D. Dragan. 2023 · 2023
Closest in time.
Definition of manipulation
APA Dictionary of Psychology. 2023 · 2023
Closest in time.
Alexander Pan, Chan Jun Shern, Andy Zou, Nathaniel Li, Steven Basart, Thomas Woodside, Jonathan Ng, Hanlin Zhang, Scott Emmons, and Dan Hendrycks. 2023 · 2023
Closest in time.
AI Deception: A Survey of Examples, Risks, and Potential Solutions
Peter S. Park, Simon Goldstein, Aidan O’Gara, Michael Chen, and Dan Hendrycks. 2023 · 2023
Closest in time.
The Amplification Paradox in Recommender Systems
Manoel Horta Ribeiro, Veniamin Veselovsky, and Robert West. 2023 · 2023
Closest in time.
Toolformer: Language Models Can Teach Themselves to Use Tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023 · 2023
Closest in time.
Emergent Deception and Emergent Optimization
Jacob Steinhardt. 2023 · 2023
Closest in time.
Twitter’s Recommendation Algorithm
Twitter. 2023 · 2023
Closest in time.
Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Tomer Ullman. 2023 · 2023
Closest in time.
Microsoft’s Bing is an emotionally manipulative liar, and people love it
James Vincent. 2023 · 2023
Closest in time.
Chatbots Shouldn’t Use Emojis
Carissa Véliz. 2023 · 2023
Closest in time.
Honesty Is the Best Policy: Defining and Mitigating AI Deception
Francis Rhys Ward, Tom Everitt, Francesca Toni, and Francesco Belardinelli. 2023 · 2023
Closest in time.