Aligning to Social Norms and Values in Interactive Narratives
Prithviraj Ammanabrolu, Liwei Jiang, Maarten Sap, Hannaneh Hajishirzi, and Yejin Choi. 2022 · 2022
Later among the works it cites.
Language Models as Agent Models
Original
Jacob Andreas. 2022 · 2022
Later among the works it cites.
Out of One, Many: Using Language Models to Simulate Human Samples
Original
Lisa P. Argyle, Ethan C. Busby, Nancy Fulda, Joshua Gubler, Christopher Rytting, and David Wingate. 2022 · 2022
Later among the works it cites.
Fine-tuning language models to find agreement among humans with diverse preferences
Original
Michiel A. Bakker, Martin J. Chadwick, Hannah R. Sheahan, Michael Henry Tessler, Lucy Campbell-Gillingham, Jan Balaguer, Nat McAleese, Amelia Glaese, John Aslanides, Matthew M. Botvinick, and Christopher Summerfield. 2022 · 2022
Later among the works it cites.
Enabling Classifiers to Make Judgements Explicitly Aligned with Human Values
Original
Yejin Bang, Tiezheng Yu, Andrea Madotto, Zhaojiang Lin, Mona Diab, and Pascale Fung. 2022 · 2022
Later among the works it cites.
Artificial Intelligence Enabled Personalised Assistive Tools to Enhance Education of Children with Neurodevelopmental Disorders—A Review
Prabal Datta Barua, Jahmunah Vicnesh, Raj Gururajan, Shu Lih Oh, Elizabeth Palmer, Muhammad Mokhzaini Azizan, Nahrizul Adib Kadri, and U. Rajendra Acharya. 2022 · 2022
Later among the works it cites.
Power to the People? Opportunities and Challenges for Participatory AI
Abeba Birhane, William Isaac, Vinodkumar Prabhakaran, Mark Diaz, Madeleine Clare Elish, Iason Gabriel, and Shakir Mohamed. 2022 · 2022
Later among the works it cites.
Measuring Progress on Scalable Oversight for Large Language Models
Original
Samuel R. Bowman, Jeeyoon Hyun, Ethan Perez, Edwin Chen, Craig Pettit, Scott Heiner, Kamilė Lukošiūtė, Amanda Askell, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Christopher Olah, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Jackson Kernion, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, Liane Lovitt, Nelson Elhage, Nicholas Schiefer, Nicholas Joseph, Noemí Mercado, Nova DasSarma, Robin Larson, Sam McCandlish, Sandipan Kundu, Scott Johnston, Shauna Kravec, Sheer El Showk, Stanislav Fort, Timothy Telleen-Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Ben Mann, and Jared Kaplan. 2022 · 2022
Later among the works it cites.
A personalized dialogue generator with implicit user persona detection
Itsugun Cho, Dongyang Wang, Ryota Takahashi, and Hiroaki Saito. 2022 · 2022
Later among the works it cites.
AI liability directive
European Commission. 2022 · 2022
Later among the works it cites.
Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
Original
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, Andy Jones, Sam Bowman, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Nelson Elhage, Sheer El-Showk, Stanislav Fort, Zac Hatfield-Dodds, Tom Henighan, Danny Hernandez, Tristan Hume, Josh Jacobson, Scott Johnston, Shauna Kravec, Catherine Olsson, Sam Ringer, Eli Tran-Johnson, Dario Amodei, Tom Brown, Nicholas Joseph, Sam McCandlish, Chris Olah, Jared Kaplan, and Jack Clark. 2022 · 2022
Later among the works it cites.
AI Trained on 4Chan Becomes ‘Hate Speech Machine’
Matthew Gault. 2022 · 2022
Later among the works it cites.
Do Not Recommend? Reduction as a Form of Content Moderation
Tarleton Gillespie. 2022 · 2022
Later among the works it cites.
Improving alignment of dialogue agents via targeted human judgements
Original
Amelia Glaese, Nat McAleese, Maja Trębacz, John Aslanides, Vlad Firoiu, Timo Ewalds, Maribeth Rauh, Laura Weidinger, Martin Chadwick, Phoebe Thacker, Lucy Campbell-Gillingham, Jonathan Uesato, Po-Sen Huang, Ramona Comanescu, Fan Yang, Abigail See, Sumanth Dathathri, Rory Greig, Charlie Chen, Doug Fritz, Jaume Sanchez Elias, Richard Green, Soňa Mokrá, Nicholas Fernando, Boxi Wu, Rachel Foley, Susannah Young, Iason Gabriel, William Isaac, John Mellor, Demis Hassabis, Koray Kavukcuoglu, Lisa Anne Hendricks, and Geoffrey Irving. 2022 · 2022
Later among the works it cites.
Algorithmically Embodied Emissions: The Environmental Harm of Everyday Life Information in Digital Culture
Jutta Haider, Malte Rödl, and Sofie Joosse. 2022 · 2022
Later among the works it cites.
Training Compute-Optimal Large Language Models
Original
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals, and Laurent Sifre. 2022 · 2022
Later among the works it cites.
Unnatural Instructions: Tuning Language Models with (Almost) No Human Labor
Original
Or Honovich, Thomas Scialom, Omer Levy, and Timo Schick. 2022 · 2022
Later among the works it cites.
Large Language Models as Simulated Economic Agents: What Can We Learn from Homo Silicus?
John J Horton. 2022 · 2022
Later among the works it cites.
Learning to Adapt Domain Shifts of Moral Values via Instance Weighting
Xiaolei Huang, Alexandra Wormley, and Adam Cohen. 2022 · 2022
Later among the works it cites.
Exploring AI chatbot affordances in the EFL classroom: young learners’ experiences and perspectives
Jaeho Jeon. 2022 · 2022
Later among the works it cites.
Survey of Hallucination in Natural Language Generation
Original
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Yejin Bang, Andrea Madotto, and Pascale Fung. 2022 · 2022
Later among the works it cites.
Can Machines Learn Morality? The Delphi Experiment
Original
Liwei Jiang, Jena D. Hwang, Chandra Bhagavatula, Ronan Le Bras, Jenny Liang, Jesse Dodge, Keisuke Sakaguchi, Maxwell Forbes, Jon Borchardt, Saadia Gabriel, Yulia Tsvetkov, Oren Etzioni, Maarten Sap, Regina Rini, and Yejin Choi. 2022 · 2022
Later among the works it cites.
When to Make Exceptions: Exploring Language Models as Accounts of Human Moral Judgment
Original
Zhijing Jin, Sydney Levine, Fernando Gonzalez, Ojasv Kamal, Maarten Sap, Mrinmaya Sachan, Rada Mihalcea, Josh Tenenbaum, and Bernhard Schölkopf. 2022 · 2022
Later among the works it cites.
Language Models (Mostly) Know What They Know
Original
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield Dodds, Nova DasSarma, Eli Tran-Johnson, Scott Johnston, Sheer El-Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yuntao Bai, Sam Bowman, Stanislav Fort, Deep Ganguli, Danny Hernandez, Josh Jacobson, Jackson Kernion, Shauna Kravec, Liane Lovitt, Kamal Ndousse, Catherine Olsson, Sam Ringer, Dario Amodei, Tom Brown, Jack Clark, Nicholas Joseph, Ben Mann, Sam McCandlish, Chris Olah, and Jared Kaplan. 2022 · 2022
Later among the works it cites.
What if ground truth is subjective? Personalized deep neural hate speech detection
Kamil Kanclerz, Marcin Gruza, Konrad Karanowski, Julita Bielaniewicz, Piotr Milkowski, Jan Kocon, and Przemyslaw Kazienko. 2022 · 2022
Later among the works it cites.
Building a personalized dialogue system with prompt-tuning
Tomohito Kasahara, Daisuke Kawahara, Nguyen Tung, Shengzhe Li, Kenta Shinzato, and Toshinori Sato. 2022 · 2022
Later among the works it cites.
In conversation with Artificial Intelligence: aligning language models with human values
Original
Atoosa Kasirzadeh and Iason Gabriel. 2022 · 2022
Later among the works it cites.
Can Pretrained Language Models Generate Persuasive, Faithful, and Informative Ad Text for Product Descriptions?
Fajri Koto, Jey Han Lau, and Timothy Baldwin. 2022 · 2022
Later among the works it cites.
PERSONACHATGEN: Generating personalized dialogues using GPT-3
Young-Jun Lee, Chae-Gyun Lim, Yunsu Choi, Ji-Hui Lm, and Ho-Jin Choi. 2022 · 2022
Later among the works it cites.
TruthfulQA: Measuring How Models Mimic Human Falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans. 2022 · 2022
Later among the works it cites.
Aligning Generative Language Models with Human Values
Ruibo Liu, Ge Zhang, Xinyu Feng, and Soroush Vosoughi. 2022 · 2022
Later among the works it cites.
Artificial empathy in marketing interactions: Bridging the human-AI gap in affective and social customer experience
Yuping Liu-Thompkins, Shintaro Okazaki, and Hairong Li. 2022 · 2022
Later among the works it cites.
Internet addiction in young adults: A meta-analysis and systematic review
Raquel Lozano-Blasco, Alberto Quilez Robres, and Alberto Soto Sánchez. 2022 · 2022
Later among the works it cites.
Towards Boosting the Open-Domain Chatbot with Human Feedback
Original
Hua Lu, Siqi Bao, Huang He, Fan Wang, Hua Wu, and Haifeng Wang. 2022 · 2022
Later among the works it cites.
The digital divide: A review and future research agenda
Sophie Lythreatis, Sanjay Kumar Singh, and Abdul-Nasser El-Kassar. 2022 · 2022
Later among the works it cites.
“Alexa, let’s talk about my productivity”: The impact of digital assistants on work productivity
Davit Marikyan, Savvas Papagiannidis, Omer F. Rana, Rajiv Ranjan, and Graham Morgan. 2022 · 2022
Later among the works it cites.
UserIdentifier: Implicit user representations for simple and effective personalized sentiment analysis
Fatemehsadat Mireshghallah, Vaishnavi Shrivastava, Milad Shokouhi, Taylor Berg-Kirkpatrick, Robert Sim, and Dimitrios Dimitriadis. 2022 · 2022
Later among the works it cites.
Personalizing weekly diet reports
Elena Monfroglio, Lucas Anselma, and Alessandro Mazzei. 2022 · 2022
Later among the works it cites.
“This is a political movement, friend”: Why “incels” support violence
Catharina O’Donnell and Eran Shor. 2022 · 2022
Later among the works it cites.
Introducing ChatGPT
OpenAI. 2022 · 2022
Later among the works it cites.
The impact of recommender systems on competition between music companies
Peter L Ormosi and Rahul Savani. 2022 · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Original
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022 · 2022
Later among the works it cites.
Discovering Language Model Behaviors with Model-Written Evaluations
Ethan Perez, Sam Ringer, Kamilė Lukošiūtė, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, Andy Jones, Anna Chen, Ben Mann, Brian Israel, Bryan Seethor, Cameron McKinnon, Christopher Olah, Da Yan, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Guro Khundadze, Jackson Kernion, James Landis, Jamie Kerr, Jared Mueller, Jeeyoon Hyun, Joshua Landau, Kamal Ndousse, Landon Goldberg, Liane Lovitt, Martin Lucas, Michael Sellitto, Miranda Zhang, Neerav Kingsland, Nelson Elhage, Nicholas Joseph, Noemí Mercado, Nova DasSarma, Oliver Rausch, Robin Larson, Sam McCandlish, Scott Johnston, Shauna Kravec, Sheer El Showk, Tamera Lanham, Timothy Telleen-Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Jack Clark, Samuel R. Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan. 2022 · 2022
Later among the works it cites.
Perturbation Augmentation for Fairer NLP
Original
Rebecca Qian, Candace Ross, Jude Fernandes, Eric Smith, Douwe Kiela, and Adina Williams. 2022 · 2022
Later among the works it cites.
In(cel)doctrination: How technologically facilitated misogyny moves violence off screens and on to streets
Kaitlyn Regehr. 2022 · 2022
Later among the works it cites.
Two Contrasting Data Annotation Paradigms for Subjective NLP Tasks
Paul Rottger, Bertie Vidgen, Dirk Hovy, and Janet Pierrehumbert. 2022 · 2022
Later among the works it cites.
Situating Search
Chirag Shah and Emily M. Bender. 2022 · 2022
Later among the works it cites.
"I’m sorry to hear that": finding bias in language models with a holistic descriptor dataset
Original
Eric Michael Smith, Melissa Hall Melanie Kambadur, Eleonora Presani, and Adina Williams. 2022 · 2022
Later among the works it cites.
Methods for Estimating and Improving Robustness of Language Models
Michal Stefanik. 2022 · 2022
Later among the works it cites.
LaMDA: Language Models for Dialog Applications
Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, YaGuang Li, Hongrae Lee, Huaixiu Steven Zheng, Amin Ghafouri, Marcelo Menegali, Yanping Huang, Maxim Krikun, Dmitry Lepikhin, James Qin, Dehao Chen, Yuanzhong Xu, Zhifeng Chen, Adam Roberts, Maarten Bosma, Vincent Zhao, Yanqi Zhou, Chung-Ching Chang, Igor Krivokon, Will Rusch, Marc Pickett, Pranesh Srinivasan, Laichee Man, Kathleen Meier-Hellstern, Meredith Ringel Morris, Tulsee Doshi, Renelito Delos Santos, Toju Duke, Johnny Soraker, Ben Zevenbergen, Vinodkumar Prabhakaran, Mark Diaz, Ben Hutchinson, Kristen Olson, Alejandra Molina, Erin Hoffman-John, Josh Lee, Lora Aroyo, Ravi Rajakumar, Alena Butryna, Matthew Lamm, Viktoriya Kuzmina, Joe Fenton, Aaron Cohen, Rachel Bernstein, Ray Kurzweil, Blaise Aguera-Arcas, Claire Cui, Marian Croak, Ed Chi, and Quoc Le. 2022 · 2022
Later among the works it cites.
Persona or context? Towards building context adaptive personalized persuasive virtual sales assistant
Abhisek Tiwari, Sriparna Saha, Shubhashis Sengupta, Anutosh Maitra, Roshni Ramnani, and Pushpak Bhattacharyya. 2022 · 2022
Later among the works it cites.
Inside a White Power echo chamber: Why fringe digital spaces are polarizing politics
Petter Törnberg and Anton Törnberg. 2022 · 2022
Later among the works it cites.
A Study of Implicit Bias in Pretrained Language Models against People with Disabilities
Pranav Narayanan Venkit, Mukund Srinath, and Shomir Wilson. 2022 · 2022
Later among the works it cites.
Self-Instruct: Aligning Language Model with Self Generated Instructions
Original
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2022 · 2022
Later among the works it cites.
Taxonomy of Risks posed by Language Models
Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa Kasirzadeh, Courtney Biles, Sasha Brown, Zac Kenton, Will Hawkins, Tom Stepleton, Abeba Birhane, Lisa Anne Hendricks, Laura Rimell, William Isaac, Julia Haas, Sean Legassick, Geoffrey Irving, and Iason Gabriel. 2022 · 2022
Later among the works it cites.
Leveraging similar users for personalized language modeling with limited data
Charles Welch, Chenxi Gu, Jonathan K. Kummerfeld, Veronica Perez-Rosas, and Rada Mihalcea. 2022 · 2022
Later among the works it cites.
Less is more: Learning to refine dialogue history for personalized dialogue generation
Hanxun Zhong, Zhicheng Dou, Yutao Zhu, Hongjin Qian, and Ji-Rong Wen. 2022 · 2022
Later among the works it cites.
The Moral Integrity Corpus: A Benchmark for Ethical Dialogue Systems
Caleb Ziems, Jane Yu, Yi-Chia Wang, Alon Halevy, and Diyi Yang. 2022 · 2022
Later among the works it cites.
Participatory Design of AI Systems: Opportunities and Challenges Across Diverse Users, Relationships, and Application Domains
Douglas Zytko, Pamela J. Wisniewski, Shion Guha, Eric P. S. Baumer, and Min Kyung Lee. 2022 · 2022
Later among the works it cites.
Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies
Original
Gati Aher, Rosa I. Arriaga, and Adam Tauman Kalai. 2023 · 2023
Closest in time.
Even kids are worried ChatGPT will make them lazy plagiarists, says a linguist who studies tech’s effect on reading, writing and thinking
Naomi S. Baron. 2023 · 2023
Closest in time.
Why People Are Confessing Their Love For AI Chatbots
Andrew R. Chow. 2023 · 2023
Closest in time.
ChatGPT Revenue and Usage Statistics (2023)
David Curry. 2023 · 2023
Closest in time.
Meet Claude: Anthropic’s Rival to ChatGPT | Blog | Scale AI
Riley Goodside and Spencer Papay. 2023 · 2023
Closest in time.
The impact of algorithmically driven recommendation systems on music consumption and production - a literature review
David Hesmondhalgh, Raquel Campos Valverde, Bondy Valdovinos Kaye D., and Li Zhongwei. 2023 · 2023
Closest in time.
ChatGPT Burns Millions Every Day. Can Computer Scientists Make AI One Million Times More Efficient?
John Koetsier. 2023 · 2023
Closest in time.
A participatory initiative to include LGBT+ voices in AI for mental health
Andrey Kormilitzin, Nenad Tomasev, Kevin R. McKee, and Dan W. Joyce. 2023 · 2023
Closest in time.
How AI encourages consumers to share their secrets? The role of anthropomorphism, personalisation, and privacy concerns and avenues for future research
Bianca Kronemann, Hatice Kizgin, Nripendra Rana, and Yogesh K. Dwivedi. 2023 · 2023
Closest in time.
Selective Explanations: Leveraging Human Input to Align Explainable AI
Original
Vivian Lai, Yiming Zhang, Chacha Chen, Q. Vera Liao, and Chenhao Tan. 2023 · 2023
Closest in time.
Second Thoughts are Best: Learning to Re-Align With Human Values from Text Edits
Original
Ruibo Liu, Chenyan Jia, Ge Zhang, Ziyu Zhuang, Tony X. Liu, and Soroush Vosoughi. 2023 · 2023
Closest in time.
The ChatGPT Addiction: 3 Reasons Why ChatGPT Will Make You Obsessed! Medium (CyberSec_sai)
Medium. 2023 · 2023
Closest in time.
Reinventing search with a new AI-powered Microsoft Bing and Edge, your copilot for the web
Yusuf Mehdi. 2023 · 2023
Closest in time.
Introducing LLaMA: A foundational, 65-billion-parameter language model
MetaAI. 2023 · 2023
Closest in time.
Auditing large language models: a three-layered approach
Original
Jakob Mökander, Jonas Schuett, Hannah Rose Kirk, and Luciano Floridi. 2023 · 2023
Closest in time.
Online Safety Bill
UK Parliament. 2023 · 2023
Closest in time.
An important next step on our AI journey
Sundar Pichai. 2023 · 2023
Closest in time.
Identifying Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction
Original
Renee Shelby, Shalaleh Rismani, Kathryn Henne, AJung Moon, Negar Rostamzadeh, Paul Nicholas, N’Mah Yilla, Jess Gallegos, Andrew Smart, Emilio Garcia, and Gurleen Virk. 2023 · 2023
Closest in time.
ChatGB: Tony Blair backs push for taxpayer-funded ‘sovereign AI’ to rival ChatGPT
James Titcomb and Matthew Field. 2023 · 2023
Closest in time.
Personalized machine translation: Predicting translational preferences
Shachar Mirkin and Jean-Luc Meunier. 2015 · 2025
Closest in time.
Dialogue Learning with Human Teaching and Feedback in End-to-End Trainable Task-Oriented Dialogue Systems
Bing Liu, Gokhan Tür, Dilek Hakkani-Tür, Pararth Shah, and Larry Heck. 2018 · 2069
Closest in time.