Fetching the paper…
Reading the bibliography…
How can we build AI systems that can learn any set of individual human values both quickly and safely, avoiding causing harm or violating societal standards for acceptable behavior during the learning process? We explore the effects of representational alignment between humans and AI agents on learning human values.
Representational content in humans and machines
Mark H Bickhard · 1993
Earlier work this paper cites.
Gaussian Processes for Machine Learning
Carl Edward Rasmussen and Christopher K. I. Williams, editors · 2005
Earlier work this paper cites.
Analysis of Thompson sampling for the multi-armed bandit problem
Shipra Agrawal and Navin Goyal · 2012
Earlier work this paper cites.
Safe exploration of state and action spaces in reinforcement learning
Javier Garcia and Fernando Fernández · 2012
Earlier work this paper cites.
Distributed representations of sentences and documents
Quoc Le and Tomas Mikolov · 2014
Earlier work this paper cites.
Support Vector Regression
Mariette Awad and Rahul Khanna · 2015
Earlier work this paper cites.
A tutorial on thompson sampling, November 2017
Daniel J. Russo, Benjamin Van Roy, Abbas Kazerouni, Ian Osband, and Zheng Wen · 2017
Earlier work this paper cites.
The moral machine experiment
Edmond Awad, Sohan Dsouza, Richard Kim, Jonathan Schulz, Joseph Henrich, Azim Shariff, Jean-François Bonnefon, and Iyad Rahwan · 2018
Earlier work this paper cites.
Adaptive sensitive reweighting to mitigate bias in fairness-aware classification
Emmanouil Krasanakis, Eleftherios Spyromitros-Xioufis, Symeon Papadopoulos, and Yiannis Kompatsiaris · 2018
Earlier work this paper cites.
Evaluating (and improving) the correspondence between deep neural networks and human representations
Joshua C. Peterson, Joshua T. Abbott, and Thomas L. Griffiths · 2018
Earlier work this paper cites.
Benchmarking safe exploration in deep reinforcement learning
Joshua Achiam and Dario Amodei · 2019
Earlier work this paper cites.
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Nils Reimers and Iryna Gurevych · 2019
Earlier work this paper cites.
What does the mind learn? a comparison of human and machine learning representations
Jake Spicer and Adam N Sanborn · 2019
Earlier work this paper cites.
A survey of inverse reinforcement learning: Challenges, methods and progress
Saurabh Arora and Prashant Doshi · 2020
Earlier work this paper cites.
Aligning AI with shared human values
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt · 2020
Earlier work this paper cites.
The curse of dense low-dimensional information retrieval for large index sizes
Nils Reimers and Iryna Gurevych · 2020
Earlier work this paper cites.
Making monolingual sentence embeddings multilingual using knowledge distillation
Nils Reimers and Iryna Gurevych · 2020
Earlier work this paper cites.
Nandan Thakur, Nils Reimers, Johannes Daxenberger, and Iryna Gurevych · 2020
Earlier work this paper cites.
Training value-aligned reinforcement learning agents using a normative prior
Md Sultan Al Nahian, Spencer Frazier, Brent Harrison, and Mark Riedl · 2021
Cited alongside, same era.
What would Jiminy Cricket do? towards agents that behave morally
Dan Hendrycks, Mantas Mazeika, Andy Zou, Sahil Patel, Christine Zhu, Jesus Navarro, Dawn Song, Bo Li, and Jacob Steinhardt · 2021
Cited alongside, same era.
Exploring alignment of representations with human perception
Vedant Nanda, Ayan Majumdar, Camila Kolling, John P. Dickerson, Krishna P. Gummadi, Bradley C. Love, and Adrian Weller · 2021
Cited alongside, same era.
BEIR: A heterogenous benchmark for zero-shot evaluation of information retrieval models
Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych · 2021
Cited alongside, same era.
Bias runs deep: Implicit reasoning biases in persona-assigned LLMs
Shashank Gupta, Vaishnavi Shrivastava, Ameet Deshpande, Ashwin Kalyan, Peter Clark, Ashish Sabharwal, and Tushar Khot · 2023
Closest in time.
Few-shot preference learning for human-in-the-loop RL
Donald Joseph Hejna III and Dorsa Sadigh · 2023
Closest in time.
AI alignment: A comprehensive survey
Jiaming Ji, Tianyi Qiu, Boyuan Chen, Borong Zhang, Hantao Lou, Kaile Wang, Yawen Duan, Zhonghao He, Jiayi Zhou, Zhaowei Zhang, et al · 2023
Closest in time.
Raja Marjieh, Ilia Sucholutsky, Pol van Rijn, Nori Jacoby, and Thomas L Griffiths · 2023
Closest in time.
What language reveals about perception: Distilling psychophysical knowledge from large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kexin Wang, Nils Reimers, and Iryna Gurevych · 2021
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Cited alongside, same era.
The challenge of value alignment
Iason Gabriel and Vafa Ghazavi · 2022
Cited alongside, same era.
Bias mitigation for machine learning classifiers: A comprehensive survey
Max Hort, Zhenpeng Chen, Jie M Zhang, Federica Sarro, and Mark Harman · 2022
Cited alongside, same era.
Predicting human similarity judgments using large language models
Raja Marjieh, Ilia Sucholutsky, Ted Sumers, Nori Jacoby, and Tom Griffiths · 2022
Cited alongside, same era.
Words are all you need? language as an approximation for human similarity judgments
Raja Marjieh, Pol Van Rijn, Ilia Sucholutsky, Theodore Sumers, Harin Lee, Thomas L Griffiths, and Nori Jacoby · 2022
Cited alongside, same era.
Language models and brain alignment: beyond word-level semantics and prediction
Gabriele Merlin and Mariya Toneva · 2022
Cited alongside, same era.
Human alignment of neural network representations
Lukas Muttenthaler, Jonas Dippel, Lorenz Linhardt, Robert A Vandermeulen, and Simon Kornblith · 2022
Cited alongside, same era.
Raja Marjieh, Ilia Sucholutsky, Pol van Rijn, Nori Jacoby, and Tom Griffiths · 2023
Closest in time.
Words are all you need? language as an approximation for human similarity judgments, 2023
Raja Marjieh, Pol van Rijn, Ilia Sucholutsky, Theodore R. Sumers, Harin Lee, Thomas L. Griffiths, and Nori Jacoby · 2023
Closest in time.
Human alignment of neural network representations, 2023
Lukas Muttenthaler, Jonas Dippel, Lorenz Linhardt, Robert A. Vandermeulen, and Simon Kornblith · 2023
Closest in time.
Improving neural network representations using human similarity judgments, 2023
Lukas Muttenthaler, Lorenz Linhardt, Jonas Dippel, Robert A. Vandermeulen, Katherine Hermann, Andrew K. Lampinen, and Simon Kornblith · 2023
Closest in time.
Kerem Oktar, Ilia Sucholutsky, Tania Lombrozo, and Thomas L Griffiths · 2023
Closest in time.
Concept alignment as a prerequisite for value alignment
Sunayana Rane, Mark Ho, Ilia Sucholutsky, and Thomas L Griffiths · 2023
Closest in time.
Alignment with human representations supports robust few-shot learning
Ilia Sucholutsky and Thomas L. Griffiths · 2023
Closest in time.
Getting aligned on representational alignment
Ilia Sucholutsky, Lukas Muttenthaler, Adrian Weller, Andi Peng, Andreea Bobu, Been Kim, Bradley C Love, Erin Grant, Jascha Achterberg, Joshua B Tenenbaum, et al · 2023
Closest in time.
Implicit bias in large language models: Experimental proof and implications for education
Melissa Warr, Nicole Jakubczyk Oster, and Roger Isaac · 2023
Closest in time.
A survey of imitation learning: Algorithms, recent developments, and challenges
Maryam Zare, Parham Kebria, Abbas Khoshravi, and Saeid Nahavandi · 2023
Closest in time.
Representation engineering: A top-down approach to AI transparency
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J. Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, J. Zico Kolter, and Dan Hendrycks · 2023
Closest in time.
Preventing harm from non-conscious bias in medical generative AI
Janna Hastings · 2024
Closest in time.
Studying the effect of globalization on color perception using multilingual online recruitment and large language models, 2024
Jakob Niedermann, Ilia Sucholutsky, Raja Marjieh, Elif Celen, Thomas L Griffiths, Nori Jacoby, and Pol van Rijn · 2024
Closest in time.
Sunayana Rane, Polyphony J Bruna, Ilia Sucholutsky, Christopher Kello, and Thomas L Griffiths · 2024
Closest in time.