Neural probabilistic motor primitives for humanoid control
Original
Josh Merel, Leonard Hasenclever, Alexandre Galashov, Arun Ahuja, Vu Pham, Greg Wayne, Yee Whye Teh, and Nicolas Heess · 2018
Later among the works it cites.
Unity, 2018
Unity · 2018
Later among the works it cites.
Embedded agency
Abram Demski and Scott Garrabrant · 2019
Later among the works it cites.
Reward tampering problems and solutions in reinforcement learning
Tom Everitt and Marcus Hutter · 2019
Later among the works it cites.
An investigation of model-free planning
Original
Arthur Guez, Mehdi Mirza, Karol Gregor, Rishabh Kabra, Sébastien Racaniere, Théophane Weber, David Raposo, Adam Santoro, Laurent Orseau, Tom Eccles, et al · 2019
Later among the works it cites.
Risks from learned optimization in advanced machine learning systems
Original
Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant · 2019
Later among the works it cites.
Categorizing wireheading in partially embedded agents
Original
Arushi Majha, Sayan Sarkar, and Davide Zagami · 2019
Later among the works it cites.
Detecting spiky corruption in Markov decision processes
Jason Mancuso, Tomasz Kisielewski, David Lindner, and Alok Singh · 2019
Later among the works it cites.
Solving rubik’s cube with a robot hand
Original
OpenAI, Ilge Akkaya, Marcin Andrychowicz, Maciek Chociej, Mateusz Litwin, Bob McGrew, Arthur Petron, Alex Paino, Matthias Plappert, Glenn Powell, Raphael Ribas, et al · 2019
Later among the works it cites.
Human compatible: Artificial intelligence and the problem of control
Stuart Russell · 2019
Later among the works it cites.
Baba is you, 2019
Arvi Teikari · 2019
Later among the works it cites.
Pitfalls of learning a reward function online
Stuart Armstrong, Jan Leike, Laurent Orseau, and Shane Legg · 2020
Closest in time.
Social media, echo chambers, and political polarization
Pablo Barberá · 2020
Closest in time.
Asymptotically unambitious artificial general intelligence
Michael K Cohen, Badri N Vellambi, and Marcus Hutter · 2020
Closest in time.
Rl unplugged: Benchmarks for offline reinforcement learning
Original
Caglar Gulcehre, Ziyu Wang, Alexander Novikov, Tom Le Paine, Sergio Gómez Colmenarejo, Konrad Zolna, Rishabh Agarwal, Josh Merel, Daniel Mankowitz, Cosmin Paduraru, et al · 2020
Closest in time.
Reward-rational (implicit) choice: A unifying formalism for reward learning
Original
Hong Jun Jeon, Smitha Milli, and Anca D Dragan · 2020
Closest in time.
Specification gaming: the flip side of AI ingenuity
Victoria Krakovna, Jonathan Uesato, Vladimir Mikulik, Matthew Rahtz, Tom Everitt, Ramana Kumar, Zac Kenton, Jan Leike, and Shane Legg · 2020
Closest in time.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Original
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Closest in time.
Effects of persuasive dialogues: Testing bot identities and inquiry strategies
Weiyan Shi, Xuewei Wang, Yoo Jung Oh, Jingwen Zhang, Saurav Sahay, and Zhou Yu · 2020
Closest in time.
Aligned tampering incentives for deep reinforcement learning
Jonathan Uesato, Ramana Kumar, Victoria Krakovna, Tom Everitt, Richard Ngo, and Shane Legg · 2020
Closest in time.
Learning to interactively learn and assist
Mark Woodward, Chelsea Finn, and Karol Hausman · 2020
Closest in time.