Fetching the paper…
Reading the bibliography…
Providing Reinforcement Learning (RL) agents with human feedback can dramatically improve various aspects of learning.
Robust tests for equality of variances.(1960)
Howard Levene, II Olkin, and H Hotelling · 1960
Earlier work this paper cites.
Two varieties of long-latency positive waves evoked by unpredictable auditory stimuli in man
Hillyard SA. Squires NK, Squires KC · 1975
Earlier work this paper cites.
Prefrontal and cingulate activity during timing behavior in the monkey
Hiroaki Niki and Masataka Watanabe · 1979
Earlier work this paper cites.
Event-related potentials, lexical decision and semantic priming
Shlomo Bentin, Gregory McCarthy, and Charles C Wood · 1985
Earlier work this paper cites.
‘error’ potentials in limbic cortex (anterior cingulate area 24) of monkeys during motor learning
H Gemba, K Sasaki, and V.B. Brooks · 1986
Earlier work this paper cites.
Effects of crossmodal divided attention on late erp components. ii. error processing in choice reaction tasks
Michael Falkenstein, Joachim Hohnsbein, Joerg Hoormann, and L. Blanke · 1991
Earlier work this paper cites.
A brain potential manifestation of error-related processing [supplement]
William Gehring, Michael Coles, David Meyer, and Emanuel Donchin · 1995
Earlier work this paper cites.
Event-related brain potentials and error-related processing: An analysis of incorrect responses to go and no-go stimuli
Marten K Scheffers, Michael GH Coles, Peter Bernstein, William J Gehring, and Emanuel Donchin · 1996
Earlier work this paper cites.
Event-related brain potentials following incorrect feedback in a time-estimation task: Evidence for a “generic” neural system for error detection
Wolfgang H. R. Miltner, Christoph H. Braun, and Michael G. H. Coles · 1997
Earlier work this paper cites.
Anterior cingulate cortex, error detection, and the online monitoring of performance
Cameron S Carter, Todd S Braver, Deanna M Barch, Matthew M Botvinick, Douglas Noll, and Jonathan D Cohen · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y Ng, Daishi Harada, and Stuart Russell · 1999
Earlier work this paper cites.
Erp components on reaction errors and their functional significance: a tutorial
Michael Falkenstein, Jörg Hoormann, Stefan Christ, and Joachim Hohnsbein · 2000
Earlier work this paper cites.
Erp components on reaction errors and their functional significance: A tutorial
Michael Falkenstein, Jörg Hoormann, Stefan Christ, and Joachim Hohnsbein · 2000
Earlier work this paper cites.
Eeg-based communication: presence of an error potential
Gerwin Schalk, Jonathan R Wolpaw, Dennis J McFarland, and Gert Pfurtscheller · 2000
Earlier work this paper cites.
A social reinforcement learning agent
LI Charles and RS Christian · 2001
Earlier work this paper cites.
Error-related brain potentials are differentially related to awareness of response errors: Evidence from an antisaccade task
Sander Nieuwenhuis, K Richard Ridderinkhof, Jos Blom, Guido PH Band, and Albert Kok · 2001
Earlier work this paper cites.
The neural basis of human error processing: reinforcement learning, dopamine, and the error-related negativity
Clay B. Holroyd and Michael G. H. Coles · 2002
Earlier work this paper cites.
The neural basis of human error processing: reinforcement learning, dopamine, and the error-related negativity
Clay B Holroyd and Michael GH Coles · 2002
Earlier work this paper cites.
Brain potentials elicited by prose-embedded linguistic anomalies
Lee Osterhout, Mark D. Allen, Judith Mclaughlin, and Kayo Inoue · 2002
Earlier work this paper cites.
Boosting bit rates and error detection for the classification of fast-paced motor commands based on single-trial eeg analysis
Benjamin Blankertz, Guido Dornhege, Christin Schafer, Roman Krepki, Jens Kohlmorgen, K-R Muller, Volker Kunzmann, Florian Losch, and Gabriel Curio · 2003
Earlier work this paper cites.
Errors in reward prediction are reflected in the event-related brain potential
Clay B Holroyd, Sander Nieuwenhuis, Nick Yeung, and Jonathan D Cohen · 2003
Earlier work this paper cites.
Response error correction-a demonstration of improved human-machine performance using real-time eeg monitoring
Lucas C Parra, Clay D Spence, Adam D Gerson, and Paul Sajda · 2003
Earlier work this paper cites.
Principled methods for advising reinforcement learning agents
Eric Wiewiora, Garrison W Cottrell, and Charles Elkan · 2003
Earlier work this paper cites.
You are wrong!—automatic detection of interaction errors from brain waves
Pierre W Ferrez and José del R Millán · 2005
Cited alongside, same era.
Reinforcement learning with supervision by combining multiple learnings and expert advices
Hyeong Soo Chang · 2006
Cited alongside, same era.
Bandit based monte-carlo planning
Levente Kocsis and Csaba Szepesvári · 2006
Cited alongside, same era.
Evaluating the robustness of learning from implicit feedback
Filip Radlinski and Thorsten Joachims · 2006
Cited alongside, same era.
Reinforcement learning with human teachers: Evidence of feedback and guidance with implications for learning performance
Andrea Lockerd Thomaz, Cynthia Breazeal, et al · 2006
Cited alongside, same era.
Error-related eeg potentials generated during simulated brain–computer interaction
Pierre W Ferrez and José del R Millán · 2008
Think: Toward practical general-purpose brain-computer communication
Mohit Agarwal and Raghupathy Sivakumar · 2015
Later among the works it cites.
Reinforcement learning from demonstration through shaping
Tim Brys, Anna Harutyunyan, Halit Bener Suay, Sonia Chernova, Matthew E Taylor, and Ann Nowé · 2015
Later among the works it cites.
Active reward learning with a novel acquisition function
Christian Daniel, Oliver Kroemer, Malte Viering, Jan Metz, and Jan Peters · 2015
Later among the works it cites.
Score-based inverse reinforcement learning
Layla El Asri, Bilal Piot, Matthieu Geist, Romain Laroche, and Olivier Pietquin · 2016
Later among the works it cites.
Learning language games through interaction
Sida I Wang, Percy Liang, and Christopher D Manning · 2016
Later among the works it cites.
Model-free preference-based reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Influence of cognitive control and mismatch on the n2 component of the erp: A review
Jonathan R Folstein and Cyma Van Petten · 2008
Cited alongside, same era.
Interactively shaping agents via human reinforcement: The tamer framework
W Bradley Knox and Peter Stone · 2009
Cited alongside, same era.
xdawn algorithm to enhance evoked potentials: application to brain–computer interface
Bertrand Rivet, Antoine Souloumiac, Virginie Attina, and Guillaume Gibert · 2009
Cited alongside, same era.
Learning from eeg error-related potentials in noninvasive brain-computer interfaces
Ricardo Chavarriaga and José del R Millán · 2010
Cited alongside, same era.
Robot reinforcement learning using eeg-based reward signals
Iñaki Iturrate, Luis Montesano, and Javier Minguez · 2010
Cited alongside, same era.
Openvibe: An open-source software platform to design, test, and use brain–computer interfaces in real and virtual environments
Yann Renard, Fabien Lotte, Guillaume Gibert, Marco Congedo, Emmanuel Maby, Vincent Delannoy, Olivier Bertrand, and Anatole Lécuyer · 2010
Cited alongside, same era.
Christian Wirth, Johannes Fürnkranz, and Gerhard Neumann · 2016
Later among the works it cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Later among the works it cites.
Reinforcement learning with deep energy-based policies
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine · 2017
Later among the works it cites.
Intrinsic interactive reinforcement learning–using error-related potentials for real world human-robot interaction
Su Kyoung Kim, Elsa Andrea Kirchner, Arne Stefes, and Frank Kirchner · 2017
Later among the works it cites.
Correcting robot mistakes in real time using eeg signals
Andres F Salazar-Gomez, Joseph DelPreto, Stephanie Gil, Frank H Guenther, and Daniela Rus · 2017
Later among the works it cites.
Improving reinforcement learning with confidence-based demonstrations
Zhaodong Wang and Matthew E Taylor · 2017
Later among the works it cites.
Dqn-tamer: Human-in-the-loop reinforcement learning with intractable feedback
Riku Arakawa, Sosuke Kobayashi, Yuya Unno, Yuta Tsuboi, and Shin-ichi Maeda · 2018
Later among the works it cites.
Efficient exploration through bayesian deep q-networks
Kamyar Azizzadenesheli, Emma Brunskill, and Animashree Anandkumar · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Deep tamer: Interactive agent shaping in high-dimensional state spaces
Garrett Warnell, Nicholas Waytowich, Vernon Lawhern, and Peter Stone · 2018
Later among the works it cites.
Cerebro: A wearable solution to detect and track user preferences using brainwaves
Mohit Agarwal and Raghupathy Sivakumar · 2019
Later among the works it cites.
Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations
Daniel S Brown, Wonjoon Goo, Prabhat Nagarajan, and Scott Niekum · 2019
Later among the works it cites.
Potential-based advice for stochastic policy learning
Baicen Xiao, Bhaskar Ramasubramanian, Andrew Clark, Hannaneh Hajishirzi, Linda Bushnell, and Radha Poovendran · 2019
Later among the works it cites.
Leveraging human guidance for deep reinforcement learning tasks
Ruohan Zhang, Faraz Torabi, Lin Guan, Dana H Ballard, and Peter Stone · 2019
Later among the works it cites.
Charge for a whole day: Extending battery life for bci wearables using a lightweight wake-up command
Mohit Agarwal and Raghupathy Sivakumar · 2020
Closest in time.
Human-in-the-loop rl with an eeg wearable headset: On effective use of brainwaves to accelerate learning
Mohit Agarwal, Shyam Krishnan Venkateswaran, and Raghupathy Sivakumar · 2020
Closest in time.
Blink to get in: Biometric authentication for mobile devices using eeg signals
E. Gupta, M. Agarwal, and R. Sivakumar · 2020
Closest in time.
Fresh: Interactive reward shaping in high-dimensional state spaces using human feedback
Baicen Xiao, Qifan Lu, Bhaskar Ramasubramanian, Andrew Clark, Linda Bushnell, and Radha Poovendran · 2020
Closest in time.