Fetching the paper…
Reading the bibliography…
The MineRL BASALT competition has served to catalyze advances in learning from human feedback through four hard-to-specify tasks in Minecraft, such as create and photograph a waterfall.
The rating of chessplayers, past and present
A. E. Elo · 1978
Earlier work this paper cites.
Trueskill™: a bayesian skill rating system
R. Herbrich, T. Minka, and T. Graepel · 2006
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Earlier work this paper cites.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Earlier work this paper cites.
Benchmarking deep reinforcement learning for continuous control
Y. Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel · 2016
Earlier work this paper cites.
The malmo platform for artificial intelligence experimentation
M. Johnson, K. Hofmann, T. Hutton, and D. Bignell · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Risk-aware active inverse reinforcement learning
D. S. Brown, Y. Cui, and S. Niekum · 2018
Earlier work this paper cites.
Large-scale study of curiosity-driven learning
Y. Burda, H. Edwards, D. Pathak, A. Storkey, T. Darrell, and A. A. Efros · 2018
Earlier work this paper cites.
Reward learning from human preferences and demonstrations in atari
B. Ibarz, J. Leike, T. Pohlen, G. Irving, S. Legg, and D. Amodei · 2018
Earlier work this paper cites.
textblob documentation
S. Loria et al · 2018
Earlier work this paper cites.
Generative design in minecraft (gdmc) settlement generation competition
C. Salge, M. C. Green, R. Canaan, and J. Togelius · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Earlier work this paper cites.
Y. Tassa, Y. Doron, A. Muldal, T. Erez, Y. Li, D. de Las Casas, D. Budden, A. Abdolmaleki, J. Merel, A. Lefrancq, et al · 2018
Earlier work this paper cites.
Deep reinforcement learning from policy-dependent human feedback
D. Arumugam, J. K. Lee, S. Saskin, and M. L. Littman · 2019
Earlier work this paper cites.
Leveraging procedural generation to benchmark reinforcement learning, 2019
K. Cobbe, C. Hesse, J. Hilton, and J. Schulman · 2019
Earlier work this paper cites.
Craftassist: A framework for dialogue-enabled interactive agents
J. Gray, K. Srinet, Y. Jernite, H. Yu, Z. Chen, D. Guo, S. Goyal, C. L. Zitnick, and A. Szlam · 2019
Cited alongside, same era.
Neurips 2019 competition: the minerl competition on sample efficient reinforcement learning using human priors
W. H. Guss, C. Codel, K. Hofmann, B. Houghton, N. Kuno, S. Milani, S. Mohanty, D. P. Liebana, R. Salakhutdinov, N. Topin, et al · 2019
Cited alongside, same era.
Collaborative dialogue in Minecraft
A. Narayan-Chen, P. Jayannavar, and J. Hockenmaier · 2019
Cited alongside, same era.
The multi-agent reinforcement learning in Malmö (MARLÖ) competition
D. Perez-Liebana, K. Hofmann, S. P. Mohanty, N. Kuno, A. Kramer, S. Devlin, R. D. Gaina, and D. Ionita · 2019
Cited alongside, same era.
Reward-rational (implicit) choice: A unifying formalism for reward learning
H. J. Jeon, S. Milli, and A. Dragan · 2020
Video pretraining (vpt): Learning to act by watching unlabeled online videos
B. Baker, I. Akkaya, P. Zhokov, J. Huizinga, J. Tang, A. Ecoffet, B. Houghton, R. Sampedro, and J. Clune · 2022
Later among the works it cites.
Use-case-grounded simulations for explanation evaluation
V. Chen, N. Johnson, N. Topin, G. Plumb, and A. Talwalkar · 2022
Later among the works it cites.
Minedojo: Building open-ended embodied agents with internet-scale knowledge
L. Fan, G. Wang, Y. Jiang, A. Mandlekar, Y. Yang, H. Zhu, A. Tang, D. A. Huang, Y. Zhu, and A. Anandkumar · 2022
Later among the works it cites.
imitation: Clean imitation learning implementations
A. Gleave, M. Taufeeque, J. Rocamonde, E. Jenner, S. H. Wang, S. Toyer, M. Ernestus, N. Belrose, S. Emmons, and S. Russell · 2022
Later among the works it cites.
The MineRL BASALT competition on learning from human feedback
A. Kanervisto, S. Milani, K. Ramanauskas, B. V. Galbraith, S. H. Wang, B. Houghton, S. Mohanty, and R. Shah · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Action space shaping in deep reinforcement learning
A. Kanervisto, C. Scheller, and V. Hautamäki · 2020
Cited alongside, same era.
Specification gaming: the flip side of ai ingenuity
V. Krakovna, J. Uesato, V. Mikulik, M. Rahtz, T. Everitt, R. Kumar, Z. Kenton, J. Leike, and S. Legg · 2020
Cited alongside, same era.
Learning to summarize with human feedback
N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano · 2020
Cited alongside, same era.
Navigation turing test (ntt): Learning to evaluate human-like navigation
S. Devlin, R. Georgescu, I. Momennejad, J. Rzepecki, E. Zuniga, G. Costello, G. Leroy, A. Shaw, and K. Hofmann · 2021
Cited alongside, same era.
Datasheets for datasets
T. Gebru, J. Morgenstern, B. Vecchione, J. W. Vaughan, H. Wallach, H. Daumé III, and K. Crawford · 2021
Cited alongside, same era.
Datasheets for datasets
T. Gebru, J. Morgenstern, B. Vecchione, J. W. Vaughan, H. Wallach, H. Daumé III, and K. Crawford · 2021
Cited alongside, same era.
Evocraft: A new challenge for open-endedness
D. Grbic, R. B. Palm, E. Najarro, C. Glanois, and S. Risi · 2021
Cited alongside, same era.
J. Kiseleva, A. Skrynnik, A. Zholus, S. Mohanty, N. Arabzadeh, M.-A. Côté, M. Aliannejadi, M. Teruel, Z. Li, M. Burtsev, M. ter Hoeve, Z. Volovikova, A. Panov, Y. Sun, K. Srinet, A. Szlam, and A. Awadallah · 2022
Later among the works it cites.
Behavioral cloning via search in video pretraining latent space
F. Malato, F. Leopold, A. Raut, V. Hautamäki, and A. Melnik · 2022
Later among the works it cites.
Retrospective on the 2021 minerl basalt competition on learning from human feedback
R. Shah, S. H. Wang, C. Wild, S. Milani, A. Kanervisto, V. G. Goecks, N. Waytowich, D. Watkins-Valls, B. Prakash, E. Mills, et al · 2022
Later among the works it cites.
Mastering diverse domains through world models
D. Hafner, J. Pasukonis, J. Ba, and T. Lillicrap · 2023
Closest in time.
Steve-1: A generative model for text-to-behavior in minecraft
S. Lifshitz, K. Paster, H. Chan, J. Ba, and S. McIlraith · 2023
Closest in time.
S. Milani, A. Kanervisto, K. Ramanauskas, S. Schulhoff, B. Houghton, S. Mohanty, B. Galbraith, K. Chen, Y. Song, T. Zhou, et al · 2023
Closest in time.
Empirical design in reinforcement learning
A. Patterson, S. Neumann, M. White, and A. White · 2023
Closest in time.
Game-based learning-teaching artificial intelligence to play minecraft: a systematic
R. Smit and H. Smuts · 2023
Closest in time.
Voyager: An open-ended embodied agent with large language models
G. Wang, Y. Xie, Y. Jiang, A. Mandlekar, C. Xiao, Y. Zhu, L. Fan, and A. Anandkumar · 2023
Closest in time.
Human-in-the-loop behavior modeling via an integral concurrent adaptive inverse reinforcement learning
H.-N. Wu and M. Wang · 2023
Closest in time.
Ghost in the minecraft: Generally capable agents for open-world environments via large language models with text-based knowledge and memory
X. Zhu, Y. Chen, H. Tian, C. Tao, W. Su, C. Yang, G. Huang, B. Li, L. Lu, X. Wang, Y. Qiao, Z. Zhang, and J. Dai · 2023
Closest in time.