Fetching the paper…
Reading the bibliography…
Assistance games are a promising alternative to reinforcement learning from human feedback (RLHF) for training AI assistants.
CraftAssist: A Framework for Dialogue-enabled Interactive Agents, July 2019
Gray, J., Srinet, K., Jernite, Y., Yu, H., Chen, Z., Guo, D., Goyal, S., Zitnick, C. L., and Szlam, A · 1907
Earlier work this paper cites.
Why Build an Assistant in Minecraft?, July 2019
Szlam, A., Gray, J., Srinet, K., Jernite, Y., Joulin, A., Synnaeve, G., Kiela, D., Yu, H., Chen, Z., Goyal, S., Guo, D., Rothermel, D., Zitnick, C. L., and Weston, J · 1907
Earlier work this paper cites.
Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., Lillicrap, T., and Silver, D · 1911
Earlier work this paper cites.
Dota 2 with Large Scale Deep Reinforcement Learning, December 2019
OpenAI, Berner, C., Brockman, G., Chan, B., Cheung, V., Debiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., Józefowicz, R., Gray, S., Olsson, C., Pachocki, J., Petrov, M., Pinto, H. P. d. O., Raiman, J., Salimans, T., Schlatter, J., Schneider, J., Sidor, S., Sutskever, I., Tang, J., Wolski, F., and Zhang, S · 1912
Earlier work this paper cites.
A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning
Ross, S., Gordon, G., and Bagnell, D · 1938
Earlier work this paper cites.
Individual choice behavior
Luce, R. D · 1959
Earlier work this paper cites.
The Choice Axiom After Twenty Years
Luce, R. D · 1977
Earlier work this paper cites.
The Complexity of Markov Decision Processes
Papadimitriou, C. H. and Tsitsiklis, J. N · 1987
Earlier work this paper cites.
Long Short-Term Memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
On the undecidability of probabilistic planning and related stochastic optimization problems
Madani, O., Hanks, S., and Condon, A · 2003
Earlier work this paper cites.
Bandit Based Monte-Carlo Planning
Kocsis, L. and Szepesvári, C · 2006
Earlier work this paper cites.
Monte-Carlo Planning in Large POMDPs
Silver, D. and Veness, J · 2010
Earlier work this paper cites.
Ad Hoc Autonomous Agent Teams: Collaboration without Pre-Coordination
Stone, P., Kaminka, G., Kraus, S., and Rosenschein, J · 2010
Earlier work this paper cites.
Modeling interaction via the principle of maximum causal entropy
Ziebart, B. D., Bagnell, J. A., and Dey, A. K · 2010
Earlier work this paper cites.
A policy-blending formalism for shared control
Dragan, A. D. and Srinivasa, S. S · 2013
Earlier work this paper cites.
A Decision-Theoretic Model of Assistance
Fern, A., Natarajan, S., Judah, K., and Tadepalli, P · 2014
Earlier work this paper cites.
Shared Autonomy via Hindsight Optimization
Javdani, S., Srinivasa, S., and Bagnell, A · 2015
Earlier work this paper cites.
Superintelligence: Paths, Dangers, Strategies
Bostrom, N · 2016
Earlier work this paper cites.
Cooperative Inverse Reinforcement Learning
Hadfield-Menell, D., Russell, S. J., Abbeel, P., and Dragan, A · 2016
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
The Malmo platform for artificial intelligence experimentation
Johnson, M., Hofmann, K., Hutton, T., and Bignell, D · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P., Leike, J., Brown, T. B., Martic, M., Legg, S., and Amodei, D · 2017
Cited alongside, same era.
Proximal Policy Optimization Algorithms, August 2017
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., Lillicrap, T., Simonyan, K., and Hassabis, D · 2017
Cited alongside, same era.
RLlib: Abstractions for Distributed Reinforcement Learning
Liang, E., Liaw, R., Nishihara, R., Moritz, P., Fox, R., Goldberg, K., Gonzalez, J., Jordan, M. I., and Stoica, I · 2018
Cited alongside, same era.
An Efficient, Generalized Bellman Update For Cooperative Inverse Reinforcement Learning
Malik, D., Palaniappan, M., Fisac, J. F., Hadfield-Menell, D., Russell, S., and Dragan, A. D · 2018
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., Joseph, N., Kadavath, S., Kernion, J., Conerly, T., El-Showk, S., Elhage, N., Hatfield-Dodds, Z., Hernandez, D., Hume, T., Johnston, S., Kravec, S., Lovitt, L., Nanda, N., Olsson, C., Amodei, D., Brown, T., Clark, J., McCandlish, S., Olah, C., Mann, B., and Kaplan, J · 2022
Later among the works it cites.
Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos
Baker, B., Akkaya, I., Zhokhov, P., Huizinga, J., Tang, J., Ecoffet, A., Houghton, B., Sampedro, R., and Clune, J · 2022
Later among the works it cites.
MineDojo: Building Open-Ended Embodied Agents with Internet-Scale Knowledge
Fan, L., Wang, G., Jiang, Y., Mandlekar, A., Yang, Y., Zhu, H., Tang, A., Huang, D.-A., Zhu, Y., and Anandkumar, A · 2022
Later among the works it cites.
Modeling Strong and Human-Like Gameplay with KL-Regularized Search
Jacob, A. P., Wu, D. J., Farina, G., Lerer, A., Hu, H., Bakhtin, A., Andreas, J., and Brown, N · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On the Utility of Learning about Humans for Human-AI Coordination
Carroll, M., Shah, R., Ho, M. K., Griffiths, T., Seshia, S. A., Abbeel, P., and Dragan, A. D · 2019
Cited alongside, same era.
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Cited alongside, same era.
Human Compatible: Artificial Intelligence and the Problem of Control
Russell, S · 2019
Cited alongside, same era.
Combining Deep Reinforcement Learning and Search for Imperfect-Information Games
Brown, N., Bakhtin, A., Lerer, A., and Gong, Q · 2020
Cited alongside, same era.
Pragmatic-Pedagogic Value Alignment
Fisac, J. F., Gates, M. A., Hamrick, J. B., Liu, C., Hadfield-Menell, D., Palaniappan, M., Malik, D., Sastry, S. S., Griffiths, T. L., and Dragan, A. D · 2020
Cited alongside, same era.
Monte-Carlo Tree Search as Regularized Policy Optimization
Grill, J.-B., Altché, F., Tang, Y., Hubert, T., Valko, M., Antonoglou, I., and Munos, R · 2020
Cited alongside, same era.
“Other-Play” for Zero-Shot Coordination
Hu, H., Lerer, A., Peysakhovich, A., and Foerster, J · 2020
Cited alongside, same era.
Kiseleva, J., Li, Z., Aliannejadi, M., Mohanty, S., ter Hoeve, M., Burtsev, M., Skrynnik, A., Zholus, A., Panov, A., and Srinet, K · 2022
Later among the works it cites.
Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPs
Ni, T., Eysenbach, B., and Salakhutdinov, R · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P. F., Leike, J., and Lowe, R · 2022
Later among the works it cites.
Skrynnik, A., Volovikova, Z., Côté, M.-A., Voronov, A., Zholus, A., Arabzadeh, N., Mohanty, S., Teruel, M., Awadallah, A., Panov, A., Burtsev, M., and Kiseleva, J · 2022
Later among the works it cites.
Yang, M., Carroll, M., and Dragan, A · 2022
Later among the works it cites.
The Surprising Effectiveness of PPO in Cooperative, Multi-Agent Games, November 2022
Yu, C., Velu, A., Vinitsky, E., Gao, J., Wang, Y., Bayen, A., and Wu, Y · 2022
Later among the works it cites.
IGLU Gridworld: Simple and Fast Environment for Embodied Dialog Agents, May 2022
Zholus, A., Skrynnik, A., Mohanty, S., Volovikova, Z., Kiseleva, J., Szlam, A., Coté, M.-A., and Panov, A. I · 2022
Later among the works it cites.
Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and Planning
Bakhtin, A., Wu, D. J., Lerer, A., Gray, J., Jacob, A. P., Farina, G., Miller, A. H., and Brown, N · 2023
Later among the works it cites.
Milani, S., Kanervisto, A., Ramanauskas, K., Schulhoff, S., Houghton, B., Mohanty, S., Galbraith, B., Chen, K., Song, Y., Zhou, T., Yu, B., Liu, H., Guan, K., Hu, Y., Lv, T., Malato, F., Leopold, F., Raut, A., Hautamäki, V., Melnik, A., Ishida, S., Henriques, J. F., Klassert, R., Laurito, W., Novoseller, E., Goecks, V. G., Waytowich, N., Watkins, D., Miller, J., and Shah, R · 2023
Later among the works it cites.
LIMA: Less Is More for Alignment
Zhou, C., Liu, P., Xu, P., Iyer, S., Sun, J., Mao, Y., Ma, X., Efrat, A., Yu, P., Yu, L., Zhang, S., Ghosh, G., Lewis, M., Zettlemoyer, L., and Levy, O · 2023
Later among the works it cites.
Cornelisse, D. and Vinitsky, E · 2024
Later among the works it cites.
Lang, L., Foote, D., Russell, S., Dragan, A., Jenner, E., and Emmons, S · 2024
Later among the works it cites.
Improving Grounded Language Understanding in a Collaborative Environment by Interacting with Agents Through Help Feedback
Mehta, N., Teruel, M., Deng, X., Sanz, S. F., Awadallah, A., and Kiseleva, J · 2024
Later among the works it cites.
Multi-turn Reinforcement Learning with Preference Human Feedback
Shani, L., Rosenberg, A., Cassel, A., Lang, O., Calandriello, D., Zipori, A., Noga, H., Keller, O., Piot, B., Szpektor, I., Hassidim, A., Matias, Y., and Munos, R · 2024
Later among the works it cites.
Voyager: An Open-Ended Embodied Agent with Large Language Models
Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., and Anandkumar, A · 2024
Later among the works it cites.
Pragmatic Instruction Following and Goal Assistance via Cooperative Language-Guided Inverse Planning
Zhi-Xuan, T., Ying, L., Mansinghka, V., and Tenenbaum, J. B · 2024
Later among the works it cites.
On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback
Williams, M., Carroll, M., Narang, A., Weisser, C., Murphy, B., and Dragan, A. D · 2025
Closest in time.