Fetching the paper…
Reading the bibliography…
Traditionally, reinforcement learning (RL) agents learn to solve new tasks by updating their neural network parameters through interactions with the task environment.
Learning to predict by the methods of temporal differences
Sutton, R. S. (1988) · 1988
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S. (1988) · 1988
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
Cybenko, G. (1989) · 1989
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Hornik, K., Stinchcombe, M., and White, H. (1989) · 1989
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
Cybenko, G. (1989) · 1989
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Hornik, K., Stinchcombe, M., and White, H. (1989) · 1989
Earlier work this paper cites.
Finding structure in time
Elman, J. L. (1990) · 1990
Earlier work this paper cites.
Finding structure in time
Elman, J. L. (1990) · 1990
Earlier work this paper cites.
On the computational power of neural nets
Siegelmann, H. T. and Sontag, E. D. (1992) · 1992
Earlier work this paper cites.
On the computational power of neural nets
Siegelmann, H. T. and Sontag, E. D. (1992) · 1992
Earlier work this paper cites.
Multilayer feedforward networks with a nonpolynomial activation function can approximate any function
Leshno, M., Lin, V. Y., Pinkus, A., and Schocken, S. (1993) · 1993
Earlier work this paper cites.
Multilayer feedforward networks with a nonpolynomial activation function can approximate any function
Leshno, M., Lin, V. Y., Pinkus, A., and Schocken, S. (1993) · 1993
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Baird, L. C. (1995) · 1995
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Baird, L. C. (1995) · 1995
Earlier work this paper cites.
Least-squares temporal difference learning
Boyan, J. A. (1999) · 1999
Earlier work this paper cites.
Average cost temporal-difference learning
Tsitsiklis, J. N. and Roy, B. V. (1999) · 1999
Earlier work this paper cites.
Least-squares temporal difference learning
Boyan, J. A. (1999) · 1999
Earlier work this paper cites.
Average cost temporal-difference learning
Tsitsiklis, J. N. and Roy, B. V. (1999) · 1999
Earlier work this paper cites.
The ode method for convergence of stochastic approximation and reinforcement learning
Borkar, V. S. and Meyn, S. P. (2000) · 2000
Earlier work this paper cites.
The ode method for convergence of stochastic approximation and reinforcement learning
Borkar, V. S. and Meyn, S. P. (2000) · 2000
Earlier work this paper cites.
Learning to learn using gradient descent
Hochreiter, S., Younger, A. S., and Conwell, P. R. (2001) · 2001
Earlier work this paper cites.
Learning to learn using gradient descent
Hochreiter, S., Younger, A. S., and Conwell, P. R. (2001) · 2001
Earlier work this paper cites.
Matplotlib: A 2d graphics environment
Hunter, J. D. (2007) · 2007
Earlier work this paper cites.
Matplotlib: A 2d graphics environment
Hunter, J. D. (2007) · 2007
Earlier work this paper cites.
Preconditioned temporal difference learning
Yao, H. and Liu, Z.-Q. (2008) · 2008
Earlier work this paper cites.
Preconditioned temporal difference learning
Yao, H. and Liu, Z.-Q. (2008) · 2008
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M. (2013) · 2013
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M. (2013) · 2013
Earlier work this paper cites.
Neural turing machines
Graves, A., Wayne, G., and Danihelka, I. (2014) · 2014
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L. (2014) · 2014
Earlier work this paper cites.
Neural turing machines
Graves, A., Wayne, G., and Danihelka, I. (2014) · 2014
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L. (2014) · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2015) · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M. A., Fidjeland, A., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D. (2015) · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2015) · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M. A., Fidjeland, A., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D. (2015) · 2015
Earlier work this paper cites.
Deepmind lab
Beattie, C., Leibo, J. Z., Teplyashin, D., Ward, T., Wainwright, M., Küttler, H., Lefrancq, A., Green, S., Valdés, V., Sadik, A., et al. (2016) · 2016
Earlier work this paper cites.
OpenAI Gym
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W. (2016) · 2016
Earlier work this paper cites.
Rl2: nforcement learning via slow reinforcement learning
Duan, Y., Schulman, J., Chen, X., Bartlett, P. L., Sutskever, I., and Abbeel, P. (2016) · 2016
Earlier work this paper cites.
Learning to reinforcement learn
Wang, J. X., Kurth-Nelson, Z., Tirumala, D., Soyer, H., Leibo, J. Z., Munos, R., Blundell, C., Kumaran, D., and Botvinick, M. (2016) · 2016
Earlier work this paper cites.
Deepmind lab
Beattie, C., Leibo, J. Z., Teplyashin, D., Ward, T., Wainwright, M., Küttler, H., Lefrancq, A., Green, S., Valdés, V., Sadik, A., et al. (2016) · 2016
Earlier work this paper cites.
OpenAI Gym
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W. (2016) · 2016
Earlier work this paper cites.
Rl2: nforcement learning via slow reinforcement learning
Duan, Y., Schulman, J., Chen, X., Bartlett, P. L., Sutskever, I., and Abbeel, P. (2016) · 2016
Earlier work this paper cites.
Learning to reinforcement learn
Wang, J. X., Kurth-Nelson, Z., Tirumala, D., Soyer, H., Leibo, J. Z., Munos, R., Blundell, C., Kumaran, D., and Botvinick, M. (2016) · 2016
Earlier work this paper cites.
Deep learning
Bengio, Y., Goodfellow, I., and Courville, A. (2017) · 2017
Earlier work this paper cites.
Distilling a neural network into a soft decision tree
Frosst, N. and Hinton, G. (2017) · 2017
Earlier work this paper cites.
Residual connections encourage iterative inference
Jastrzębski, S., Arpit, D., Ballas, N., Verma, V., Che, T., and Bengio, Y. (2017) · 2017
Earlier work this paper cites.
The expressive power of neural networks: A view from the width
Lu, Z., Pu, H., Wang, F., Hu, Z., and Wang, L. (2017) · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I. (2017) · 2017
Earlier work this paper cites.
Deep learning
Bengio, Y., Goodfellow, I., and Courville, A. (2017) · 2017
Earlier work this paper cites.
Distilling a neural network into a soft decision tree
Frosst, N. and Hinton, G. (2017) · 2017
Earlier work this paper cites.
Residual connections encourage iterative inference
Jastrzębski, S., Arpit, D., Ballas, N., Verma, V., Che, T., and Bengio, Y. (2017) · 2017
Earlier work this paper cites.
The expressive power of neural networks: A view from the width
Lu, Z., Pu, H., Wang, F., Hu, Z., and Wang, L. (2017) · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I. (2017) · 2017
Earlier work this paper cites.
Towards robust interpretability with self-explaining neural networks
Alvarez Melis, D. and Jaakkola, T. (2018) · 2018
Earlier work this paper cites.
A finite time analysis of temporal difference learning with linear function approximation
Bhandari, J., Russo, D., and Singal, R. (2018) · 2018
Cited alongside, same era.
Reinforcement Learning: An Introduction (2nd Edition)
Sutton, R. S. and Barto, A. G. (2018) · 2018
Cited alongside, same era.
Towards robust interpretability with self-explaining neural networks
Alvarez Melis, D. and Jaakkola, T. (2018) · 2018
Cited alongside, same era.
A finite time analysis of temporal difference learning with linear function approximation
Bhandari, J., Russo, D., and Singal, R. (2018) · 2018
Cited alongside, same era.
Reinforcement Learning: An Introduction (2nd Edition)
Sutton, R. S. and Barto, A. G. (2018) · 2018
Cited alongside, same era.
Neural temporal-difference and q-learning provably converge to global optima
Cai, Q., Yang, Z., Lee, J. D., and Wang, Z. (2019) · 2019
Do transformers parse while predicting the masked word?
Zhao, H., Panigrahi, A., Ge, R., and Arora, S. (2023) · 2023
Later among the works it cites.
Emergence of in-context reinforcement learning from noise distillation
Zisman, I., Kurenkov, V., Nikulin, A., Sinii, V., and Kolesnikov, S. (2023) · 2023
Later among the works it cites.
Linear attention is (maybe) all you need (to understand transformer optimization)
Ahn, K., Cheng, X., Song, M., Yun, C., Jadbabaie, A., and Sra, S. (2023) · 2023
Later among the works it cites.
What learning algorithm is in-context learning? investigations with linear models
Akyürek, E., Schuurmans, D., Andreas, J., Ma, T., and Zhou, D. (2023) · 2023
Later among the works it cites.
Physics of language models: Part 1, context-free grammar
Allen-Zhu, Z. and Li, Y. (2023) · 2023
Later among the works it cites.
Human-timescale adaptation in an open-ended task space
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Improving generalization in meta reinforcement learning using learned objectives
Kirsch, L., van Steenkiste, S., and Schmidhuber, J. (2019) · 2019
Cited alongside, same era.
Neural temporal-difference and q-learning provably converge to global optima
Cai, Q., Yang, Z., Lee, J. D., and Wang, Z. (2019) · 2019
Cited alongside, same era.
Improving generalization in meta reinforcement learning using learned objectives
Kirsch, L., van Steenkiste, S., and Schmidhuber, J. (2019) · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. (2020) · 2020
Cited alongside, same era.
Rethinking attention with performers
Choromanski, K., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J., Mohiuddin, A., Kaiser, L., et al. (2020) · 2020
Cited alongside, same era.
Array programming with NumPy
Harris, C. R., Millman, K. J., van der Walt, S. J., Gommers, R., Virtanen, P., Cournapeau, D., Wieser, E., Taylor, J., Berg, S., Smith, N. J., Kern, R., Picus, M., Hoyer, S., van Kerkwijk, M. H., Brett, M., Haldane, A., del Río, J. F., Wiebe, M., Peterson, P., Gérard-Marchant, P., Sheppard, K., Reddy, T., Weckesser, W., Abbasi, H., Gohlke, C., and Oliphant, T. E. (2020) · 2020
Cited alongside, same era.
Bauer, J., Baumli, K., Behbahani, F., Bhoopchand, A., Bradley-Schmieg, N., Chang, M., Clay, N., Collister, A., Dasagi, V., Gonzalez, L., et al. (2023) · 2023
Later among the works it cites.
A survey of meta-reinforcement learning
Beck, J., Vuorio, R., Liu, E. Z., Xiong, Z., Zintgraf, L., Finn, C., and Whiteson, S. (2023) · 2023
Later among the works it cites.
Amago: Scalable in-context reinforcement learning for adaptive agents
Grigsby, J., Fan, L., and Zhu, Y. (2023) · 2023
Later among the works it cites.
Towards general-purpose in-context learning agents
Kirsch, L., Harrison, J., Freeman, C., Sohl-Dickstein, J., and Schmidhuber, J. (2023) · 2023
Later among the works it cites.
Transformers as decision makers: Provable in-context reinforcement learning via supervised pretraining
Lin, L., Bai, Y., and Mei, S. (2023) · 2023
Later among the works it cites.
Structured state space models for in-context reinforcement learning
Lu, C., Schroecker, Y., Gu, A., Parisotto, E., Foerster, J., Singh, S., and Behbahani, F. (2023) · 2023
Later among the works it cites.
One step of gradient descent is provably the optimal in-context learner with one layer of linear self-attention
Mahankali, A., Hashimoto, T. B., and Ma, T. (2023) · 2023
Later among the works it cites.
Generalization to new sequential decision making tasks with in-context learning
Raparthy, S. C., Hambro, E., Kirk, R., Henaff, M., and Raileanu, R. (2023) · 2023
Later among the works it cites.
In-context reinforcement learning for variable action spaces
Sinii, V., Nikulin, A., Kurenkov, V., Zisman, I., and Kolesnikov, S. (2023) · 2023
Later among the works it cites.
How many pretraining tasks are needed for in-context learning of linear regression?
Wu, J., Zou, D., Chen, Z., Braverman, V., Gu, Q., and Bartlett, P. L. (2023) · 2023
Later among the works it cites.
White-box transformers via sparse rate reduction
Yu, Y., Buchanan, S., Pai, D., Chu, T., Wu, Z., Tong, S., Haeffele, B., and Ma, Y. (2023) · 2023
Later among the works it cites.
Do transformers parse while predicting the masked word?
Zhao, H., Panigrahi, A., Ge, R., and Arora, S. (2023) · 2023
Later among the works it cites.
Emergence of in-context reinforcement learning from noise distillation
Zisman, I., Kurenkov, V., Nikulin, A., Sinii, V., and Kolesnikov, S. (2023) · 2023
Later among the works it cites.
Transformers learn to implement preconditioned gradient descent for in-context learning
Ahn, K., Cheng, X., Daneshmand, H., and Sra, S. (2024) · 2024
Closest in time.
PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph Compilation
Ansel, J., Yang, E., He, H., Gimelshein, N., Jain, A., Voznesensky, M., Bao, B., Bell, P., Berard, D., Burovski, E., Chauhan, G., Chourdia, A., Constable, W., Desmaison, A., DeVito, Z., Ellison, E., Feng, W., Gong, J., Gschwind, M., Hirsh, B., Huang, S., Kalambarkar, K., Kirsch, L., Lazos, M., Lezcano, M., Liang, Y., Liang, J., Lu, Y., Luk, C., Maher, B., Pan, Y., Puhrsch, C., Reso, M., Saroufim, M., Siraichi, M. Y., Suk, H., Suo, M., Tillet, P., Wang, E., Wang, X., Wen, W., Zhang, S., Zhao, X., Zhou, K., Zou, R., Mathews, A., Chanan, G., Wu, P., and Chintala, S. (2024) · 2024
Closest in time.
Artificial generational intelligence: Cultural accumulation in reinforcement learning
Cook, J., Lu, C., Hughes, E., Leibo, J. Z., and Foerster, J. (2024) · 2024
Closest in time.
In-context exploration-exploitation for reinforcement learning
Dai, Z., Tomasi, F., and Ghiassian, S. (2024) · 2024
Closest in time.
Can looped transformers learn to implement multi-step gradient descent for in-context learning?
Gatmiry, K., Saunshi, N., Reddi, S. J., Jegelka, S., and Kumar, S. (2024) · 2024
Closest in time.
Amago-2: Breaking the multi-task barrier in meta-reinforcement learning with transformers
Grigsby, J., Sasek, J., Parajuli, S., Adebi, I. D., Zhang, A., and Zhu, Y. (2024) · 2024
Closest in time.
Can large language models explore in-context?
Krishnamurthy, A., Harris, K., Foster, D. J., Zhang, C., and Slivkins, A. (2024) · 2024
Closest in time.
Supervised pretraining can learn in-context reinforcement learning
Lee, J., Xie, A., Pacchiano, A., Chandak, Y., Finn, C., Nachum, O., and Brunskill, E. (2024) · 2024
Closest in time.
Do llm agents have regret? a case study in online learning and games
Park, C., Liu, X., Ozdaglar, A., and Zhang, K. (2024) · 2024
Closest in time.
Almost sure convergence rates and concentration of stochastic approximation and reinforcement learning with markovian noise
Qian, X., Xie, Z., Liu, X., and Zhang, S. (2024) · 2024
Closest in time.
How do transformers perform in-context autoregressive learning?
Sander, M. E., Giryes, R., Suzuki, T., Blondel, M., and Peyré, G. (2024) · 2024
Closest in time.
Cross-episodic curriculum for transformer agents
Shi, L. X., Jiang, Y., Grigsby, J., Fan, L., and Zhu, Y. (2024) · 2024
Closest in time.
Hierarchical prompt decision transformer: Improving few-shot policy generalization with global and adaptive
Wang, Z., Wang, H., and Qi, Y. (2024) · 2024
Closest in time.
Meta-reinforcement learning robust to distributional shift via performing lifelong in-context learning
Xu, T., Li, Z., and Ren, Q. (2024) · 2024
Closest in time.
Trained transformers learn linear models in-context
Zhang, R., Frei, S., and Bartlett, P. L. (2024) · 2024
Closest in time.
On mesa-optimization in autoregressively trained transformers: Emergence and capability
Zheng, C., Huang, W., Wang, R., Wu, G., Zhu, J., and Li, C. (2024) · 2024
Closest in time.
Transformers learn to implement preconditioned gradient descent for in-context learning
Ahn, K., Cheng, X., Daneshmand, H., and Sra, S. (2024) · 2024
Closest in time.
PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph Compilation
Ansel, J., Yang, E., He, H., Gimelshein, N., Jain, A., Voznesensky, M., Bao, B., Bell, P., Berard, D., Burovski, E., Chauhan, G., Chourdia, A., Constable, W., Desmaison, A., DeVito, Z., Ellison, E., Feng, W., Gong, J., Gschwind, M., Hirsh, B., Huang, S., Kalambarkar, K., Kirsch, L., Lazos, M., Lezcano, M., Liang, Y., Liang, J., Lu, Y., Luk, C., Maher, B., Pan, Y., Puhrsch, C., Reso, M., Saroufim, M., Siraichi, M. Y., Suk, H., Suo, M., Tillet, P., Wang, E., Wang, X., Wen, W., Zhang, S., Zhao, X., Zhou, K., Zou, R., Mathews, A., Chanan, G., Wu, P., and Chintala, S. (2024) · 2024
Closest in time.
Artificial generational intelligence: Cultural accumulation in reinforcement learning
Cook, J., Lu, C., Hughes, E., Leibo, J. Z., and Foerster, J. (2024) · 2024
Closest in time.
In-context exploration-exploitation for reinforcement learning
Dai, Z., Tomasi, F., and Ghiassian, S. (2024) · 2024
Closest in time.
Can looped transformers learn to implement multi-step gradient descent for in-context learning?
Gatmiry, K., Saunshi, N., Reddi, S. J., Jegelka, S., and Kumar, S. (2024) · 2024
Closest in time.
Amago-2: Breaking the multi-task barrier in meta-reinforcement learning with transformers
Grigsby, J., Sasek, J., Parajuli, S., Adebi, I. D., Zhang, A., and Zhu, Y. (2024) · 2024
Closest in time.
Can large language models explore in-context?
Krishnamurthy, A., Harris, K., Foster, D. J., Zhang, C., and Slivkins, A. (2024) · 2024
Closest in time.
Supervised pretraining can learn in-context reinforcement learning
Lee, J., Xie, A., Pacchiano, A., Chandak, Y., Finn, C., Nachum, O., and Brunskill, E. (2024) · 2024
Closest in time.
Do llm agents have regret? a case study in online learning and games
Park, C., Liu, X., Ozdaglar, A., and Zhang, K. (2024) · 2024
Closest in time.
Almost sure convergence rates and concentration of stochastic approximation and reinforcement learning with markovian noise
Qian, X., Xie, Z., Liu, X., and Zhang, S. (2024) · 2024
Closest in time.
How do transformers perform in-context autoregressive learning?
Sander, M. E., Giryes, R., Suzuki, T., Blondel, M., and Peyré, G. (2024) · 2024
Closest in time.
Cross-episodic curriculum for transformer agents
Shi, L. X., Jiang, Y., Grigsby, J., Fan, L., and Zhu, Y. (2024) · 2024
Closest in time.
Hierarchical prompt decision transformer: Improving few-shot policy generalization with global and adaptive
Wang, Z., Wang, H., and Qi, Y. (2024) · 2024
Closest in time.
Meta-reinforcement learning robust to distributional shift via performing lifelong in-context learning
Xu, T., Li, Z., and Ren, Q. (2024) · 2024
Closest in time.
Trained transformers learn linear models in-context
Zhang, R., Frei, S., and Bartlett, P. L. (2024) · 2024
Closest in time.
On mesa-optimization in autoregressively trained transformers: Emergence and capability
Zheng, C., Huang, W., Wang, R., Wu, G., Zhu, J., and Li, C. (2024) · 2024
Closest in time.
The ODE method for stochastic approximation and reinforcement learning with markovian noise
Liu, S., Chen, S., and Zhang, S. (2025) · 2025
Closest in time.
A survey of in-context reinforcement learning
Moeini, A., Wang, J., Beck, J., Blaser, E., Whiteson, S., Chandra, R., and Zhang, S. (2025) · 2025
Closest in time.
The ODE method for stochastic approximation and reinforcement learning with markovian noise
Liu, S., Chen, S., and Zhang, S. (2025) · 2025
Closest in time.
A survey of in-context reinforcement learning
Moeini, A., Wang, J., Beck, J., Blaser, E., Whiteson, S., Chandra, R., and Zhang, S. (2025) · 2025
Closest in time.