Fetching the paper…
Reading the bibliography…
In this work, from a theoretical lens, we aim to understand why large language model (LLM) empowered agents are able to solve decision-making problems in the physical world.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020) · 1901
Earlier work this paper cites.
Grammar variational autoencoder
Kusner, M. J., Paige, B., and Hernández-Lobato, J. M. (2017) · 1954
Earlier work this paper cites.
Asymptotic evaluation of certain markov process expectations for large time—iii
Donsker, M. D. and Varadhan, S. S. (1976) · 1976
Earlier work this paper cites.
Learning to act using real-time dynamic programming
Barto, A. G., Bradtke, S. J., and Singh, S. P. (1995) · 1995
Earlier work this paper cites.
Dynamic noncooperative game theory
Başar, T. and Olsder, G. J. (1998) · 1998
Earlier work this paper cites.
Bayesian model averaging: a tutorial
Hoeting, J. A., Madigan, D., Raftery, A. E., and Volinsky, C. T. (1999) · 1999
Earlier work this paper cites.
Error bounds for approximate value iteration
Munos, R. (2005) · 1999
Earlier work this paper cites.
Empirical Processes in M-estimation
Geer, S. A. (2000) · 2000
Earlier work this paper cites.
Planning as heuristic search
Bonet, B. and Geffner, H. (2001) · 2001
Earlier work this paper cites.
Recent advances in hierarchical reinforcement learning
Barto, A. G. and Mahadevan, S. (2003) · 2003
Earlier work this paper cites.
Latent dirichlet allocation
Blei, D. M., Ng, A. Y., and Jordan, M. I. (2003) · 2003
Earlier work this paper cites.
Automated Planning: theory and practice
Ghallab, M., Nau, D., and Traverso, P. (2004) · 2004
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Gutmann, M. and Hyvärinen, A. (2010) · 2010
Earlier work this paper cites.
Efficient reductions for imitation learning
Ross, S. and Bagnell, D. (2010) · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G., and Bagnell, D. (2011) · 2011
Earlier work this paper cites.
Value-difference based exploration: adaptive control between epsilon-greedy and softmax
Tokic, M. and Palm, G. (2011) · 2011
Earlier work this paper cites.
A survey of monte carlo tree search methods
Browne, C. B., Powley, E., Whitehouse, D., Lucas, S. M., Cowling, P. I., Rohlfshagen, P., Tavener, S., Perez, D., Samothrakis, S., and Colton, S. (2012) · 2012
Earlier work this paper cites.
Probability in high dimension
Van Handel, R. (2014) · 2014
Earlier work this paper cites.
Concentration inequalities for markov chains by marton couplings and spectral methods
Paulin, D. (2015) · 2015
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R. (2016) · 2016
Earlier work this paper cites.
Error bounds for approximations with deep relu networks
Yarotsky, D. (2017) · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2018) · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Earlier work this paper cites.
Imitating latent policies from observation
Edwards, A., Sahni, H., Schroecker, Y., and Isbell, C. (2019) · 2019
Earlier work this paper cites.
Are sixteen heads really better than one?
Michel, P., Levy, O., and Neubig, G. (2019) · 2019
Earlier work this paper cites.
Flambe: Structural complexity and representation learning of low rank mdps
Agarwal, A., Kakade, S., Krishnamurthy, A., and Sun, W. (2020) · 2020
Cited alongside, same era.
Minimax-optimal off-policy evaluation with linear function approximation
Duan, Y., Jia, Z., and Wang, M. (2020) · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. (2020) · 2020
Cited alongside, same era.
Goal-conditioned reinforcement learning with imagined subgoals
Chane-Sane, E., Schmid, C., and Laptev, I. (2021) · 2021
Cited alongside, same era.
The statistical complexity of interactive decision making
Foster, D. J., Kakade, S. M., Qian, J., and Rakhlin, A. (2021) · 2021
Cited alongside, same era.
Making linear mdps practical via contrastive representation learning
Zhang, T., Ren, T., Yang, M., Gonzalez, J., Schuurmans, D., and Dai, B. (2022) · 2022
Later among the works it cites.
In-context learning through the bayesian prism
Ahuja, K., Panwar, M., and Goyal, N. (2023) · 2023
Later among the works it cites.
Do as i can, not as i say: Grounding language in robotic affordances
Brohan, A., Chebotar, Y., Finn, C., Hausman, K., Herzog, A., Ho, D., Ibarz, J., Irpan, A., Jang, E., Julian, R., et al. (2023) · 2023
Later among the works it cites.
Du, Y., Yang, M., Florence, P., Xia, F., Wahid, A., Ichter, B., Sermanet, P., Yu, T., Abbeel, P., Tenenbaum, J. B., et al. (2023) · 2023
Later among the works it cites.
Amago: Scalable in-context reinforcement learning for adaptive agents
Grigsby, J., Fan, L., and Zhu, Y. (2023) · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scaling up visual and vision-language representation learning with noisy text supervision
Jia, C., Yang, Y., Xia, Y., Chen, Y.-T., Parekh, Z., Pham, H., Le, Q., Sung, Y.-H., Li, Z., and Duerig, T. (2021) · 2021
Cited alongside, same era.
Mobile: Model-based imitation learning from observation alone
Kidambi, R., Chang, J., and Sun, W. (2021) · 2021
Cited alongside, same era.
What makes good in-context examples for gpt- 3 3 ?
Liu, J., Shen, D., Zhang, Y., Dolan, B., Carin, L., and Chen, W. (2021) · 2021
Cited alongside, same era.
Transformers can do bayesian inference
Müller, S., Hollmann, N., Arango, S. P., Grabocka, J., and Hutter, F. (2021) · 2021
Cited alongside, same era.
Hierarchical reinforcement learning: A comprehensive survey
Pateria, S., Subagdja, B., Tan, A.-h., and Quek, C. (2021) · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. (2021) · 2021
Cited alongside, same era.
An explanation of in-context learning as implicit bayesian inference
Xie, S. M., Raghunathan, A., Liang, P., and Ma, T. (2021) · 2021
Cited alongside, same era.
Later among the works it cites.
A theory of emergent in-context learning as implicit structure induction
Hahn, M. and Goyal, N. (2023) · 2023
Later among the works it cites.
Reasoning with language model is planning with world model
Hao, S., Gu, Y., Ma, H., Hong, J. J., Wang, Z., Wang, D. Z., and Hu, Z. (2023) · 2023
Later among the works it cites.
Tree-planner: Efficient close-loop task planning with large language models
Hu, M., Mu, Y., Yu, X., Ding, M., Wu, S., Shao, W., Chen, Q., Wang, B., Qiao, Y., and Luo, P. (2023) · 2023
Later among the works it cites.
Language models, agent models, and world models: The law for machine reasoning and planning
Hu, Z. and Shu, T. (2023) · 2023
Later among the works it cites.
A latent space theory for emergent abilities in large language models
Jiang, H. (2023) · 2023
Later among the works it cites.
Supervised pretraining can learn in-context reinforcement learning
Lee, J. N., Xie, A., Pacchiano, A., Chandak, Y., Finn, C., Nachum, O., and Brunskill, E. (2023) · 2023
Later among the works it cites.
Liu, Z., Hu, H., Zhang, S., Guo, H., Ke, S., Liu, B., and Wang, Z. (2023) · 2023
Later among the works it cites.
Roco: Dialectic multi-robot collaboration with large language models
Mandi, Z., Jain, S., and Song, S. (2023) · 2023
Later among the works it cites.
Nottingham, K., Ammanabrolu, P., Suhr, A., Choi, Y., Hajishirzi, H., Singh, S., and Fox, R. (2023) · 2023
Later among the works it cites.
Gpt-4 technical report
OpenAI, R. (2023) · 2023
Later among the works it cites.
Self-supervised learning for videos: A survey
Schiappa, M. C., Rawat, Y. S., and Shah, M. (2023) · 2023
Later among the works it cites.
Progprompt: Generating situated robot task plans using large language models
Singh, I., Blukis, V., Mousavian, A., Goyal, A., Xu, D., Tremblay, J., Fox, D., Thomason, J., and Garg, A. (2023) · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Team, G., Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al. (2023) · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. (2023) · 2023
Later among the works it cites.
The learnability of in-context learning
Wies, N., Levine, Y., and Shashua, A. (2023) · 2023
Later among the works it cites.
Multimodal learning with transformers: A survey
Xu, P., Zhu, X., and Clifton, D. A. (2023) · 2023
Later among the works it cites.
Mathematical analysis of machine learning algorithms
Zhang, T. (2023) · 2023
Later among the works it cites.
Zhang, Y., Zhang, F., Yang, Z., and Wang, Z. (2023) · 2023
Later among the works it cites.
Drive like a human: Rethinking autonomous driving with large language models
Fu, D., Li, X., Wen, L., Dou, M., Cai, P., Shi, B., and Qiao, Y. (2024) · 2024
Closest in time.