Fetching the paper…
Reading the bibliography…
In this paper, we present a benchmark to pressure-test today's frontier models' multimodal decision-making capabilities in the very long-context regime (up to one million tokens) and investigate whether these models can learn from large numbers of expert demonstrations in their context.
The concept of mind
Ryle, G · 1949
Earlier work this paper cites.
Encyclopaedia of Chess Openings
Matanović, A · 1978
Earlier work this paper cites.
Stockfish, 2008
Romstad, T., Costalba, M., Kiiski, J., Linscott, G., Nasu, Y., Isozaki, M., Noda, H., and et al · 2008
Earlier work this paper cites.
genxword, 2011
Whitlock, D · 2011
Earlier work this paper cites.
python-chess, 2012
Fiekas, N · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. A · 2013
Earlier work this paper cites.
One-shot imitation learning
Duan, Y., Andrychowicz, M., Stadie, B. C., Ho, J., Schneider, J., Sutskever, I., Abbeel, P., and Zaremba, W · 2017
Earlier work this paper cites.
Distributed distributional deterministic policy gradients
Barth-Maron, G., Hoffman, M. W., Budden, D., Dabney, W., Horgan, D., TB, D., Muldal, A., Heess, N., and Lillicrap, T. P · 2018
Earlier work this paper cites.
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., de Las Casas, D., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., Lillicrap, T. P., and Riedmiller, M. A · 2018
Earlier work this paper cites.
Babyai: A platform to study the sample efficiency of grounded language learning
Chevalier-Boisvert, M., Bahdanau, D., Lahlou, S., Willems, L., Saharia, C., Nguyen, T. H., and Bengio, Y · 2019
Earlier work this paper cites.
Meta-learning of sequential strategies
Ortega, P. A., Wang, J. X., Rowland, M., Genewein, T., Kurth-Nelson, Z., Pascanu, R., Heess, N., Veness, J., Pritzel, A., Sprechmann, P., Jayakumar, S. M., McGrath, T., Miller, K. J., Azar, M. G., Osband, I., Rabinowitz, N. C., György, A., Chiappa, S., Osindero, S., Teh, Y. W., van Hasselt, H., de Freitas, N., Botvinick, M. M., and Legg, S · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
The nethack learning environment
Küttler, H., Nardelli, N., Miller, A. H., Raileanu, R., Selvatici, M., Grefenstette, E., and Rocktäschel, T · 2020
Earlier work this paper cites.
Meta-trained agents implement bayes-optimal agents
Mikulik, V., Delétang, G., McGrath, T., Genewein, T., Martic, M., Legg, S., and Ortega, P. A · 2020
Earlier work this paper cites.
In-context learning enables robot action prediction in llms
Yin, Y., Wang, Z., Sharma, Y., Niu, D., Darrell, T., and Herzig, R · 2020
Earlier work this paper cites.
Decision transformer: Reinforcement learning via sequence modeling
Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I · 2021
Earlier work this paper cites.
Muesli: Combining improvements in policy optimization
Hessel, M., Danihelka, I., Viola, F., Guez, A., Schmitt, S., Sifre, L., Weber, T., Silver, D., and van Hasselt, H · 2021
Earlier work this paper cites.
Shaking the foundations: delusions in sequence models for interaction and control
Ortega, P. A., Kunesch, M., Delétang, G., Genewein, T., Grau-Moya, J., Veness, J., Buchli, J., Degrave, J., Piot, B., Pérolat, J., Everitt, T., Tallec, C., Parisotto, E., Erez, T., Chen, Y., Reed, S. E., Hutter, M., de Freitas, N., and Legg, S · 2021
Earlier work this paper cites.
Recurrent memory transformer
Bulatov, A., Kuratov, Y., and Burtsev, M · 2022
Earlier work this paper cites.
Benchmarking the spectrum of agent capabilities
Hafner, D · 2022
Earlier work this paper cites.
Language models as zero-shot planners: Extracting actionable knowledge for embodied agents
Huang, W., Abbeel, P., Pathak, D., and Mordatch, I · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y · 2022
Earlier work this paper cites.
Pre-trained language models for interactive decision-making
Li, S., Puig, X., Paxton, C., Du, Y., Wang, C., Fan, L., Chen, T., Huang, D., Akyürek, E., Anandkumar, A., Andreas, J., Mordatch, I., Torralba, A., and Zhu, Y · 2022
Earlier work this paper cites.
A generalist agent
Reed, S. E., Zolna, K., Parisotto, E., Colmenarejo, S. G., Novikov, A., Barth-Maron, G., Gimenez, M., Sulsky, Y., Kay, J., Springenberg, J. T., Eccles, T., Bruce, J., Razavi, A., Edwards, A., Heess, N., Chen, Y., Hadsell, R., Vinyals, O., Bordbar, M., and de Freitas, N · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E. H., Le, Q. V., and Zhou, D · 2022
Cited alongside, same era.
Prompting decision transformer for few-shot policy generalization
Xu, M., Shen, Y., Zhang, S., Lu, Y., Zhao, D., Tenenbaum, J. B., and Gan, C · 2022
Cited alongside, same era.
Atari-5: Distilling the arcade learning environment down to five games
Aitchison, M., Sweetser, P., and Hutter, M · 2023
Cited alongside, same era.
Gemini: A family of highly capable multimodal models
Anil, R., Borgeaud, S., Wu, Y., Alayrac, J., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., Millican, K., Silver, D., Petrov, S., Johnson, M., Antonoglou, I., Schrittwieser, J., Glaese, A., Chen, J., Pitler, E., Lillicrap, T. P., Lazaridou, A., Firat, O., Molloy, J., Isard, M., Barham, P. R., Hennigan, T., Lee, B., Viola, F., Reynolds, M., Xu, Y., Doherty, R., Collins, E., Meyer, C., Rutherford, E., Moreira, E., Ayoub, K., Goel, M., Tucker, G., Piqueras, E., Krikun, M., Barr, I., Savinov, N., Danihelka, I., Roelofs, B., White, A., Andreassen, A., von Glehn, T., Yagati, L., Kazemi, M., Gonzalez, L., Khalman, M., Sygnowski, J., and et al · 2023
Cited alongside, same era.
BABILong: Testing the limits of llms with long context reasoning-in-a-haystack
Kuratov, Y., Bulatov, A., Anokhin, P., Rodkin, I., Sorokin, D., Sorokin, A. Y., and Burtsev, M · 2024
Closest in time.
Hello gpt-4o, 2024a
OpenAI · 2024
Closest in time.
Models, 2024b
OpenAI · 2024
Closest in time.
Introducing openai o1-preview, 2024c
OpenAI · 2024
Closest in time.
Reasoning models, 2024d
OpenAI · 2024
Closest in time.
Vision, 2024e
OpenAI · 2024
Closest in time.
Keypoint action tokens enable in-context imitation learning in robotics
Palo, N. D. and Johns, E · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Human-timescale adaptation in an open-ended task space
Bauer, J., Baumli, K., Behbahani, F. M. P., Bhoopchand, A., Bradley-Schmieg, N., Chang, M., Clay, N., Collister, A., Dasagi, V., Gonzalez, L., Gregor, K., Hughes, E., Kashem, S., Loks-Thompson, M., Openshaw, H., Parker-Holder, J., Pathak, S., Nieves, N. P., Rakicevic, N., Rocktäschel, T., Schroecker, Y., Singh, S., Sygnowski, J., Tuyls, K., York, S., Zacherl, A., and Zhang, L. M · 2023
Cited alongside, same era.
Large language models can implement policy iteration
Brooks, E. A., Walls, L., Lewis, R. L., and Singh, S · 2023
Cited alongside, same era.
Playing chess with large language models, 2023
Carlini, N · 2023
Cited alongside, same era.
Memory-based meta-learning on non-stationary distributions
Genewein, T., Delétang, G., Ruoss, A., Wenliang, L. K., Catt, E., Dutordoir, V., Grau-Moya, J., Orseau, L., Hutter, M., and Veness, J · 2023
Cited alongside, same era.
Supervised pretraining can learn in-context reinforcement learning
Lee, J., Xie, A., Pacchiano, A., Chandak, Y., Finn, C., Nachum, O., and Brunskill, E · 2023
Cited alongside, same era.
Large language models as general pattern machines
Mirchandani, S., Xia, F., Florence, P., Ichter, B., Driess, D., Arenas, M. G., Rao, K., Sadigh, D., and Zeng, A · 2023
Cited alongside, same era.
OpenAI · 2023
Cited alongside, same era.
React: Synergizing reasoning and acting in language models
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K. R., and Cao, Y · 2023
Cited alongside, same era.
Generalization to new sequential decision making tasks with in-context learning
Raparthy, S. C., Hambro, E., Kirk, R., Henaff, M., and Raileanu, R · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reid, M., Savinov, N., Teplyashin, D., Lepikhin, D., Lillicrap, T. P., Alayrac, J., Soricut, R., Lazaridou, A., Firat, O., Schrittwieser, J., Antonoglou, I., Anil, R., Borgeaud, S., Dai, A. M., Millican, K., Dyer, E., Glaese, M., Sottiaux, T., Lee, B., Viola, F., Reynolds, M., Xu, Y., Molloy, J., Chen, J., Isard, M., Barham, P., Hennigan, T., McIlroy, R., Johnson, M., Schalkwyk, J., Collins, E., Rutherford, E., Moreira, E., Ayoub, K., Goel, M., Meyer, C., Thornton, G., Yang, Z., Michalewski, H., Abbas, Z., Schucher, N., Anand, A., Ives, R., Keeling, J., Lenc, K., Haykal, S., Shakeri, S., Shyam, P., Chowdhery, A., Ring, R., Spencer, S., Sezener, E., and et al · 2024
Closest in time.
Mathematical discoveries from program search with large language models
Romera-Paredes, B., Barekatain, M., Novikov, A., Balog, M., Kumar, M. P., Dupont, E., Ruiz, F. J. R., Ellenberg, J. S., Wang, P., Fawzi, O., Kohli, P., and Fawzi, A · 2024
Closest in time.
Amortized planning with large-scale transformers: A case study on chess
Ruoss, A., Delétang, G., Medapati, S., Grau-Moya, J., Wenliang, L. K., Catt, E., Reid, J., Lewis, C. A., Veness, J., and Genewein, T · 2024
Closest in time.
TaskBench: Benchmarking large language models for task automation
Shen, Y., Song, K., Tan, X., Zhang, W., Ren, K., Yuan, S., Lu, W., Li, D., and Zhuang, Y · 2024
Closest in time.
Scaling instructable agents across many simulated worlds
SIMA Team, Raad, M. A., Ahuja, A., Barros, C., Besse, F., Bolt, A., Bolton, A., Brownfield, B., Buttimore, G., Cant, M., Chakera, S., Chan, S. C. Y., Clune, J., Collister, A., Copeman, V., Cullum, A., Dasgupta, I., de Cesare, D., Trapani, J. D., Donchev, Y., Dunleavy, E., Engelcke, M., Faulkner, R., Garcia, F., Gbadamosi, C., Gong, Z., Gonzalez, L., Gupta, K., Gregor, K., Hallingstad, A. O., Harley, T., Haves, S., Hill, F., Hirst, E., Hudson, D. A., Hudson, J., Hughes-Fitt, S., Rezende, D. J., Jasarevic, M., Kampis, L., Ke, N. R., Keck, T., Kim, J., Knagg, O., Kopparapu, K., Lampinen, A. K., Legg, S., Lerchner, A., Limont, M., Liu, Y., Loks-Thompson, M., Marino, J., Cussons, K. M., Matthey, L., Mcloughlin, S., Mendolicchio, P., Merzic, H., Mitenkova, A., Moufarek, A., Oliveira, V., Oliveira, Y. G., Openshaw, H., Pan, R., Pappu, A., Platonov, A., Purkiss, O., Reichert, D. P., Reid, J., Richemond, P. H., Roberts, T., Ruscoe, G., Elias, J. S., Sandars, T., Sawyer, D. P., Scholtes, T., Simmons, G., Slater, D., Soyer, H., Strathmann, H., Stys, P., Tam, A. C., Teplyashin, D., Terzi, T., Vercelli, D., Vujatovic, B., Wainwright, M., Wang, J. X., Wang, Z., Wierstra, D., Williams, D., Wong, N., York, S., and Young, N · 2024
Closest in time.
REGENT: A retrieval-augmented generalist agent that can act in-context in new environments
Sridhar, K., Dutta, S., Jayaraman, D., and Lee, I · 2024
Closest in time.
Prompt a robot to walk with large language models
Wang, Y., Zhang, B., Chen, J., and Sreenath, K · 2024
Closest in time.
Waytowich, N. R., White, D., Sunbeam, M., and Goecks, V. G · 2024
Closest in time.
Large language models as optimizers
Yang, C., Wang, X., Lu, Y., Liu, H., Le, Q. V., Zhou, D., and Chen, X · 2024
Closest in time.
Understanding prompt tuning and in-context learning via meta-learning
Genewein, T., Li, K. W., Grau-Moya, J., Ruoss, A., Orseau, L., and Hutter, M · 2025
Closest in time.
Gemini 2.0 flash experimental, 2024a
Google DeepMind · 2025
Closest in time.
Introducing gemini 2.0: our new ai model for the agentic era, 2024b
Google DeepMind · 2025
Closest in time.
Vision language models are in-context value learners
Ma, Y. J., Hejna, J., Fu, C., Shah, D., Liang, J., Xu, Z., Kirmani, S., Xu, P., Driess, D., Xiao, T., Bastani, O., Jayaraman, D., Yu, W., Zhang, T., Sadigh, D., and Xia, F · 2025
Closest in time.
Balrog: Benchmarking agentic llm and vlm reasoning on games
Paglieri, D., Cupiał, B., Coward, S., Piterbarg, U., Wolczyk, M., Khan, A., Pignatelli, E., Łukasz Kuciński, Pinto, L., Fergus, R., Foerster, J. N., Parker-Holder, J., and Rocktäschel, T · 2025
Closest in time.
Mastering board games by external and internal planning with language models
Schultz, J., Adamek, J., Jusup, M., Lanctot, M., Kaisers, M., Perrin, S., Hennes, D., Shar, J., Lewis, C., Ruoss, A., Zahavy, T., Veličković, P., Prince, L., Singh, S., Malmi, E., and Tomašev, N · 2025
Closest in time.
Why is prompting hard? understanding prompts on binary sequence predictors
Wenliang, L. K., Ruoss, A., Grau-Moya, J., Hutter, M., and Genewein, T · 2025
Closest in time.
LiveBench: A challenging, contamination-limited LLM benchmark
White, C., Dooley, S., Roberts, M., Pal, A., Feuer, B., Jain, S., Shwartz-Ziv, R., Jain, N., Saifullah, K., Dey, S., Shubh-Agrawal, Sandha, S. S., Naidu, S. V., Hegde, C., LeCun, Y., Goldstein, T., Neiswanger, W., and Goldblum, M · 2025
Closest in time.
The rise and potential of large language model based agents: a survey
Xi, Z., Chen, W., Guo, X., He, W., Ding, Y., Hong, B., Zhang, M., Wang, J., Jin, S., Zhou, E., Zheng, R., Fan, X., Wang, X., Xiong, L., Zhou, Y., Wang, W., Jiang, C., Zou, Y., Liu, X., Yin, Z., Dou, S., Weng, R., Qin, W., Zheng, Y., Qiu, X., Huang, X., Zhang, Q., and Gui, T · 2025
Closest in time.