Fetching the paper…
Reading the bibliography…
Recent advancements in large language models (LLMs) have expanded their capabilities beyond traditional text-based tasks to multimodal domains, integrating visual, auditory, and textual data.
Goecks, V. G.; Gremillion, G. M.; Lawhern, V. J.; Valasek, J.; and Waytowich, N. R. 2019 · 1910
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y.; Bottou, L.; Bengio, Y.; and Haffner, P. 1998 · 1998
Earlier work this paper cites.
Agent57: Outperforming the Atari Human Benchmark
Badia, A. P.; Piot, B.; Kapturowski, S.; Sprechmann, P.; Vitvitskyi, A.; Guo, D.; and Blundell, C. 2020 · 2003
Earlier work this paper cites.
Recurrent neural network based language model
Mikolov, T.; Karafiát, M.; Burget, L.; Cernockỳ, J.; and Khudanpur, S. 2010 · 2010
Earlier work this paper cites.
Keep calm and explore: Language models for action generation in text-based games
Yao, S.; Rao, R.; Hausknecht, M.; and Narasimhan, K. 2020 · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A.; Sutskever, I.; and Hinton, G. E. 2012 · 2012
Earlier work this paper cites.
The Arcade Learning Environment: An Evaluation Platform for General Agents
Bellemare, M. G.; Naddaf, Y.; Veness, J.; and Bowling, M. 2013 · 2013
Earlier work this paper cites.
Playing Atari with Deep Reinforcement Learning
Mnih, V.; Kavukcuoglu, K.; Silver, D.; Graves, A.; Antonoglou, I.; Wierstra, D.; and Riedmiller, M. 2013 · 2013
Earlier work this paper cites.
Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN)
Mao, J.; Xu, W.; Yang, Y.; Wang, J.; Huang, Z.; and Yuille, A. 2015 · 2015
Earlier work this paper cites.
Imitation learning: A survey of learning methods
Hussein, A.; Gaber, M. M.; Elyan, E.; and Jayne, C. 2017 · 2017
Earlier work this paper cites.
Sensor data acquisition and multimodal sensor fusion for human activity recognition using deep learning
Chung, S.; Lim, J.; Noh, K. J.; Kim, G.; and Jeong, H. 2019 · 2019
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Earlier work this paper cites.
Human-level play in the game of Diplomacy by combining language models with strategic reasoning
(FAIR)†, M. F. A. R. D. T.; Bakhtin, A.; Brown, N.; Dinan, E.; Farina, G.; Flaherty, C.; Fried, D.; Goff, A.; Gray, J.; Hu, H.; et al. 2022 · 2022
Earlier work this paper cites.
Atari Agents
Gogianu, F.; Berariu, T.; Bușoniu, L.; and Burceanu, E. 2022 · 2022
Cited alongside, same era.
Emergent world representations: Exploring a sequence model trained on a synthetic task
Li, K.; Hopkins, A. K.; Bau, D.; Viégas, F.; Pfister, H.; and Wattenberg, M. 2022 · 2022
Cited alongside, same era.
Reed, S.; Zolna, K.; Parisotto, E.; Colmenarejo, S. G.; Novikov, A.; Barth-Maron, G.; Gimenez, M.; Sulsky, Y.; Kay, J.; Springenberg, J. T.; et al. 2022 · 2022
Cited alongside, same era.
Playing repeated games with large language models
Akata, E.; Schulz, L.; Coda-Forno, J.; Oh, S. J.; Bethge, M.; and Schulz, E. 2023 · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Chiang, W.-L.; Li, Z.; Lin, Z.; Sheng, Y.; Wu, Z.; Zhang, H.; Zheng, L.; Zhuang, S.; Zhuang, Y.; Gonzalez, J. E.; et al. 2023 · 2023
Ferret: Refer and Ground Anything Anywhere at Any Granularity
You, H.; Zhang, H.; Gan, Z.; Du, X.; Zhang, B.; Wang, Z.; Cao, L.; Chang, S.-F.; and Yang, Y. 2023 · 2023
Later among the works it cites.
Scaling instructable agents across many simulated worlds
Abi Raad, M.; Ahuja, A.; Barros, C.; Besse, F.; Bolt, A.; Bolton, A.; Brownfield, B.; Buttimore, G.; Cant, M.; Chakera, S.; et al. 2024 · 2024
Closest in time.
Introducing the next generation of Claude
Anthropic. 2024 · 2024
Closest in time.
Gemini Flash
DeepMind, G. 2024 · 2024
Closest in time.
Large Language Models and Games: A Survey and Roadmap
Gallotta, R.; Todd, G.; Zammit, M.; Earle, S.; Liapis, A.; Togelius, J.; and Yannakakis, G. N. 2024 · 2024
Closest in time.
Introducing ChatGPT
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Interactive Task Planning with Language Models
Li, B.; Wu, P.; Abbeel, P.; and Malik, J. 2023 · 2023
Cited alongside, same era.
Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2023 · 2023
Cited alongside, same era.
Eureka: Human-Level Reward Design via Coding Large Language Models
Ma, Y. J.; Liang, W.; Wang, G.; Huang, D.-A.; Bastani, O.; Jayaraman, D.; Zhu, Y.; Fan, L.; and Anandkumar, A. 2023 · 2023
Cited alongside, same era.
SayPlan: Grounding Large Language Models using 3D Scene Graphs for Scalable Robot Task Planning
Rana, K.; Haviland, J.; Garg, S.; Abou-Chakra, J.; Reid, I.; and Suenderhauf, N. 2023 · 2023
Cited alongside, same era.
Can large language models play text games well? current state-of-the-art and open questions
Tsai, C. F.; Zhou, X.; Liu, S. S.; Li, J.; Yu, M.; and Mei, H. 2023 · 2023
Cited alongside, same era.
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2023 · 2023
Cited alongside, same era.
Voyager: An open-ended embodied agent with large language models
Wang, G.; Xie, Y.; Jiang, Y.; Mandlekar, A.; Xiao, C.; Zhu, Y.; Fan, L.; and Anandkumar, A. 2023 · 2023
Cited alongside, same era.
OpenAI. 2022 · 2024
Closest in time.
GPT-4o mini: advancing cost-efficient intelligence
OpenAI. 2024a · 2024
Closest in time.
Hello GPT-4o
OpenAI. 2024b · 2024
Closest in time.
OpenAI; Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2024 · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reid, M.; Savinov, N.; Teplyashin, D.; Lepikhin, D.; Lillicrap, T.; baptiste Alayrac, J.; Soricut, R.; Lazaridou, A.; Firat, O.; Schrittwieser, J.; Antonoglou, I.; et al. 2024 · 2024
Closest in time.
Gemini: A Family of Highly Capable Multimodal Models
Team, G.; Anil, R.; Borgeaud, S.; Alayrac, J.-B.; Yu, J.; Soricut, R.; Schalkwyk, J.; Dai, A. M.; Hauth, A.; Millican, K.; and Silver, D. 2024 · 2024
Closest in time.
Read and reap the rewards: Learning to play atari with the help of instruction manuals
Wu, Y.; Fan, Y.; Liang, P. P.; Azaria, A.; Li, Y.; and Mitchell, T. M. 2024 · 2024
Closest in time.
A Survey on Robotics with Foundation Models: toward Embodied AI
Xu, Z.; Wu, K.; Wen, J.; Li, J.; Liu, N.; Che, Z.; and Tang, J. 2024 · 2024
Closest in time.