Fetching the paper…
Reading the bibliography…
This paper describes our research on AI agents embodied in visual, virtual or physical forms, enabling them to interact with both users and their environments.
Phenomenology of Perception
Maurice Merleau-Ponty · 1945
Earlier work this paper cites.
The problem of temporal order in stimulation and perception
James J Gibson · 1966
Earlier work this paper cites.
Facial action coding system
Paul Ekman and Wallace V Friesen · 1978
Earlier work this paper cites.
Mental models: Towards a cognitive science of language, inference, and consciousness
Philip Nicholas Johnson-Laird · 1983
Earlier work this paper cites.
Learning to generate sub-goals for action sequences
Jürgen Schmidhuber · 1991
Earlier work this paper cites.
Optimization of computer simulation models with rare events
Reuven Y. Rubinstein · 1997
Earlier work this paper cites.
Multiagent systems: a modern approach to distributed artificial intelligence
Gerhard Weiss · 1999
Earlier work this paper cites.
Machines and mindlessness: Social responses to computers
Clifford Nass and Youngme Moon · 2000
Earlier work this paper cites.
Embodied conversational agents: representation and intelligence in user interfaces
Justine Cassell · 2001
Earlier work this paper cites.
Establishing and maintaining long-term human-computer relationships
Timothy W. Bickmore and Rosalind W. Picard · 2005
Earlier work this paper cites.
The development of embodied cognition: Six lessons from babies
Linda Smith and Michael Gasser · 2005
Earlier work this paper cites.
Calibrating noise to sensitivity in private data analysis
Cynthia Dwork, Frank McSherry, Kobbi Nissim, and Adam Smith · 2006
Earlier work this paper cites.
Cognitive tutors: technology bringing learning science to the classroom
Kenneth R Koedinger and Albert T Corbett · 2006
Earlier work this paper cites.
Planning Algorithms
Steven M. LaValle · 2006
Earlier work this paper cites.
Automatic rigging and animation of 3d characters
Ilya Baran and Jovan Popović · 2007
Earlier work this paper cites.
On seeing human: a three-factor theory of anthropomorphism
Nicholas Epley, Adam Waytz, and John T Cacioppo · 2007
Earlier work this paper cites.
Multimodal interfaces
Sharon Oviatt · 2007
Earlier work this paper cites.
Designing the user interface: strategies for effective human-computer interaction
Ben Shneiderman and Catherine Plaisant · 2010
Earlier work this paper cites.
Differentially private empirical risk minimization
Kamalika Chaudhuri, Claire Monteleoni, and Anand D Sarwate · 2011
Earlier work this paper cites.
Early childhood development and later outcome , chapter Core cognition and beyond, page 33–65
R. Baillargeon and S. Carey · 2012
Earlier work this paper cites.
3D Video and its Applications
Takashi Matsuyama, Sho Nobuhara, Takaaki Takai, and Tony Tung · 2012
Earlier work this paper cites.
Representation learning: A review and new perspectives
Yoshua Bengio, Aaron Courville, and Pascal Vincent · 2013
Earlier work this paper cites.
Multimodal interaction: A review
Matthew Turk · 2014
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Earlier work this paper cites.
Action-conditional video prediction using deep networks in atari games
Junhyuk Oh, Xiaoxiao Guo, Honglak Lee, Richard Lewis, and Satinder Singh · 2015
Earlier work this paper cites.
Variational inference with normalizing flows
Danilo Rezende and Shakir Mohamed · 2015
Earlier work this paper cites.
Learning to communicate with deep multi-agent reinforcement learning
Jakob Foerster, Ioannis Alexandros Assael, Nando De Freitas, and Shimon Whiteson · 2016
Earlier work this paper cites.
Anymal - a highly mobile and dynamic quadrupedal robot
Marco Hutter, Christian Gehring, Dominic Jud, Andreas Lauber, C. Dario Bellicoso, Vassilios Tsounis, Jemin Hwangbo, Karen Bodie, Peter Fankhauser, Michael Bloesch, Remo Diethelm, Samuel Bachmann, Amir Melzer, and Mark Hoepflinger · 2016
Earlier work this paper cites.
Video pixel networks
Nal Kalchbrenner, Aaron van den Oord, Karen Simonyan, Ivo Danihelka, Oriol Vinyals, Alex Graves, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Predictive Control for Linear and Hybrid Systems
Francesco Borrelli, Alberto Bemporad, and Manfred Morari · 2017
Earlier work this paper cites.
Recurrent environment simulators
Silvia Chiappa, Sébastien Racaniere, Daan Wierstra, and Shakir Mohamed · 2017
Earlier work this paper cites.
AI2-THOR: An Interactive 3D Environment for Visual AI
Eric Kolve, Roozbeh Mottaghi, Winson Han, Eli VanderBilt, Luca Weihs, Alvaro Herrasti, Daniel Gordon, Yuke Zhu, Abhinav Gupta, and Ali Farhadi · 2017
Earlier work this paper cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi I Wu, Aviv Tamar, Jean Harb, OpenAI Pieter Abbeel, and Igor Mordatch · 2017
Earlier work this paper cites.
Dynamic walking on randomly-varying discrete terrain with one-step preview
Quan Nguyen, Ayush Agrawal, Xingye Da, William Martin, Hartmut Geyer, Jessy Grizzle, and Koushil Sreenath · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell · 2017
Earlier work this paper cites.
Learning multiple visual domains with residual adapters
Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Nora the empathetic psychologist
Genta Indra Winata, Onno Kampman, Yang Yang, Anik Dey, and Pascale Fung · 2017
Earlier work this paper cites.
Stochastic variational video prediction
Mohammad Babaeizadeh, Chelsea Finn, Dumitru Erhan, Roy H. Campbell, and Sergey Levine · 2018
Earlier work this paper cites.
Emergent communication through negotiation
Kris Cao, Angeliki Lazaridou, Marc Lanctot, Joel Z Leibo, Karl Tuyls, and Stephen Clark · 2018
Earlier work this paper cites.
Stochastic video generation with a learned prior
Remi Denton and Rob Fergus · 2018
Earlier work this paper cites.
Cognitive science in the era of artificial intelligence: A roadmap for reverse-engineering the infant language-learner
Emmanuel Dupoux · 2018
Earlier work this paper cites.
Learning actionable representations from visual observations
Debidatta Dwibedi, Jonathan Tompson, Corey Lynch, and Pierre Sermanet · 2018
Earlier work this paper cites.
Towards empathetic human-robot interactions
Pascale Fung, Dario Bertero, Yan Wan, Anik Dey, Ricky Ho Yin Chan, Farhad Bin Siddique, Yang Yang, Chien-Sheng Wu, and Ruixi Lin · 2018
Earlier work this paper cites.
Deep appearance models for face rendering
Stephen Lombardi, Jason Saragih, Tomas Simon, and Yaser Sheikh · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Earlier work this paper cites.
Deepmind control suite, 2018
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, Timothy Lillicrap, and Martin Riedmiller · 2018
Earlier work this paper cites.
Gibson Env: real-world perception for embodied agents
Fei Xia, Amir R. Zamir, Zhi-Yang He, Alexander Sax, Jitendra Malik, and Silvio Savarese · 2018
Cited alongside, same era.
Openpose: Realtime multi-person 2d pose estimation using part affinity fields
Zhe Cao, Gines Hidalgo, Tomas Simon, Shih-En Wei, and Yaser Sheikh · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Revisiting the evaluation of theory of mind through question answering
Matthew Le, Y-Lan Boureau, and Maximilian Nickel · 2019
Cited alongside, same era.
Exploring perceived emotional intelligence of personality-driven virtual agents in handling user challenges
Xiaojuan Ma, Emily Yang, and Pascale Fung · 2019
Cited alongside, same era.
Expressive body capture: 3D hands, face, and body from a single image
Homerobot: Open-vocabulary mobile manipulation
Sriram Yenamandra, Arun Ramachandran, Karmesh Yadav, Austin Wang, Mukul Khanna, Theophile Gervet, Tsung-Yen Yang, Vidhi Jain, Alexander William Clegg, John Turner, et al · 2023
Later among the works it cites.
Magvit: Masked generative video transformer
Lijun Yu, Yong Cheng, Kihyuk Sohn, José Lezama, Han Zhang, Huiwen Chang, Alexander G. Hauptmann, Ming-Hsuan Yang, Yuan Hao, Irfan Essa, and Lu Jiang · 2023
Later among the works it cites.
Lima: Less is more for alignment
Chunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, et al · 2023
Later among the works it cites.
Neuro-motor controlled wearable augmentations: current research and emerging trends
Haneen Alsuradi, Joseph Hong, Helin Mazi, and Mohamad Eid · 2024
Later among the works it cites.
Genie: Generative interactive environments
Jake Bruce, Michael D Dennis, Ashley Edwards, Jack Parker-Holder, Yuge Shi, Edward Hughes, Matthew Lai, Aditi Mavalankar, Richie Steigerwald, Chris Apps, et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A. A. Osman, Dimitrios Tzionas, and Michael J. Black · 2019
Cited alongside, same era.
Cooperative heterogeneous multi-robot systems: A survey
Yara Rizk, Mariette Awad, and Edward W Tunstel · 2019
Cited alongside, same era.
Exploring virtual agents for augmented reality
Isaac Wang, Jesse Smith, and Jaime Ruiz · 2019
Cited alongside, same era.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning, 2019
Tianhe Yu, Deirdre Quillen, Zhanpeng He, Ryan Julian, Avnish Narayan, Hayden Shively, Adithya Bellathur, Karol Hausman, Chelsea Finn, and Sergey Levine · 2019
Cited alongside, same era.
Inverting gradients-how easy is it to break privacy in federated learning?
Jonas Geiping, Hartmut Bauermeister, Hannah Dröge, and Michael Moeller · 2020
Cited alongside, same era.
Human instruction-following with deep reinforcement learning via transfer-learning from text
Felix Hill, Sona Mokra, Nathaniel Wong, and Tim Harley · 2020
Cited alongside, same era.
Arch: Animatable reconstruction of clothed humans
Zeng Huang, Yuanlu Xu, Christoph Lassner, Hao Li, and Tony Tung · 2020
Cited alongside, same era.
High-dimension human value representation in large language models
Samuel Cahyawijaya, Delong Chen, Yejin Bang, Leila Khalatbari, Bryan Wilie, Ziwei Ji, Etsuko Ishii, and Pascale Fung · 2024
Later among the works it cites.
Ai robots and humanoid ai: Review, perspectives and directions
Longbing Cao · 2024
Later among the works it cites.
Next-generation intelligent assistants for wearable devices
Xin Luna Dong · 2024
Later among the works it cites.
Moshi: a speech-text foundation model for real-time dialogue, 2024
Alexandre Défossez, Laurent Mazaré, Manu Orsini, Amélie Royer, Patrick Pérez, Hervé Jégou, Edouard Grave, and Neil Zeghidour · 2024
Later among the works it cites.
Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives
Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, et al · 2024
Later among the works it cites.
Mastering diverse domains through world models, 2024
Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap · 2024
Later among the works it cites.
Sparsh: Self-supervised touch representations for vision-based tactile sensing
Carolina Higuera, Akash Sharma, Chaithanya Krishna Bodduluri, Taosha Fan, Patrick Lancaster, Mrinal Kalakrishnan, Michael Kaess, Byron Boots, Mike Lambeta, Tingfan Wu, and Mustafa Mukadam · 2024
Later among the works it cites.
Dino-foresight: Looking into the future with dino
Efstathios Karypidis, Ioannis Kakogeorgiou, Spyros Gidaris, and Nikos Komodakis · 2024
Later among the works it cites.
Openvla: An open-source vision-language-action model
Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan P Foster, Pannag R Sanketi, Quan Vuong, et al · 2024
Later among the works it cites.
Evaluating real-world robot manipulation policies in simulation
Xuanlin Li, Kyle Hsu, Jiayuan Gu, Oier Mees, Karl Pertsch, Homer Rich Walke, Chuyuan Fu, Ishikaa Lunawat, Isabel Sieh, Sean Kirmani, Sergey Levine, Jiajun Wu, Chelsea Finn, Hao Su, Quan Vuong, and Ted Xiao · 2024
Later among the works it cites.
A decade of dcase: Achievements, practices, evaluations and future challenges, 2024
Annamaria Mesaros, Romain Serizel, Toni Heittola, Tuomas Virtanen, and Mark D. Plumbley · 2024
Later among the works it cites.
Robocasa: Large-scale simulation of everyday tasks for generalist robots
Soroush Nasiriany, Abhiram Maddukuri, Lance Zhang, Adeet Parikh, Aaron Lo, Abhishek Joshi, Ajay Mandlekar, and Yuke Zhu · 2024
Later among the works it cites.
Ellma-t: an embodied llm-agent for supporting english language learning in social vr
Mengxu Pan, Alexandra Kitson, Hongyu Wan, and Mirjana Prpa · 2024
Later among the works it cites.
Genie 2: A large-scale foundation world model
Jack Parker-Holder, Philip Ball, Jake Bruce, Vibhavari Dasagi, Kristian Holsheimer, Christos Kaplanis, Alexandre Moufarek, Guy Scully, Jeremy Shar, Jimmy Shi, Stephen Spencer, Jessica Yung, Michael Dennis, Sultan Kenjeyev, Shangbang Long, Vlad Mnih, Harris Chan, Maxime Gazeau, Bonnie Li, Fabio Pardo, Luyu Wang, Lei Zhang, Frederic Besse, Tim Harley, Anna Mitenkova, Jane Wang, Jeff Clune, Demis Hassabis, Raia Hadsell, Adrian Bolton, Satinder Singh, and Tim Rocktäschel · 2024
Later among the works it cites.
Habitat 3.0: A co-habitat for humans, avatars, and robots
Xavier Puig, Eric Undersander, Andrew Szot, Mikael Dallaire Cote, Tsung-Yen Yang, Ruslan Partsey, Ruta Desai, Alexander Clegg, Michal Hlavac, So Yeon Min, Vladimír Vondruš, Theophile Gervet, Vincent-Pierre Berges, John M Turner, Oleksandr Maksymets, Zsolt Kira, Mrinal Kalakrishnan, Jitendra Malik, Devendra Singh Chaplot, Unnat Jain, Dhruv Batra, Akshara Rai, and Roozbeh Mottaghi · 2024
Later among the works it cites.
Melanie Sclar, Jane Yu, Maryam Fazel-Zarandi, Yulia Tsvetkov, Yonatan Bisk, Yejin Choi, and Asli Celikyilmaz · 2024
Later among the works it cites.
The art of llm refinement: Ask, refine, trust
Kumar Shridhar, Koustuv Sinha, Andrew Cohen, Tianlu Wang, Ping Yu, Ramakanth Pasunuru, Mrinmaya Sachan, Jason E Weston, and Asli Celikyilmaz · 2024
Later among the works it cites.
Neutral residues: revisiting adapters for model extension
Franck Signe Talla, Herve Jegou, and Edouard Grave · 2024
Later among the works it cites.
What do we learn from a large-scale study of pre-trained visual representations in sim and real environments?
Sneha Silwal, Karmesh Yadav, Tingfan Wu, Jay Vakil, Arjun Majumdar, Sergio Arnaud, Claire Chen, Vincent-Pierre Berges, Dhruv Batra, Aravind Rajeswaran, Mrinal Kalakrishnan, Franziska Meier, and Oleksandr Maksymets · 2024
Later among the works it cites.
Advancements in humanoid robots: A comprehensive review and future prospects
Yuchuang Tong, Haotian Liu, and Zhengtao Zhang · 2024
Later among the works it cites.
Meta smart glasses—large language models and the future for assistive glasses for individuals with vision impairments
Ethan Waisberg, Joshua Ong, Mouayad Masalkhi, Nasif Zaman, Prithul Sarker, Andrew G Lee, and Alireza Tavakkoli · 2024
Later among the works it cites.
M-best-rq: A multi-channel speech foundation model for smart glasses, 2024
Yufeng Yang, Desh Raj, Ju Lin, Niko Moritz, Junteng Jia, Gil Keren, Egor Lakomkin, Yiteng Huang, Jacob Donley, Jay Mahadeokar, and Ozlem Kalinli · 2024
Later among the works it cites.
Dino-wm: World models on pre-trained visual features enable zero-shot planning, 2024
Gaoyue Zhou, Hengkai Pan, Yann LeCun, and Lerrel Pinto · 2024
Later among the works it cites.
The chime-8 mmcsg challenge: Multi-modal conversations in smart glasses
Katerina Zmolikova, Simone Merello, Kaustubh Kalgaonkar, Ju Lin, Niko Moritz, Pingchuan Ma, Ming Sun, Honglie Chen, Antoine Saliou, Stavros Petridis, Christian Fuegen, and Michael Mandel · 2024
Later among the works it cites.
Seamless interaction: Dyadic audiovisual motion modeling and large-scale dataset
Vasu Agrawal, Akinniyi Akinyemi, Kathryn Alvero, Morteza Behrooz, Julia Buffalini, Fabio Maria Carlucci, Joy Chen, Junming Chen, Zhang Chen, Shiyang Cheng, Praveen Chowdary, Joe Chuang, Antony D’Avirro, Jon Daly, Ning Dong, Mark Duppenthaler, Cynthia Gao, Jeff Girard, Martin Gleize, Sahir Gomez, Hongyu Gong, Srivathsan Govindarajan, Brandon Han, Sen He, Denise Hernandez, Yordan Hristov, Rongjie Huang, Hirofumi Inaguma, Somya Jain, Raj Janardhan, Qingyao Jia, Christopher Klaiber, Dejan Kovachev, Moneish Kumar, Hang Li, Yilei Li, Pavel Litvin, Wei Liu, Guangyao Ma, Jing Ma, Martin Ma, Xutai Ma, Lucas Mantovani, Sagar Miglani, Sreyas Mohan, Louis-Philippe Morency, Evonne Ng, Kam-Woh Ng, Tu Anh Nguyen, Amia Oberai, Benjamin Peloquin, Juan Pino, Jovan Popovic, Omid Poursaeed, Fabian Prada, Alice Rakotoarison, Alexander Richard, Christophe Ropers, Safiyyah Saleem, Vasu Sharma, Alex Shcherbyna, Jia Shen, Jie Shen, Anastasis Stathopoulos, Anna Sun, Paden Tomasello, Tuan Tran, Arina Turkatenko, Bo Wan, Chao Wang, Jeff Wang, Mary Williamson, Carleigh Wood, Tao Xiang, Yilin Yang, Julien Yao, Chen Zhang, Jiemin Zhang, Xinyue Zhang, Jason Zheng, Pavlo Zhyzheria, Jan Zikes, and Michael Zollhoefer · 2025
Closest in time.
V-jepa 2: Self-supervised video models enable understanding, prediction and planning, 2025
Mahmoud Assran, Adrien Bardes, David Fan, Quentin Garrido, Russell Howes, Mojtaba Komeili, Matthew Muckley, Ammar Rizvi, Claire Roberts, Koustuv Sinha, Artem Zholus, Sergio Arnaud, Abha Gejji, Ada Martin, Francois Robert Hogan, Daniel Dugas, Piotr Bojanowski, Vasil Khalidov, Patrick Labatut, Francisco Massa, Marc Szafraniec, Kapil Krishnakumar, Yong Li, Xiaodong Ma, Sarath Chandar, Franziska Meier, Yann LeCun, Michael Rabbat, and Nicolas Ballas · 2025
Closest in time.
Roboarena: Distributed real-world evaluation of generalist robot policies
Pranav Atreya, Karl Pertsch, Tony Lee, Moo Jin Kim, Arhan Jain, Artur Kuramshin, Clemens Eppner, Cyrus Neary, Edward Hu, Fabio Ramos, et al · 2025
Closest in time.
Hallulens: Llm hallucination benchmark
Yejin Bang, Ziwei Ji, Alan Schelten, Anthony Hartshorn, Tara Fowler, Cheng Zhang, Nicola Cancedda, and Pascale Fung · 2025
Closest in time.
Gr00t n1: An open foundation model for generalist humanoid robots
Johan Bjorck, Fernando Castañeda, Nikita Cherniadev, Xingye Da, Runyu Ding, Linxi Fan, Yu Fang, Dieter Fox, Fengyuan Hu, Spencer Huang, et al · 2025
Closest in time.
Intphys 2: Benchmarking intuitive physics understanding in complex synthetic environments
Florian Bordes, Quentin Garrido, Justine T Kao, Adina Williams, Michael Rabbat, and Emmanuel Dupoux · 2025
Closest in time.
Delong Chen, Willy Chung, Yejin Bang, Ziwei Ji, and Pascale Fung · 2025
Closest in time.
Plan-and-act: Improving planning of agents for long-horizon tasks
Lutfi Eren Erdogan, Nicholas Lee, Sehoon Kim, Suhong Moon, Hiroki Furuta, Gopala Anumanchipalli, Kurt Keutzer, and Amir Gholami · 2025
Closest in time.
Causalvqa: A physically grounded causal reasoning benchmark for video models
Aaron Foss, Chloe Evans, Sasha Mitts, Koustuv Sinha, Ammar Rizvi, and Justine T Kao · 2025
Closest in time.
pi0.5: a vision-language-action model with open-world generalization
Physical Intelligence, Kevin Black, Noah Brown, James Darpinian, Karan Dhabalia, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, et al · 2025
Closest in time.
Calibrating verbal uncertainty as a linear feature to reduce hallucinations
Ziwei Ji, Lei Yu, Yeskendir Koishekenov, Yejin Bang, Anthony Hartshorn, Alan Schelten, Cheng Zhang, Pascale Fung, and Nicola Cancedda · 2025
Closest in time.
A shortcut-aware video-qa benchmark for physical understanding via minimal video pairs
Benno Krojer, Mojtaba Komeili, Candace Ross, Quentin Garrido, Koustuv Sinha, Nicolas Ballas, and Mahmoud Assran · 2025
Closest in time.
Embodied large language models enable robots to complete complex tasks in unpredictable environments
Ruaridh Mon-Williams, Gen Li, Ran Long, Wenqian Du, and Christopher G Lucas · 2025
Closest in time.
Position: Episodic memory is the missing piece for long-term LLM agents
Mathis Pink, Qinyuan Wu, Vy Ai Vo, Javier Turek, Jianing Mu, Alexander Huth, and Mariya Toneva · 2025
Closest in time.
Smolvla: A vision-language-action model for affordable and efficient robotics
Mustafa Shukor, Dana Aubakirova, Francesco Capuano, Pepijn Kooijmans, Steven Palma, Adil Zouitine, Michel Aractingi, Caroline Pascal, Martino Russi, Andres Marafioti, et al · 2025
Closest in time.
Learning from reward-free offline data: A case for planning with latent dynamics models, 02 2025
Vlad Sobal, Wancong Zhang, Kynghyun Cho, Randall Balestriero, Tim Rudner, and Yann Lecun · 2025
Closest in time.
Gemini robotics: Bringing ai into the physical world
Gemini Robotics Team, Saminda Abeyruwan, Joshua Ainslie, Jean-Baptiste Alayrac, Montserrat Gonzalez Arenas, Travis Armstrong, Ashwin Balakrishna, Robert Baruch, Maria Bauza, Michiel Blokzijl, et al · 2025
Closest in time.
Agentdam: Privacy leakage evaluation for autonomous web agents
Arman Zharmagambetov, Chuan Guo, Ivan Evtimov, Maya Pavlova, Ruslan Salakhutdinov, and Kamalika Chaudhuri · 2025
Closest in time.
Autoeval: Autonomous evaluation of generalist robot manipulation policies in the real world
Zhiyuan Zhou, Pranav Atreya, You Liang Tan, Karl Pertsch, and Sergey Levine · 2025
Closest in time.