Fetching the paper…
Reading the bibliography…
The integration of Generative Artificial Intelligence (AI) into autonomous machines represents a major paradigm shift in how these systems operate and unlocks new solutions to problems once deemed intractable.
Motion planning diffusion: Learning and planning of robot motions with diffusion models. In 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 1916–1923
Joao Carvalho, An T Le, Mark Baierl, Dorothea Koert, and Jan Peters. 2023 · 1923
Earlier work this paper cites.
Learning representations by back-propagating errors
David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams. 1986 · 1986
Earlier work this paper cites.
The utility driven dynamic error propagation network . Vol. 11
Anthony J Robinson and Frank Fallside. 1987 · 1987
Earlier work this paper cites.
The problem of learning long-term dependencies in recurrent networks. In IEEE international conference on neural networks . IEEE, 1183–1188
Yoshua Bengio, Paolo Frasconi, and Patrice Simard. 1993 · 1993
Earlier work this paper cites.
Long Short-term Memory
S Hochreiter. 1997 · 1997
Earlier work this paper cites.
Unsupervised learning of invariant feature hierarchies with applications to object recognition. In 2007 IEEE conference on computer vision and pattern recognition . IEEE, 1–8
Marc’Aurelio Ranzato, Fu Jie Huang, Y-Lan Boureau, and Yann LeCun. 2007 · 2007
Earlier work this paper cites.
Stacked convolutional auto-encoders for hierarchical feature extraction. In Artificial Neural Networks and Machine Learning–ICANN 2011: 21st International Conference on Artificial Neural Networks, Espoo, Finland, June 14-17, 2011, Proceedings, Part I 21 . Springer, 52–59
Jonathan Masci, Ueli Meier, Dan Cireşan, and Jürgen Schmidhuber. 2011 · 2011
Earlier work this paper cites.
Autoencoders, unsupervised learning, and deep architectures. In Proceedings of ICML workshop on unsupervised and transfer learning . JMLR Workshop and Conference Proceedings, 37–49
Pierre Baldi. 2012 · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma. 2013 · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
R Pascanu. 2013 · 2013
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Conditional generative adversarial nets
Mehdi Mirza and Simon Osindero. 2014 · 2014
Earlier work this paper cites.
Sequence to Sequence Learning with Neural Networks
I Sutskever. 2014 · 2014
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation. In Medical image computing and computer-assisted intervention–MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, proceedings, part III 18 . Springer, 234–241
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. 2015 · 2015
Earlier work this paper cites.
Learning structured output representation using deep conditional generative models
Kihyuk Sohn, Honglak Lee, and Xinchen Yan. 2015 · 2015
Earlier work this paper cites.
Generative adversarial text to image synthesis. In International conference on machine learning . PMLR, 1060–1069
Scott Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, and Honglak Lee. 2016 · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu. 2016 · 2016
Earlier work this paper cites.
Towards principled methods for training generative adversarial networks
Martin Arjovsky and Léon Bottou. 2017 · 2017
Earlier work this paper cites.
Image-to-image translation with conditional adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition . 1125–1134
Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros. 2017 · 2017
Earlier work this paper cites.
Factorization tricks for LSTM networks
Oleksii Kuchaiev and Boris Ginsburg. 2017 · 2017
Earlier work this paper cites.
Photo-realistic single image super-resolution using a generative adversarial network. In Proceedings of the IEEE conference on computer vision and pattern recognition . 4681–4690
Christian Ledig, Lucas Theis, Ferenc Huszár, Jose Caballero, Andrew Cunningham, Alejandro Acosta, Andrew Aitken, Alykhan Tejani, Johannes Totz, Zehan Wang, et al · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
A Vaswani. 2017 · 2017
Earlier work this paper cites.
Unpaired image-to-image translation using cycle-consistent adversarial networks. In Proceedings of the IEEE international conference on computer vision . 2223–2232
Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A Efros. 2017 · 2017
Earlier work this paper cites.
Large Scale GAN Training for High Fidelity Natural Image Synthesis
Andrew Brock. 2018 · 2018
Earlier work this paper cites.
Generating wikipedia by summarizing long sequences
Peter J Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, and Noam Shazeer. 2018 · 2018
Earlier work this paper cites.
A comparison of ARIMA and LSTM in forecasting time series. In 2018 17th IEEE international conference on machine learning and applications (ICMLA) . Ieee, 1394–1401
Sima Siami-Namini, Neda Tavakoli, and Akbar Siami Namin. 2018 · 2018
Earlier work this paper cites.
A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 4401–4410
Tero Karras, Samuli Laine, and Timo Aila. 2019 · 2019
Earlier work this paper cites.
Model cards for model reporting. In Proceedings of the conference on fairness, accountability, and transparency . 220–229
Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. 2019 · 2019
Earlier work this paper cites.
Recurrent neural networks (rnns): A gentle introduction and overview
Robin M Schmidt. 2019 · 2019
Earlier work this paper cites.
Understanding LSTM–a tutorial into long short-term memory recurrent neural networks
Ralf C Staudemeyer and Eric Rothstein Morris. 2019 · 2019
Earlier work this paper cites.
A hierarchical recurrent neural network for symbolic melody generation
Jian Wu, Changran Hu, Yulong Wang, Xiaolin Hu, and Jun Zhu. 2019 · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom B Brown. 2020 · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy. 2020 · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020 · 2020
Earlier work this paper cites.
Multilingual denoising pre-training for neural machine translation
Y Liu. 2020 · 2020
Earlier work this paper cites.
Denoising diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon. 2020 · 2020
Earlier work this paper cites.
Deep learning-based stock price prediction using LSTM and bi-directional LSTM model. In 2020 2nd novel intelligent and leading emerging sciences conference (NILES) . IEEE, 87–92
Md Arif Istiake Sunny, Mirza Mohd Shahriar Maswood, and Abdullah G Alharbi. 2020 · 2020
Earlier work this paper cites.
Intellicode compose: Code generation using transformer. In Proceedings of the 28th ACM joint meeting on European software engineering conference and symposium on the foundations of software engineering . 1433–1443
Alexey Svyatkovskiy, Shao Kun Deng, Shengyu Fu, and Neel Sundaresan. 2020 · 2020
Earlier work this paper cites.
mt5: A massively multilingual pre-trained text-to-text transformer
L Xue. 2020 · 2020
Earlier work this paper cites.
Tfix: Learning to fix coding errors with a text-to-text transformer. In International Conference on Machine Learning . PMLR, 780–791
Berkay Berabi, Jingxuan He, Veselin Raychev, and Martin Vechev. 2021 · 2021
Earlier work this paper cites.
Multimodal datasets: misogyny, pornography, and malignant stereotypes
Abeba Birhane, Vinay Uday Prabhu, and Emmanuel Kahembwe. 2021 · 2021
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol. 2021 · 2021
Earlier work this paper cites.
LongT5: Efficient text-to-text transformer for long sequences
Mandy Guo, Joshua Ainslie, David Uthus, Santiago Ontanon, Jianmo Ni, Yun-Hsuan Sung, and Yinfei Yang. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision. In International conference on machine learning . PMLR, 8748–8763
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Zero-shot text-to-image generation. In International conference on machine learning . Pmlr, 8821–8831
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. 2021 · 2021
Earlier work this paper cites.
Do as i can, not as i say: Grounding language in robotic affordances
Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, et al · 2022
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al · 2022
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Joseph Dabis, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Jasmine Hsu, et al · 2022
Earlier work this paper cites.
What can transformers learn in-context? a case study of simple function classes
Shivam Garg, Dimitris Tsipras, Percy S Liang, and Gregory Valiant. 2022 · 2022
Earlier work this paper cites.
Imagen video: High definition video generation with diffusion models
Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P Kingma, Ben Poole, Mohammad Norouzi, David J Fleet, et al · 2022
Earlier work this paper cites.
Video diffusion models
Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet. 2022b · 2022
Earlier work this paper cites.
Modular robot design optimization with generative adversarial networks. In 2022 International Conference on Robotics and Automation (ICRA) . IEEE, 4282–4288
Jiaheng Hu, Julian Whitman, Matthew Travers, and Howie Choset. 2022 · 2022
Earlier work this paper cites.
Inner monologue: Embodied reasoning through planning with language models
Wenlong Huang, Fei Xia, Ted Xiao, Harris Chan, Jacky Liang, Pete Florence, Andy Zeng, Jonathan Tompson, Igor Mordatch, Yevgen Chebotar, et al · 2022
Earlier work this paper cites.
Planning with diffusion for flexible behavior synthesis
Michael Janner, Yilun Du, Joshua B Tenenbaum, and Sergey Levine. 2022 · 2022
Earlier work this paper cites.
TAMOLS: Terrain-aware motion optimization for legged systems
Fabian Jenelten, Ruben Grandia, Farbod Farshidian, and Marco Hutter. 2022 · 2022
Earlier work this paper cites.
Coderl: Mastering code generation through pretrained models and deep reinforcement learning
Hung Le, Yue Wang, Akhilesh Deepak Gotmare, Silvio Savarese, and Steven Chu Hong Hoi. 2022 · 2022
Earlier work this paper cites.
Competition-level code generation with alphacode
Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, et al · 2022
Earlier work this paper cites.
StructDiffusion: Language-guided creation of physically-valid structures using unseen objects
Weiyu Liu, Yilun Du, Tucker Hermans, Sonia Chernova, and Chris Paxton. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Earlier work this paper cites.
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. 2022 · 2022
Earlier work this paper cites.
Plant: Explainable planning transformers via object-level representations
Katrin Renz, Kashyap Chitta, Otniel-Bogdan Mercea, A Koepke, Zeynep Akata, and Andreas Geiger. 2022 · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 10684–10695
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022 · 2022
Earlier work this paper cites.
Photorealistic text-to-image diffusion models with deep language understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al · 2022
Earlier work this paper cites.
Correcting robot plans with natural language feedback
Pratyusha Sharma, Balakumar Sundaralingam, Valts Blukis, Chris Paxton, Tucker Hermans, Antonio Torralba, Jacob Andreas, and Dieter Fox. 2022 · 2022
Earlier work this paper cites.
Git: A generative image-to-text transformer for vision and language
Jianfeng Wang, Zhengyuan Yang, Xiaowei Hu, Linjie Li, Kevin Lin, Zhe Gan, Zicheng Liu, Ce Liu, and Lijuan Wang. 2022 · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Earlier work this paper cites.
Taxonomy of risks posed by language models. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency . 214–229
Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa Kasirzadeh, et al · 2022
Earlier work this paper cites.
Robotic skill acquisition via instruction augmentation with vision-language models
Ted Xiao, Harris Chan, Pierre Sermanet, Ayzaan Wahid, Anthony Brohan, Karol Hausman, Sergey Levine, and Jonathan Tompson. 2022 · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Earlier work this paper cites.
Drag-guided diffusion models for vehicle image generation
Nikos Arechiga, Frank Permenter, Binyang Song, and Chenyang Yuan. 2023 · 2023
Earlier work this paper cites.
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al · 2023
Earlier work this paper cites.
Autoencoders
Dor Bank, Noam Koenigstein, and Raja Giryes. 2023 · 2023
Cited alongside, same era.
Tell me where to go: A composable framework for context-aware embodied robot navigation
Harel Biggie, Ajay Narasimha Mopidevi, Dusty Woods, and Christoffer Heckman. 2023 · 2023
Cited alongside, same era.
Align your latents: High-resolution video synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 22563–22575
Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. 2023 · 2023
Cited alongside, same era.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Xi Chen, Krzysztof Choromanski, Tianli Ding, Danny Driess, Avinava Dubey, Chelsea Finn, et al · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
NVIDIA Jetson AGX Orin Technical Brief
[n. d.] · 2024
Closest in time.
ANYmal Technical Specifications
ANYbotics. 2023 · 2024
Closest in time.
Demonstrating Event-Triggered Investigation and Sample Collection for Human Scientists using Field Robots and Large Foundation Models. In Robotics: Science and Systems (RSS)
Tirthankar Bandyopadhyay, Fletcher Talbot, Callum Bennie, Hashini Senaratne, Xun Li, Brendan Tidd, Mingze Xi, Jan Stiefel, Volkan Dedeoglu, Rod Taylor, et al · 2024
Closest in time.
Rt-h: Action hierarchies using language
Suneel Belkhale, Tianli Ding, Ted Xiao, Pierre Sermanet, Quon Vuong, Jonathan Tompson, Yevgen Chebotar, Debidatta Dwibedi, and Dorsa Sadigh. 2024 · 2024
Closest in time.
AutoGPT+ P: Affordance-based Task Planning with Large Language Models
Timo Birr, Christoph Pohl, Abdelrahman Younes, and Tamim Asfour. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al · 2023
Cited alongside, same era.
Making large multimodal models understand arbitrary visual prompts
Mu Cai, Haotian Liu, Siva Karthik Mustikovela, Gregory P Meyer, Yuning Chai, Dennis Park, and Yong Jae Lee. 2023 · 2023
Cited alongside, same era.
Matthew Chang, Theophile Gervet, Mukul Khanna, Sriram Yenamandra, Dhruv Shah, So Yeon Min, Kavit Shah, Chris Paxton, Saurabh Gupta, Dhruv Batra, et al · 2023
Cited alongside, same era.
Open-vocabulary queryable scene representations for real world planning. In 2023 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 11509–11522
Boyuan Chen, Fei Xia, Brian Ichter, Kanishka Rao, Keerthana Gopalakrishnan, Michael S Ryoo, Austin Stone, and Daniel Kappler. 2023d · 2023
Cited alongside, same era.
Polarnet: 3d point clouds for language-guided robotic manipulation
Shizhe Chen, Ricardo Garcia, Cordelia Schmid, and Ivan Laptev. 2023b · 2023
Cited alongside, same era.
Genaug: Retargeting behaviors to unseen situations via generative augmentation
Zoey Chen, Sho Kiami, Abhishek Gupta, and Vikash Kumar. 2023c · 2023
Cited alongside, same era.
Diffusion policy: Visuomotor policy learning via action diffusion
Cheng Chi, Siyuan Feng, Yilun Du, Zhenjia Xu, Eric Cousineau, Benjamin Burchfiel, and Shuran Song. 2023 · 2023
Cited alongside, same era.
Dall-eval: Probing the reasoning skills and social biases of text-to-image generation models. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 3043–3054
Jaemin Cho, Abhay Zala, and Mohit Bansal. 2023 · 2023
Cited alongside, same era.
The (r) evolution of multimodal large language models: A survey
Davide Caffagni, Federico Cocchi, Luca Barsellotti, Nicholas Moratelli, Sara Sarto, Lorenzo Baraldi, Marcella Cornia, and Rita Cucchiara. 2024 · 2024
Closest in time.
Creation of Novel Soft Robot Designs using Generative AI
Wee Kiat Chan, PengWei Wang, and Raye Chen-Hua Yeow. 2024 · 2024
Closest in time.
URDFormer: A Pipeline for Constructing Articulated Simulation Environments from Real-World Images
Zoey Chen, Aaron Walsman, Marius Memmel, Kaichun Mo, Alex Fang, Karthikeya Vemuri, Alan Wu, Dieter Fox, and Abhishek Gupta. 2024 · 2024
Closest in time.
AI Safety in Generative AI Large Language Models: A Survey
Jaymari Chua, Yun Li, Shiyi Yang, Chen Wang, and Lina Yao. 2024 · 2024
Closest in time.
Generative modeling of residuals for real-time risk-sensitive safety with discrete-time control barrier functions. In 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, i–viii
Ryan K Cosner, Igor Sadalski, Jana K Woo, Preston Culbertson, and Aaron D Ames. 2024 · 2024
Closest in time.
Receive, reason, and react: Drive as you say, with large language models in autonomous vehicles
Can Cui, Yunsheng Ma, Xu Cao, Wenqian Ye, and Ziran Wang. 2024 · 2024
Closest in time.
Keypoint Action tokens enable in-context imitation learning in robotics
Norman Di Palo and Edward Johns. 2024 · 2024
Closest in time.
Longrope: Extending llm context window beyond 2 million tokens
Yiran Ding, Li Lyna Zhang, Chengruidong Zhang, Yuanyuan Xu, Ning Shang, Jiahang Xu, Fan Yang, and Mao Yang. 2024 · 2024
Closest in time.
DiffuserLite: Towards Real-time Diffusion Planning
Zibin Dong, Jianye Hao, Yifu Yuan, Fei Ni, Yitian Wang, Pengyi Li, and Yan Zheng. 2024 · 2024
Closest in time.
SAGE: Bridging Semantic and Actionable Parts for Generalizable Manipulation of Articulated Objects. In ICLR 2024 Workshop on Large Language Model (LLM) Agents
Haoran Geng, Songlin Wei, Congyue Deng, Bokui Shen, He Wang, and Leonidas Guibas. 2024 · 2024
Closest in time.
SEEK: Semantic Reasoning for Object Goal Navigation in Real World Inspection Tasks
Muhammad Fadhil Ginting, Sung-Kyun Kim, David D Fan, Matteo Palieri, Mykel J Kochenderfer, and Ali-akbar Agha-Mohammadi. 2024 · 2024
Closest in time.
RVT-2: Learning Precise Manipulation from Few Demonstrations
Ankit Goyal, Valts Blukis, Jie Xu, Yijie Guo, Yu-Wei Chao, and Dieter Fox. 2024 · 2024
Closest in time.
InterPreT: Interactive Predicate Learning from Language Feedback for Generalizable Task Planning
Muzhi Han, Yifeng Zhu, Song-Chun Zhu, Ying Nian Wu, and Yuke Zhu. 2024 · 2024
Closest in time.
Fast LiDAR Upsampling using Conditional Diffusion Models
Sander Elias Magnussen Helgesen, Kazuto Nakashima, Jim Tørresen, and Ryo Kurazume. 2024 · 2024
Closest in time.
EfficientDreamer: High-Fidelity and Robust 3D Creation via Orthogonal-view Diffusion Priors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 4949–4958
Zhipeng Hu, Minda Zhao, Chaoyi Zhao, Xinyue Liang, Lincheng Li, Zeng Zhao, Changjie Fan, Xiaowei Zhou, and Xin Yu. 2024 · 2024
Closest in time.
New Solutions on LLM Acceleration, Optimization, and Application
Yingbing Huang, Lily Jiaxin Wan, Hanchen Ye, Manvi Jha, Jinghua Wang, Yuhong Li, Xiaofan Zhang, and Deming Chen. 2024 · 2024
Closest in time.
Vid2robot: End-to-end video-conditioned policy learning with cross-attention transformers
Vidhi Jain, Maria Attarian, Nikhil J Joshi, Ayzaan Wahid, Danny Driess, Quan Vuong, Pannag R Sanketi, Pierre Sermanet, Stefan Welker, Christine Chan, et al · 2024
Closest in time.
Gen2sim: Scaling up robot learning in simulation with generative models. In 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 6672–6679
Pushkal Katara, Zhou Xian, and Katerina Fragkiadaki. 2024 · 2024
Closest in time.
Constraint-Aware Diffusion Models for Trajectory Optimization
Anjian Li, Zihan Ding, Adji Bousso Dieng, and Ryne Beeson. 2024 · 2024
Closest in time.
Learning to learn faster from human feedback with language model predictive control
Jacky Liang, Fei Xia, Wenhao Yu, Andy Zeng, Montserrat Gonzalez Arenas, Maria Attarian, Maria Bauza, Matthew Bennice, Alex Bewley, Adil Dostmohamed, et al · 2024
Closest in time.
SocialGAIL: Faithful Crowd Simulation for Social Robot Navigation. In 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 16873–16880
Bo Ling, Yan Lyu, Dongxiao Li, Guanyu Gao, Yi Shi, Xueyong Xu, and Weiwei Wu. 2024 · 2024
Closest in time.
Moka: Open-vocabulary robotic manipulation through mark-based visual prompting
Fangchen Liu, Kuan Fang, Pieter Abbeel, and Sergey Levine. 2024a · 2024
Closest in time.
Dipper: Diffusion-based 2d path planner applied on legged robots. In 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 9264–9270
Jianwei Liu, Maria Stamatopoulou, and Dimitrios Kanoulas. 2024e · 2024
Closest in time.
Ok-robot: What really matters in integrating open-knowledge models for robotics
Peiqi Liu, Yaswanth Orru, Chris Paxton, Nur Muhammad Mahi Shafiullah, and Lerrel Pinto. 2024d · 2024
Closest in time.
Composable part-based manipulation
Weiyu Liu, Jiayuan Mao, Joy Hsu, Tucker Hermans, Animesh Garg, and Jiajun Wu. 2024c · 2024
Closest in time.
Wonder3d: Single image to 3d using cross-domain diffusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 9970–9980
Xiaoxiao Long, Yuan-Chen Guo, Cheng Lin, Yuan Liu, Zhiyang Dou, Lingjie Liu, Yuexin Ma, Song-Hai Zhang, Marc Habermann, Christian Theobalt, et al · 2024
Closest in time.
DrEureka: Language Model Guided Sim-To-Real Transfer
Yecheng Jason Ma, William Liang, Hung-Ju Wang, Sam Wang, Yuke Zhu, Linxi Fan, Osbert Bastani, and Dinesh Jayaraman. 2024 · 2024
Closest in time.
Reorientdiff: Diffusion model based reorientation for object manipulation. In 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 10867–10873
Utkarsh A Mishra and Yongxin Chen. 2024 · 2024
Closest in time.
CBF-LLM: Safe Control for LLM Alignment
Yuya Miyaoka and Masaki Inoue. 2024 · 2024
Closest in time.
Kazuki Mizuta and Karen Leung. 2024 · 2024
Closest in time.
AI Safety Benchmarks
MLCommons. 2024 · 2024
Closest in time.
CppFlow: Generative Inverse Kinematics for Efficient and Robust Cartesian Path Planning. In 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 12279–12785
Jeremy Morgan, David Millard, and Gaurav S Sukhatme. 2024 · 2024
Closest in time.
LiDAR Data Synthesis with Denoising Diffusion Probabilistic Models. In 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 14724–14731
Kazuto Nakashima and Ryo Kurazume. 2024 · 2024
Closest in time.
RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots
Soroush Nasiriany, Abhiram Maddukuri, Lance Zhang, Adeet Parikh, Aaron Lo, Abhishek Joshi, Ajay Mandlekar, and Yuke Zhu. 2024 · 2024
Closest in time.
Consistency policy: Accelerated visuomotor policies via consistency distillation
Aaditya Prasad, Kevin Lin, Jimmy Wu, Linqi Zhou, and Jeannette Bohg. 2024 · 2024
Closest in time.
Mobile Edge Intelligence for Large Language Models: A Contemporary Survey
Guanqiao Qu, Qiyuan Chen, Wei Wei, Zheng Lin, Xianhao Chen, and Kaibin Huang. 2024 · 2024
Closest in time.
Adversarial Nibbler: An Open Red-Teaming Method for Identifying Diverse Harms in Text-to-Image Generation. In The 2024 ACM Conference on Fairness, Accountability, and Transparency . 388–406
Jessica Quaye, Alicia Parrish, Oana Inel, Charvi Rastogi, Hannah Rose Kirk, Minsuk Kahng, Erin Van Liemt, Max Bartolo, Jess Tsang, Justin White, et al · 2024
Closest in time.
Explore until Confident: Efficient Exploration for Embodied Question Answering
Allen Z Ren, Jaden Clark, Anushri Dixit, Masha Itkina, Anirudha Majumdar, and Dorsa Sadigh. 2024 · 2024
Closest in time.
OmniLRS: A Photorealistic Simulator for Lunar Robotics. In 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 16901–16907
Antoine Richard, Junnosuke Kamohara, Kentaro Uno, Shreya Santra, Dave van der Meer, Miguel Olivares-Mendez, and Kazuya Yoshida. 2024 · 2024
Closest in time.
Generative AI and deepfakes: a human rights approach to tackling harmful content
Felipe Romero Moreno. 2024 · 2024
Closest in time.
Fast and Modular Autonomy Software for Autonomous Racing Vehicles
Andrew Saba, Aderotimi Adetunji, Adam Johnson, Aadi Kothari, Matthew Sivaprakasam, Joshua Spisak, Prem Bharatia, Arjun Chauhan, Brendan Duff Jr, Noah Gasparro, et al · 2024
Closest in time.
Tool Use Leaderboard
Scale AI. 2024 · 2024
Closest in time.
Yell at your robot: Improving on-the-fly from language corrections
Lucy Xiaoyang Shi, Zheyuan Hu, Tony Z Zhao, Archit Sharma, Karl Pertsch, Jianlan Luo, Sergey Levine, and Chelsea Finn. 2024 · 2024
Closest in time.
Real-Time Anomaly Detection and Reactive Planning with Large Language Models
Rohan Sinha, Amine Elhafsi, Christopher Agia, Matthew Foutter, Edward Schmerling, and Marco Pavone. 2024 · 2024
Closest in time.
Nomad: Goal masked diffusion policies for navigation and exploration. In 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 63–70
Ajay Sridhar, Dhruv Shah, Catherine Glossop, and Sergey Levine. 2024 · 2024
Closest in time.
DiPPeST: Diffusion-based Path Planner for Synthesizing Trajectories Applied on Quadruped Robots
Maria Stamatopoulou, Jianwei Liu, and Dimitrios Kanoulas. 2024 · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. 2024 · 2024
Closest in time.
Prompt, plan, perform: Llm-based humanoid control via quantized imitation learning. In 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 16236–16242
Jingkai Sun, Qiang Zhang, Yiqun Duan, Xiaoyang Jiang, Chong Cheng, and Renjing Xu. 2024 · 2024
Closest in time.
Octo: An open-source generalist robot policy
Octo Model Team, Dibya Ghosh, Homer Walke, Karl Pertsch, Kevin Black, Oier Mees, Sudeep Dasari, Joey Hejna, Tobias Kreiman, Charles Xu, et al · 2024
Closest in time.
A comprehensive survey of hallucination mitigation techniques in large language models
SM Tonmoy, SM Zaman, Vinija Jain, Anku Rani, Vipula Rawte, Aman Chadha, and Amitava Das. 2024 · 2024
Closest in time.
Language models don’t always say what they think: unfaithful explanations in chain-of-thought prompting
Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman. 2024 · 2024
Closest in time.
Render and Diffuse: Aligning Image and Action Spaces for Diffusion-based Behaviour Cloning
Vitalis Vosylius, Younggyo Seo, Jafar Uruç, and Stephen James. 2024 · 2024
Closest in time.
Diffusebot: Breeding soft robots with physics-augmented generative diffusion models
Tsun-Hsuan Johnson Wang, Juntian Zheng, Pingchuan Ma, Yilun Du, Byungchul Kim, Andrew Spielberg, Josh Tenenbaum, Chuang Gan, and Daniela Rus. 2024 · 2024
Closest in time.
Hierarchical Open-Vocabulary 3D Scene Graphs for Language-Grounded Robot Navigation. In First Workshop on Vision-Language Models for Navigation and Manipulation at ICRA 2024
Abdelrhman Werby, Chenguang Huang, Martin Büchner, Abhinav Valada, and Wolfram Burgard. 2024 · 2024
Closest in time.
The learnability of in-context learning
Noam Wies, Yoav Levine, and Amnon Shashua. 2024 · 2024
Closest in time.
Dynamics-Guided Diffusion Model for Robot Manipulator Design
Xiaomeng Xu, Huy Ha, and Shuran Song. 2024 · 2024
Closest in time.
Vlfm: Vision-language frontier maps for zero-shot semantic navigation. In 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 42–48
Naoki Yokoyama, Sehoon Ha, Dhruv Batra, Jiuguang Wang, and Bernadette Bucher. 2024 · 2024
Closest in time.
Natural Language Can Help Bridge the Sim2Real Gap
Albert Yu, Adeline Foote, Raymond Mooney, and Roberto Martín-Martín. 2024a · 2024
Closest in time.
Octopi: Object Property Reasoning with Large Tactile-Language Models
Samson Yu, Kelvin Lin, Anxing Xiao, Jiafei Duan, and Harold Soh. 2024b · 2024
Closest in time.
Jianhao Yuan, Shuyang Sun, Daniel Omeiza, Bo Zhao, Paul Newman, Lars Kunze, and Matthew Gadd. 2024 · 2024
Closest in time.
3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations. In ICRA 2024 Workshop on 3D Visual Representations for Robot Manipulation
Yanjie Ze, Gu Zhang, Kangning Zhang, Chenyuan Hu, Muhan Wang, and Huazhe Xu. 2024 · 2024
Closest in time.
NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation
Jiazhao Zhang, Kunyu Wang, Rongtao Xu, Gengze Zhou, Yicong Hong, Xiaomeng Fang, Qi Wu, Zhizheng Zhang, and Wang He. 2024b · 2024
Closest in time.
Diffusion Meets DAgger: Supercharging Eye-in-hand Imitation Learning
Xiaoyu Zhang, Matthew Chang, Pranav Kumar, and Saurabh Gupta. 2024a · 2024
Closest in time.
VLMPC: Vision-Language Model Predictive Control for Robotic Manipulation
Wentao Zhao, Jiaming Chen, Ziyu Meng, Donghui Mao, Ran Song, and Wei Zhang. 2024 · 2024
Closest in time.
3d-oae: Occlusion auto-encoders for self-supervised learning on point clouds. In 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 15416–15423
Junsheng Zhou, Xin Wen, Baorui Ma, Yu-Shen Liu, Yue Gao, Yi Fang, and Zhizhong Han. 2024 · 2024
Closest in time.
Playfusion: Skill acquisition via diffusion from language-annotated play. In Conference on Robot Learning . PMLR, 2012–2029
Lili Chen, Shikhar Bahl, and Deepak Pathak. 2023a · 2029
Closest in time.