Fetching the paper…
Reading the bibliography…
Foundation models pretrained on diverse data at scale have demonstrated extraordinary capabilities in a wide range of vision and language tasks.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning. In International conference on machine learning . PMLR, 1928–1937
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu. 2016 · 1937
Earlier work this paper cites.
A multiple shooting algorithm for direct solution of optimal control problems
Hans Georg Bock and Karl-Josef Plitt. 1984 · 1984
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
Dean A Pomerleau. 1988 · 1988
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
Dean A Pomerleau. 1989 · 1989
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Richard S Sutton. 1990 · 1990
Earlier work this paper cites.
Direct and indirect methods for trajectory optimization
Oskar Von Stryk and Roland Bulirsch. 1992 · 1992
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan. 1992 · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams. 1992 · 1992
Earlier work this paper cites.
Numerical solution of optimal control problems by direct collocation
Oskar Von Stryk. 1993 · 1993
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L Puterman. 1994 · 1994
Earlier work this paper cites.
TD-Gammon, a self-teaching backgammon program, achieves master-level play
Gerald Tesauro. 1994 · 1994
Earlier work this paper cites.
Abstraction and approximate decision-theoretic planning
Richard Dearden and Craig Boutilier. 1997 · 1997
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David McAllester, Satinder Singh, and Yishay Mansour. 1999 · 1999
Earlier work this paper cites.
A natural policy gradient
Sham M Kakade. 2001 · 2001
Earlier work this paper cites.
State abstraction for programmable reinforcement learning agents. In Aaai/iaai . 119–125
David Andre and Stuart J Russell. 2002 · 2002
Earlier work this paper cites.
Multiple model-based reinforcement learning
Kenji Doya, Kazuyuki Samejima, Ken-ichi Katagiri, and Mitsuo Kawato. 2002 · 2002
Earlier work this paper cites.
Metrics for Finite Markov Decision Processes.. In UAI , Vol. 4. 162–169
Norm Ferns, Prakash Panangaden, and Doina Precup. 2004 · 2004
Earlier work this paper cites.
Dynamic abstraction in reinforcement learning via clustering. In Proceedings of the twenty-first international conference on Machine learning . 71
Shie Mannor, Ishai Menache, Amit Hoze, and Uri Klein. 2004 · 2004
Earlier work this paper cites.
Indri: A language model-based search engine for complex queries. In Proceedings of the international conference on intelligent analysis , Vol. 2. Washington, DC., 2–6
Trevor Strohman, Donald Metzler, Howard Turtle, and W Bruce Croft. 2005 · 2005
Earlier work this paper cites.
Improved monte-carlo search
Levente Kocsis, Csaba Szepesvári, and Jan Willemson. 2006 · 2006
Earlier work this paper cites.
A tutorial on energy-based learning
Yann LeCun, Sumit Chopra, Raia Hadsell, M Ranzato, and Fujie Huang. 2006 · 2006
Earlier work this paper cites.
Outliers: The story of success
Malcolm Gladwell. 2008 · 2008
Earlier work this paper cites.
Using bisimulation for policy transfer in MDPs. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 24
Pablo Castro and Doina Precup. 2010 · 2010
Earlier work this paper cites.
Relative entropy policy search. In Twenty-Fourth AAAI Conference on Artificial Intelligence
Jan Peters, Katharina Mulling, and Yasemin Altun. 2010 · 2010
Earlier work this paper cites.
PILCO: A model-based and data-efficient approach to policy search. In Proceedings of the 28th International Conference on machine learning (ICML-11) . 465–472
Marc Deisenroth and Carl E Rasmussen. 2011 · 2011
Earlier work this paper cites.
Synthesis and stabilization of complex behaviors through online trajectory optimization. In 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems . IEEE, 4906–4913
Yuval Tassa, Tom Erez, and Emanuel Todorov. 2012 · 2012
Earlier work this paper cites.
Model predictive control
Eduardo F Camacho and Carlos Bordons Alba. 2013 · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling. 2013 · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller. 2013 · 2013
Earlier work this paper cites.
Deterministic policy gradient algorithms. In International conference on machine learning . PMLR, 387–395
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller. 2014 · 2014
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. 2015 · 2015
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel. 2015b · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning . PMLR, 2256–2265
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. 2015 · 2015
Earlier work this paper cites.
Rl 2 : Fast reinforcement learning via slow reinforcement learning
Yan Duan, John Schulman, Xi Chen, Peter L Bartlett, Ilya Sutskever, and Pieter Abbeel. 2016 · 2016
Earlier work this paper cites.
Q-prop: Sample-efficient policy gradient with an off-policy critic
Shixiang Gu, Timothy Lillicrap, Zoubin Ghahramani, Richard E Turner, and Sergey Levine. 2016 · 2016
Earlier work this paper cites.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon. 2016 · 2016
Earlier work this paper cites.
Loss is its own reward: Self-supervision for reinforcement learning
Evan Shelhamer, Parsa Mahmoudieh, Max Argus, and Trevor Darrell. 2016 · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Earlier work this paper cites.
Learning to reinforcement learn
Jane X Wang, Zeb Kurth-Nelson, Dhruva Tirumala, Hubert Soyer, Joel Z Leibo, Remi Munos, Charles Blundell, Dharshan Kumaran, and Matt Botvinick. 2016 · 2016
Earlier work this paper cites.
The" something something" video database for learning and evaluating visual common sense. In Proceedings of the IEEE international conference on computer vision . 5842–5850
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al · 2017
Earlier work this paper cites.
Sequence Tutor: Conservative Fine-Tuning of Sequence Generation Models with KL-control
N. Jaques, S. Gu, D. Bahdanau, J. M. Hernandez-Lobato, R. E. Turner, and D. Eck. 2017 · 2017
Earlier work this paper cites.
Bridging the gap between value and policy based reinforcement learning
Ofir Nachum, Mohammad Norouzi, Kelvin Xu, and Dale Schuurmans. 2017 · 2017
Earlier work this paper cites.
Value prediction network
Junhyuk Oh, Satinder Singh, and Honglak Lee. 2017 · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction. In International conference on machine learning . PMLR, 2778–2787
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell. 2017 · 2017
Earlier work this paper cites.
Imagination-augmented agents for deep reinforcement learning
Sébastien Racanière, Théophane Weber, David Reichert, Lars Buesing, Arthur Guez, Danilo Jimenez Rezende, Adrià Puigdomènech Badia, Oriol Vinyals, Nicolas Heess, Yujia Li, et al · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 · 2017
Earlier work this paper cites.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2017
Earlier work this paper cites.
Domain randomization for transferring deep neural networks from simulation to the real world. In 2017 IEEE/RSJ international conference on intelligent robots and systems (IROS) . IEEE, 23–30
Josh Tobin, Rachel Fong, Alex Ray, Jonas Schneider, Wojciech Zaremba, and Pieter Abbeel. 2017 · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aaron Van Den Oord, Oriol Vinyals, et al · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
State abstractions for lifelong reinforcement learning. In International Conference on Machine Learning . PMLR, 10–19
David Abel, Dilip Arumugam, Lucas Lehnert, and Michael Littman. 2018 · 2018
Earlier work this paper cites.
Playing hard exploration games by watching youtube
Yusuf Aytar, Tobias Pfaff, David Budden, Thomas Paine, Ziyu Wang, and Nando De Freitas. 2018 · 2018
Earlier work this paper cites.
Babyai: A platform to study the sample efficiency of grounded language learning
Maxime Chevalier-Boisvert, Dzmitry Bahdanau, Salem Lahlou, Lucas Willems, Chitwan Saharia, Thien Huu Nguyen, and Yoshua Bengio. 2018 · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Model-based value estimation for efficient model-free reinforcement learning
Vladimir Feinberg, Alvin Wan, Ion Stoica, Michael I Jordan, Joseph E Gonzalez, and Sergey Levine. 2018 · 2018
Earlier work this paper cites.
Scalable deep reinforcement learning for vision-based robotic manipulation. In Conference on Robot Learning . PMLR, 651–673
Dmitry Kalashnikov, Alex Irpan, Peter Pastor, Julian Ibarz, Alexander Herzog, Eric Jang, Deirdre Quillen, Ethan Holly, Mrinal Kalakrishnan, Vincent Vanhoucke, et al · 2018
Earlier work this paper cites.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning. In 2018 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 7559–7566
Anusha Nagabandi, Gregory Kahn, Ronald S Fearing, and Sergey Levine. 2018 · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. 2018 · 2018
Earlier work this paper cites.
Virtualhome: Simulating household activities via programs. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 8494–8502
Xavier Puig, Kevin Ra, Marko Boben, Jiaman Li, Tingwu Wang, Sanja Fidler, and Antonio Torralba. 2018 · 2018
Earlier work this paper cites.
Time-contrastive networks: Self-supervised learning from video. In 2018 IEEE international conference on robotics and automation (ICRA) . IEEE, 1134–1141
Pierre Sermanet, Corey Lynch, Yevgen Chebotar, Jasmine Hsu, Eric Jang, Stefan Schaal, Sergey Levine, and Google Brain. 2018 · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto. 2018 · 2018
Earlier work this paper cites.
Reinforcement and imitation learning for diverse visuomotor skills
Yuke Zhu, Ziyu Wang, Josh Merel, Andrei Rusu, Tom Erez, Serkan Cabi, Saran Tunyasuvunakool, János Kramár, Raia Hadsell, Nando de Freitas, et al · 2018
Earlier work this paper cites.
Solving rubik’s cube with a robot hand
Ilge Akkaya, Marcin Andrychowicz, Maciek Chociej, Mateusz Litwin, Bob McGrew, Arthur Petron, Alex Paino, Matthias Plappert, Glenn Powell, Raphael Ribas, et al · 2019
Earlier work this paper cites.
Learning from demonstration in the wild. In 2019 International Conference on Robotics and Automation (ICRA) . IEEE, 775–781
Feryal Behbahani, Kyriacos Shiarlis, Xi Chen, Vitaly Kurin, Sudhanshu Kasewa, Ciprian Stirbu, Joao Gomes, Supratik Paul, Frans A Oliehoek, Joao Messias, et al · 2019
Earlier work this paper cites.
Dota 2 with large scale deep reinforcement learning
Christopher Berner, Greg Brockman, Brooke Chan, Vicki Cheung, Przemysław Dębiak, Christy Dennison, David Farhi, Quirin Fischer, Shariq Hashme, Chris Hesse, et al · 2019
Earlier work this paper cites.
Robonet: Large-scale multi-robot learning
Sudeep Dasari, Frederik Ebert, Stephen Tian, Suraj Nair, Bernadette Bucher, Karl Schmeckpeper, Siddharth Singh, Sergey Levine, and Chelsea Finn. 2019 · 2019
Earlier work this paper cites.
Model Based Planning with Energy Based Models
Yilun Du, Toru Lin, and Igor Mordatch. 2019 · 2019
Earlier work this paper cites.
Implicit generation and generalization in energy-based models
Yilun Du and Igor Mordatch. 2019 · 2019
Earlier work this paper cites.
Task-Agnostic Dynamics Priors for Deep Reinforcement Learning. In International Conference on Machine Learning
Yilun Du and Karthik Narasimhan. 2019 · 2019
Earlier work this paper cites.
Deepmdp: Learning continuous latent space models for representation learning. In International Conference on Machine Learning . PMLR, 2170–2179
Carles Gelada, Saurabh Kumar, Jacob Buckman, Ofir Nachum, and Marc G Bellemare. 2019 · 2019
Earlier work this paper cites.
Dream to control: Learning behaviors by latent imagination
Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi. 2019 · 2019
Earlier work this paper cites.
Model-based reinforcement learning for atari
Lukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski, Roy H Campbell, Konrad Czechowski, Dumitru Erhan, Chelsea Finn, Piotr Kozakowski, Sergey Levine, et al · 2019
Earlier work this paper cites.
Aviral Kumar, Xue Bin Peng, and Sergey Levine. 2019 · 2019
Earlier work this paper cites.
Reinforcement learning applications
Yuxi Li. 2019 · 2019
Cited alongside, same era.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 2630–2640
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. 2019 · 2019
Cited alongside, same era.
Reinforcement Learning Upside Down: Don’t Predict Rewards–Just Map Them to Actions
Juergen Schmidhuber. 2019 · 2019
Cited alongside, same era.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al · 2019
Cited alongside, same era.
Understanding HTML with Large Language Models
Izzeddin Gur, Ofir Nachum, Yingjie Miao, Mustafa Safdari, Austin Huang, Aakanksha Chowdhery, Sharan Narang, Noah Fiedel, and Aleksandra Faust. 2022 · 2022
Later among the works it cites.
Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 16000–16009
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. 2022 · 2022
Later among the works it cites.
Imagen video: High definition video generation with diffusion models
Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P Kingma, Ben Poole, Mohammad Norouzi, David J Fleet, et al · 2022
Later among the works it cites.
Classifier-free diffusion guidance
Jonathan Ho and Tim Salimans. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yifan Wu, George Tucker, and Ofir Nachum. 2019 · 2019
Cited alongside, same era.
Towards a human-like open-domain chatbot
Daniel Adiwardana, Minh-Thang Luong, David R So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, et al · 2020
Cited alongside, same era.
Opal: Offline primitive discovery for accelerating offline reinforcement learning
Anurag Ajay, Aviral Kumar, Pulkit Agrawal, Sergey Levine, and Ofir Nachum. 2020 · 2020
Cited alongside, same era.
A simple framework for contrastive learning of visual representations. In International conference on machine learning . PMLR, 1597–1607
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020 · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Cited alongside, same era.
Compositional Visual Generation with Energy Based Models. In Advances in Neural Information Processing Systems
Yilun Du, Shuang Li, and Igor Mordatch. 2020 · 2020
Cited alongside, same era.
Rl unplugged: A suite of benchmarks for offline reinforcement learning
Caglar Gulcehre, Ziyu Wang, Alexander Novikov, Thomas Paine, Sergio Gómez, Konrad Zolna, Rishabh Agarwal, Josh S Merel, Daniel J Mankowitz, Cosmin Paduraru, et al · 2020
Cited alongside, same era.
Mastering atari with discrete world models
Danijar Hafner, Timothy Lillicrap, Mohammad Norouzi, and Jimmy Ba. 2020 · 2020
Cited alongside, same era.
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Later among the works it cites.
Language models as zero-shot planners: Extracting actionable knowledge for embodied agents
Wenlong Huang, Pieter Abbeel, Deepak Pathak, and Igor Mordatch. 2022a · 2022
Later among the works it cites.
Inner monologue: Embodied reasoning through planning with language models
Wenlong Huang, Fei Xia, Ted Xiao, Harris Chan, Jacky Liang, Pete Florence, Andy Zeng, Jonathan Tompson, Igor Mordatch, Yevgen Chebotar, et al · 2022
Later among the works it cites.
Planning with Diffusion for Flexible Behavior Synthesis
Michael Janner, Yilun Du, Joshua B Tenenbaum, and Sergey Levine. 2022 · 2022
Later among the works it cites.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Yejin Bang, Andrea Madotto, and Pascale Fung. 2022 · 2022
Later among the works it cites.
Vima: General robot manipulation with multimodal prompts
Yunfan Jiang, Agrim Gupta, Zichen Zhang, Guanzhi Wang, Yongqiang Dou, Yanjun Chen, Li Fei-Fei, Anima Anandkumar, Yuke Zhu, and Linxi Fan. 2022 · 2022
Later among the works it cites.
Simple but effective: Clip embeddings for embodied ai. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 14829–14838
Apoorv Khandelwal, Luca Weihs, Roozbeh Mottaghi, and Aniruddha Kembhavi. 2022 · 2022
Later among the works it cites.
Offline Q-Learning on Diverse Multi-Task Data Both Scales And Generalizes
Aviral Kumar, Rishabh Agarwal, Xinyang Geng, George Tucker, and Sergey Levine. 2022 · 2022
Later among the works it cites.
In-context reinforcement learning with algorithm distillation
Michael Laskin, Luyu Wang, Junhyuk Oh, Emilio Parisotto, Stephen Spencer, Richie Steigerwald, DJ Strouse, Steven Hansen, Angelos Filos, Ethan Brooks, et al · 2022
Later among the works it cites.
Internet-augmented language models through few-shot prompting for open-domain question answering
Angeliki Lazaridou, Elena Gribovskaya, Wojciech Stokowiec, and Nikolai Grigorev. 2022 · 2022
Later among the works it cites.
Multi-Game Decision Transformers
Kuang-Huei Lee, Ofir Nachum, Mengjiao Yang, Lisa Lee, Daniel Freeman, Winnie Xu, Sergio Guadarrama, Ian Fischer, Eric Jang, Henryk Michalewski, et al · 2022
Later among the works it cites.
Solving quantitative reasoning problems with language models
Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, et al · 2022
Later among the works it cites.
Composing Ensembles of Pre-trained Models via Iterative Consensus
Shuang Li, Yilun Du, Joshua B Tenenbaum, Antonio Torralba, and Igor Mordatch. 2022a · 2022
Later among the works it cites.
Pre-trained language models for interactive decision-making
Shuang Li, Xavier Puig, Yilun Du, Clinton Wang, Ekin Akyurek, Antonio Torralba, Jacob Andreas, and Igor Mordatch. 2022b · 2022
Later among the works it cites.
Masked Autoencoding for Scalable and Generalizable Decision Making
Fangchen Liu, Hao Liu, Aditya Grover, and Pieter Abbeel. 2022c · 2022
Later among the works it cites.
Instruction-Following Agents with Jointly Pre-Trained Vision-Language Models
Hao Liu, Lisa Lee, Kimin Lee, and Pieter Abbeel. 2022a · 2022
Later among the works it cites.
Compositional Visual Generation with Composable Diffusion Models
Nan Liu, Shuang Li, Yilun Du, Antonio Torralba, and Joshua B Tenenbaum. 2022b · 2022
Later among the works it cites.
Mind’s Eye: Grounded Language Model Reasoning through Simulation
Ruibo Liu, Jason Wei, Shixiang Shane Gu, Te-Yen Wu, Soroush Vosoughi, Claire Cui, Denny Zhou, and Andrew M Dai. 2022d · 2022
Later among the works it cites.
Zero-Shot Reward Specification via Grounded Natural Language. In ICLR 2022 Workshop on Generalizable Policy Learning in Physical World
Parsa Mahmoudieh, Deepak Pathak, and Trevor Darrell. 2022 · 2022
Later among the works it cites.
CACTI: A Framework for Scalable Multi-Task Multi-Scene Visual Imitation Learning
Zhao Mandi, Homanga Bharadhwaj, Vincent Moens, Shuran Song, Aravind Rajeswaran, and Vikash Kumar. 2022 · 2022
Later among the works it cites.
Contrastive Value Learning: Implicit Models for Simple Offline RL
Bogdan Mazoure, Benjamin Eysenbach, Ofir Nachum, and Jonathan Tompson. 2022 · 2022
Later among the works it cites.
Transformers are sample efficient world models
Vincent Micheli, Eloi Alonso, and François Fleuret. 2022 · 2022
Later among the works it cites.
Learning language-conditioned robot behavior from offline data and crowd-sourced annotation. In Conference on Robot Learning . PMLR, 1303–1315
Suraj Nair, Eric Mitchell, Kevin Chen, Silvio Savarese, Chelsea Finn, et al · 2022
Later among the works it cites.
CHATGPT: Optimizing language models for dialogue
OpenAI. 2022 · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
Joint Representation Training in Sequential Tasks with Shared Structure
Aldo Pacchiano, Ofir Nachum, Nilseh Tripuraneni, and Peter Bartlett. 2022 · 2022
Later among the works it cites.
Talm: Tool augmented language models
Aaron Parisi, Yao Zhao, and Noah Fiedel. 2022 · 2022
Later among the works it cites.
You Can’t Count on Luck: Why Decision Transformers Fail in Stochastic Environments
Keiran Paster, Sheila McIlraith, and Jimmy Ba. 2022 · 2022
Later among the works it cites.
Measuring and Narrowing the Compositionality Gap in Language Models
Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah A Smith, and Mike Lewis. 2022 · 2022
Later among the works it cites.
Planning with Large Language Models via Corrective Re-prompting
Shreyas Sundara Raman, Vanya Cohen, Eric Rosen, Ifrah Idrees, David Paulius, and Stefanie Tellex. 2022 · 2022
Later among the works it cites.
Scott Reed, Konrad Zolna, Emilio Parisotto, Sergio Gomez Colmenarejo, Alexander Novikov, Gabriel Barth-Maron, Mai Gimenez, Yury Sulsky, Jackie Kay, Jost Tobias Springenberg, et al · 2022
Later among the works it cites.
Can Wikipedia Help Offline Reinforcement Learning?
Machel Reid, Yutaro Yamada, and Shixiang Shane Gu. 2022 · 2022
Later among the works it cites.
Latent Variable Representation for Reinforcement Learning
Tongzheng Ren, Chenjun Xiao, Tianjun Zhang, Na Li, Zhaoran Wang, Sujay Sanghavi, Dale Schuurmans, and Bo Dai. 2022 · 2022
Later among the works it cites.
Laion-5b: An open large-scale dataset for training next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al · 2022
Later among the works it cites.
Masked world models for visual control
Younggyo Seo, Danijar Hafner, Hao Liu, Fangchen Liu, Stephen James, Kimin Lee, and Pieter Abbeel. 2022a · 2022
Later among the works it cites.
Behavior Transformers: Cloning k k modes with one stone
Nur Muhammad Mahi Shafiullah, Zichen Jeff Cui, Ariuntuya Altanzaya, and Lerrel Pinto. 2022 · 2022
Later among the works it cites.
Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action
Dhruv Shah, Blazej Osinski, Brian Ichter, and Sergey Levine. 2022 · 2022
Later among the works it cites.
VideoDex: Learning Dexterity from Internet Videos
Kenneth Shaw, Shikhar Bahl, and Deepak Pathak. 2022 · 2022
Later among the works it cites.
Cliport: What and where pathways for robotic manipulation. In Conference on Robot Learning . PMLR, 894–906
Mohit Shridhar, Lucas Manuelli, and Dieter Fox. 2022 · 2022
Later among the works it cites.
BlenderBot 3: a deployed conversational agent that continually learns to responsibly engage
Kurt Shuster, Jing Xu, Mojtaba Komeili, Da Ju, Eric Michael Smith, Stephen Roller, Megan Ung, Moỹa Chen, Kushal Arora, Joshua Lane, Morteza Behrooz, William Ngan, Spencer Poff, Naman Goyal, Arthur Szlam, Ỹ-Lan Boureau, Melanie Kambadur, and Jason Weston. 2022 · 2022
Later among the works it cites.
Offline rl for natural language generation with implicit language q learning
Charlie Snell, Ilya Kostrikov, Yi Su, Mengjiao Yang, and Sergey Levine. 2022a · 2022
Later among the works it cites.
Context-aware language modeling for goal-oriented dialogue systems
Charlie Snell, Sherry Yang, Justin Fu, Yi Su, and Sergey Levine. 2022b · 2022
Later among the works it cites.
Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al · 2022
Later among the works it cites.
PlaTe: Visually-grounded planning with transformers in procedural tasks
Jiankai Sun, De-An Huang, Bo Lu, Yun-Hui Liu, Bolei Zhou, and Animesh Garg. 2022 · 2022
Later among the works it cites.
Semantic exploration from language abstractions and pretrained representations
Allison C Tam, Neil C Rabinowitz, Andrew K Lampinen, Nicholas A Roy, Stephanie CY Chan, DJ Strouse, Jane X Wang, Andrea Banino, and Felix Hill. 2022 · 2022
Later among the works it cites.
Evaluating Vision Transformer Methods for Deep Reinforcement Learning from Pixels
Tianxin Tao, Daniele Reda, and Michiel van de Panne. 2022 · 2022
Later among the works it cites.
Lamda: Language models for dialog applications
Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al · 2022
Later among the works it cites.
Multi-Environment Pretraining Enables Transfer to Action Limited Datasets
David Venuto, Sherry Yang, Pieter Abbeel, Doina Precup, Igor Mordatch, and Ofir Nachum. 2022 · 2022
Later among the works it cites.
Chai: A chatbot ai for task-oriented dialogue with offline reinforcement learning
Siddharth Verma, Justin Fu, Mengjiao Yang, and Sergey Levine. 2022 · 2022
Later among the works it cites.
Phenaki: Variable length video generation from open domain textual description
Ruben Villegas, Mohammad Babaeizadeh, Pieter-Jan Kindermans, Hernan Moraldo, Han Zhang, Mohammad Taghi Saffar, Santiago Castro, Julius Kunze, and Dumitru Erhan. 2022 · 2022
Later among the works it cites.
Emergent abilities of large language models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. 2022b · 2022
Later among the works it cites.
Masked visual pre-training for motor control
Tete Xiao, Ilija Radosavovic, Trevor Darrell, and Jitendra Malik. 2022 · 2022
Later among the works it cites.
Chain of thought imitation with procedure cloning
Mengjiao Yang, Dale Schuurmans, Pieter Abbeel, and Ofir Nachum. 2022a · 2022
Later among the works it cites.
Dichotomy of control: Separating what you can control from what you cannot
Mengjiao Yang, Dale Schuurmans, Pieter Abbeel, and Ofir Nachum. 2022b · 2022
Later among the works it cites.
ReAct: Synergizing Reasoning and Acting in Language Models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2022 · 2022
Later among the works it cites.
Socratic models: Composing zero-shot multimodal reasoning with language
Andy Zeng, Adrian Wong, Stefan Welker, Krzysztof Choromanski, Federico Tombari, Aveek Purohit, Michael Ryoo, Vikas Sindhwani, Johnny Lee, Vincent Vanhoucke, et al · 2022
Later among the works it cites.
Least-to-most prompting enables complex reasoning in large language models
Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Olivier Bousquet, Quoc Le, and Ed Chi. 2022 · 2022
Later among the works it cites.
Efficient Online Reinforcement Learning with Offline Data
Philip J Ball, Laura Smith, Ilya Kostrikov, and Sergey Levine. 2023 · 2023
Closest in time.
PaLM-E: An Embodied Multimodal Language Model. In arXiv preprint arXiv:2302.11111
Danny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, Wenlong Huang, Yevgen Chebotar, Pierre Sermanet, Daniel Duckworth, Sergey Levine, Vincent Vanhoucke, Karol Hausman, Marc Toussaint, Klaus Greff, Andy Zeng, Igor Mordatch, and Pete Florence. 2023 · 2023
Closest in time.
Guiding Pretraining in Reinforcement Learning with Large Language Models
Yuqing Du, Olivia Watkins, Zihan Wang, Cédric Colas, Trevor Darrell, Pieter Abbeel, Abhishek Gupta, and Jacob Andreas. 2023a · 2023
Closest in time.
Learning Universal Policies via Text-Guided Video Generation
Yilun Du, Mengjiao Yang, Bo Dai, Hanjun Dai, Ofir Nachum, Joshua B Tenenbaum, Dale Schuurmans, and Pieter Abbeel. 2023b · 2023
Closest in time.
Looped Transformers as Programmable Computers
Angeliki Giannou, Shashank Rajput, Jy-yong Sohn, Kangwook Lee, Jason D Lee, and Dimitris Papailiopoulos. 2023 · 2023
Closest in time.
Languages are Rewards: Hindsight Finetuning using Human Feedback
Hao Liu, Carmelo Sferrazza, and Pieter Abbeel. 2023a · 2023
Closest in time.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2023b · 2023
Closest in time.
Toolformer: Language Models Can Teach Themselves to Use Tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023 · 2023
Closest in time.
Memory Augmented Large Language Models are Computationally Universal
Dale Schuurmans. 2023 · 2023
Closest in time.
Large Language Models Can Be Easily Distracted by Irrelevant Context
Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed Chi, Nathanael Schärli, and Denny Zhou. 2023 · 2023
Closest in time.
Zihao Wang, Shaofei Cai, Anji Liu, Xiaojian Ma, and Yitao Liang. 2023 · 2023
Closest in time.