Fetching the paper…
Reading the bibliography…
Recent works have shown how the reasoning capabilities of Large Language Models (LLMs) can be applied to domains beyond natural language processing, such as planning and interaction for robots.
Once for all: Train one network and specialize it for efficient deployment
H. Cai, C. Gan, T. Wang, Z. Zhang, and S. Han · 1908
Earlier work this paper cites.
Play and its role in the mental development of the child
L. S. Vygotsky · 1967
Earlier work this paper cites.
Strips: A new approach to the application of theorem proving to problem solving
R. E. Fikes and N. J. Nilsson · 1971
Earlier work this paper cites.
A structure for plans and behavior
E. D. Sacerdoti · 1975
Earlier work this paper cites.
Tool and symbol in child development
L. Vygotsky · 1994
Earlier work this paper cites.
Thinking in language?: evolution and a modularist possibility
P. Carruthers · 1998
Earlier work this paper cites.
Shop: Simple hierarchical ordered planner
D. Nau, Y. Cao, A. Lotem, and H. Munoz-Avila · 1999
Earlier work this paper cites.
Recent advances in hierarchical reinforcement learning
A. G. Barto and S. Mahadevan · 2003
Earlier work this paper cites.
Language conditioned imitation learning over unstructured data
C. Lynch and P. Sermanet · 2005
Earlier work this paper cites.
Planning algorithms
S. M. LaValle · 2006
Earlier work this paper cites.
Hierarchical planning in the now
L. P. Kaelbling and T. Lozano-Pérez · 2010
Earlier work this paper cites.
Toward understanding natural language directions
T. Kollar, S. Tellex, D. Roy, and N. Roy · 2010
Earlier work this paper cites.
Understanding natural language commands for robotic navigation and mobile manipulation
S. Tellex, T. Kollar, S. Dickerson, M. Walter, A. Banerjee, S. Teller, and N. Roy · 2011
Earlier work this paper cites.
Thought and language
L. S. Vygotsky · 2012
Earlier work this paper cites.
Integrated task and motion planning in belief space
L. P. Kaelbling and T. Lozano-Pérez · 2013
Earlier work this paper cites.
Interpreting and executing recipes with a cooking robot
M. Bollini, S. Tellex, T. Thompson, N. Roy, and D. Rus · 2013
Earlier work this paper cites.
Combined task and motion planning through an extensible planner-independent interface layer
S. Srivastava, E. Fang, L. Riano, R. Chitnis, S. Russell, and P. Abbeel · 2014
Earlier work this paper cites.
Asking for help using inverse semantics
S. Tellex, R. Knepper, A. Li, D. Rus, and N. Roy · 2014
Earlier work this paper cites.
Grounding verbs of motion in natural language commands to robots
T. Kollar, S. Tellex, D. Roy, and N. Roy · 2014
Earlier work this paper cites.
Logic-geometric programming: An optimization-based approach to combined task and motion planning
M. Toussaint · 2015
Earlier work this paper cites.
Deep learning for detecting robotic grasps
I. Lenz, H. Lee, and A. Saxena · 2015
Earlier work this paper cites.
Recurrent convolutional neural network for object recognition
M. Liang and X. Hu · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Earlier work this paper cites.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Differentiable physics and stable modes for tool-use and manipulation planning
M. A. Toussaint, K. R. Allen, K. A. Smith, and J. B. Tenenbaum · 2018
Earlier work this paper cites.
Neural task programming: Learning to generalize across hierarchical tasks
D. Xu, S. Nair, Y. Zhu, J. Gao, A. Garg, L. Fei-Fei, and S. Savarese · 2018
Earlier work this paper cites.
Universal planning networks: Learning generalizable representations for visuomotor control
A. Srinivas, A. Jabri, P. Abbeel, S. Levine, and C. Finn · 2018
Earlier work this paper cites.
Learning plannable representations with causal infogan
T. Kurutach, A. Tamar, G. Yang, S. J. Russell, and P. Abbeel · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Real-world multiobject, multigrasp detection
F.-J. Chu, R. Xu, and P. A. Vela · 2018
Earlier work this paper cites.
Language models as knowledge bases?
F. Petroni, T. Rocktäschel, P. Lewis, A. Bakhtin, Y. Wu, A. H. Miller, and S. Riedel · 2019
Earlier work this paper cites.
Commonsense knowledge mining from pretrained models
J. Davison, J. Feldman, and A. M. Rush · 2019
Cited alongside, same era.
Search on the replay buffer: Bridging planning and reinforcement learning
B. Eysenbach, R. R. Salakhutdinov, and S. Levine · 2019
Cited alongside, same era.
Regression planning networks
D. Xu, R. Martín-Martín, D.-A. Huang, Y. Zhu, S. Savarese, and L. F. Fei-Fei · 2019
Cited alongside, same era.
Learning to map natural language instructions to physical quadcopter control using simulated flight
V. Blukis, Y. Terme, E. Niklasson, R. A. Knepper, and Y. Artzi · 2019
Cited alongside, same era.
Language as an abstraction for hierarchical deep reinforcement learning
Y. Jiang, S. Gu, K. Murphy, and C. Finn · 2019
Cited alongside, same era.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. d. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al · 2021
Later among the works it cites.
Finetuned language models are zero-shot learners
J. Wei, M. Bosma, V. Y. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le · 2021
Later among the works it cites.
A joint network for grasp detection conditioned on natural language commands
Y. Chen, R. Xu, Y. Lin, and P. A. Vela · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Later among the works it cites.
Simvlm: Simple visual language model pretraining with weak supervision
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov · 2019
Cited alongside, same era.
Sentence-bert: Sentence embeddings using siamese bert-networks
N. Reimers and I. Gurevych · 2019
Cited alongside, same era.
Prospection: Interpretable plans from language by predicting the future
C. Paxton, Y. Bisk, J. Thomason, A. Byravan, and D. Foxl · 2019
Cited alongside, same era.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
J. Lu, D. Batra, D. Parikh, and S. Lee · 2019
Cited alongside, same era.
Zero-shot anticipation for instructional activities
F. Sener and A. Yao · 2019
Cited alongside, same era.
Object detection in 20 years: A survey
Z. Zou, Z. Shi, Y. Guo, and J. Ye · 2019
Cited alongside, same era.
How can we know what language models know?
Z. Jiang, F. F. Xu, J. Araki, and G. Neubig · 2020
Cited alongside, same era.
Z. Wang, J. Yu, A. W. Yu, Z. Dai, Y. Tsvetkov, and Y. Cao · 2021
Later among the works it cites.
Embodied bert: A transformer model for embodied, language-guided visual task completion
A. Suglia, Q. Gao, J. Thomason, G. Thattai, and G. Sukhatme · 2021
Later among the works it cites.
Mural: multimodal, multitask retrieval across languages
A. Jain, M. Guo, K. Srinivasan, T. Chen, S. Kudugunta, C. Jia, Y. Yang, and J. Baldridge · 2021
Later among the works it cites.
Open-vocabulary object detection via vision and language knowledge distillation
X. Gu, T.-Y. Lin, W. Kuo, and Y. Cui · 2021
Later among the works it cites.
Mt-opt: Continuous multi-task robotic reinforcement learning at scale
D. Kalashnikov, J. Varley, Y. Chebotar, B. Swanson, R. Jonschkowski, C. Finn, S. Levine, and K. Hausman · 2021
Later among the works it cites.
Grounding predicates through actions
T. Migimatsu and J. Bohg · 2021
Later among the works it cites.
Mdetr-modulated detection for end-to-end multi-modal understanding
A. Kamath, M. Singh, Y. LeCun, G. Synnaeve, I. Misra, and N. Carion · 2021
Later among the works it cites.
Palm: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al · 2022
Closest in time.
Chain of thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, E. Chi, Q. Le, and D. Zhou · 2022
Closest in time.
Large language models are zero-shot reasoners
T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa · 2022
Closest in time.
Can language models learn from explanations in context?
A. K. Lampinen, I. Dasgupta, S. C. Chan, K. Matthewson, M. H. Tessler, A. Creswell, J. L. McClelland, J. X. Wang, and F. Hill · 2022
Closest in time.
Vygotskian autotelic artificial intelligence: Language and culture internalization for human-like ai
C. Colas, T. Karch, C. Moulin-Frier, and P.-Y. Oudeyer · 2022
Closest in time.
Socratic models: Composing zero-shot multimodal reasoning with language
A. Zeng, A. Wong, S. Welker, K. Choromanski, F. Tombari, A. Purohit, M. Ryoo, V. Sindhwani, J. Lee, V. Vanhoucke, et al · 2022
Closest in time.
Language models as zero-shot planners: Extracting actionable knowledge for embodied agents
W. Huang, P. Abbeel, D. Pathak, and I. Mordatch · 2022
Closest in time.
Do as i can and not as i say: Grounding language in robotic affordances
M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, D. Ho, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, E. Jang, R. J. Ruano, K. Jeffrey, S. Jesmonth, N. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, K.-H. Lee, S. Levine, Y. Lu, L. Luu, C. Parada, P. Pastor, J. Quiambao, K. Rao, J. Rettinghouse, D. Reyes, P. Sermanet, N. Sievers, C. Tan, A. Toshev, V. Vanhoucke, F. Xia, T. Xiao, P. Xu, S. Xu, and M. Yan · 2022
Closest in time.
Inventing relational state and action abstractions for effective and efficient bilevel planning
T. Silver, R. Chitnis, N. Kumar, W. McClinton, T. Lozano-Perez, L. P. Kaelbling, and J. Tenenbaum · 2022
Closest in time.
Value function spaces: Skill-centric state abstractions for long-horizon reasoning
D. Shah, P. Xu, Y. Lu, T. Xiao, A. Toshev, S. Levine, and B. Ichter · 2022
Closest in time.
Deep hierarchical planning from pixels
D. Hafner, K.-H. Lee, I. Fischer, and P. Abbeel · 2022
Closest in time.
Pre-trained language models for interactive decision-making
S. Li, X. Puig, Y. Du, C. Wang, E. Akyurek, A. Torralba, J. Andreas, and I. Mordatch · 2022
Closest in time.
What matters in language conditioned robotic imitation learning
O. Mees, L. Hermann, and W. Burgard · 2022
Closest in time.
Intra-agent speech permits zero-shot task acquisition
C. Yan, F. Carnevale, P. Georgiev, A. Santoro, A. Guy, A. Muldal, C.-C. Hung, J. Abramson, T. Lillicrap, and G. Wayne · 2022
Closest in time.
Plate: Visually-grounded planning with transformers in procedural tasks
J. Sun, D.-A. Huang, B. Lu, Y.-H. Liu, B. Zhou, and A. Garg · 2022
Closest in time.
Simple but effective: Clip embeddings for embodied ai
A. Khandelwal, L. Weihs, R. Mottaghi, and A. Kembhavi · 2022
Closest in time.
Cliport: What and where pathways for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2022
Closest in time.
Can foundation models perform zero-shot task specification for robot manipulation?
Y. Cui, S. Niekum, A. Gupta, V. Kumar, and A. Rajeswaran · 2022
Closest in time.
Flamingo: a visual language model for few-shot learning
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al · 2022
Closest in time.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Closest in time.