Fetching the paper…
Reading the bibliography…
Language is compositional; an instruction can express multiple relation constraints to hold among objects in a scene that a robot is tasked to rearrange.
Skew-fit: State-covering self-supervised reinforcement learning
Vitchyr H. Pong, Murtaza Dalal, Steven Lin, Ashvin Nair, Shikhar Bahl, and Sergey Levine · 1903
Earlier work this paper cites.
Model based planning with energy based models
Yilun Du, Toru Lin, and Igor Mordatch · 1909
Earlier work this paper cites.
Danfei Xu, Roberto Martín-Martín, De-An Huang, Yuke Zhu, Silvio Savarese, and Li Fei-Fei · 1909
Earlier work this paper cites.
Object-centric task and motion planning in dynamic environments
Toki Migimatsu and Jeannette Bohg · 1911
Earlier work this paper cites.
Your classifier is secretly an energy based model and you should treat it like one
Will Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud, Mohammad Norouzi, and Kevin Swersky · 1912
Earlier work this paper cites.
Optimization by simulated annealing: Quantitative studies
Scott Kirkpatrick · 1984
Earlier work this paper cites.
Language as a cognitive tool to imagine goals in curiosity-driven exploration
Cédric Colas, Tristan Karch, Nicolas Lair, Jean-Michel Dussoux, Clément Moulin-Frier, Peter Ford Dominey, and Pierre-Yves Oudeyer · 2002
Earlier work this paper cites.
Language models are few-shot learners, 2020
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2005
Earlier work this paper cites.
Hierarchical task and motion planning in the now
Leslie Pack Kaelbling and Tomás Lozano-Pérez · 2011
Earlier work this paper cites.
Interactive furniture layout using interior design guidelines
Paul Merrell, Eric Schkufza, Zeyang Li, Maneesh Agrawala, and Vladlen Koltun · 2011
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
Max Welling and Yee W Teh · 2011
Earlier work this paper cites.
Make it home: automatic optimization of furniture arrangement
Lap Fai Yu, Sai Kit Yeung, Chi Keung Tang, Demetri Terzopoulos, Tony F Chan, and Stanley J Osher · 2011
Earlier work this paper cites.
Improved contrastive divergence training of energy based models
Yilun Du, Shuang Li, Joshua B. Tenenbaum, and Igor Mordatch · 2012
Earlier work this paper cites.
Example-based synthesis of 3d object arrangements
Matthew Fisher, Daniel Ritchie, Manolis Savva, Thomas Funkhouser, and Pat Hanrahan · 2012
Earlier work this paper cites.
Hierarchical planning for long-horizon manipulation with geometric and symbolic scene graphs
Yifeng Zhu, Jonathan Tremblay, Stan Birchfield, and Yuke Zhu · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Microsoft COCO: common objects in context
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, Lubomir D. Bourdev, Ross B. Girshick, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick · 2014
Earlier work this paper cites.
Generating images from captions with attention, 2015
Elman Mansimov, Emilio Parisotto, Jimmy Lei Ba, and Ruslan Salakhutdinov · 2015
Earlier work this paper cites.
Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models
Bryan A. Plummer, Liwei Wang, Christopher M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, and Svetlana Lazebnik · 2015
Earlier work this paper cites.
Logic-geometric programming: An optimization-based approach to combined task and motion planning
Marc Toussaint · 2015
Earlier work this paper cites.
Building a semantic parser overnight
Yushi Wang, Jonathan Berant, and Percy Liang · 2015
Earlier work this paper cites.
Language to logical form with neural attention
Li Dong and Mirella Lapata · 2016
Earlier work this paper cites.
Incorporating copying mechanism in sequence-to-sequence learning
Jiatao Gu, Zhengdong Lu, Hang Li, and Victor OK Li · 2016
Cited alongside, same era.
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei-Fei · 2016
Cited alongside, same era.
A theory of generative convnet
Jianwen Xie, Yang Lu, Song-Chun Zhu, and Ying Nian Wu · 2016
Cited alongside, same era.
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick · 2017
Cited alongside, same era.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick · 2017
Cited alongside, same era.
Grounding Language to Autonomously-Acquired Skills via Goal Generation
Ahmed Akakzia, Cédric Colas, Pierre-Yves Oudeyer, Mohamed Chetouani, and Olivier Sigaud · 2021
Later among the works it cites.
Plasticinelab: A soft-body manipulation benchmark with differentiable physics
Zhiao Huang, Yuanming Hu, Tao Du, Siyuan Zhou, Hao Su, Joshua B Tenenbaum, and Chuang Gan · 2021
Later among the works it cites.
MDETR - modulated detection for end-to-end multi-modal understanding
Aishwarya Kamath, Mannat Singh, Yann LeCun, Ishan Misra, Gabriel Synnaeve, and Nicolas Carion · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Synthesizing dynamic patterns by spatial-temporal generative convnet
Jianwen Xie, Song-Chun Zhu, and Ying Nian Wu · 2017
Cited alongside, same era.
Stripstream: Integrating symbolic planners and blackbox samplers
Caelan Reed Garrett, Tomás Lozano-Pérez, and Leslie Pack Kaelbling · 2018
Cited alongside, same era.
Image generation from scene graphs
Justin Johnson, Agrim Gupta, and Li Fei-Fei · 2018
Cited alongside, same era.
SDRL: interpretable and data-efficient deep reinforcement learning leveraging symbolic planning
Daoming Lyu, Fangkai Yang, Bo Liu, and Steven Gustafson · 2018
Cited alongside, same era.
Concept learning with energy-based models
Igor Mordatch · 2018
Cited alongside, same era.
Visual reinforcement learning with imagined goals
Ashvin Nair, Vitchyr Pong, Murtaza Dalal, Shikhar Bahl, Steven Lin, and Sergey Levine · 2018
Cited alongside, same era.
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever · 2021
Later among the works it cites.
Learning to rearrange deformable cables, fabrics, and bags with goal-conditioned transporter networks
Daniel Seita, Pete Florence, Jonathan Tompson, Erwin Coumans, Vikas Sindhwani, Ken Goldberg, and Andy Zeng · 2021
Later among the works it cites.
Cliport: What and where pathways for robotic manipulation
Mohit Shridhar, Lucas Manuelli, and Dieter Fox · 2021
Later among the works it cites.
Hyperdynamics: Meta-learning object and agent dynamics with hypernetworks
Zhou Xian, Shamit Lal, Hsiao-Yu Tung, Emmanouil Antonios Platanios, and Katerina Fragkiadaki · 2021
Later among the works it cites.
Cédric Colas, Tristan Karch, Clément Moulin-Frier, and Pierre-Yves Oudeyer · 2022
Later among the works it cites.
Bottom up top down detection transformers for language grounding in images and point clouds
Ayush Jain, Nikolaos Gkanatsios, Ishita Mediratta, and Katerina Fragkiadaki · 2022
Later among the works it cites.
Dall-e-bot: Introducing web-scale diffusion models to robotics
Ivan Kapelyukh, Vitalis Vosylius, and Edward Johns · 2022
Later among the works it cites.
Code as policies: Language model programs for embodied control, 2022
Jacky Liang, Wenlong Huang, Fei Xia, Peng Xu, Karol Hausman, Brian Ichter, Pete Florence, and Andy Zeng · 2022
Later among the works it cites.
On grounded planning for embodied tasks with language models, 2022
Bill Yuchen Lin, Chengsong Huang, Qian Liu, Wenda Gu, Sam Sommerer, and Xiang Ren · 2022
Later among the works it cites.
kpam: Keypoint affordances for category-level robotic manipulation
Lucas Manuelli, Wei Gao, Peter Florence, and Russ Tedrake · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Later among the works it cites.
Photorealistic text-to-image diffusion models with deep language understanding, 2022
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J Fleet, and Mohammad Norouzi · 2022
Later among the works it cites.
Guiding multi-step rearrangement tasks with natural language instructions
Elias Stengel-Eskin, Andrew Hundt, Zhuohong He, Aditya Murali, Nakul Gopalan, Matthew Gombolay, and Gregory Hager · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed H. Chi, Quoc Le, and Denny Zhou · 2022
Later among the works it cites.
Transporters with visual foresight for solving unseen rearrangement tasks, 2022
Hongtao Wu, Jikai Ye, Xin Meng, Chris Paxton, and Gregory Chirikjian · 2022
Later among the works it cites.
Scaling autoregressive models for content-rich text-to-image generation, 2022
Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gunjan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yinfei Yang, Burcu Karagol Ayan, Ben Hutchinson, Wei Han, Zarana Parekh, Xin Li, Han Zhang, Jason Baldridge, and Yonghui Wu · 2022
Later among the works it cites.
Fluidlab: A differentiable environment for benchmarking complex fluid manipulation
Zhou Xian, Bo Zhu, Zhenjia Xu, Hsiao-Yu Tung, Antonio Torralba, Katerina Fragkiadaki, and Chuang Gan · 2023
Closest in time.