Fetching the paper…
Reading the bibliography…
Current NLP techniques have been greatly applied in different domains.
Retrieval, analogy, and composition: A framework for compositional generalization in image captioning
Zhan Shi, Hui Liu, Martin Renqiang Min, Christopher Malon, Li Erran Li, and Xiaodan Zhu. 2021 · 2000
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2015 · 2015
Earlier work this paper cites.
Natural language object retrieval
Ronghang Hu, Huazhe Xu, Marcus Rohrbach, Jiashi Feng, Kate Saenko, and Trevor Darrell. 2016 · 2016
Earlier work this paper cites.
Faster r-cnn: towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Learning multi-modal grounded linguistic semantics by playing" i spy"
Jesse Thomason, Jivko Sinapov, Maxwell Svetlik, Peter Stone, and Raymond J Mooney. 2016 · 2016
Earlier work this paper cites.
Self-critical sequence training for image captioning
Steven J Rennie, Etienne Marcheret, Youssef Mroueh, Jerret Ross, and Vaibhava Goel. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Scene graph generation by iterative message passing
Danfei Xu, Yuke Zhu, Christopher B Choy, and Li Fei-Fei. 2017 · 2017
Earlier work this paper cites.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sünderhauf, Ian Reid, Stephen Gould, and Anton Van Den Hengel. 2018 · 2018
Earlier work this paper cites.
Monotonic chunkwise attention
Chung-Cheng Chiu and Colin Raffel. 2018 · 2018
Earlier work this paper cites.
Real-world multiobject, multigrasp detection
Fu-Jen Chu, Ruinian Xu, and Patricio A Vela. 2018 · 2018
Earlier work this paper cites.
Yolov3: An incremental improvement
Joseph Redmon and Ali Farhadi. 2018 · 2018
Earlier work this paper cites.
Modeling relational data with graph convolutional networks
Michael Schlichtkrull, Thomas N Kipf, Peter Bloem, Rianne Van Den Berg, Ivan Titov, and Max Welling. 2018 · 2018
Earlier work this paper cites.
Learning from implicit information in natural language instructions for robotic manipulations
Ozan Arkan Can, Pedro Zuidberg Dos Martires, Andreas Persson, Julian Gaal, Amy Loutfi, Luc De Raedt, Deniz Yuret, and Alessandro Saffiotti. 2019 · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Learning ambidextrous robot grasping policies
Jeffrey Mahler, Matthew Matl, Vishal Satish, Michael Danielczuk, Bill DeRose, Stephen McKinley, and Ken Goldberg. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Improving robot success detection using static object data
Rosario Scalise, Jesse Thomason, Yonatan Bisk, and Siddhartha Srinivasa. 2019 · 2019
Cited alongside, same era.
Interactive visual grounding of referring expressions for human-robot interaction
Mohit Shridhar and David Hsu. 2020 · 2020
Later among the works it cites.
Ai sensing for robotics using deep learning based visual and language modeling
Yuvaram Singh and Kameshwar Rao JV. 2020 · 2020
Later among the works it cites.
Bridge the gap: High-level semantic planning for image captioning
Chenxi Yuan, Yang Bai, and Chun Yuan. 2020 · 2020
Later among the works it cites.
Avplug: Approach vector planning for unicontact grasping amid clutter
Yahav Avigal, Vishal Satish, Zachary Tam, Huang Huang, Harry Zhang, Michael Danielczuk, Jeffrey Ichnowski, and Ken Goldberg. 2021 · 2021
Later among the works it cites.
A joint network for grasp detection conditioned on natural language commands
Yiye Chen, Ruinian Xu, Yunzhi Lin, and Patricio A Vela. 2021b · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhengyuan Yang, Boqing Gong, Liwei Wang, Wenbing Huang, Dong Yu, and Jiebo Luo. 2019 · 2019
Cited alongside, same era.
A multi-task convolutional neural network for autonomous robotic grasping in object stacking scenes
Hanbo Zhang, Xuguang Lan, Site Bai, Lipeng Wan, Chenjie Yang, and Nanning Zheng. 2019a · 2019
Cited alongside, same era.
Roi-based robotic grasp detection for object overlapping scenes
Hanbo Zhang, Xuguang Lan, Site Bai, Xinwen Zhou, Zhiqiang Tian, and Nanning Zheng. 2019b · 2019
Cited alongside, same era.
Meshed-memory transformer for image captioning
Marcella Cornia, Matteo Stefanini, Lorenzo Baraldi, and Rita Cucchiara. 2020 · 2020
Cited alongside, same era.
Room-across-room: Multilingual vision-and-language navigation with dense spatiotemporal grounding
Alexander Ku, Peter Anderson, Roma Patel, Eugene Ie, and Jason Baldridge. 2020 · 2020
Cited alongside, same era.
Antipodal robotic grasping using generative residual convolutional neural network
Sulabh Kumra, Shirin Joshi, and Ferat Sahin. 2020 · 2020
Cited alongside, same era.
Oscar: Object-semantics aligned pre-training for vision-language tasks
Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et al. 2020 · 2020
Cited alongside, same era.
Sonia Raychaudhuri, Saim Wani, Shivansh Patel, Unnat Jain, and Angel X Chang. 2021 · 2021
Later among the works it cites.
Behavior: Benchmark for everyday household activities in virtual, interactive, and ecological environments
Sanjana Srivastava, Chengshu Li, Michael Lingelbach, Roberto Martín-Martín, Fei Xia, Kent Elliott Vainio, Zheng Lian, Cem Gokmen, Shyamal Buch, Karen Liu, et al. 2021 · 2021
Later among the works it cites.
Ocid-ref: A 3d robotic dataset with embodied language for clutter scene grounding
Ke-Jyun Wang, Yun-Hsuan Liu, Hung-Ting Su, Jen-Wei Wang, Yu-Siang Wang, Winston Hsu, and Wen-Chin Chen. 2021 · 2021
Later among the works it cites.
Piglet: Language grounding through neuro-symbolic interaction in a 3d world
Rowan Zellers, Ari Holtzman, Matthew Peters, Roozbeh Mottaghi, Aniruddha Kembhavi, Ali Farhadi, and Yejin Choi. 2021 · 2021
Later among the works it cites.
Hierarchical planning for long-horizon manipulation with geometric and symbolic scene graphs
Yifeng Zhu, Jonathan Tremblay, Stan Birchfield, and Yuke Zhu. 2021 · 2021
Later among the works it cites.
Do as i can, not as i say: Grounding language in robotic affordances
Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, et al. 2022 · 2022
Closest in time.
A rationale-centric framework for human-in-the-loop machine learning
Jinghui Lu, Linyi Yang, Brian Namee, and Yue Zhang. 2022 · 2022
Closest in time.
In situ bidirectional human-robot value alignment
Luyao Yuan, Xiaofeng Gao, Zilong Zheng, Mark Edmonds, Ying Nian Wu, Federico Rossano, Hongjing Lu, Yixin Zhu, and Song-Chun Zhu. 2022 · 2022
Closest in time.