Fetching the paper…
Reading the bibliography…
This paper investigates robot manipulation based on human instruction with ambiguous requests.
C. Finucane, G. Jing, and H. Kress-Gazit, “Ltlmop: Experimenting with language, temporal logic and robot control,” in IROS . IEEE, 2010, pp. 1988–1993
1993
Earlier work this paper cites.
H. Kress-Gazit, G. E. Fainekos, and G. J. Pappas, “From structured english to robot motion,” in IROS . IEEE, 2007, pp. 2717–2722
2007
Earlier work this paper cites.
S. Tellex, T. Kollar, S. Dickerson, M. Walter, A. Banerjee, S. Teller, and N. Roy, “Understanding natural language commands for robotic navigation and mobile manipulation,” in AAAI , vol. 25, no. 1, 2011, pp. 1507–1514
2011
Earlier work this paper cites.
M. Tenorth and M. Beetz, “Knowrob: A knowledge processing infrastructure for cognition-enabled robots,” IJRR , vol. 32, no. 5, pp. 566–590, 2013
2013
Earlier work this paper cites.
S. Guadarrama, L. Riano, D. Golland, D. Go, Y. Jia, D. Klein, P. Abbeel, T. Darrell et al. , “Grounding spatial relations for human-robot interaction,” in IROS . IEEE, 2013, pp. 1640–1647
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
J. Donahue, Y. Jia, O. Vinyals, J. Hoffman, N. Zhang, E. Tzeng, and T. Darrell, “Decaf: A deep convolutional activation feature for generic visual recognition,” in ICML . PMLR, 2014, pp. 647–655
2014
Earlier work this paper cites.
R. Girshick, J. Donahue, T. Darrell, and J. Malik, “Rich feature hierarchies for accurate object detection and semantic segmentation,” in IEEE CVPR , 2014, pp. 580–587
2014
Earlier work this paper cites.
B. Hariharan, P. Arbeláez, R. Girshick, and J. Malik, “Simultaneous detection and segmentation,” in ECCV . Springer, 2014, pp. 297–312
2014
Earlier work this paper cites.
J. Pennington, R. Socher, and C. D. Manning, “Glove: Global vectors for word representation,” in EMNLP , 2014, pp. 1532–1543
2014
Earlier work this paper cites.
K. Cho, B. Van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase representations using RNN encoder-decoder for statistical machine translation,” in EMNLP . ACL, 2014, pp. 1724–1734
2014
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” NIPS , vol. 28, 2015
2015
Earlier work this paper cites.
J. Long, E. Shelhamer, and T. Darrell, “Fully convolutional networks for semantic segmentation,” in IEEE CVPR , 2015, pp. 3431–3440
2015
Earlier work this paper cites.
R. Kiros, Y. Zhu, R. R. Salakhutdinov, R. Zemel, R. Urtasun, A. Torralba, and S. Fidler, “Skip-thought vectors,” NIPS , vol. 28, 2015
2015
Earlier work this paper cites.
A. Antunes, L. Jamone, G. Saponaro, A. Bernardino, and R. Ventura, “From human instructions to robot actions: Formulation of goals, affordances and probabilistic planning,” in ICRA . IEEE, 2016, pp. 5449–5454
2016
Earlier work this paper cites.
D. K. Misra, J. Sung, K. Lee, and A. Saxena, “Tell me dave: Context-sensitive grounding of natural language to manipulation instructions,” IJRR , vol. 35, no. 1-3, pp. 281–300, 2016
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in IEEE CVPR , 2016, pp. 770–778
2016
Earlier work this paper cites.
E. Kolve, R. Mottaghi, W. Han, E. VanderBilt, L. Weihs, A. Herrasti, D. Gordon, Y. Zhu, A. Gupta, and A. Farhadi, “AI2-THOR: An Interactive 3D Environment for Visual AI,” arXiv , 2017
2017
Cited alongside, same era.
P. Lindes, A. Mininger, J. R. Kirk, and J. E. Laird, “Grounding language for interactive task learning,” in Proceedings of the First Workshop on Language Grounding for Robotics , 2017, pp. 1–9
2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” NIPS , vol. 30, 2017
2017
Cited alongside, same era.
K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask r-cnn,” in ICCV , 2017, pp. 2961–2969
2017
Cited alongside, same era.
K. M. Dipendra, B. Andrew, B. Valts, N. Eyvind, S. Max, and A. Yoav, “Mapping instructions to actions in 3d environments with visual goal prediction,” in EMNLP . Association for Computational Linguistics, 2018, pp. 2667–2678
N. Reimers and I. Gurevych, “Sentence-bert: Sentence embeddings using siamese bert-networks,” in EMNLP . Association for Computational Linguistics, 11 2019
2019
Later among the works it cites.
M. Nazarczuk and K. Mikolajczyk, “V2a-vision to action: Learning robotic arm actions based on vision and language,” in ACCV , 2020
2020
Later among the works it cites.
M. Shridhar, J. Thomason, D. Gordon, Y. Bisk, W. Han, R. Mottaghi, L. Zettlemoyer, and D. Fox, “Alfred: A benchmark for interpreting grounded instructions for everyday tasks,” in IEEE CVPR , 2020, pp. 10 740–10 749
2020
Later among the works it cites.
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang, “Bottom-up and top-down attention for image captioning and visual question answering,” in IEEE CVPR , 2018, pp. 6077–6086
2018
Cited alongside, same era.
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever, “Improving language understanding by generative pre-training,” 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
R. Zellers, M. Yatskar, S. Thomson, and Y. Choi, “Neural motifs: Scene graph parsing with global context,” in CVPR , 2018, pp. 5831–5840
2018
Cited alongside, same era.
P. Pramanick, C. Sarkar, and I. Bhattacharya, “Your instruction may be crisp, but not clear to me!” in IEEE RO-MAN , 2019, pp. 1–8
2019
Cited alongside, same era.
F.-J. Chu, R. Xu, L. Seguin, and P. A. Vela, “Toward affordance detection and ranking on novel objects for real-world robotic manipulation,” RA-L , vol. 4, no. 4, pp. 4070–4077, 2019
2019
Cited alongside, same era.
C. Sun, A. Myers, C. Vondrick, K. Murphy, and C. Schmid, “Videobert: A joint model for video and language representation learning,” in IEEE ICCV , 2019, pp. 7464–7473
2019
Cited alongside, same era.
2020
Later among the works it cites.
S. Tellex, N. Gopalan, H. Kress-Gazit, and C. Matuszek, “Robots that use language: A survey,” 2020
2020
Later among the works it cites.
H. Jiang, I. Misra, M. Rohrbach, E. Learned-Miller, and X. Chen, “In defense of grid features for visual question answering,” in IEEE CVPR , 2020, pp. 10 267–10 276
2020
Later among the works it cites.
2020
Later among the works it cites.
L. Liu, X. Liu, J. Gao, W. Chen, and J. Han, “Understanding the difficulty of training transformers,” in EMNLP . Association for Computational Linguistics, 2020, pp. 5747–5763
2020
Later among the works it cites.
K. P. Singh, S. Bhambri, B. Kim, R. Mottaghi, and J. Choi, “Factorizing perception and policy for interactive instruction following,” in IEEE ICCV , 2021, pp. 1888–1897
2021
Later among the works it cites.
2021
Later among the works it cites.
A. Mogadala, M. Kalimuthu, and D. Klakow, “Trends in integration of vision and language research: A survey of tasks, datasets, and methods,” JAIR , vol. 71, pp. 1183–1317, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
P. Damodaran, “Parrot: Paraphrase generation for nlu.” 2021
2021
Later among the works it cites.
V. Blukis, C. Paxton, D. Fox, A. Garg, and Y. Artzi, “A persistent spatial semantic representation for high-level natural language instruction execution,” in CoRL . PMLR, 2022, pp. 706–717
2022
Closest in time.
R. Xu, H. Chen, Y. Lin, and P. A. Vela, “IVALAb: Vision and Language Symbolic Goal Learning,” https://github.com/ivalab/mmf , 2022
2022
Closest in time.