Fetching the paper…
Reading the bibliography…
Humans interpret scenes by recognizing both the identities and positions of objects in their observations.
M. A. Goodale and A. D. Milner, “Separate visual pathways for perception and action,” Trends in neurosciences , vol. 15, no. 1, pp. 20–25, 1992
1992
Earlier work this paper cites.
L. G. Ungerleider and J. V. Haxby, “‘what’and ‘where’in the human brain,” Current opinion in neurobiology , vol. 4, no. 2, pp. 157–165, 1994
1994
Earlier work this paper cites.
E. H. de Haan and A. Cowey, “On the usefulness of ‘what’and ‘where’pathways in vision,” Trends in cognitive sciences , vol. 15, no. 10, pp. 460–466, 2011
2011
Earlier work this paper cites.
V. Ordonez, G. Kulkarni, and T. Berg, “Im2text: Describing images using 1 million captioned photographs,” Advances in neural information processing systems , vol. 24, 2011
2011
Earlier work this paper cites.
M. N. Hebart and G. Hesselmann, “What visual information is processed in the human dorsal stream?” Journal of Neuroscience , vol. 32, no. 24, pp. 8107–8109, 2012
2012
Earlier work this paper cites.
B. R. Sheth and R. Young, “Two visual pathways in primates based on sampling of space: exploitation and exploration of visual information,” Frontiers in integrative neuroscience , vol. 10, p. 37, 2016
2016
Earlier work this paper cites.
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg, “Modeling context in referring expressions,” in Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part II 14 . Springer, 2016, pp. 69–85
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
E. Freud, J. C. Culham, D. C. Plaut, and M. Behrmann, “The large-scale organization of shape processing in the ventral and dorsal pathways,” elife , vol. 6, p. e27576, 2017
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
D. Wang, C. Devin, Q.-Z. Cai, F. Yu, and T. Darrell, “Deep object-centric policies for autonomous driving,” in 2019 International Conference on Robotics and Automation (ICRA) . IEEE, 2019, pp. 8853–8859
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
R. Ma, L. Lam, B. A. Spiegel, A. Ganeshan, R. Patel, B. Abbatematteo, D. Paulius, S. Tellex, and G. Konidaris, “Skill generalization with verbs,” arXiv preprint , 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
M. Laskin, K. Lee, A. Stooke, L. Pinto, P. Abbeel, and A. Srinivas, “Reinforcement learning with augmented data,” Advances in neural information processing systems , vol. 33, pp. 19 884–19 895, 2020
2020
Earlier work this paper cites.
T. Migimatsu and J. Bohg, “Object-centric task and motion planning in dynamic environments,” IEEE Robotics and Automation Letters , vol. 5, no. 2, pp. 844–851, 2020
2020
Earlier work this paper cites.
F. Locatello, D. Weissenborn, T. Unterthiner, A. Mahendran, G. Heigold, J. Uszkoreit, A. Dosovitskiy, and T. Kipf, “Object-centric learning with slot attention,” Advances in Neural Information Processing Systems , vol. 33, pp. 11 525–11 538, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Z. Lin, Z. Zhang, L.-Z. Chen, M.-M. Cheng, and S.-P. Lu, “Interactive image segmentation with first click attention,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2020, pp. 13 339–13 348
2020
Earlier work this paper cites.
M. Hao, Y. Liu, X. Zhang, and J. Sun, “Labelenc: A new intermediate supervision method for object detection,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXV 16 . Springer, 2020, pp. 529–545
2020
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International conference on machine learning . PMLR, 2021, pp. 8748–8763
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
J. Shi, J. Qian, Y. J. Ma, and D. Jayaraman, “Plug-and-play object-centric representations from “what” and “where” foundation models,” arXiv preprint , 2021
2021
Earlier work this paper cites.
C. Wang, R. Wang, A. Mandlekar, L. Fei-Fei, S. Savarese, and D. Xu, “Generalization through hand-eye coordination: An action space for learning spatially-invariant visuomotor control,” in 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2021, pp. 8913–8920
2021
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou et al. , “Chain-of-thought prompting elicits reasoning in large language models,” Advances in Neural Information Processing Systems , vol. 35, pp. 24 824–24 837, 2022
C. Huang, O. Mees, A. Zeng, and W. Burgard, “Visual language maps for robot navigation,” in 2023 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2023, pp. 10 608–10 615
2023
Later among the works it cites.
J. Liang, W. Huang, F. Xia, P. Xu, K. Hausman, B. Ichter, P. Florence, and A. Zeng, “Code as policies: Language model programs for embodied control,” in 2023 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2023, pp. 9493–9500
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Cui, S. Karamcheti, R. Palleti, N. Shivakumar, P. Liang, and D. Sadigh, “No, to the right: Online language corrections for robotic manipulation via shared autonomy,” in Proceedings of the 2023 ACM/IEEE International Conference on Human-Robot Interaction , 2023, pp. 93–101
2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa, “Large language models are zero-shot reasoners,” Advances in neural information processing systems , vol. 35, pp. 22 199–22 213, 2022
2022
Cited alongside, same era.
L. Fan, G. Wang, Y. Jiang, A. Mandlekar, Y. Yang, H. Zhu, A. Tang, D.-A. Huang, Y. Zhu, and A. Anandkumar, “Minedojo: Building open-ended embodied agents with internet-scale knowledge,” Advances in Neural Information Processing Systems , vol. 35, pp. 18 343–18 362, 2022
2022
Cited alongside, same era.
M. Shridhar, L. Manuelli, and D. Fox, “Cliport: What and where pathways for robotic manipulation,” in Conference on Robot Learning . PMLR, 2022, pp. 894–906
2022
Cited alongside, same era.
W. Huang, P. Abbeel, D. Pathak, and I. Mordatch, “Language models as zero-shot planners: Extracting actionable knowledge for embodied agents,” in International Conference on Machine Learning . PMLR, 2022, pp. 9118–9147
2022
Cited alongside, same era.
2022
Cited alongside, same era.
E. Jang, A. Irpan, M. Khansari, D. Kappler, F. Ebert, C. Lynch, S. Levine, and C. Finn, “Bc-z: Zero-shot task generalization with robotic imitation learning,” in Conference on Robot Learning . PMLR, 2022, pp. 991–1002
2022
Cited alongside, same era.
2022
Cited alongside, same era.
S. Nair, E. Mitchell, K. Chen, S. Savarese, C. Finn et al. , “Learning language-conditioned robot behavior from offline data and crowd-sourced annotation,” in Conference on Robot Learning . PMLR, 2022, pp. 1303–1315
2022
Cited alongside, same era.
Later among the works it cites.
2023
Later among the works it cites.
M. Shridhar, L. Manuelli, and D. Fox, “Perceiver-actor: A multi-task transformer for robotic manipulation,” in Conference on Robot Learning . PMLR, 2023, pp. 785–799
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
B. Chen, F. Xia, B. Ichter, K. Rao, K. Gopalakrishnan, M. S. Ryoo, A. Stone, and D. Kappler, “Open-vocabulary queryable scene representations for real world planning,” in 2023 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2023, pp. 11 509–11 522
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Zhu, Z. Jiang, P. Stone, and Y. Zhu, “Learning generalizable manipulation policies with object-centric 3d representations,” in 7th Annual Conference on Robot Learning , 2023
2023
Later among the works it cites.
N. Heravi, A. Wahid, C. Lynch, P. Florence, T. Armstrong, J. Tompson, P. Sermanet, J. Bohg, and D. Dwibedi, “Visuomotor control in multi-object scenes using object-aware representations,” in 2023 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2023, pp. 9515–9522
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Q. Zhou and Y. Zhu, “Make a long image short: Adaptive token length for vision transformers,” in Joint European Conference on Machine Learning and Knowledge Discovery in Databases . Springer, 2023, pp. 69–85
2023
Later among the works it cites.
Y. Zhu, Q. Zhou, N. Liu, Z. Xu, Z. Ou, X. Mou, and J. Tang, “Scalekd: Distilling scale-aware knowledge in small object detector,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 19 723–19 733
2023
Later among the works it cites.