Fetching the paper…
Reading the bibliography…
Embodied AI is one of the most popular studies in artificial intelligence and robotics, which can effectively improve the intelligence of real-world agents (i.e.
S. Wang, X. Huang, C. Chen, L. Wu, and J. Li, “Reform: Error-aware few-shot knowledge graph completion,” in Proceedings of the 30th ACM International Conference on Information & Knowledge Management , 2021, pp. 1979–1988
1988
Earlier work this paper cites.
G. A. Miller, “Wordnet: a lexical database for english,” Communications of the ACM , vol. 38, no. 11, pp. 39–41, 1995
1995
Earlier work this paper cites.
M. C. Jensen and W. H. Heckling, “Specific and general knowledge, and organizational structure,” Journal of applied corporate finance , vol. 8, no. 2, pp. 4–18, 1995
1995
Earlier work this paper cites.
D. B. Lenat, “Cyc: A large-scale investment in knowledge infrastructure,” Communications of the ACM , vol. 38, no. 11, pp. 33–38, 1995
1995
Earlier work this paper cites.
O. Khatib, K. Yokoi, O. Brock, K. Chang, and A. Casal, “Robots in human environments: Basic autonomous capabilities,” The International Journal of Robotics Research , vol. 18, no. 7, pp. 684–696, 1999
1999
Earlier work this paper cites.
M. T. Mason, Mechanics of robotic manipulation . MIT press, 2001
2001
Earlier work this paper cites.
X. Tang, X. Gao, J. Liu, and H. Zhang, “A spatial-temporal approach for video caption detection and recognition,” IEEE transactions on neural networks , vol. 13, no. 4, pp. 961–971, 2002
2002
Earlier work this paper cites.
J. M. Henderson, “Human gaze control during real-world scene perception,” Trends in cognitive sciences , vol. 7, no. 11, pp. 498–504, 2003
2003
Earlier work this paper cites.
H. Liu and P. Singh, “Conceptnet—a practical commonsense reasoning tool-kit,” BT technology journal , vol. 22, no. 4, pp. 211–226, 2004
2004
Earlier work this paper cites.
K. Patterson, P. J. Nestor, and T. T. Rogers, “Where do you know what you know? the representation of semantic knowledge in the human brain,” Nature reviews neuroscience , vol. 8, no. 12, pp. 976–987, 2007
2007
Earlier work this paper cites.
B. Siciliano, O. Khatib, and T. Kröger, Springer handbook of robotics . Springer, 2008, vol. 200
2008
Earlier work this paper cites.
K. Bollacker, C. Evans, P. Paritosh, T. Sturge, and J. Taylor, “Freebase: a collaboratively created graph database for structuring human knowledge,” in Proceedings of the 2008 ACM SIGMOD international conference on Management of data , 2008, pp. 1247–1250
2008
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in IEEE Conference on Computer Vision and Pattern Recognition . IEEE, 2009, pp. 248–255
2009
Earlier work this paper cites.
M. Waibel, M. Beetz, J. Civera, R. d’Andrea, J. Elfring, D. Galvez-Lopez, K. Häussermann, R. Janssen, J. Montiel, A. Perzylo et al. , “Roboearth,” IEEE Robotics & Automation Magazine , vol. 18, no. 2, pp. 69–82, 2011
2011
Earlier work this paper cites.
W. Wu, H. Li, H. Wang, and K. Q. Zhu, “Probase: A probabilistic taxonomy for text understanding,” in Proceedings of the 2012 ACM SIGMOD international conference on management of data , 2012, pp. 481–492
2012
Earlier work this paper cites.
D. R. Hofstadter and E. Sander, Surfaces and essences: Analogy as the fuel and fire of thinking . Basic books, 2013
2013
Earlier work this paper cites.
D. Vrandečić and M. Krötzsch, “Wikidata: a free collaborative knowledgebase,” Communications of the ACM , vol. 57, no. 10, pp. 78–85, 2014
2014
Earlier work this paper cites.
N. Tandon, G. De Melo, F. Suchanek, and G. Weikum, “Webchild: Harvesting and organizing commonsense knowledge from the web,” in Proceedings of the 7th ACM international conference on Web search and data mining , 2014, pp. 523–532
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
Y. Zhu, A. Fathi, and L. Fei-Fei, “Reasoning about object affordances in a knowledge base representation,” in European conference on computer vision . Springer, 2014, pp. 408–424
2014
Earlier work this paper cites.
S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg, “Referitgame: Referring to objects in photographs of natural scenes,” in Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP) , 2014, pp. 787–798
2014
Earlier work this paper cites.
J. Johnson, R. Krishna, M. Stark, L.-J. Li, D. Shamma, M. Bernstein, and L. Fei-Fei, “Image retrieval using scene graphs,” in IEEE Conference on Computer Vision and Pattern Recognition , 2015, pp. 3668–3678
2015
Earlier work this paper cites.
H. Su, C. R. Qi, Y. Li, and L. J. Guibas, “Render for cnn: Viewpoint estimation in images using cnns trained with rendered 3d model views,” in IEEE International Conference on Computer Vision , 2015, pp. 2686–2694
2015
Earlier work this paper cites.
J. Lehmann, R. Isele, M. Jakob, A. Jentzsch, D. Kontokostas, P. N. Mendes, S. Hellmann, M. Morsey, P. Van Kleef, S. Auer et al. , “Dbpedia–a large-scale, multilingual knowledge base extracted from wikipedia,” Semantic web , vol. 6, no. 2, pp. 167–195, 2015
2015
Earlier work this paper cites.
M. Tenorth, J. Winkler, D. Beßler, and M. Beetz, “Open-ease: a cloud-based knowledge service for autonomous learning,” KI-Künstliche Intelligenz , vol. 29, pp. 407–411, 2015
2015
Earlier work this paper cites.
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik, “Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models,” in IEEE International Conference on Computer Vision , 2015, pp. 2641–2649
2015
Earlier work this paper cites.
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg, “Modeling context in referring expressions,” in European conference on computer vision . Springer, 2016, pp. 69–85
2016
Earlier work this paper cites.
J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, “You only look once: Unified, real-time object detection,” in IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 779–788
2016
Earlier work this paper cites.
S.-E. Wei, V. Ramakrishna, T. Kanade, and Y. Sheikh, “Convolutional pose machines,” in IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 4724–4732
2016
Earlier work this paper cites.
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li et al. , “Visual genome: Connecting language and vision using crowdsourced dense image annotations,” International journal of computer vision , vol. 123, no. 1, pp. 32–73, 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
J. Liu, G. Wang, L.-Y. Duan, K. Abdiyeva, and A. C. Kot, “Skeleton-based human action recognition with global context-aware attention lstm networks,” IEEE Transactions on Image Processing , vol. 27, no. 4, pp. 1586–1599, 2017
2017
Earlier work this paper cites.
D. Xu, Y. Zhu, C. B. Choy, and L. Fei-Fei, “Scene graph generation by iterative message passing,” in IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 5410–5419
2017
Earlier work this paper cites.
P. Anderson, Q. Wu, D. Teney, J. Bruce, M. Johnson, N. Sünderhauf, I. Reid, S. Gould, and A. Van Den Hengel, “Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments,” in IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 3674–3683
2018
Earlier work this paper cites.
H. Paulheim, “How much is a triple?” in IEEE International Semantic Web Conference , 2018
2018
Cited alongside, same era.
T. Mitchell, W. Cohen, E. Hruschka, P. Talukdar, B. Yang, J. Betteridge, A. Carlson, B. Dalvi, M. Gardner, B. Kisiel et al. , “Never-ending learning,” Communications of the ACM , vol. 61, no. 5, pp. 103–115, 2018
2018
Cited alongside, same era.
C.-Y. Chuang, J. Li, A. Torralba, and S. Fidler, “Learning to act properly: Predicting and explaining affordances from images,” in IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 975–983
2018
Cited alongside, same era.
Z. Li, X. Ding, and T. Liu, “Constructing narrative event evolutionary graph for script event prediction,” in the 27th International Joint Conference on Artificial Intelligence , 2018, pp. 4201–4207
2018
Cited alongside, same era.
2022
Later among the works it cites.
J. Thomason, M. Shridhar, Y. Bisk, C. Paxton, and L. Zettlemoyer, “Language grounding with 3d objects,” in Conference on Robot Learning . PMLR, 2022, pp. 1691–1701
2022
Later among the works it cites.
R. Corona, S. Zhu, D. Klein, and T. Darrell, “Voxel-informed language grounding,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , 2022, pp. 54–60
2022
Later among the works it cites.
R. Gao, Y.-Y. Chang, S. Mall, L. Fei-Fei, and J. Wu, “Objectfolder: A dataset of objects with implicit visual, auditory, and tactile representations,” in Conference on Robot Learning . PMLR, 2022, pp. 466–476
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
N. Zhang, S. Deng, Z. Sun, G. Wang, X. Chen, W. Zhang, and H. Chen, “Long-tail relation extraction via knowledge graph embeddings and graph convolution networks,” in Proceedings of NAACL-HLT , 2019, pp. 3016–3025
2019
Cited alongside, same era.
T. Ebisu and R. Ichise, “Generalized translation-based embedding of knowledge graph,” IEEE Transactions on Knowledge and Data Engineering , vol. 32, no. 5, pp. 941–951, 2019
2019
Cited alongside, same era.
J. Lu, D. Batra, D. Parikh, and S. Lee, “Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,” Advances in neural information processing systems , vol. 32, 2019
2019
Cited alongside, same era.
L. Ke, X. Li, Y. Bisk, A. Holtzman, Z. Gan, J. Liu, J. Gao, Y. Choi, and S. Srinivasa, “Tactical rewind: Self-correction via backtracking in vision-and-language navigation,” in IEEE Conference on Computer Vision and Pattern Recognition , 2019, pp. 6741–6749
2019
Cited alongside, same era.
I. Armeni, Z.-Y. He, J. Gwak, A. R. Zamir, M. Fischer, J. Malik, and S. Savarese, “3d scene graph: A structure for unified semantics, 3d space, and camera,” in IEEE/CVF International Conference on Computer Vision , 2019, pp. 5664–5673
2019
Cited alongside, same era.
U.-H. Kim, J.-M. Park, T.-J. Song, and J.-H. Kim, “3-d scene graph: A sparse and semantic representation of physical environments for intelligent agents,” IEEE transactions on cybernetics , vol. 50, no. 12, pp. 4921–4933, 2019
2019
Cited alongside, same era.
M. Sap, R. Le Bras, E. Allaway, C. Bhagavatula, N. Lourie, H. Rashkin, B. Roof, N. A. Smith, and Y. Choi, “Atomic: An atlas of machine commonsense for if-then reasoning,” in AAAI conference on artificial intelligence , vol. 33, no. 01, 2019, pp. 3027–3035
2019
Cited alongside, same era.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in North American Chapter of the Association for Computational Linguistics , 2019, pp. 4171–4186
2019
Cited alongside, same era.
R. Gao, Z. Si, Y.-Y. Chang, S. Clarke, J. Bohg, L. Fei-Fei, W. Yuan, and J. Wu, “Objectfolder 2.0: A multisensory object dataset for sim2real transfer,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 10 598–10 608
2022
Later among the works it cites.
A. Khandelwal, L. Weihs, R. Mottaghi, and A. Kembhavi, “Simple but effective: Clip embeddings for embodied ai,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 14 829–14 838
2022
Later among the works it cites.
X. Liang, F. Zhu, L. Lingling, H. Xu, and X. Liang, “Visual-language navigation pretraining via prompt-based environmental self-exploration,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2022, pp. 4837–4851
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
S. Nair, A. Rajeswaran, V. Kumar, C. Finn, and A. Gupta, “R3m: A universal visual representation for robot manipulation,” in 6th Annual Conference on Robot Learning , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
S. Deng, C. Wang, Z. Li, N. Zhang, Z. Dai, H. Chen, F. Xiong, M. Yan, Q. Chen, M. Chen et al. , “Construction and applications of billion-scale pre-trained multimodal business knowledge graph,” in 2023 IEEE 39th International Conference on Data Engineering (ICDE) . IEEE, 2023, pp. 2988–3002
2023
Closest in time.
W. Liu, A. Daruna, M. Patel, K. Ramachandruni, and S. Chernova, “A survey of semantic reasoning frameworks for robotic systems,” Robotics and Autonomous Systems , vol. 159, p. 104294, 2023
2023
Closest in time.
S. Vemprala, R. Bonatti, A. Bucker, and A. Kapoor, “Chatgpt for robotics: Design principles and model abilities,” Microsoft, Tech. Rep. MSR-TR-2023-8, February 2023.[Online]. Available … , 2023
2023
Closest in time.
“Chatgpt: Optimizing language models for dialogue,” https://openai.com/blog/chatgpt/
2023
Closest in time.
OpenAI, “Gpt-4 technical report,” 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
P. Liu, W. Yuan, J. Fu, Z. Jiang, H. Hayashi, and G. Neubig, “Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing,” ACM Computing Surveys , vol. 55, no. 9, pp. 1–35, 2023
2023
Closest in time.
S. Liang, “Knowledge graph embedding based on graph neural network,” in 2023 IEEE 39th International Conference on Data Engineering (ICDE) . IEEE, 2023, pp. 3908–3912
2023
Closest in time.
2023
Closest in time.
W. Liu, D. Bansal, A. Daruna, and S. Chernova, “Learning instance-level n-ary semantic knowledge at scale for robots operating in everyday environments,” Autonomous Robots , vol. 47, no. 5, pp. 529–547, 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto, “Stanford alpaca: An instruction-following llama model,” https://github.com/tatsu-lab/stanford_alpaca
2023
Closest in time.
2023
Closest in time.
Q. Ye, H. Xu, G. Xu, J. Ye, M. Yan, Y. Zhou, J. Wang, A. Hu, P. Shi, Y. Shi, C. Li, Y. Xu, H. Chen, J. Tian, Q. Qi, J. Zhang, and F. Huang, “mplug-owl: Modularization empowers large language models with multimodality,” 2023
2023
Closest in time.
2023
Closest in time.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhudinov, R. Zemel, and Y. Bengio, “Show, attend and tell: Neural image caption generation with visual attention,” in International Conference on Machine Learning . PMLR, 2015, pp. 2048–2057
2057
Closest in time.