Fetching the paper…
Reading the bibliography…
The ability to accurately locate and navigate to a specific object is a crucial capability for embodied agents that operate in the real world and interact with objects to complete tasks.
Distributed optimization and statistical learning via the alternating direction method of multipliers
Boyd, S., Parikh, N., Chu, E., Peleato, B., and Eckstein, J · 1935
Earlier work this paper cites.
A frontier-based approach for autonomous exploration
Yamauchi, B · 1997
Earlier work this paper cites.
Objectnav revisited: On evaluation of embodied agents navigating to objects
Batra, D., Gokaslan, A., Kembhavi, A., Maksymets, O., Mottaghi, R., Savva, M., Toshev, A., and Wijmans, E · 2006
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Hinge-loss markov random fields and probabilistic soft logic
Bach, S. H., Broecheler, M., Huang, B., and Getoor, L · 2017
Earlier work this paper cites.
Matterport3D: Learning from RGB-D data in indoor environments
Chang, A., Dai, A., Funkhouser, T., Halber, M., Niessner, M., Savva, M., Song, S., Zeng, A., and Zhang, Y · 2017
Earlier work this paper cites.
Mask r-cnn
He, K., Gkioxari, G., Dollár, P., and Girshick, R · 2017
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., and Batra, D · 2017
Earlier work this paper cites.
On evaluation of embodied navigation agents
Anderson, P., Chang, A. X., Chaplot, D. S., Dosovitskiy, A., Gupta, S., Koltun, V., Kosecka, J., Malik, J., Mottaghi, R., Savva, M., and Zamir, A. R · 2018
Earlier work this paper cites.
Visual semantic navigation using scene priors
Yang, W., Wang, X., Farhadi, A., Gupta, A., and Mottaghi, R · 2019
Earlier work this paper cites.
Integrating egocentric localization for more realistic point-goal navigation agents
Datta, S., Maksymets, O., Hoffman, J., Lee, S., Batra, D., and Parikh, D · 2020
Earlier work this paper cites.
Robothor: An open simulation-to-real embodied ai platform
Deitke, M., Han, W., Herrasti, A., Kembhavi, A., Kolve, E., Mottaghi, R., Salvador, J., Schwenk, D., VanderBilt, E., Wallingford, M., Weihs, L., Yatskar, M., and Farhadi, A · 2020
Earlier work this paper cites.
Debertav3: Improving deberta using electra-style pre-training with gradient-disentangled embedding sharing, 2021
He, P., Gao, J., and Chen, W · 2021
Earlier work this paper cites.
Thda: Treasure hunt data augmentation for semantic navigation
Maksymets, O., Cartillier, V., Gokaslan, A., Wijmans, E., Galuba, W., Lee, S., and Batra, D · 2021
Cited alongside, same era.
Memory-augmented reinforcement learning for image-goal navigation
Mezghani, L., Sukhbaatar, S., Lavril, T., Maksymets, O., Batra, D., Bojanowski, P., and Alahari, K · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Cited alongside, same era.
Habitat-matterport 3d dataset (HM3d): 1000 large-scale 3d environments for embodied AI
Ramakrishnan, S. K., Gokaslan, A., Wijmans, E., Maksymets, O., Clegg, A., Turner, J. M., Undersander, E., Galuba, W., Westbury, A., Chang, A. X., Savva, M., Zhao, Y., and Batra, D · 2021
Cited alongside, same era.
Large language models are zero-shot reasoners
Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y · 2022
Later among the works it cites.
Grounded language-image pre-training
Li*, L. H., Zhang*, P., Zhang*, H., Yang, J., Li, C., Zhong, Y., Wang, L., Yuan, L., Zhang, L., Hwang, J.-N., Chang, K.-W., and Gao, J · 2022
Later among the works it cites.
Neuro-symbolic causal language planning with commonsense prompting
Lu, Y., Feng, W., Zhu, W., Xu, W., Wang, X. E., Eckstein, M., and Wang, W. Y · 2022
Later among the works it cites.
ZSON: Zero-shot object-goal navigation using multimodal goal embeddings
Majumdar, A., Aggarwal, G., Devnani, B. S., Hoffman, J., and Batra, D · 2022
Later among the works it cites.
FILM: Following instructions in language with modular methods
Min, S. Y., Chaplot, D. S., Ravikumar, P. K., Bisk, Y., and Salakhutdinov, R · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shen, S., Li, L. H., Tan, H., Bansal, M., Rohrbach, A., Chang, K.-W., Yao, Z., and Keutzer, K · 2021
Cited alongside, same era.
Auxiliary tasks and exploration enable objectgoal navigation
Ye, J., Batra, D., Das, A., and Wijmans, E · 2021
Cited alongside, same era.
Semantic linking maps for active visual object search (extended abstract)
Zeng, Z., Röfer, A., and Jenkins, O. C · 2021
Cited alongside, same era.
Do as i can and not as i say: Grounding language in robotic affordances
Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Fu, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Ho, D., Hsu, J., Ibarz, J., Ichter, B., Irpan, A., Jang, E., Ruano, R. J., Jeffrey, K., Jesmonth, S., Joshi, N., Julian, R., Kalashnikov, D., Kuang, Y., Lee, K.-H., Levine, S., Lu, Y., Luu, L., Parada, C., Pastor, P., Quiambao, J., Rao, K., Rettinghouse, J., Reyes, D., Sermanet, P., Sievers, N., Tan, C., Toshev, A., Vanhoucke, V., Xia, F., Xiao, T., Xu, P., Xu, S., Yan, M., and Zeng, A · 2022
Cited alongside, same era.
Zero experience required: Plug & play modular transfer learning for semantic visual navigation
Al-Halah, Z., Ramakrishnan, S. K., and Grauman, K · 2022
Cited alongside, same era.
A persistent spatial semantic representation for high-level natural language instruction execution
Blukis, V., Paxton, C., Fox, D., Garg, A., and Artzi, Y · 2022
Cited alongside, same era.
Procthor: Large-scale embodied ai using procedural generation, 2022
Deitke, M., VanderBilt, E., Herrasti, A., Weihs, L., Salvador, J., Ehsani, K., Han, W., Kolve, E., Farhadi, A., Kembhavi, A., and Mottaghi, R · 2022
Cited alongside, same era.
Simple but effective: Clip embeddings for embodied ai
Khandelwal, A., Weihs, L., Mottaghi, R., and Kembhavi, A · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Gray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R · 2022
Later among the works it cites.
Poni: Potential functions for objectgoal navigation with interaction-free learning
Ramakrishnan, S. K., Chaplot, D. S., Al-Halah, Z., Malik, J., and Grauman, K · 2022
Later among the works it cites.
Habitat-web: Learning embodied object-search strategies from human demonstrations at scale
Ramrakhya, R., Undersander, E., Batra, D., and Das, A · 2022
Later among the works it cites.
Tidee: Tidying up novel rooms using visuo-semantic commonsense priors
Sarch, G., Fang, Z., Harley, A. W., Schydlo, P., Tarr, M. J., Gupta, S., and Fragkiadaki, K · 2022
Later among the works it cites.
Robotic Navigation with Large Pre-Trained Models of Language, Vision, and Action
Shah, D., Osinski, B., Ichter, B., and Levine, S · 2022
Later among the works it cites.
Skill induction and planning with latent language
Sharma, P., Torralba, A., and Andreas, J · 2022
Later among the works it cites.
Jarvis: A neuro-symbolic commonsense reasoning framework for conversational embodied agents
Zheng, K., Zhou, K., Gu, J., Fan, Y., Wang, J., Li, Z., He, X., and Wang, X. E · 2022
Later among the works it cites.
Cows on pasture: Baselines and benchmarks for language-driven zero-shot object navigation
Gadre, S. Y., Wortsman, M., Ilharco, G., Schmidt, L., and Song, S · 2023
Closest in time.