Fetching the paper…
Reading the bibliography…
We present PoliFormer (Policy Transformer), an RGB-only indoor navigation agent trained end-to-end with reinforcement learning at scale that generalizes to the real-world without adaptation despite being trained purely in simulation.
A frontier-based approach for autonomous exploration
B. Yamauchi · 1997
Earlier work this paper cites.
ObjectNav revisited: On evaluation of embodied agents navigating to objects
D. Batra, A. Gokaslan, A. Kembhavi, O. Maksymets, R. Mottaghi, M. Savva, A. Toshev, and E. Wijmans · 2006
Earlier work this paper cites.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. A. Riedmiller · 2013
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
K. Cho, B. van Merrienboer, Ç. Gülçehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
J. Chung, C. Gulcehre, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2015
Earlier work this paper cites.
C. Beattie, J. Z. Leibo, D. Teplyashin, T. Ward, M. Wainwright, H. Küttler, A. Lefrancq, S. Green, V. Valdés, A. Sadik, J. Schrittwieser, K. Anderson, S. York, M. Cant, A. Cain, A. Bolton, S. Gaffney, H. King, D. Hassabis, S. Legg, and S. Petersen · 2016
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
AI2-THOR: An Interactive 3D Environment for Visual AI
E. Kolve, R. Mottaghi, W. Han, E. VanderBilt, L. Weihs, A. Herrasti, M. Deitke, K. Ehsani, D. Gordon, Y. Zhu, A. Kembhavi, A. K. Gupta, and A. Farhadi · 2017
Earlier work this paper cites.
Robot learning in homes: Improving generalization and reducing dataset bias
A. Gupta, A. Murali, D. P. Gandhi, and L. Pinto · 2018
Earlier work this paper cites.
VirtualHome: Simulating Household Activities via Programs
X. Puig, K. Ra, M. Boben, J. Li, T. Wang, S. Fidler, and A. Torralba · 2018
Earlier work this paper cites.
Learning to walk via deep reinforcement learning
T. Haarnoja, A. Zhou, S. Ha, J. Tan, G. Tucker, and S. Levine · 2018
Earlier work this paper cites.
Visual semantic navigation using scene priors
W. Yang, X. Wang, A. Farhadi, A. K. Gupta, and R. Mottaghi · 2018
Earlier work this paper cites.
Neural predictive belief representations
Z. D. Guo, M. G. Azar, B. Piot, B. A. Pires, and R. Munos · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, et al · 2019
Earlier work this paper cites.
DD-PPO: learning near-perfect pointgoal navigators from 2.5 billion frames
E. Wijmans, A. Kadian, A. Morcos, S. Lee, I. Essa, D. Parikh, M. Savva, and D. Batra · 2020
Earlier work this paper cites.
MultiON: Benchmarking Semantic Map Memory using Multi-Object Navigation
S. Wani, S. Patel, U. Jain, A. X. Chang, and M. Savva · 2020
Earlier work this paper cites.
Learning to explore using active neural slam
D. S. Chaplot, D. Gandhi, S. Gupta, A. Gupta, and R. Salakhutdinov · 2020
Earlier work this paper cites.
Neural topological slam for visual navigation
D. S. Chaplot, R. Salakhutdinov, A. Gupta, and S. Gupta · 2020
Earlier work this paper cites.
SAPIEN: A SimulAted Part-based Interactive ENvironment
F. Xiang, Y. Qin, K. Mo, Y. Xia, H. Zhu, F. Liu, M. Liu, H. Jiang, Y. Yuan, H. Wang, L. Yi, A. X. Chang, L. J. Guibas, and H. Su · 2020
Earlier work this paper cites.
Bootstrap latent-predictive representations for multitask reinforcement learning
Z. D. Guo, B. A. Pires, B. Piot, J.-B. Grill, F. Altché, R. Munos, and M. G. Azar · 2020
Earlier work this paper cites.
D4RL: Datasets for deep data-driven reinforcement learning
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine · 2020
Earlier work this paper cites.
Stabilizing transformers for reinforcement learning
E. Parisotto, F. Song, J. Rae, R. Pascanu, C. Gulcehre, S. Jayakumar, M. Jaderberg, R. L. Kaufman, A. Clark, S. Noury, et al · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Earlier work this paper cites.
AllenAct: A framework for embodied ai research
L. Weihs, J. Salvador, K. Kotar, U. Jain, K.-H. Zeng, R. Mottaghi, and A. Kembhavi · 2020
Earlier work this paper cites.
Auxiliary tasks and exploration enable objectnav
J. Ye, D. Batra, A. Das, and E. Wijmans · 2021
Earlier work this paper cites.
Provably breaking the quadratic error compounding barrier in imitation learning, optimally
N. Rajaraman, Y. Han, L. F. Yang, K. Ramchandran, and J. Jiao · 2021
Earlier work this paper cites.
Auxiliary tasks and exploration enable objectnav
J. Ye, D. Batra, A. Das, and E. Wijmans · 2021
Earlier work this paper cites.
BEHAVIOR: Benchmark for Everyday Household Activities in Virtual, Interactive, and Ecological Environments
S. Srivastava, C. Li, M. Lingelbach, R. Mart’in-Mart’in, F. Xia, K. Vainio, Z. Lian, C. Gokmen, S. Buch, C. K. Liu, S. Savarese, H. Gweon, J. Wu, and L. Fei-Fei · 2021
Earlier work this paper cites.
igibson 2.0: Object-centric simulation for robot learning of everyday household tasks
C. Li, F. Xia, R. Martín-Martín, M. Lingelbach, S. Srivastava, B. Shen, K. E. Vainio, C. Gokmen, G. Dharan, T. Jain, A. Kurenkov, C. K. Liu, H. Gweon, J. Wu, L. Fei-Fei, and S. Savarese · 2021
Earlier work this paper cites.
iGibson 1.0: A Simulation Environment for Interactive Tasks in Large Realistic Scenes
B. Shen, F. Xia, C. Li, R. Martín-Martín, L. Fan, G. Wang, C. Pérez-D’Arpino, S. Buch, S. Srivastava, L. Tchapmi, M. Tchapmi, K. Vainio, J. Wong, L. Fei-Fei, and S. Savarese · 2021
Cited alongside, same era.
Habitat 2.0: Training Home Assistants to Rearrange their Habitat
A. Szot, A. Clegg, E. Undersander, E. Wijmans, Y. Zhao, J. M. Turner, N. Maestre, M. Mukadam, D. S. Chaplot, O. Maksymets, A. Gokaslan, V. Vondrus, S. Dharur, F. Meier, W. Galuba, A. X. Chang, Z. Kira, V. Koltun, J. Malik, M. Savva, and D. Batra · 2021
Cited alongside, same era.
ThreeDWorld: A Platform for Interactive Multi-Modal Physical Simulation
C. Gan, J. Schwartz, S. Alter, D. Mrowca, M. Schrimpf, J. Traer, J. D. Freitas, J. Kubilius, A. Bhandwaldar, N. Haber, M. Sano, K. Kim, E. Wang, M. Lingelbach, A. Curtis, K. T. Feigelis, D. Bear, D. Gutfreund, D. D. Cox, A. Torralba, J. J. DiCarlo, J. Tenenbaum, J. H. McDermott, and D. Yamins · 2021
Cited alongside, same era.
Pushing it out of the way: Interactive visual navigation
K.-H. Zeng, L. Weihs, A. Farhadi, and R. Mottaghi · 2021
Cited alongside, same era.
Decision transformer: Reinforcement learning via sequence modeling
L. Chen, K. Lu, A. Rajeswaran, K. Lee, A. Grover, M. Laskin, P. Abbeel, A. Srinivas, and I. Mordatch · 2021
Adaptive skill coordination for robotic mobile manipulation
N. Yokoyama, A. Clegg, E. Undersander, S. Ha, D. Batra, and A. Rai · 2023
Later among the works it cites.
RVT: Robotic View Transformer for 3D Object Manipulation
A. Goyal, J. Xu, Y. Guo, V. Blukis, Y. Chao, and D. Fox · 2023
Later among the works it cites.
Multi-skill Mobile Manipulation for Object Rearrangement
J. Gu, D. S. Chaplot, H. Su, and J. Malik · 2023
Later among the works it cites.
M. Chang, T. Gervet, M. Khanna, S. Yenamandra, D. Shah, S. Y. Min, K. Shah, C. Paxton, S. Gupta, D. Batra, et al · 2023
Later among the works it cites.
Navigating to objects in the real world
T. Gervet, S. Chintala, D. Batra, J. Malik, and D. S. Chaplot · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Rethinking attention with performers
K. Choromanski, V. Likhosherstov, D. Dohan, X. Song, A. Gane, T. Sarlos, P. Hawkins, J. Davis, A. Mohiuddin, L. Kaiser, et al · 2021
Cited alongside, same era.
ProcTHOR: Large-scale embodied AI using procedural generation
M. Deitke, E. VanderBilt, A. Herrasti, L. Weihs, K. Ehsani, J. Salvador, W. Han, E. Kolve, A. Kembhavi, and R. Mottaghi · 2022
Cited alongside, same era.
The Design of Stretch: A Compact, Lightweight Mobile Manipulator for Indoor Human Environments
C. C. Kemp, A. Edsinger, H. M. Clever, and B. Matulevich · 2022
Cited alongside, same era.
Phone2Proc: Bringing robust robots into our chaotic world, 2022
M. Deitke, R. Hendrix, L. Weihs, A. Farhadi, K. Ehsani, and A. Kembhavi · 2022
Cited alongside, same era.
Detecting twenty-thousand classes using image-level supervision
X. Zhou, R. Girdhar, A. Joulin, P. Krähenbühl, and I. Misra · 2022
Cited alongside, same era.
ZSON: zero-shot object-goal navigation using multimodal goal embeddings
A. Majumdar, G. Aggarwal, B. Devnani, J. Hoffman, and D. Batra · 2022
Cited alongside, same era.
Habitat-web: Learning embodied object-search strategies from human demonstrations at scale
R. Ramrakhya, E. Undersander, D. Batra, and A. Das · 2022
Cited alongside, same era.
Where are we in the search for an artificial visual cortex for embodied intelligence?
A. Majumdar, K. Yadav, S. Arnaud, Y. J. Ma, C. Chen, S. Silwal, A. Jain, V. Berges, P. Abbeel, J. Malik, D. Batra, Y. Lin, O. Maksymets, A. Rajeswaran, and F. Meier · 2023
Later among the works it cites.
RT-1: robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. J. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, I. Leal, K. Lee, S. Levine, Y. Lu, U. Malla, D. Manjunath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Pertsch, J. Quiambao, K. Rao, M. S. Ryoo, G. Salazar, P. R. Sanketi, K. Sayed, J. Singh, S. Sontakke, A. Stone, C. Tan, H. T. Tran, V. Vanhoucke, S. Vega, Q. Vuong, F. Xia, T. Xiao, P. Xu, S. Xu, T. Yu, and B. Zitkovich · 2023
Later among the works it cites.
Open X-Embodiment: Robotic Learning Datasets and RT-X Models
O. X. Collaboration, A. Padalkar, A. Pooley, A. Jain, A. Bewley, A. Herzog, A. Irpan, A. Khazatsky, A. Raj, A. Singh, A. Brohan, A. Raffin, A. Wahid, B. Burgess-Limerick, B. Kim, B. Schölkopf, B. Ichter, C. Lu, C. Xu, C. Finn, C. Xu, C. Chi, C. Huang, C. Chan, C. Pan, C. Fu, C. Devin, D. Driess, D. Pathak, D. Shah, D. Büchler, D. Kalashnikov, D. Sadigh, E. Johns, F. Ceola, F. Xia, F. Stulp, G. Zhou, G. S. Sukhatme, G. Salhotra, G. Yan, G. Schiavi, G. Kahn, H. Su, H. Fang, H. Shi, H. B. Amor, H. I. Christensen, H. Furuta, H. Walke, H. Fang, I. Mordatch, I. Radosavovic, and et al · 2023
Later among the works it cites.
ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills
J. Gu, F. Xiang, X. Li, Z. Ling, X. Liu, T. Mu, Y. Tang, S. Tao, X. Wei, Y. Yao, X. Yuan, P. Xie, Z. Huang, R. Chen, and H. Su · 2023
Later among the works it cites.
When learning is out of reach, reset: Generalization in autonomous visuomotor reinforcement learning
Z. Zhang and L. Weihs · 2023
Later among the works it cites.
Scene Graph Contrastive Learning for Embodied Navigation
K. P. Singh, J. Salvador, L. Weihs, and A. Kembhavi · 2023
Later among the works it cites.
Offline pre-trained multi-agent decision transformer
L. Meng, M. Wen, C. Le, X. Li, D. Xing, W. Zhang, Y. Wen, H. Zhang, J. Wang, Y. Yang, et al · 2023
Later among the works it cites.
Skill transformer: A monolithic policy for mobile manipulation
X. Huang, D. Batra, A. Rai, and A. Szot · 2023
Later among the works it cites.
Learning model predictive controllers with real-time attention for real-world navigation
X. Xiao, T. Zhang, K. Choromanski, E. Lee, A. Francis, J. Varley, S. Tu, S. Singh, P. Xu, F. Xia, et al · 2023
Later among the works it cites.
Spatial-language attention policies for efficient robot learning
P. Parashar, V. Jain, X. Zhang, J. Vakil, S. Powers, Y. Bisk, and C. Paxton · 2023
Later among the works it cites.
GNM: A general navigation model to drive any robot
D. Shah, A. Sridhar, A. Bhorkar, N. Hirose, and S. Levine · 2023
Later among the works it cites.
Nomad: Goal masked diffusion policies for navigation and exploration
A. Sridhar, D. Shah, C. Glossop, and S. Levine · 2023
Later among the works it cites.
Vint: A foundation model for visual navigation
D. Shah, A. Sridhar, N. Dashora, K. Stachowicz, K. Black, N. Hirose, and S. Levine · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. Canton-Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V. Goswami, N. Goyal, A. Hartshorn, S. Hosseini, R. Hou, H. Inan, M. Kardas, V. Kerkez, M. Khabsa, I. Kloumann, A. Korenev, P. S. Koura, M. Lachaux, T. Lavril, J. Lee, D. Liskovich, Y. Lu, Y. Mao, X. Martinet, T. Mihaylov, P. Mishra, I. Molybog, Y. Nie, A. Poulton, J. Reizenstein, R. Rungta, K. Saladi, A. Schelten, R. Silva, E. M. Smith, R. Subramanian, X. E. Tan, B. Tang, R. Taylor, A. Williams, J. X. Kuan, P. Xu, Z. Yan, I. Zarov, Y. Zhang, A. Fan, M. Kambadur, S. Narang, A. Rodriguez, R. Stojnic, S. Edunov, and T. Scialom · 2023
Later among the works it cites.
Selective visual representations improve convergence and generalization for embodied ai
A. Eftekhar, K.-H. Zeng, J. Duan, A. Farhadi, A. Kembhavi, and R. Krishna · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Imitating shortest paths in simulation enables effective navigation and manipulation in the real world
K. Ehsani, T. Gupta, R. Hendrix, J. Salvador, L. Weihs, K.-H. Zeng, K. P. Singh, Y. Kim, W. Han, A. Herrasti, et al · 2024
Closest in time.
ObjaTHOR: Python package for importing and loading external assets into ai2thor
A. I. for AI · 2024
Closest in time.
Pushing the limits of cross-embodiment learning for manipulation and navigation
J. Yang, C. Glossop, A. Bhorkar, D. Shah, Q. Vuong, C. Finn, D. Sadigh, and S. Levine · 2024
Closest in time.
PIVOT: Iterative visual prompting elicits actionable knowledge for vlms
S. Nasiriany, F. Xia, W. Yu, T. Xiao, J. Liang, I. Dasgupta, A. Xie, D. Driess, A. Wahid, Z. Xu, et al · 2024
Closest in time.
DROID: A large-scale in-the-wild robot manipulation dataset
A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y. Chen, K. Ellis, et al · 2024
Closest in time.
Universal manipulation interface: In-the-wild robot teaching without in-the-wild robots
C. Chi, Z. Xu, C. Pan, E. Cousineau, B. Burchfiel, S. Feng, R. Tedrake, and S. Song · 2024
Closest in time.
Habitat synthetic scenes dataset (hssd-200): An analysis of 3d scene scale and realism tradeoffs for objectgoal navigation
M. Khanna, Y. Mao, H. Jiang, S. Haresh, B. Schacklett, D. Batra, A. Clegg, E. Undersander, A. X. Chang, and M. Savva · 2024
Closest in time.
PDiT: Interleaving perception and decision-making transformers for deep reinforcement learning
H. Mao, R. Zhao, Z. Li, Z. Xu, H. Chen, Y. Chen, B. Zhang, Z. Xiao, J. Zhang, and J. Yin · 2024
Closest in time.
Real-world humanoid locomotion with reinforcement learning
I. Radosavovic, T. Xiao, B. Zhang, T. Darrell, J. Malik, and K. Sreenath · 2024
Closest in time.