Fetching the paper…
Reading the bibliography…
Transformer, originally devised for natural language processing, has also attested significant success in computer vision.
J. MacQueen, “Classification and analysis of multivariate observations,” in 5th Berkeley Symp. Math. Statist. Probability , 1967
1967
Earlier work this paper cites.
R. Reddy, “Speech understanding systems. summary of results of the five-year research effort at carnegie-mellon university,” 1977
1977
Earlier work this paper cites.
A. Pnueli, “The temporal logic of programs,” in 18th Annual Symposium on Foundations of Computer Science . ieee, 1977
1977
Earlier work this paper cites.
R. G. Morris, “Spatial localization does not require the presence of local cues,” Learning and motivation , 1981
1981
Earlier work this paper cites.
R. D. Shachter, “Probabilistic inference and influence diagrams,” Operations research , 1988
1988
Earlier work this paper cites.
D. A. Pomerleau, “Efficient training of artificial neural networks for autonomous navigation,” Neural computation , 1991
1991
Earlier work this paper cites.
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine learning , 1992
1992
Earlier work this paper cites.
J. Schmidhuber, “Learning to control fast-weight memories: An alternative to dynamic recurrent networks,” Neural Computation , 1992
1992
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , 1997
1997
Earlier work this paper cites.
L. P. Kaelbling, M. L. Littman, and A. R. Cassandra, “Planning and acting in partially observable stochastic domains,” Artificial intelligence , 1998
1998
Earlier work this paper cites.
L. P. Kaelbling, M. L. Littman, and A. R. Cassandra, “Planning and acting in partially observable stochastic domains,” Artificial intelligence , 1998
1998
Earlier work this paper cites.
R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour, “Policy gradient methods for reinforcement learning with function approximation,” NeurIPS , 1999
1999
Earlier work this paper cites.
S. M. Kakade, “A natural policy gradient,” NeurIPS , 2001
2001
Earlier work this paper cites.
C. Boutilier, R. Reiter, and B. Price, “Symbolic dynamic programming for first-order mdps,” in IJCAI , 2001
2001
Earlier work this paper cites.
R. Coulom, “Efficient selectivity and backup operators in monte-carlo tree search,” in International conference on computers and games . Springer, 2006
2006
Earlier work this paper cites.
C. Diuk, A. Cohen, and M. L. Littman, “An object-oriented representation for efficient reinforcement learning,” in ICML , 2008
2008
Earlier work this paper cites.
A. Saxena and K. Goebel, “Turbofan engine degradation simulation data set,” NASA Ames Prognostics Data Repository , 2008
2008
Earlier work this paper cites.
M. Toussaint, “Robot trajectory optimization using approximate inference,” in ICML , 2009
2009
Earlier work this paper cites.
J. Pearl, Causality . Cambridge university press, 2009
2009
Earlier work this paper cites.
S. J. Russell, Artificial intelligence a modern approach . Pearson Education, Inc., 2010
2010
Earlier work this paper cites.
P. M. Nadkarni, L. Ohno-Machado, and W. W. Chapman, “Natural language processing: an introduction,” Journal of the American Medical Informatics Association , 2011
2011
Earlier work this paper cites.
E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model-based control,” in 2012 IEEE/RSJ international conference on intelligent robots and systems . IEEE, 2012
2012
Earlier work this paper cites.
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling, “The arcade learning environment: An evaluation platform for general agents,” Journal of Artificial Intelligence Research , 2013
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski et al. , “Human-level control through deep reinforcement learning,” nature , 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
R. Girshick, “Fast r-cnn,” in ICCV , 2015
2015
Earlier work this paper cites.
M. Hausknecht and P. Stone, “Deep recurrent q-learning for partially observable mdps,” in 2015 aaai fall symposium series , 2015
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
Z. Wang, T. Schaul, M. Hessel, H. Hasselt, M. Lanctot, and N. Freitas, “Dueling network architectures for deep reinforcement learning,” in ICML . PMLR, 2016
2016
Earlier work this paper cites.
H. Van Hasselt, A. Guez, and D. Silver, “Deep reinforcement learning with double q-learning,” in AAAI , 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” in ICML . PMLR, 2016
2016
Earlier work this paper cites.
A. Van den Oord, N. Kalchbrenner, L. Espeholt, O. Vinyals, A. Graves et al. , “Conditional image generation with pixelcnn decoders,” NeurIPS , 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
M. G. Bellemare, G. Ostrovski, A. Guez, P. Thomas, and R. Munos, “Increasing the action gap: New operators for reinforcement learning,” in AAAI , 2016
2016
Earlier work this paper cites.
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” in ICML . PMLR, 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR , 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
T. Trouillon, J. Welbl, S. Riedel, É. Gaussier, and G. Bouchard, “Complex embeddings for simple link prediction,” in ICML . PMLR, 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” NeurIPS , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
J. Gehring, M. Auli, D. Grangier, D. Yarats, and Y. N. Dauphin, “Convolutional sequence to sequence learning,” in ICML . PMLR, 2017
2017
Earlier work this paper cites.
C. Finn, P. Abbeel, and S. Levine, “Model-agnostic meta-learning for fast adaptation of deep networks,” in ICML . PMLR, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, and V. Koltun, “Carla: An open urban driving simulator,” in CoRL . PMLR, 2017
2017
Earlier work this paper cites.
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” in ICCV , 2017
2017
Earlier work this paper cites.
S. Zheng, K. Ristovski, A. Farahat, and C. Gupta, “Long short-term memory network for remaining useful life estimation,” in ICPHM . IEEE, 2017
2017
Earlier work this paper cites.
S. Omidshafiei, J. Pazis, C. Amato, J. P. How, and J. Vian, “Deep decentralized multi-task multi-agent reinforcement learning under partial observability,” in ICML . PMLR, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction . MIT press, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
L. Espeholt, H. Soyer, R. Munos, K. Simonyan, V. Mnih, T. Ward, Y. Doron, V. Firoiu, T. Harley, I. Dunning et al. , “Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures,” in ICML . PMLR, 2018
2018
Earlier work this paper cites.
S. Kapturowski, G. Ostrovski, J. Quan, R. Munos, and W. Dabney, “Recurrent experience replay in distributed reinforcement learning,” in ICLR , 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
S. Kapturowski, G. Ostrovski, J. Quan, R. Munos, and W. Dabney, “Recurrent experience replay in distributed reinforcement learning,” in ICLR , 2018
2018
Earlier work this paper cites.
M. Hessel, J. Modayil, H. Van Hasselt, T. Schaul, G. Ostrovski, W. Dabney, D. Horgan, B. Piot, M. Azar, and D. Silver, “Rainbow: Combining improvements in deep reinforcement learning,” in AAAI , 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
X. Puig, K. Ra, M. Boben, J. Li, T. Wang, S. Fidler, and A. Torralba, “Virtualhome: Simulating household activities via programs,” in CVPR , 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
O. Bastani, Y. Pu, and A. Solar-Lezama, “Verifiable reinforcement learning via policy extraction,” NeurIPS , 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
D. Ha and J. Schmidhuber, “World models,” arXiv preprint arXiv:1803.10122 , 2018
2018
Earlier work this paper cites.
M. Chevalier-Boisvert, L. Willems, and S. Pal, “Minimalistic gridworld environment for openai gym,” 2018
2018
Earlier work this paper cites.
X. Wang, R. Girshick, A. Gupta, and K. He, “Non-local neural networks,” in CVPR , 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
M.-A. Côté, A. Kádár, X. Yuan, B. Kybartas, T. Barnes, E. Fine, J. Moore, M. Hausknecht, L. E. Asri, M. Adada et al. , “Textworld: A learning environment for text-based games,” in Workshop on Computer Games . Springer, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
M. Schlichtkrull, T. N. Kipf, P. Bloem, R. v. d. Berg, I. Titov, and M. Welling, “Modeling relational data with graph convolutional networks,” in ESWC . Springer, 2018
2018
Earlier work this paper cites.
P. Anderson, Q. Wu, D. Teney, J. Bruce, M. Johnson, N. Sünderhauf, I. Reid, S. Gould, and A. Van Den Hengel, “Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments,” in CVPR , 2018
2018
Earlier work this paper cites.
X. Wang, W. Xiong, H. Wang, and W. Y. Wang, “Look before you leap: Bridging model-free and model-based reinforcement learning for planned-ahead vision-and-language navigation,” in ECCV , 2018
2018
Earlier work this paper cites.
Z. Wei, Q. Liu, B. Peng, H. Tou, T. Chen, X.-J. Huang, K.-F. Wong, and X. Dai, “Task-oriented dialogue system for automatic diagnosis,” in ACL , 2018
2018
Earlier work this paper cites.
T. Rashid, M. Samvelyan, C. Schroeder, G. Farquhar, J. Foerster, and S. Whiteson, “Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning,” in ICML . PMLR, 2018
2018
Earlier work this paper cites.
J. Xu, X. Sun, Z. Zhang, G. Zhao, and J. Lin, “Understanding and improving layer normalization,” NeurIPS , 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. R. Salakhutdinov, and Q. V. Le, “Xlnet: Generalized autoregressive pretraining for language understanding,” NeurIPS , 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
S. Dasari and A. Gupta, “Transformers for one-shot visual imitation,” in CoRL . PMLR, 2021
2021
Later among the works it cites.
S. Chen, P.-L. Guhur, C. Schmid, and I. Laptev, “History aware multimodal transformer for vision-and-language navigation,” NeurIPS , 2021
2021
Later among the works it cites.
A. Prakash, K. Chitta, and A. Geiger, “Multi-modal fusion transformer for end-to-end autonomous driving,” in CVPR , 2021
2021
Later among the works it cites.
T. Yu, A. Kumar, R. Rafailov, A. Rajeswaran, S. Levine, and C. Finn, “Combo: Conservative offline model-based policy optimization,” NeurIPS , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
S. Fujimoto, D. Meger, and D. Precup, “Off-policy deep reinforcement learning without exploration,” in ICML . PMLR, 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Y. Ding, C. Florensa, P. Abbeel, and M. Phielipp, “Goal-conditioned imitation learning,” NeurIPS , 2019
2019
Cited alongside, same era.
N. Rhinehart, R. McAllister, K. Kitani, and S. Levine, “Precog: Prediction conditioned on goals in visual multi-agent settings,” in ICCV , 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
I. Schlag, K. Irie, and J. Schmidhuber, “Linear transformers are secretly fast weight programmers,” in ICML . PMLR, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
W. Ye, S. Liu, T. Kurutach, P. Abbeel, and Y. Gao, “Mastering atari games with limited data,” NeurIPS , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
Z. Liu, Z. Fan, Y. Wang, and P. S. Yu, “Augmenting sequential recommendation with pseudo-prior items via reversely pre-training transformer,” in Proceedings of the 44th international ACM SIGIR conference on Research and development in information retrieval , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
E. Mitchell, R. Rafailov, X. B. Peng, S. Levine, and C. Finn, “Offline meta-reinforcement learning with advantage weighting,” in ICML . PMLR, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
E. Mitchell, R. Rafailov, X. B. Peng, S. Levine, and C. Finn, “Offline meta-reinforcement learning with advantage weighting,” in ICML . PMLR, 2021
2021
Later among the works it cites.
Y. Tang and D. Ha, “The sensory neuron as a transformer: Permutation-invariant neural networks for reinforcement learning,” NeurIPS , 2021
2021
Later among the works it cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in ICML . PMLR, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
H. Kim, Y. Ohmura, and Y. Kuniyoshi, “Transformer-based deep imitation learning for dual-arm robot manipulation,” in IROS) . IEEE, 2021
2021
Later among the works it cites.
Y. Hong, Q. Wu, Y. Qi, C. Rodriguez-Opazo, and S. Gould, “Vln bert: A recurrent vision-and-language bert for navigation,” in CVPR , 2021
2021
Later among the works it cites.
A. Arnab, M. Dehghani, G. Heigold, C. Sun, M. Lučić, and C. Schmid, “Vivit: A video vision transformer,” in ICCV , 2021
2021
Later among the works it cites.
A. Pashevich, C. Schmid, and C. Sun, “Episodic transformer for vision-and-language navigation,” in ICCV , 2021
2021
Later among the works it cites.
C. Chen, Z. Al-Halah, and K. Grauman, “Semantic audio-visual navigation,” in CVPR , 2021
2021
Later among the works it cites.
A. Zhao, T. He, Y. Liang, H. Huang, G. Van den Broeck, and S. Soatto, “Sam: Squeeze-and-mimic networks for conditional visual driving policy learning,” in CoRL . PMLR, 2021
2021
Later among the works it cites.
K. Chitta, A. Prakash, and A. Geiger, “Neat: Neural attention fields for end-to-end autonomous driving,” in ICCV , 2021
2021
Later among the works it cites.
Z. Zhang, A. Liniger, D. Dai, F. Yu, and L. Van Gool, “End-to-end urban driving by imitating a reinforcement learning coach,” in ICCV , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
Y. Liu, J. Zhang, L. Fang, Q. Jiang, and B. Zhou, “Multimodal motion prediction with stacked transformers,” in CVPR , 2021
2021
Later among the works it cites.
A. Quintanar, D. Fernández-Llorca, I. Parra, R. Izquierdo, and M. Sotelo, “Predicting vehicles trajectories in urban scenarios with transformer networks and augmented information,” in 2021 IEEE Intelligent Vehicles Symposium (IV) . IEEE, 2021
2021
Later among the works it cites.
Y. Yuan, X. Weng, Y. Ou, and K. M. Kitani, “Agentformer: Agent-aware transformers for socio-temporal multi-agent forecasting,” in ICCV , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
M. S. Teixeira, V. Maran, and M. Dragoni, “The interplay of a conversational ontology and ai planning for health dialogue management,” in Proceedings of the 36th annual ACM symposium on applied computing , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
J. G. Kuba, M. Wen, L. Meng, H. Zhang, D. Mguni, J. Wang, Y. Yang et al. , “Settling the variance of multi-agent policy gradients,” NeurIPS , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2022
Closest in time.
2022
Closest in time.
B. Singh, R. Kumar, and V. P. Singh, “Reinforcement learning in robotic applications: a comprehensive survey,” Artificial Intelligence Review , 2022
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
J. Sun, D.-A. Huang, B. Lu, Y.-H. Liu, B. Zhou, and A. Garg, “Plate: Visually-grounded planning with transformers in procedural tasks,” RAL , 2022
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
J. Shang, K. Kahatapitiya, X. Li, and M. S. Ryoo, “Starformer: Transformer with state-action-reward representations for visual reinforcement learning,” in ECCV . Springer, 2022
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
S. Takagi, “On the effect of pre-training for transformer in different modality on offline reinforcement learning,” NeurIPS , 2022
2022
Closest in time.
F. Liu, H. Liu, A. Grover, and P. Abbeel, “Masked autoencoding for scalable and generalizable decision making,” NeurIPS , 2022
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
M. Xu, Y. Shen, S. Zhang, Y. Lu, D. Zhao, J. Tenenbaum, and C. Gan, “Prompting decision transformer for few-shot policy generalization,” in ICML . PMLR, 2022
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
G. Liu, A. Adhikari, A.-m. Farahmand, and P. Poupart, “Learning object-oriented dynamics for planning from text,” in ICLR , 2022
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
A. R. Villaflor, Z. Huang, S. Pande, J. M. Dolan, and J. Schneider, “Addressing optimism bias in sequence modeling for reinforcement learning,” in ICML . PMLR, 2022
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
J. Chen, D. Li, Q. Chen, W. Zhou, and X. Liu, “Diaformer: Automatic diagnosis via symptoms sequence generation,” in AAAI , 2022
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
Y. Li, T. Gao, J. Yang, H. Xu, and Y. Wu, “Phasic self-imitative reduction for sparse-reward goal-conditioned reinforcement learning,” in ICML . PMLR, 2022
2022
Closest in time.
2022
Closest in time.
B. Trabucco, M. Phielipp, and G. Berseth, “Anymorph: Learning transferable polices by inferring agent morphology,” in ICML . PMLR, 2022
2022
Closest in time.
Z. Mandi, F. Liu, K. Lee, and P. Abbeel, “Towards more generalizable one-shot visual imitation learning,” in ICRA . IEEE, 2022
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
M. Shridhar, L. Manuelli, and D. Fox, “Cliport: What and where pathways for robotic manipulation,” in CoRL . PMLR, 2022
2022
Closest in time.
E. Jang, A. Irpan, M. Khansari, D. Kappler, F. Ebert, C. Lynch, S. Levine, and C. Finn, “Bc-z: Zero-shot task generalization with robotic imitation learning,” in CoRL . PMLR, 2022
2022
Closest in time.
R. Jangir, N. Hansen, S. Ghosal, M. Jain, and X. Wang, “Look closer: Bridging egocentric and third-person views with transformers for robotic manipulation,” RAL , 2022
2022
Closest in time.
S. Hu, L. Chen, P. Wu, H. Li, J. Yan, and D. Tao, “St-p3: End-to-end vision-based autonomous driving via spatial-temporal feature learning,” in ECCV . Springer, 2022
2022
Closest in time.
R. Sanjaya, J. Wang, and Y. Yang, “Measuring the non-transitivity in chess,” Algorithms , 2022
2022
Closest in time.
2022
Closest in time.
Y. Yang, G. Chen, W. Wang, X. Hao, J. Hao, and P. A. Heng, “Transformer-based working memory for multiagent reinforcement learning with action parsing,” NeurIPS , 2022
2022
Closest in time.