Fetching the paper…
Reading the bibliography…
Reinforcement learning algorithms require many samples when solving complex hierarchical tasks with sparse and delayed rewards.
SQIL: imitation learning via regularized behavioral cloning
Reddy, S., Dragan, A. D., and Levine, S · 1905
Earlier work this paper cites.
A general method applicable to the search for similarities in the amino acid sequence of two proteins
Needleman, S. B. and Wunsch, C. D · 1970
Earlier work this paper cites.
A linear space algorithm for computing maximal common subsequences
Hirschberg, D. S · 1975
Earlier work this paper cites.
Atlas of Protein Sequence and Structure , volume 3
Dayhoff, M. O · 1978
Earlier work this paper cites.
Cases in which parsimony or compatibility methods will be positively misleading
Felsenstein, J · 1978
Earlier work this paper cites.
Identification of common molecular subsequences
Smith, T. F. and Waterman, M. S · 1981
Earlier work this paper cites.
An improved algorithm for matching biological sequences
Gotoh, O · 1982
Earlier work this paper cites.
Use of the ’Perceptron’ algorithm to distinguish translational initiation sites in E. coli
Stormo, G. D., Schneider, T. D., Gold, L., and Ehrenfeucht, A · 1982
Earlier work this paper cites.
Multiple sequence alignment with hierarchical clustering
Corpet, F · 1988
Earlier work this paper cites.
Learning from Delayed Rewards
Watkins, C. J. C. H · 1989
Earlier work this paper cites.
Basic local alignment search tool
Altschul, S. F., Gish, W., Miller, W., Myers, E. W., and Lipman, D. J · 1990
Earlier work this paper cites.
Methods for assessing the statistical significance of molecular sequence features by using general scoring schemes
Karlin, S. and Altschul, S. F · 1990
Earlier work this paper cites.
Statistical composition of high-scoring segments from molecular sequences
Karlin, S., Dembo, A., and Kawabata, T · 1990
Earlier work this paper cites.
Cognitive models from subcognitive skills , pp. 71–99
Michie, D., Bain, M., and Hayes-Michie, J · 1990
Earlier work this paper cites.
Untersuchungen zu dynamischen neuronalen Netzen
Hochreiter, S · 1991
Earlier work this paper cites.
Efficient training of artificial neural networks for autonomous navigation
Pomerleau, D. A · 1991
Earlier work this paper cites.
Amino acid substitution matrices from protein blocks
Henikoff, S. and Henikoff, J. G · 1992
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
Dayan, P · 1993
Earlier work this paper cites.
PROSITE: recent developments
Bairoch, A. and Bucher, P · 1994
Earlier work this paper cites.
Building Symbolic Representations of Intuitive Real-Time Skills from Performance Data , pp. 385–418
Michie, D. and Camacho, R · 1994
Earlier work this paper cites.
CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice
Thompson, J. D., Higgins, D. G., and Gibson, T. J · 1994
Earlier work this paper cites.
On the Complexity of Multiple Sequence Alignment
Wang, L. and Jiang, T · 1994
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1995
Earlier work this paper cites.
Learning from demonstration
Schaal, S · 1996
Earlier work this paper cites.
Gapped BLAST and PSI-BLAST: a new generation of protein database search programs
Altschul, S. F., Madden, T. L., Schäffer, A. A., Zhang, J., Zhang, Z., Miller, W., and Lipman, D. J · 1997
Earlier work this paper cites.
Highly specific protein sequence motifs for genome analysis
Mccammon, J. A. and Wolynes, P. G · 1998
Earlier work this paper cites.
Learning to take actions
Khardon, R · 1999
Earlier work this paper cites.
Between MDPs and Semi-MDPs: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S. P · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the MAXQ value function decomposition
Dietterich, T. G · 2000
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, A. Y. and Russell, S. J · 2000
Earlier work this paper cites.
T-coffee: a novel method for fast and accurate multiple sequence alignment
Notredame, C., Higgins, D. G., and Heringa, J · 2000
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. and Langford, J · 2002
Earlier work this paper cites.
Learning options in reinforcement learning
Stolle, M. and Precup, D · 2002
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P. and Ng, A. Y · 2004
Earlier work this paper cites.
MUSCLE: multiple sequence alignment with high accuracy and high throughput
Edgar, R. C · 2004
Earlier work this paper cites.
DIALIGN: Multiple DNA and protein sequence alignment at BiBiServ
Morgenstern, B · 2004
Earlier work this paper cites.
Markov Decision Processes
Puterman, M. L · 2005
Earlier work this paper cites.
Large-scale kernel machines , chapter Scaling learning algorithms towards AI, pp. 321–359
Bengio, Y. and Lecun, Y · 2007
Earlier work this paper cites.
Clustering by passing messages between data points
Frey, B. J. and Dueck, D · 2007
Earlier work this paper cites.
Matplotlib: A 2d graphics environment
Hunter, J. D · 2007
Earlier work this paper cites.
A game-theoretic approach to apprenticeship learning
Syed, U. and Schapire, R. E · 2007
Cited alongside, same era.
Robot programming by demonstration
Billard, A., Calinon, S., Dillmann, R., and Schaal, S · 2008
Cited alongside, same era.
Maximum entropy inverse reinforcement learning
Ziebart, B. D., Maas, A., Bagnell, J. A., and Dey, A. K · 2008
Cited alongside, same era.
Sequence comparison: theory and methods
Chao, K. and Zhang, L · 2009
Cited alongside, same era.
Biopython: freely available Python tools for computational molecular biology and bioinformatics
Cock, P. J. A., Antao, T., Chang, J. T., Chapman, B. A., Cox, C. J., Dalke, A., Friedberg, I., Hamelryck, T., Kauff, F., Wilczynski, B., and de Hoon, M. J. L · 2009
Cited alongside, same era.
Search-based structured prediction, 2009
III, H. D., Langford, J., and Marcu, D · 2009
Cited alongside, same era.
FeUdal networks for hierarchical reinforcement learning
Vezhnevets, A. S., Osindero, S., Schaul, T., Heess, N., Jaderberg, M., Silver, D., and Kavukcuoglu, K · 2017
Later among the works it cites.
RUDDER: return decomposition for delayed rewards
Arjona-Medina, J. A., Gillhofer, M., Widrich, M., Unterthiner, T., and Hochreiter, S · 2018
Later among the works it cites.
IMPALA: Scalable distributed Deep-RL with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., Legg, S., and Kavukcuoglu, K · 2018
Later among the works it cites.
Meta learning shared hierarchies
Frans, K., Ho, J., Chen, X., Abbeel, P., and Schulman, J · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Active learning for reward estimation in inverse reinforcement learning
Lopes, M., Melo, F. S., and Montesano, L · 2009
Cited alongside, same era.
Effects of feedback delay on learning
Rahmandad, H., Repenning, N., and Sterman, J · 2009
Cited alongside, same era.
Cross-domain few-shot learning by representation fusion, 2020
Adler, T., Brandstetter, J., Widrich, M., Mayr, A., Kreil, D., Kopp, M., Klambauer, G., and Hochreiter, S · 2010
Cited alongside, same era.
Optimal policy switching algorithms for reinforcement learning
Comanici, G. and Precup, D · 2010
Cited alongside, same era.
Efficient reductions for imitation learning
Ross, S. and Bagnell, D · 2010
Cited alongside, same era.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G., and Bagnell, D · 2011
Cited alongside, same era.
Deep q-learning from demonstrations
Hester, T., Vecerík, M., Pietquin, O., Lanctot, M., Schaul, T., Piot, B., Horgan, D., Quan, J., Sendonaris, A., Osband, I., Dulac-Arnold, G., Agapiou, J., Leibo, J. Z., and Gruslys, A · 2018
Later among the works it cites.
Overcoming exploration in reinforcement learning with demonstrations
Nair, A., McGrew, B., Andrychowicz, M., Zaremba, W., and Abbeel, P · 2018
Later among the works it cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2018
Later among the works it cites.
Truncated horizon policy search: Combining reinforcement learning & imitation learning
Sun, W., Bagnell, J., and Boots, B · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G · 2018
Later among the works it cites.
Pretraining deep actor-critic reinforcement learning algorithms with expert demonstrations
Zhang, X. and Ma, H · 2018
Later among the works it cites.
mazelab: A customizable framework to create maze and gridworld environments
Zuo, X · 2018
Later among the works it cites.
RUDDER: return decomposition for delayed rewards
Arjona-Medina, J. A., Gillhofer, M., Widrich, M., Unterthiner, T., Brandstetter, J., and Hochreiter, S · 2019
Later among the works it cites.
Go-Explore: A new approach for hard-exploration problems
Ecoffet, A., Huizinga, J., Lehman, J., Stanley, K. O., and Clune, J · 2019
Later among the works it cites.
Search on the replay buffer: Bridging planning and reinforcement learning
Eysenbach, B., Salakhutdinov, R., and Levine, S · 2019
Later among the works it cites.
Reinforcement learning from imperfect demonstrations under soft expert guidance
Jing, M., Ma, X., Huang, W., Sun, F., Yang, C., Fang, B., and Liu, H · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Later among the works it cites.
Successor options: An option discovery framework for reinforcement learning
Ramesh, R., Tomar, M., and Ravindran, B · 2019
Later among the works it cites.
Experience replay for continual learning
Rolnick, D., Ahuja, A., Schwarz, J., Lillicrap, T. P., and Wayne, G · 2019
Later among the works it cites.
Continuous deep maximum entropy inverse reinforcement learning using online POMDP
Silva, J. A. R., Grassi, V., and Wolf, D. F · 2019
Later among the works it cites.
Hierarchical deep q-network with forgetting from imperfect demonstrations in Minecraft
Skrynnik, A., Staroverov, A., Aitygulov, E., Aksenov, K., Davydov, V., and Panov, A. I · 2019
Later among the works it cites.
Array programming with NumPy
Harris, C. R., Millman, K. J., van der Walt, S. J., Gommers, R., Virtanen, P., Cournapeau, D., Wieser, E., Taylor, J., Berg, S., Smith, N. J., Kern, R., Picus, M., Hoyer, S., van Kerkwijk, M. H., Brett, M., Haldane, A., del Río, J. F., Wiebe, M., Peterson, P., Gérard-Marchant, P., Sheppard, K., Reddy, T., Weckesser, W., Abbasi, H., Gohlke, C., and Oliphant, T. E · 2020
Closest in time.
Convergence proof for actor-critic methods applied to PPO and RUDDER
Holzleitner, M., Gruber, L., Arjona-Medina, J. A., Brandstetter, J., and Hochreiter, S · 2020
Closest in time.
Playing Minecraft with behavioural cloning
Kanervisto, A., Karttunen, J., and Hautamäki, V · 2020
Closest in time.
Retrospective analysis of the 2019 MineRL competition on sample efficient reinforcement learning
Milani, S., Topin, N., Houghton, B., Guss, W. H., Mohanty, S. P., Nakata, K., Vinyals, O., and Kuno, N. S · 2020
Closest in time.
Sample efficient reinforcement learning through learning from demonstrations in Minecraft
Scheller, C., Schraner, Y., and Vogel, M · 2020
Closest in time.
Forgetful experience replay in hierarchical reinforcement learning from demonstrations, 2020
Skrynnik, A., Staroverov, A., Aitygulov, E., Aksenov, K., Davydov, V., and Panov, A. I · 2020
Closest in time.
Modern hopfield networks and attention for immune repertoire classification
Widrich, M., Schäfl, B., Pavlović, M., Ramsauer, H., Gruber, L., Holzleitner, M., Brandstetter, J., Sandve, G. K., Greiff, V., Hochreiter, S., and Klambauer, G · 2020
Closest in time.
Watch, try, learn: Meta-learning from demonstrations and rewards
Zhou, A., Jang, E., Kappler, D., Herzog, A., Khansari, M., Wohlhart, P., Bai, Y., Kalakrishnan, M., Levine, S., and Finn, C · 2020
Closest in time.
Convergence Proof for Actor-Critic Methods Applied to PPO and RUDDER , pp. 105–130
Holzleitner, M., Gruber, L., Arjona-Medina, J., Brandstetter, J., and Hochreiter, S · 2021
Closest in time.
Hopfield networks is all you need
Ramsauer, H., Schäfl, B., Lehner, J., Seidl, P., Widrich, M., Gruber, L., Holzleitner, M., Adler, T., Kreil, D., Kopp, M. K., Klambauer, G., Brandstetter, J., and Hochreiter, S · 2021
Closest in time.
Understanding the effects of dataset characteristics on offline reinforcement learning
Schweighofer, K., Hofmarcher, M., Dinu, M., Renz, P., Bitto-Nemling, A., Patil, V. P., and Hochreiter, S · 2021
Closest in time.
Modern hopfield networks for return decomposition for delayed rewards
Widrich, M., Hofmarcher, M., Patil, V. P., Bitto-Nemling, A., and Hochreiter, S · 2021
Closest in time.
XAI and Strategy Extraction via Reward Redistribution , pp. 177–205
Dinu, M.-C., Hofmarcher, M., Patil, V. P., Dorfer, M., Blies, P. M., Brandstetter, J., Arjona-Medina, J. A., and Hochreiter, S · 2022
Closest in time.
A globally convergent evolutionary strategy for stochastic constrained optimization with applications to reinforcement learning
Diouane, Y., Lucchi, A., and Prakash Patil, V · 2022
Closest in time.
Cloob: Modern hopfield networks with infoloob outperform clip, 2022
Fürst, A., Rumetshofer, E., Lehner, J., Tran, V., Tang, F., Ramsauer, H., Kreil, D., Kopp, M., Klambauer, G., Bitto-Nemling, A., and Hochreiter, S · 2022
Closest in time.
Few-shot learning by dimensionality reduction in gradient space, 2022
Gauch, M., Beck, M., Adler, T., Kotsur, D., Fiel, S., Eghbal-zadeh, H., Brandstetter, J., Kofler, J., Holzleitner, M., Zellinger, W., Klotz, D., Hochreiter, S., and Lehner, S · 2022
Closest in time.
History compression via language models in reinforcement learning
Paischer, F., Adler, T., Patil, V., Bitto-Nemling, A., Holzleitner, M., Lehner, S., Eghbal-zadeh, H., and Hochreiter, S · 2022
Closest in time.
Hopular: Modern hopfield networks for tabular data, 2022
Schäfl, B., Gruber, L., Bitto-Nemling, A., and Hochreiter, S · 2022
Closest in time.