Fetching the paper…
Reading the bibliography…
Reinforcement learning algorithms typically struggle in the absence of a dense, well-shaped reward function.
Sentence-bert: Sentence embeddings using siamese bert-networks
Reimers, N. and Gurevych, I · 1908
Earlier work this paper cites.
Introduction to reinforcement learning , volume 135
Sutton, R. S., Barto, A. G., et al · 1998
Earlier work this paper cites.
Language models are few-shot learners, 2020
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2005
Earlier work this paper cites.
Exploiting open-endedness to solve problems through the search for novelty
Lehman, J., Stanley, K. O., et al · 2008
Earlier work this paper cites.
What is intrinsic motivation? a typology of computational approaches
Oudeyer, P.-Y. and Kaplan, F · 2009
Earlier work this paper cites.
Active learning of inverse models with intrinsically motivated goal exploration in robots
Baranes, A. and Oudeyer, P.-Y · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Concrete problems in ai safety
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mané, D · 2016
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
Van Hasselt, H., Guez, A., and Silver, D · 2016
Earlier work this paper cites.
Dueling network architectures for deep reinforcement learning
Wang, Z., Schaul, T., Hessel, M., Hasselt, H., Lanctot, M., and Freitas, N · 2016
Earlier work this paper cites.
Grounded language learning in a simulated 3d world
Hermann, K. M., Hill, F., Green, S., Wang, F., Faulkner, R., Soyer, H., Szepesvari, D., Czarnecki, W. M., Jaderberg, M., Teplyashin, D., et al · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T · 2017
Earlier work this paper cites.
Gep-pg: Decoupling exploration and exploitation in deep reinforcement learning algorithms
Colas, C., Sigaud, O., and Oudeyer, P.-Y · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Investigating human priors for playing video games
Dubey, R., Agrawal, P., Pathak, D., Griffiths, T. L., and Efros, A. A · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al · 2018
Earlier work this paper cites.
A survey on intrinsic motivation in reinforcement learning
Aubret, A., Matignon, L., and Hassas, S · 2019
Earlier work this paper cites.
Exploration by random network distillation
Burda, Y., Edwards, H., Storkey, A., and Klimov, O · 2019
Cited alongside, same era.
Actrce: Augmenting experience via teacher’s advice for multi-goal reinforcement learning
Chan, H., Wu, Y., Kiros, J., Fidler, S., and Ba, J · 2019
Cited alongside, same era.
Environmental drivers of systematicity and generalization in a situated agent
Hill, F., Lampinen, A., Schneider, R., Clark, S., Botvinick, M., McClelland, J. L., and Santoro, A · 2019
Cited alongside, same era.
A survey of reinforcement learning informed by natural language
Luketina, J., Nardelli, N., Farquhar, G., Foerster, J., Andreas, J., Grefenstette, E., Whiteson, S., and Rocktäschel, T · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Later among the works it cites.
Skill induction and planning with latent language
Sharma, P., Torralba, A., and Andreas, J · 2021
Later among the works it cites.
Reinforcement learning with prototypical representations
Yarats, D., Fergus, R., Lazaric, A., and Pinto, L · 2021
Later among the works it cites.
Noveld: A simple yet effective exploration criterion
Zhang, T., Xu, H., Wang, X., Wu, Y., Keutzer, K., Gonzalez, J. E., and Tian, Y · 2021
Later among the works it cites.
Do as i can, not as i say: Grounding language in robotic affordances, 2022
Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Ho, D., Hsu, J., Ibarz, J., Ichter, B., Irpan, A., Jang, E., Ruano, R. J., Jeffrey, K., Jesmonth, S., Joshi, N. J., Julian, R., Kalashnikov, D., Kuang, Y., Lee, K.-H., Levine, S., Lu, Y., Luu, L., Parada, C., Pastor, P., Quiambao, J., Rao, K., Rettinghouse, J., Reyes, D., Sermanet, P., Sievers, N., Tan, C., Toshev, A., Vanhoucke, V., Xia, F., Xiao, T., Xu, P., Xu, S., and Yan, M · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Language as a cognitive tool to imagine goals in curiosity driven exploration
Colas, C., Karch, T., Lair, N., Dussoux, J.-M., Moulin-Frier, C., Dominey, P., and Oudeyer, P.-Y · 2020
Cited alongside, same era.
Human instruction-following with deep reinforcement learning via transfer-learning from text
Hill, F., Mokra, S., Wong, N., and Harley, T · 2020
Cited alongside, same era.
The Surprising Creativity of Digital Evolution: A Collection of Anecdotes from the Evolutionary Computation and Artificial Life Research Communities
Lehman, J., Clune, J., Misevic, D., Adami, C., Altenberg, L., Beaulieu, J., Bentley, P. J., Bernard, S., Beslon, G., Bryson, D. M., Cheney, N., Chrabaszcz, P., Cully, A., Doncieux, S., Dyer, F. C., Ellefsen, K. O., Feldt, R., Fischer, S., Forrest, S., Fŕenoy, A., Gagńe, C., Le Goff, L., Grabowski, L. M., Hodjat, B., Hutter, F., Keller, L., Knibbe, C., Krcah, P., Lenski, R. E., Lipson, H., MacCurdy, R., Maestre, C., Miikkulainen, R., Mitri, S., Moriarty, D. E., Mouret, J.-B., Nguyen, A., Ofria, C., Parizeau, M., Parsons, D., Pennock, R. T., Punch, W. F., Ray, T. S., Schoenauer, M., Schulte, E., Sims, K., Stanley, K. O., Taddei, F., Tarapore, D., Thibault, S., Watson, R., Weimer, W., and Yosinski, J · 2020
Cited alongside, same era.
Adapting behavior via intrinsic reward: A survey and empirical study
Linke, C., Ady, N. M., White, M., Degris, T., and White, A · 2020
Cited alongside, same era.
Language conditioned imitation learning over unstructured data
Lynch, C. and Sermanet, P · 2020
Cited alongside, same era.
Stereoset: Measuring stereotypical bias in pretrained language models
Nadeem, M., Bethke, A., and Reddy, S · 2020
Cited alongside, same era.
Keep calm and explore: Language models for action generation in text-based games
Yao, S., Rao, R., Hausknecht, M., and Narasimhan, K · 2020
Cited alongside, same era.
Later among the works it cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al · 2022
Later among the works it cites.
LMPriors: Pre-trained language models as task-specific priors
Choi, K., Cundy, C., Srivastava, S., and Ermon, S · 2022
Later among the works it cites.
Autotelic agents with intrinsically motivated goal-conditioned reinforcement learning: a short survey
Colas, C., Karch, T., Sigaud, O., and Oudeyer, P.-Y · 2022
Later among the works it cites.
Housekeep: Tidying virtual households using commonsense reasoning
Kant, Y., Ramachandran, A., Yenamandra, S., Gilitschenski, I., Batra, D., Szot, A., and Agrawal, H · 2022
Later among the works it cites.
Human action recognition and prediction: A survey
Kong, Y. and Fu, Y · 2022
Later among the works it cites.
Exploration in deep reinforcement learning: A survey
Ladosz, P., Weng, L., Kim, M., and Oh, H · 2022
Later among the works it cites.
Improving intrinsic exploration with language abstractions
Mu, J., Zhong, V., Raileanu, R., Jiang, M., Goodman, N., Rocktäschel, T., and Grefenstette, E · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Later among the works it cites.
Learning to generalize with object-centric agents in the open world survival game crafter
Stanić, A., Tang, Y., Ha, D., and Schmidhuber, J · 2022
Later among the works it cites.
From show to tell: a survey on deep learning-based image captioning
Stefanini, M., Cornia, M., Baraldi, L., Cascianelli, S., Fiameni, G., and Cucchiara, R · 2022
Later among the works it cites.
Semantic exploration from language abstractions and pretrained representations
Tam, A. C., Rabinowitz, N. C., Lampinen, A. K., Roy, N. A., Chan, S. C., Strouse, D., Wang, J. X., Banino, A., and Hill, F · 2022
Later among the works it cites.
Curiosity-driven exploration
Ten, A., Oudeyer, P.-Y., and Moulin-Frier, C · 2022
Later among the works it cites.
A survey of modern deep learning based object detection models
Zaidi, S. S. A., Ansari, M. S., Aslam, A., Kanwal, N., Asghar, M., and Lee, B · 2022
Later among the works it cites.
Reward design with language models
Kwon, M., Xie, S. M., Bullard, K., and Sadigh, D · 2023
Closest in time.