Fetching the paper…
Reading the bibliography…
Reinforcement Learning from Human Feedback (RLHF) has shown potential in qualitative tasks where easily defined performance measures are lacking.
The information bottleneck method
Tishby, N., Pereira, F. C., and Bialek, W · 2000
Earlier work this paper cites.
Learning a similarity metric discriminatively, with application to face verification
Chopra, S., Hadsell, R., and LeCun, Y · 2005
Earlier work this paper cites.
Online passive-aggressive algorithms
Crammer, K., Dekel, O., Keshet, J., Shalev-Shwartz, S., Singer, Y., and Warmuth, M. K · 2006
Earlier work this paper cites.
Dimensionality reduction by learning an invariant mapping
Hadsell, R., Chopra, S., and LeCun, Y · 2006
Earlier work this paper cites.
Large scale online learning of image similarity through ranking
Chechik, G., Sharma, V., Shalit, U., and Bengio, S · 2010
Earlier work this paper cites.
Encouraging behavioral diversity in evolutionary robotics: An empirical study
Mouret, J.-B. and Doncieux, S · 2012
Earlier work this paper cites.
Discriminative unsupervised feature learning with convolutional neural networks
Dosovitskiy, A., Springenberg, J. T., Riedmiller, M., and Brox, T · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Robots that can adapt like animals
Cully, A., Clune, J., Tarapore, D., and Mouret, J.-B · 2015
Earlier work this paper cites.
Illuminating search spaces by mapping elites
Mouret, J.-B. and Clune, J · 2015
Earlier work this paper cites.
Confronting the challenge of quality diversity
Pugh, J. K., Soros, L. B., Szerlip, P. A., and Stanley, K. O · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al · 2015
Earlier work this paper cites.
Facenet: A unified embedding for face recognition and clustering
Schroff, F., Kalenichenko, D., and Philbin, J · 2015
Earlier work this paper cites.
Learning behavior characterizations for novelty search
Meyerson, E., Lehman, J., and Miikkulainen, R · 2016
Earlier work this paper cites.
Quality diversity: A new frontier for evolutionary computation
Pugh, J. K., Soros, L. B., and Stanley, K. O · 2016
Earlier work this paper cites.
Rapid phenotypic landscape exploration through hierarchical spatial partitioning
Smith, D., Tokarchuk, L., and Wiggins, G · 2016
Earlier work this paper cites.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 2016
Earlier work this paper cites.
Deep variational information bottleneck
Alemi, A. A., Fischer, I., Dillon, J. V., and Murphy, K · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
A mathematical introduction to robotic manipulation
Murray, R. M., Li, Z., and Sastry, S. S · 2017
Cited alongside, same era.
Using centroidal voronoi tessellations to scale up the multidimensional archive of phenotypic elites algorithm
Vassiliades, V., Chatzilygeroudis, K., and Mouret, J.-B · 2017
Cited alongside, same era.
Improving exploration in evolution strategies for deep reinforcement learning via a population of novelty-seeking agents
Conti, E., Madhavan, V., Petroski Such, F., Lehman, J., Stanley, K., and Clune, J · 2018
Cited alongside, same era.
Reward learning from human preferences and demonstrations in atari
Ibarz, B., Leike, J., Pohlen, T., Irving, G., Legg, S., and Amodei, D · 2018
Cited alongside, same era.
Dynamic mutation in map-elites for robotic repertoire generation
Nordmoen, J., Samuelsen, E., Ellefsen, K. O., and Glette, K · 2018
Cited alongside, same era.
Differentiable quality diversity
Fontaine, M. and Nikolaidis, S · 2021
Later among the works it cites.
Illuminating mario scenes in the latent space of a generative adversarial network
Fontaine, M. C., Liu, R., Khalifa, A., Modi, J., Togelius, J., Hoover, A. K., and Nikolaidis, S · 2021
Later among the works it cites.
Webgpt: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Later among the works it cites.
Monte carlo elites: Quality-diversity selection as a multi-armed bandit problem
Sfikas, K., Liapis, A., and Yannakakis, G. N · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vassiliades, V. and Mouret, J.-B · 2018
Cited alongside, same era.
The unreasonable effectiveness of deep features as a perceptual metric
Zhang, R., Isola, P., Efros, A. A., Shechtman, E., and Wang, O · 2018
Cited alongside, same era.
Autonomous skill discovery with quality-diversity and unsupervised descriptors
Cully, A · 2019
Cited alongside, same era.
Mapping hearthstone deck spaces through map-elites with sliding boundaries
Fontaine, M. C., Lee, S., Soros, L. B., de Mesentier Silva, F., Togelius, J., and Hoover, A. K · 2019
Cited alongside, same era.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
Jaques, N., Ghandeharioun, A., Shen, J. H., Ferguson, C., Lapedriza, A., Jones, N., Gu, S., and Picard, R · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
A simple framework for contrastive learning of visual representations
Chen, T., Kornblith, S., Norouzi, M., and Hinton, G · 2020
Cited alongside, same era.
Wu, J., Ouyang, L., Ziegler, D. M., Stiennon, N., Lowe, R., Leike, J., and Christiano, P · 2021
Later among the works it cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al · 2022
Later among the works it cites.
Optimizing neural networks with gradient lexicase selection
Ding, L. and Spector, L · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Later among the works it cites.
Multifusion: Fusing pre-trained models for multi-lingual, multi-modal image generation
Bellagente, M., Brack, M., Teufel, H., Friedrich, F., Deiseroth, B., Eichenberg, C., Dai, A., Baldock, R., Nanda, S., Oostermeijer, K., et al · 2023
Closest in time.
Objectives are all you need: Solving deceptive problems without explicit diversity maintenance
Boldi, R., Ding, L., and Spector, L · 2023
Closest in time.
Attend-and-excite: Attention-based semantic guidance for text-to-image diffusion models
Chefer, H., Alaluf, Y., Vinker, Y., Wolf, L., and Cohen-Or, D · 2023
Closest in time.
Probabilistic lexicase selection
Ding, L., Pantridge, E., and Spector, L · 2023
Closest in time.
Dreamsim: Learning new dimensions of human visual similarity using synthetic data
Fu, S., Tamir, N., Sundaram, S., Chai, L., Zhang, R., Dekel, T., and Isola, P · 2023
Closest in time.
Kheperax: a lightweight jax-based robot control environment for benchmarking quality-diversity algorithms
Grillotti, L. and Cully, A · 2023
Closest in time.
Dalex: Lexicase-like selection via diverse aggregation
Ni, A., Ding, L., and Spector, L · 2024
Closest in time.
Particularity
Spector, L., Ding, L., and Boldi, R · 2024
Closest in time.