Fetching the paper…
Reading the bibliography…
Aligning language models with preferences can be posed as approximating a target distribution representing some desired behavior.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
Jaques, N., Ghandeharioun, A., Shen, J. H., Ferguson, C., Lapedriza, À., Jones, N., Gu, S., and Picard, R. W · 1907
Earlier work this paper cites.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 1909
Earlier work this paper cites.
Distributional reinforcement learning for energy-based sequential models
Parshakova, T., Andreoli, J.-M., and Dymetman, M · 1912
Earlier work this paper cites.
Convex analysis , volume 18
Rockafellar, R. T · 1970
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
A maximum entropy approach to natural language processing
Berger, A. L., Della Pietra, S. A., and Della Pietra, V. J · 1996
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
Baxter, J. and Bartlett, P. L · 2001
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
A tutorial on energy-based learning
LeCun, Y., Chopra, S., Hadsell, R., Ranzato, M., and Huang, F · 2006
Earlier work this paper cites.
On divergences and informations in statistics and information theory
Liese, F. and Vajda, I · 2006
Earlier work this paper cites.
Linearly-solvable markov decision problems
Todorov, E · 2006
Earlier work this paper cites.
Linearly-solvable markov decision problems
Todorov, E · 2006
Earlier work this paper cites.
Reinforcement learning by reward-weighted regression for operational space control
Peters, J. and Schaal, S · 2007
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C · 2011
Earlier work this paper cites.
Optimal control as a graphical model inference problem
Kappen, H. J., Gómez, V., and Opper, M · 2012
Earlier work this paper cites.
Convex analysis and minimization algorithms I: Fundamentals , volume 305
Hiriart-Urruty, J.-B. and Lemaréchal, C · 2013
Earlier work this paper cites.
Optimal control as a graphical model inference problem
Kappen, H. J., Gómez, V., and Opper, M · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. A · 2013
Earlier work this paper cites.
How (not) to train your generative model: Scheduled sampling, likelihood, adversary?
Huszar, F · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Neural text generation from structured data with application to the biography domain
Lebret, R., Grangier, D., and Auli, M · 2016
Earlier work this paper cites.
How NOT to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation
Liu, C.-W., Lowe, R., Serban, I., Noseworthy, M., Charlin, L., and Pineau, J · 2016
Earlier work this paper cites.
Abstractive text summarization using sequence-to-sequence RNNs and beyond
Nallapati, R., Zhou, B., dos Santos, C., Gulçehre, Ç., and Xiang, B · 2016
Earlier work this paper cites.
Reward augmented maximum likelihood for neural structured prediction
Norouzi, M., Bengio, S., Chen, Z., Jaitly, N., Schuster, M., Wu, Y., and Schuurmans, D · 2016
Earlier work this paper cites.
f-gan: Training generative neural samplers using variational divergence minimization
Nowozin, S., Cseke, B., and Tomioka, R · 2016
Earlier work this paper cites.
Sequence level training with recurrent neural networks
Ranzato, M., Chopra, S., Auli, M., and Zaremba, W · 2016
Earlier work this paper cites.
Probabilistic model for code with decision trees
Raychev, V., Bielik, P., and Vechev, M · 2016
Earlier work this paper cites.
f f -divergence inequalities
Sason, I. and Verdú, S · 2016
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M. I., and Abbeel, P · 2016
Earlier work this paper cites.
A note on the evaluation of generative models
Theis, L., van den Oord, A., and Bethge, M · 2016
Cited alongside, same era.
Wasserstein generative adversarial networks
Arjovsky, M., Chintala, S., and Bottou, L · 2017
Cited alongside, same era.
An actor-critic algorithm for sequence prediction
Bahdanau, D., Brakel, P., Xu, K., Goyal, A., Lowe, R., Pineau, J., Courville, A. C., and Bengio, Y · 2017
Cited alongside, same era.
Mode regularized generative adversarial networks
Che, T., Li, Y., Jacob, A. P., Bengio, Y., and Li, W · 2017
Cited alongside, same era.
Sequence tutor: Conservative fine-tuning of sequence generation models with kl-control
Jaques, N., Gu, S., Bahdanau, D., Hernández-Lobato, J. M., Turner, R. E., and Eck, D · 2017
Cited alongside, same era.
Reinforced video captioning with entailment rewards
A general language assistant as a laboratory for alignment
Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., DasSarma, N., et al · 2021
Later among the works it cites.
GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow, 2021
Black, S., Gao, L., Wang, P., Leahy, C., and Biderman, S · 2021
Later among the works it cites.
Imitation learning as f-divergence minimization
Ke, L., Choudhury, S., Barnes, M., Sun, W., Lee, G., and Srinivasa, S. S · 2021
Later among the works it cites.
A distributional approach to controlled text generation
Khalifa, M., Elsahar, H., and Dymetman, M · 2021
Later among the works it cites.
Entity-level factual consistency of abstractive text summarization
Nan, F., Nallapati, R., Wang, Z., Nogueira dos Santos, C., Zhu, H., Zhang, D., McKeown, K., and Xiang, B · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pasunuru, R. and Bansal, M · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Hierarchical neural story generation
Fan, A., Lewis, M., and Dauphin, Y · 2018
Cited alongside, same era.
Reinforcement learning and control as probabilistic inference: Tutorial and review
Levine, S · 2018
Cited alongside, same era.
Which training methods for gans do actually converge?
Mescheder, L. M., Geiger, A., and Nowozin, S · 2018
Cited alongside, same era.
A deep reinforced model for abstractive summarization
Paulus, R., Xiong, C., and Socher, R · 2018
Cited alongside, same era.
On f-divergences: Integral representations, local behavior, and inequalities
Sason, I · 2018
Cited alongside, same era.
Ngo, H., Raterink, C., Araújo, J. G., Zhang, I., Chen, C., Morisot, A., and Frosst, N · 2021
Later among the works it cites.
Recipes for building an open-domain chatbot
Roller, S., Dinan, E., Goyal, N., Ju, D., Williamson, M., Liu, Y., Xu, J., Ott, M., Smith, E. M., Boureau, Y.-L., and Weston, J · 2021
Later among the works it cites.
Process for adapting language models to society (palms) with values-targeted datasets
Solaiman, I. and Dennison, C · 2021
Later among the works it cites.
Ethical and social risks of harm from language models
Weidinger, L., Mellor, J., Rauh, M., Griffin, C., Uesato, J., Huang, P., Cheng, M., Glaese, M., Balle, B., Kasirzadeh, A., Kenton, Z., Brown, S., Hawkins, W., Stepleton, T., Biles, C., Birhane, A., Haas, J., Rimell, L., Hendricks, L. A., Isaac, W. S., Legassick, S., Irving, G., and Gabriel, I · 2021
Later among the works it cites.
Challenges in detoxifying language models
Welbl, J., Glaese, A., Uesato, J., Dathathri, S., Mellor, J., Hendricks, L. A., Anderson, K., Kohli, P., Coppin, B., and Huang, P.-S · 2021
Later among the works it cites.
Bot-adversarial dialogue for safe conversational agents
Xu, J., Ju, D., Li, M., Boureau, Y.-L., Weston, J., and Dinan, E · 2021
Later among the works it cites.
Ethical-advice taker: Do language models understand natural language interventions?
Zhao, J., Khashabi, D., Khot, T., Sabharwal, A., and Chang, K.-W · 2021
Later among the works it cites.
Theory-grounded measurement of U.S. social stereotypes in English language models
Cao, Y., Sotnikova, A., Daumé III, H., Rudinger, R., and Zou, L · 2022
Later among the works it cites.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., et al · 2022
Later among the works it cites.
Dohan, D., Xu, W., Lewkowycz, A., Austin, J., Bieber, D., Lopes, R. G., Wu, Y., Michalewski, H., Saurous, R. A., Sohl-Dickstein, J., et al · 2022
Later among the works it cites.
An approximate sampler for energy-based models with divergence diagnostics
Eikema, B., Kruszewski, G., Dance, C. R., Elsahar, H., and Dymetman, M · 2022
Later among the works it cites.
Improving alignment of dialogue agents via targeted human judgements
Glaese, A., McAleese, N., Trebacz, M., Aslanides, J., Firoiu, V., Ewalds, T., Rauh, M., Weidinger, L., Chadwick, M., Thacker, P., Campbell-Gillingham, L., Uesato, J., Huang, P., Comanescu, R., Yang, F., See, A., Dathathri, S., Greig, R., Chen, C., Fritz, D., Elias, J. S., Green, R., Mokrá, S., Fernando, N., Wu, B., Foley, R., Young, S., Gabriel, I., Isaac, W., Mellor, J., Hassabis, D., Kavukcuoglu, K., Hendricks, L. A., and Irving, G · 2022
Later among the works it cites.
Exposing the implicit energy networks behind masked language models via metropolis–hastings
Goyal, K., Dyer, C., and Berg-Kirkpatrick, T · 2022
Later among the works it cites.
distilbert-base-uncased-finetuned-sst-2-english (revision bfdd146), 2022
HF Canonical Model Maintainers · 2022
Later among the works it cites.
TruthfulQA: Measuring how models mimic human falsehoods
Lin, S., Hilton, J., and Evans, O · 2022
Later among the works it cites.
Teaching language models to support answers with verified quotes, 2022
Menick, J., Trebacz, M., Mikulik, V., Aslanides, J., Song, F., Chadwick, M., Glaese, M., Young, S., Campbell-Gillingham, L., Irving, G., and McAleese, N · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Gray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R · 2022
Later among the works it cites.
Cold decoding: Energy-based constrained text generation with langevin dynamics
Qin, L., Welleck, S., Khashabi, D., and Choi, Y · 2022
Later among the works it cites.
Self-critiquing models for assisting human evaluators
Saunders, W., Yeh, C., Wu, J., Bills, S., Ouyang, L., Ward, J., and Leike, J · 2022
Later among the works it cites.
Training language models with natural language feedback
Scheurer, J., Campos, J. A., Chan, J. S., Chen, A., Cho, K., and Perez, E · 2022
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., Kluska, A., Lewkowycz, A., Agarwal, A., Power, A., Ray, A., Warstadt, A., Kocurek, A. W., Safaya, A., Tazarv, A., Xiang, A., Parrish, A., Nie, A., Hussain, A., Askell, A., Dsouza, A., Rahane, A., Iyer, A. S., Andreassen, A., Santilli, A., Stuhlmüller, A., Dai, A. M., La, A., Lampinen, A. K., Zou, A., Jiang, A., Chen, A., Vuong, A., Gupta, A., Gottardi, A., Norelli, A., Venkatesh, A., Gholamidavoodi, A., Tabassum, A., Menezes, A., Kirubarajan, A., Mullokandov, A., Sabharwal, A., Herrick, A., Efrat, A., Erdem, A., Karakas, A., and et al · 2022
Later among the works it cites.
Lamda: Language models for dialog applications
Thoppilan, R., Freitas, D. D., Hall, J., Shazeer, N., Kulshreshtha, A., Cheng, H., Jin, A., Bos, T., Baker, L., Du, Y., Li, Y., Lee, H., Zheng, H. S., Ghafouri, A., Menegali, M., Huang, Y., Krikun, M., Lepikhin, D., Qin, J., Chen, D., Xu, Y., Chen, Z., Roberts, A., Bosma, M., Zhou, Y., Chang, C., Krivokon, I., Rusch, W., Pickett, M., Meier-Hellstern, K. S., Morris, M. R., Doshi, T., Santos, R. D., Duke, T., Soraker, J., Zevenbergen, B., Prabhakaran, V., Diaz, M., Hutchinson, B., Olson, K., Molina, A., Hoffman-John, E., Lee, J., Aroyo, L., Rajakumar, R., Butryna, A., Lamm, M., Kuzmina, V., Fenton, J., Cohen, A., Bernstein, R., Kurzweil, R., Aguera-Arcas, B., Cui, C., Croak, M., Chi, E. H., and Le, Q · 2022
Later among the works it cites.
STar: Bootstrapping reasoning with reasoning
Zelikman, E., Wu, Y., Mu, J., and Goodman, N · 2022
Later among the works it cites.
Adversarial training for high-stakes reliability
Ziegler, D. M., Nix, S., Chan, L., Bauman, T., Schmidt-Nielsen, P., Lin, T., Scherlis, A., Nabeshima, N., Weinstein-Raun, B., de Haas, D., Shlegeris, B., and Thomas, N · 2022
Later among the works it cites.