Fetching the paper…
Reading the bibliography…
Numerous capability and safety techniques of Large Language Models (LLMs), including RLHF, automated red-teaming, prompt engineering, and infilling, can be cast as sampling from an unnormalized target distribution defined by a given reward or potential function over the full sequence.
Sequential Monte Carlo methods in practice , volume 1
Doucet, A., De Freitas, N., Gordon, N. J., et al · 2001
Earlier work this paper cites.
On the optimality of conditional expectation as a bregman predictor
Banerjee, A., Guo, X., and Wang, H · 2005
Earlier work this paper cites.
Sequential monte carlo samplers
Del Moral, P., Doucet, A., and Jasra, A · 2006
Earlier work this paper cites.
Particle markov chain monte carlo methods
Andrieu, C., Doucet, A., and Holenstein, R · 2010
Earlier work this paper cites.
Smoothing algorithms for state–space models
Briers, M., Doucet, A., and Maskell, S · 2010
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Gutmann, M. and Hyvärinen, A · 2010
Earlier work this paper cites.
Twisted particle filters
Whiteley, N. and Lee, A · 2014
Earlier work this paper cites.
Importance weighted autoencoders
Burda, Y., Grosse, R., and Salakhutdinov, R · 2015
Earlier work this paper cites.
On extended state-space constructions for Monte Carlo methods
Finke, A · 2015
Earlier work this paper cites.
Sandwiching the marginal likelihood using bidirectional monte carlo
Grosse, R. B., Ghahramani, Z., and Adams, R. P · 2015
Earlier work this paper cites.
Neural adaptive sequential monte carlo
Gu, S. S., Ghahramani, Z., and Turner, R. E · 2015
Earlier work this paper cites.
Measuring the reliability of mcmc inference with bidirectional monte carlo
Grosse, R. B., Ancha, S., and Roy, D · 2016
Earlier work this paper cites.
Filtering variational objectives
Maddison, C. J., Lawson, J., Tucker, G., Heess, N., Norouzi, M., Mnih, A., Doucet, A., and Teh, Y · 2017
Earlier work this paper cites.
Bridging the gap between value and policy based reinforcement learning
Nachum, O., Norouzi, M., Xu, K., and Schuurmans, D · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Importance weighting and variational inference
Domke, J. and Sheldon, D. R · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Earlier work this paper cites.
Twisted variational sequential monte carlo
Lawson, D., Tucker, G., Naesseth, C. A., Maddison, C., Adams, R. P., and Teh, Y. W · 2018
Earlier work this paper cites.
Reinforcement learning and control as probabilistic inference: Tutorial and review
Levine, S · 2018
Earlier work this paper cites.
Probabilistic planning with sequential monte carlo methods
Piché, A., Thomas, V., Ibrahim, C., Bengio, Y., and Pal, C · 2018
Earlier work this paper cites.
Plug and play language models: A simple approach to controlled text generation
Dathathri, S., Madotto, A., Lan, J., Hung, J., Frank, E., Molino, P., Yosinski, J., and Liu, R · 2019
Earlier work this paper cites.
The curious case of neural text degeneration
Holtzman, A., Buys, J., Du, L., Forbes, M., and Choi, Y · 2019
Cited alongside, same era.
Distributional reinforcement learning for energy-based sequential models
Parshakova, T., Andreoli, J.-M., and Dymetman, M · 2019
Cited alongside, same era.
Importance weighted hierarchical variational inference
Sobolev, A. and Vetrov, D. P · 2019
Cited alongside, same era.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 2019
Cited alongside, same era.
An introduction to sequential Monte Carlo , volume 4
Chopin, N., Papaspiliopoulos, O., et al · 2020
Cited alongside, same era.
Critic sequential monte carlo
Lioutas, V., Lavington, J. W., Sefas, J., Niedoba, M., Liu, Y., Zwartsenberg, B., Dabiri, S., Wood, F., and Scibior, A · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Later among the works it cites.
Red teaming language models with language models
Perez, E., Huang, S., Song, F., Cai, T., Ring, R., Aslanides, J., Glaese, A., McAleese, N., and Irving, G · 2022
Later among the works it cites.
Cold decoding: Energy-based constrained text generation with langevin dynamics
Qin, L., Welleck, S., Khashabi, D., and Choi, Y · 2022
Later among the works it cites.
Offline rl for natural language generation with implicit language q learning
Snell, C. V., Kostrikov, I., Su, Y., Yang, S., and Levine, S · 2022
Later among the works it cites.
Aira, 2023
Corrêa, N. K · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Heng, J., Bishop, A., Deligiannidis, G., and Doucet, A · 2020
Cited alongside, same era.
A distributional approach to controlled text generation
Khalifa, M., Elsahar, H., and Dymetman, M · 2020
Cited alongside, same era.
Gedi: Generative discriminator guided sequence generation
Krause, B., Gotmare, A. D., McCann, B., Keskar, N. S., Joty, S., Socher, R., and Rajani, N. F · 2020
Cited alongside, same era.
Learning to summarize with human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F · 2020
Cited alongside, same era.
Learning to give checkable answers with prover-verifier games
Anil, C., Zhang, G., Wu, Y., and Grosse, R · 2021
Cited alongside, same era.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Cited alongside, same era.
Efficient (soft) q-learning for text generation with limited good data
Guo, H., Tan, B., Liu, Z., Xing, E. P., and Hu, Z · 2021
Cited alongside, same era.
Later among the works it cites.
Reward-augmented decoding: Efficient controlled text generation with a unidirectional reward model
Deng, H. and Raffel, C · 2023
Later among the works it cites.
Tinystories: How small can language models be and still speak coherent english?
Eldan, R. and Li, Y · 2023
Later among the works it cites.
Aligning foundation models for language with preferences through f f -divergence minimization
Go, D., Korbak, T., Kruszewski, G., Rozen, J., Ryu, N., and Dymetman, M · 2023
Later among the works it cites.
Amortizing intractable inference in large language models
Hu, E. J., Jain, M., Elmoznino, E., Kaddar, Y., Lajoie, G., Bengio, Y., and Malkin, N · 2023
Later among the works it cites.
Sequential monte carlo steering of large language models using probabilistic programs
Lew, A. K., Zhi-Xuan, T., Grand, G., and Mansinghka, V. K · 2023
Later among the works it cites.
Don’t throw away your value model! making ppo even better via value-guided monte-carlo tree search decoding
Liu, J., Cohen, A., Pasunuru, R., Choi, Y., Hajishirzi, H., and Celikyilmaz, A · 2023
Later among the works it cites.
Controlled decoding from language models
Mudgal, S., Lee, J., Ganapathy, H., Li, Y., Wang, T., Huang, Y., Chen, Z., Cheng, H.-T., Collins, M., Strohman, T., et al · 2023
Later among the works it cites.
Training chain-of-thought via latent-variable inference
Phan, D., Hoffman, M. D., Douglas, S., Le, T. A., Parisi, A. T., Sountsov, P., Sutton, C., Vikram, S., Saurous, R. A., et al · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C · 2023
Later among the works it cites.
Long horizon temperature scaling
Shih, A., Sadigh, D., and Ermon, S · 2023
Later among the works it cites.
Arithmetic sampling: parallel diverse decoding for large language models
Vilnis, L., Zemlyanskiy, Y., Murray, P., Passos, A. T., and Sanghai, S · 2023
Later among the works it cites.
A survey of controllable text generation using transformer-based pre-trained language models
Zhang, H., Song, H., Li, S., Zhou, M., and Song, D · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M · 2023
Later among the works it cites.
ARGS: Alignment as reward-guided search
Khanov, M., Burapacheep, J., and Li, Y · 2024
Closest in time.