Fetching the paper…
Reading the bibliography…
We present an approach called Dialogue Action Tokens (DAT) that adapts language model agents to plan goal-directed dialogues.
Target-guided open-domain conversation
Tang, J., Zhao, T., Xiong, C., Liang, X., Xing, E. P., and Hu, Z. (2019) · 1905
Earlier work this paper cites.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
Jaques, N., Ghandeharioun, A., Shen, J. H., Ferguson, C., Lapedriza, A., Jones, N., Gu, S., and Picard, R. (2019) · 1907
Earlier work this paper cites.
Recommendation as a communication game: Self-supervised bot-play for goal-oriented dialogue
Kang, D., Balakrishnan, A., Shah, P., Crook, P., Boureau, Y.-L., and Weston, J. (2019) · 1909
Earlier work this paper cites.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G. (2019) · 1909
Earlier work this paper cites.
Plug and play language models: A simple approach to controlled text generation
Dathathri, S., Madotto, A., Lan, J., Hung, J., Frank, E., Molino, P., Yosinski, J., and Liu, R. (2019) · 1912
Earlier work this paper cites.
Speech acts: An essay in the philosophy of language
Searle, J. R. (1969) · 1969
Earlier work this paper cites.
How to do things with words
Austin, J. L. (1975) · 1975
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J. (2020) · 2005
Earlier work this paper cites.
What matters in on-policy reinforcement learning? a large-scale empirical study
Andrychowicz, M., Raichuk, A., Stańczyk, P., Orsini, M., Girgin, S., Marinier, R., Hussenot, L., Geist, M., Pietquin, O., Michalski, M., et al. (2020) · 2006
Earlier work this paper cites.
Gedi: Generative discriminator guided sequence generation
Krause, B., Gotmare, A. D., McCann, B., Keskar, N. S., Joty, S., Socher, R., and Rajani, N. F. (2020) · 2009
Earlier work this paper cites.
Deep reinforcement learning for dialogue generation
Li, J., Monroe, W., Ritter, A., Galley, M., Gao, J., and Jurafsky, D. (2016) · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al. (2016) · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D. (2017) · 2017
Earlier work this paper cites.
Deal or no deal? end-to-end learning for negotiation dialogues
Lewis, M., Yarats, D., Dauphin, Y. N., Parikh, D., and Batra, D. (2017) · 2017
Earlier work this paper cites.
Neural discrete representation learning
Van Den Oord, A., Vinyals, O., et al. (2017) · 2017
Earlier work this paper cites.
Decoupling strategy and generation in negotiation dialogues
He, H., Chen, D., Balakrishnan, A., and Liang, P. (2018) · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Earlier work this paper cites.
Control prefixes for parameter-efficient text generation
Clive, J., Cao, K., and Rei, M. (2021) · 2021
Earlier work this paper cites.
A minimalist approach to offline reinforcement learning
Fujimoto, S. and Gu, S. S. (2021) · 2021
Earlier work this paper cites.
Prefix-tuning: Optimizing continuous prompts for generation
Li, X. L. and Liang, P. (2021) · 2021
Cited alongside, same era.
Webgpt: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al. (2021) · 2021
Cited alongside, same era.
Timkey, W. and Van Schijndel, M. (2021) · 2021
Cited alongside, same era.
Human-level play in the game of diplomacy by combining language models with strategic reasoning
(FAIR)†, M. F. A. R. D. T., Bakhtin, A., Brown, N., Dinan, E., Farina, G., Flaherty, C., Fried, D., Goff, A., Gray, J., Hu, H., et al. (2022) · 2022
Cited alongside, same era.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K. (2023) · 2023
Later among the works it cites.
Decision-oriented dialogue for human-ai collaboration
Lin, J., Tomlin, N., Andreas, J., and Eisner, J. (2023) · 2023
Later among the works it cites.
Generative agents: Interactive simulacra of human behavior
Park, J. S., O’Brien, J., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S. (2023) · 2023
Later among the works it cites.
Rolellm: Benchmarking, eliciting, and enhancing role-playing abilities of large language models
Wang, Z. M., Peng, Z., Que, H., Liu, J., Zhou, W., Wu, Y., Guo, H., Gan, R., Ni, Z., Zhang, M., et al. (2023) · 2023
Later among the works it cites.
Sotopia: Interactive evaluation for social intelligence in language agents
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ganguli, D., Lovitt, L., Kernion, J., Askell, A., Bai, Y., Kadavath, S., Mann, B., Perez, E., Schiefer, N., Ndousse, K., et al. (2022) · 2022
Cited alongside, same era.
Evaluating human-language model interaction
Lee, M., Srivastava, M., Hardy, A., Thickstun, J., Durmus, E., Paranjape, A., Gerard-Ursin, I., Li, X. L., Ladhak, F., Rong, F., et al. (2022) · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. (2022) · 2022
Cited alongside, same era.
Offline rl for natural language generation with implicit language q learning
Snell, C., Kostrikov, I., Su, Y., Yang, M., and Levine, S. (2022) · 2022
Cited alongside, same era.
Extracting latent steering vectors from pretrained language models
Subramani, N., Suresh, N., and Peters, M. E. (2022) · 2022
Cited alongside, same era.
Task ambiguity in humans and language models
Tamkin, A., Handa, K., Shrestha, A., and Goodman, N. (2022) · 2022
Cited alongside, same era.
Corl: Research-oriented deep offline reinforcement learning library
Tarasov, D., Nikulin, A., Akimov, D., Kurenkov, V., and Kolesnikov, S. (2022) · 2022
Cited alongside, same era.
Chai: A chatbot ai for task-oriented dialogue with offline reinforcement learning
Verma, S., Fu, J., Yang, M., and Levine, S. (2022) · 2022
Cited alongside, same era.
Zhou, X., Zhu, H., Mathur, L., Zhang, R., Yu, H., Qi, Z., Morency, L.-P., Bisk, Y., Fried, D., Neubig, G., et al. (2023) · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M. (2023) · 2023
Later among the works it cites.
Star-gate: Teaching language models to ask clarifying questions
Andukuri, C., Fränken, J.-P., Gerstenberg, T., and Goodman, N. D. (2024) · 2024
Closest in time.
What changed? converting representational interventions to natural language
Avitan, M., Cotterell, R., Goldberg, Y., and Ravfogel, S. (2024) · 2024
Closest in time.
How well can llms negotiate? negotiationarena platform and analysis
Bianchi, F., Chia, P. J., Yuksekgonul, M., Tagliabue, J., Jurafsky, D., and Zou, J. (2024) · 2024
Closest in time.
Durably reducing conspiracy beliefs through dialogues with ai
Costello, T. H., Pennycook, G., and Rand, D. G. (2024) · 2024
Closest in time.
Safe, secure, and trustworthy development and use of artificial intelligence
Executive Office of the President (2023) · 2024
Closest in time.
Finding alignments between interpretable causal variables and distributed neural representations
Geiger, A., Wu, Z., Potts, C., Icard, T., and Goodman, N. (2024) · 2024
Closest in time.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal
Mazeika, M., Phan, L., Yin, X., Zou, A., Wang, Z., Mu, N., Sakhaee, E., Li, N., Basart, S., Li, B., et al. (2024) · 2024
Closest in time.
Great, now write an article about that: The crescendo multi-turn llm jailbreak attack
Russinovich, M., Salem, A., and Eldan, R. (2024) · 2024
Closest in time.
A critical evaluation of ai feedback for aligning large language models
Sharma, A., Keh, S., Mitchell, E., Finn, C., Arora, K., and Kollar, T. (2024) · 2024
Closest in time.
A strongreject for empty jailbreaks
Souly, A., Lu, Q., Bowen, D., Trinh, T., Hsieh, E., Pandey, S., Abbeel, P., Svegliato, J., Emmons, S., Watkins, O., et al. (2024) · 2024
Closest in time.
Preference fine-tuning of llms should leverage suboptimal, on-policy data
Tajwar, F., Singh, A., Sharma, A., Rafailov, R., Schneider, J., Xie, T., Ermon, S., Finn, C., and Kumar, A. (2024) · 2024
Closest in time.
Zeng, Y., Lin, H., Zhang, J., Yang, D., Jia, R., and Shi, W. (2024) · 2024
Closest in time.
Archer: Training language model agents via hierarchical multi-turn rl
Zhou, Y., Zanette, A., Pan, J., Levine, S., and Kumar, A. (2024) · 2024
Closest in time.