DICES Dataset: Diversity in Conversational AI Evaluation for Safety
Original
Aroyo, L.; Taylor, A. S.; Diaz, M.; Homan, C. M.; Parrish, A.; Serapio-Garcia, G.; Prabhakaran, V.; and Wang, D. 2023 · 2023
Closest in time.
I2D2: Inductive Knowledge Distillation with NeuroLogic and Self-Imitation
Original
Bhagavatula, C.; Hwang, J. D.; Downey, D.; Bras, R. L.; Lu, X.; Qin, L.; Sakaguchi, K.; Swayamdipta, S.; West, P.; and Choi, Y. 2023 · 2023
Closest in time.
Pluralism — Ideology, Diversity & Tolerance — britannica.com
Britannica Editors. 2002 · 2023
Closest in time.
Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Original
Casper, S.; Davies, X.; Shi, C.; Gilbert, T. K.; Scheurer, J.; Rando, J.; Freedman, R.; Korbak, T.; Lindner, D.; Freire, P.; Wang, T.; Marks, S.; Segerie, C.-R.; Carroll, M.; Peng, A.; Christoffersen, P.; Damani, M.; Slocum, S.; Anwar, U.; Siththaranjan, A.; Nadeau, M.; Michaud, E. J.; Pfau, J.; Krasheninnikov, D.; Chen, X.; Langosco, L.; Hase, P.; Bıyık, E.; Dragan, A.; Krueger, D.; Sadigh, D.; and Hadfield-Menell, D. 2023 · 2023
Closest in time.
Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality
Chiang, W.-L.; Li, Z.; Lin, Z.; Sheng, Y.; Wu, Z.; Zhang, H.; Zheng, L.; Zhuang, S.; Zhuang, Y.; Gonzalez, J. E.; Stoica, I.; and Xing, E. P. 2023 · 2023
Closest in time.
Towards Measuring the Representation of Subjective Global Opinions in Language Models
Original
Durmus, E.; Nyugen, K.; Liao, T. I.; Schiefer, N.; Askell, A.; Bakhtin, A.; Chen, C.; Hatfield-Dodds, Z.; Hernandez, D.; Joseph, N.; et al. 2023 · 2023
Closest in time.
From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models
Original
Feng, S.; Park, C. Y.; Liu, Y.; and Tsvetkov, Y. 2023 · 2023
Closest in time.
ChatGPT outperforms crowd workers for text-annotation tasks
Gilardi, F.; Alizadeh, M.; and Kubli, M. 2023 · 2023
Closest in time.
Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes
Hsieh, C.-Y.; Li, C.-L.; Yeh, C.-k.; Nakhost, H.; Fujii, Y.; Ratner, A.; Krishna, R.; Lee, C.-Y.; and Pfister, T. 2023 · 2023
Closest in time.
Impossible Distillation: from Low-Quality Model to High-Quality Dataset & Model for Summarization and Paraphrasing
Original
Jung, J.; West, P.; Jiang, L.; Brahman, F.; Lu, X.; Fisher, J.; Sorensen, T.; and Choi, Y. 2023 · 2023
Closest in time.
SODA: Million-scale Dialogue Distillation with Social Commonsense Contextualization
Original
Kim, H.; Hessel, J.; Jiang, L.; West, P.; Lu, X.; Yu, Y.; Zhou, P.; Bras, R. L.; Alikhani, M.; Kim, G.; Sap, M.; and Choi, Y. 2023 · 2023
Closest in time.
What does a Text Classifier Learn about Morality? An Explainable Method for Cross-Domain Comparison of Moral Rhetoric
Liscio, E.; Araque, O.; Gatti, L.; Constantinescu, I.; Jonker, C.; Kalimeri, K.; and Murukannaiah, P. K. 2023 · 2023
Closest in time.
Learning Ambiguity from Crowd Sequential Annotations
Original
Lu, X. 2023 · 2023
Closest in time.
Inference-Time Policy Adapters (IPA): Tailoring Extreme-Scale LMs without Fine-tuning
Original
Lu, X.; Brahman, F.; West, P.; Jang, J.; Chandu, K.; Ravichander, A.; Qin, L.; Ammanabrolu, P.; Jiang, L.; Ramnath, S.; et al. 2023 · 2023
Closest in time.
GPT-4 Technical Report
Original
OpenAI. 2023 · 2023
Closest in time.
ClarifyDelphi: Reinforced Clarification Questions with Defeasibility Rewards for Social and Moral Situations
Pyatkin, V.; Hwang, J. D.; Srikumar, V.; Lu, X.; Jiang, L.; Choi, Y.; and Bhagavatula, C. 2023 · 2023
Closest in time.
Relationship Between Basic Human Values and Decision-Making Styles in Adolescents
Páez, J.; De-Juanas, A.; García-Castilla, F.; and Muelas, A. 2020 · 2023
Closest in time.
Towards Coding Social Science Datasets with Language Models
Original
Rytting, C. M.; Sorensen, T.; Argyle, L.; Busby, E.; Fulda, N.; Gubler, J.; and Wingate, D. 2023 · 2023
Closest in time.
Whose Opinions Do Language Models Reflect?
Original
Santurkar, S.; Durmus, E.; Ladhak, F.; Lee, C.; Liang, P.; and Hashimoto, T. 2023 · 2023
Closest in time.
NLPositionality: Characterizing Design Biases of Datasets and Models
Original
Santy, S.; Liang, J. T.; Bras, R. L.; Reinecke, K.; and Sap, M. 2023 · 2023
Closest in time.
Evaluating the Moral Beliefs Encoded in LLMs
Original
Scherrer, N.; Shi, C.; Feder, A.; and Blei, D. M. 2023 · 2023
Closest in time.
NovaCOMET: Open Commonsense Foundation Models with Symbolic Knowledge Distillation
Original
West, P.; Bras, R. L.; Sorensen, T.; Lin, B. Y.; Jiang, L.; Lu, X.; Chandu, K.; Hessel, J.; Baheti, A.; Bhagavatula, C.; and Choi, Y. 2023 · 2023
Closest in time.
Fine-Grained Human Feedback Gives Better Rewards for Language Model Training
Original
Wu, Z.; Hu, Y.; Shi, W.; Dziri, N.; Suhr, A.; Ammanabrolu, P.; Smith, N. A.; Ostendorf, M.; and Hajishirzi, H. 2023 · 2023
Closest in time.
Baize: An Open-Source Chat Model with Parameter-Efficient Tuning on Self-Chat Data
Original
Xu, C.; Guo, D.; Duan, N.; and McAuley, J. 2023 · 2023
Closest in time.
COBRA Frames: Contextual Reasoning about Effects and Harms of Offensive Statements
Zhou, X.; Zhu, H.; Yerukola, A.; Davidson, T.; Hwang, J. D.; Swayamdipta, S.; and Sap, M. 2023 · 2023
Closest in time.
Can Large Language Models Transform Computational Social Science?
Original
Ziems, C.; Held, W.; Shaikh, O.; Chen, J.; Zhang, Z.; and Yang, D. 2023 · 2023
Closest in time.