Hindutva 2.0: How a Conference on Hindu Nationalism Launches a Change in Strategy for North American Hindutva Organizations
Dheepa Sundaram. 2022 · 2022
Later among the works it cites.
Taxonomy of Risks Posed by Language Models. In Proceedings of the ACM Conference on Fairness, Accountability, and Transparency . 214–229
Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa Kasirzadeh, Courtney Biles, Sasha Brown, Zac Kenton, Will Hawkins, Tom Stepleton, Abeba Birhane, Lisa Anne Hendricks, Laura Rimell, William Isaac, Julia Haas, Sean Legassick, Geoffrey Irving, and Iason Gabriel. 2022 · 2022
Later among the works it cites.
Foundation Models: Opportunities, Risks and Mitigations
2023 · 2023
Closest in time.
Assessing LLMs for Moral Value Pluralism
Original
Noam Benkler, Drisana Mosaphir, Scott Friedman, Andrew Smart, and Sonja Schmer-Galunder. 2023 · 2023
Closest in time.
Considerations for Governing Open Foundation Models
Rishi Bommasani, Sayash Kapoor, Kevin Klyman, Shayne Longpre, Ashwin Ramaswami, Daniel Zhang, Marietje Schaake, Daniel E. Ho, Arvind Narayanan, and Percy Liang. 2023 · 2023
Closest in time.
Who Owns the Generative AI Platform?
Matt Bornstein, Guido Appenzeller, and Martin Casado. 2023 · 2023
Closest in time.
Multi-Armed Bandit Problem and Application
Djallel Bouneffouf. 2023 · 2023
Closest in time.
Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Original
Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, Jérémy Scheurer, Javier Rando, Rachel Freedman, Tomasz Korbak, David Lindner, Pedro Freire, Tony Wang, Samuel Marks, Charbel-Raphaël Segerie, Micah Carroll, Andi Peng, Phillip Christoffersen, Mehul Damani, Stewart Slocum, Usman Anwar, Anand Siththaranjan, Max Nadeau, Eric J. Michaud, Jacob Pfau, Dmitrii Krasheninnikov, Xin Chen, Lauro Langosco, Peter Hase, Erdem Bıyık, Anca Dragan, David Krueger, Dorsa Sadigh, and Dylan Hadfield-Menell. 2023 · 2023
Closest in time.
Building Socio-culturally Inclusive Stereotype Resources with Community Engagement
Original
Sunipa Dev, Jaya Goyal, Dinesh Tewari, Shachi Dave, and Vinodkumar Prabhakaran. 2023 · 2023
Closest in time.
Broadening AI Ethics Narratives: An Indic Art View. In Proceedings of the ACM Conference on Fairness, Accountability, and Transparency . 2–11
Ajay Divakaran, Aparna Sridhar, and Ramya Srinivasan. 2023 · 2023
Closest in time.
Towards Measuring the Representation of Subjective Global Opinions in Language Models
Original
Esin Durmus, Karina Nyugen, Thomas I. Liao, Nicholas Schiefer, Amanda Askell, Anton Bakhtin, Carol Chen, Zac Hatfield-Dodds, Danny Hernandez, Nicholas Joseph, Liane Lovitt, Sam McCandlish, Orowa Sikder, Alex Tamkin, Janel Thamkul, Jared Kaplan, Jack Clark, and Deep Ganguli. 2023 · 2023
Closest in time.
AI Is a Lot of Work
Josh Dzieza. 2023 · 2023
Closest in time.
From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models
Original
Shangbin Feng, Chan Young Park, Yuhan Liu, and Yulia Tsvetkov. 2023 · 2023
Closest in time.
https://twitter.com/davidfrawleyved/status/1681689995554476033?s=20
David Frawley. 2023 · 2023
Closest in time.
Atomist or Holist? A Diagnosis and Vision for More Productive Interdisciplinary AI Ethics Dialogue
Travis Greene, Amit Dhurandhar, and Galit Shmueli. 2023 · 2023
Closest in time.
Speaking Multiple Languages Affects the Moral Bias of Language Models. In Findings of the Association for Computational Linguistics . 2137–2156
Katharina Haemmerl, Bjoern Deiseroth, Patrick Schramowski, Jindřich Libovický, Constantin Rothkopf, Alexander Fraser, and Kristian Kersting. 2023 · 2023
Closest in time.
Governing Algorithms from the South: A Case Study of AI Development in Africa
Yousif Hassan. 2023 · 2023
Closest in time.
LoraHub: Efficient Cross-Task Generalization via Dynamic LoRA Composition
Original
Chengsong Huang, Qian Liu, Bill Yuchen Lin, Tianyu Pang, Chao Du, and Min Lin. 2023 · 2023
Closest in time.
Personalized Soups: Personalized Large Language Model Alignment via Post-Hoc Parameter Merging
Original
Joel Jang, Seungone Kim, Bill Yuchen Lin, Yizhong Wang, Jack Hessel, Luke Zettlemoyer, Hannaneh Hajishirzi, Yejin Choi, and Prithviraj Ammanabrolu. 2023 · 2023
Closest in time.
AI Alignment: A Comprehensive Survey
Original
Jiaming Ji, Tianyi Qiu, Boyuan Chen, Borong Zhang, Hantao Lou, Kaile Wang, Yawen Duan, Zhonghao He, Jiayi Zhou, Zhaowei Zhang, Fanzhi Zeng, Kwan Yee Ng, Juntao Dai, Xuehai Pan, Aidan O’Gara, Yingshan Lei, Hua Xu, Brian Tse, Jie Fu, Stephen McAleer, Yaodong Yang, Yizhou Wang, Song-Chun Zhu, Yike Guo, and Wen Gao. 2023 · 2023
Closest in time.
The Horrific Content a Kenyan Worker Had to See While Training ChatGPT
Alex Kantrowitz. 2023 · 2023
Closest in time.
Trustworthy AI and the Logics of Intersectional Resistance. In Proceedings of the ACM Conference on Fairness, Accountability, and Transparency . 172–182
Bran Knowles, Jasmine Fledderjohann, John T. Richards, and Kush R. Varshney. 2023 · 2023
Closest in time.
Unveiling Safety Vulnerabilities of Large Language Models. In EMNLP Workshop on Generation, Evaluation & Metrics
George Kour, Marcel Zalmanovici, Naama Zwerdling, Esther Goldbraich, Ora Nova Fandina, Ateret Anaby-Tavor, Orna Raz, and Eitan Farchi. 2023 · 2023
Closest in time.
AI Transparency in the Age of LLMs: A Human-Centered Research Roadmap
Original
Q. Vera Liao and Jennifer Wortman Vaughan. 2023 · 2023
Closest in time.
All About the Human: A Buddhist Take on AI Ethics
Chien-Te Lin. 2023 · 2023
Closest in time.
Handbook of Critical Studies of Artificial Intelligence
Simon Lindgren (Ed.). 2023 · 2023
Closest in time.
Decolonizing AI Ethics: Relational Autonomy as a Means to Counter AI Harms
Sábëlo Mhlambi and Simona Tiribelli. 2023 · 2023
Closest in time.
Auditing Large Language Models: A Three-Layered Approach
Jakob Mökander, Jonas Schuett, Hannah Rose Kirk, and Luciano Floridi. 2023 · 2023
Closest in time.
Artificial Intelligence in the Colonial Matrix of Power
James Muldoon and Boxi A. Wu. 2023 · 2023
Closest in time.
One AI Startup Wants to Tackle Bias by Teaching Black History: Equality
Jessica Nix. 2023 · 2023
Closest in time.
OpenAI Used Kenyan Workers on Less Than $2 Per Hour to Make ChatGPT Less Toxic
Billy Perrigo. 2023a · 2023
Closest in time.
The Workers Behind AI Rarely See Its Rewards. This Indian Startup Wants to Fix That
Billy Perrigo. 2023b · 2023
Closest in time.
Meta and IBM Launch ‘AI Alliance’ to Promote Open-Source AI Development
Associated Press. 2023 · 2023
Closest in time.
Fine-Tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Original
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. 2023 · 2023
Closest in time.
On AI Anthropomorphism
Ben Shneiderman and Michael Muller. 2023 · 2023
Closest in time.
Against Relationality: A Response to Abeba Birhane
Mimi St. Johns. 2023 · 2023
Closest in time.
Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human Supervision
Zhiqing Sun, Yikang Shen, Qinhong Zhou, Hongxin Zhang, Zhenfang Chen, David Cox, Yiming Yang, and Chuang Gan. 2023 · 2023
Closest in time.
AI Empire: Unraveling the Interlocking Systems of Oppression in Generative AI’s Global Order
Jasmina Tacheva and Srividya Ramasubramanian. 2023 · 2023
Closest in time.
Aligning Large Language Models with Human: A Survey
Original
Yufei Wang, Wanjun Zhong, Liangyou Li, Fei Mi, Xingshan Zeng, Wenyong Huang, Lifeng Shang, Xin Jiang, and Qun Liu. 2023 · 2023
Closest in time.
Open (For Business): Big Tech, Concentrated Power, and the Political Economy of Open AI
David Gray Widder, Sarah West, and Meredith Whittaker. 2023 · 2023
Closest in time.
Fine-Grained Human Feedback Gives Better Rewards for Language Model Training
Original
Zeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri, Alane Suhr, Prithviraj Ammanabrolu, Noah A. Smith, Mari Ostendorf, and Hannaneh Hajishirzi. 2023 · 2023
Closest in time.
On Diversified Preferences of Large Language Model Alignment
Original
Dun Zeng, Yong Dai, Pengyu Cheng, Tianhao Hu, Wanshun Chen, Nan Du, and Zenglin Xu. 2023 · 2023
Closest in time.
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Original
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023 · 2023
Closest in time.
Contextual Moral Value Alignment Through Context-Based Aggregation
Pierre Dognin, Jesus Rios, Ronny Luss, Inkit Padhi, Matthew D. Riemer, Miao Liu, Prasanna Sattigeri, Manish Nagireddy, Kush R. Varshney, and Djallel Bouneffouf. 2024 · 2024
Closest in time.
Open-Source AI Is Uniquely Dangerous
David Evan Harris. 2024 · 2024
Closest in time.
SocialStigmaQA: A Benchmark to Uncover Stigma Amplification in Generative Language Models. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 21454–21462
Manish Nagireddy, Lamogha Chiazor, Moninder Singh, and Ioana Baldini. 2024 · 2024
Closest in time.
A Minimaximalist Approach to Reinforcement Learning from Human Feedback
Original
Gokul Swamy, Christoph Dann, Rahul Kidambi, Zhiwei Steven Wu, and Alekh Agarwal. 2024 · 2024
Closest in time.
Hindu Nationalism and the New Jim Crow
Ashutosh Varshney and Connor Staggs. 2024 · 2024
Closest in time.