Fetching the paper…
Reading the bibliography…
Foundation models such as GPT-4 are fine-tuned to avoid unsafe or otherwise problematic behavior, such as helping to commit crimes or producing racist text.
Social Choice and Individual Values
Arrow, K. J · 1951
Earlier work this paper cites.
The Theory of Social Choice
Fishburn, P. C · 1973
Earlier work this paper cites.
Manipulation of voting schemes: a general result
Gibbard, A · 1973
Earlier work this paper cites.
Strategy-proofness and Arrow’s conditions: Existence and correspondence theorems for voting procedures and social welfare functions
Satterthwaite, M · 1975
Earlier work this paper cites.
Prospect theory: An analysis of decision under risk
Kahneman, D. and Tversky, A · 1979
Earlier work this paper cites.
On strategy-proofness and single peakedness
Moulin, H · 1980
Earlier work this paper cites.
Optimal decision rules in uncertain dichotomous choice situations
Nitzan, S. and Paroush, J · 1982
Earlier work this paper cites.
Social Choice and Multicriterion Decision-Making
Arrow, K. J. and Raynaud, H · 1986
Earlier work this paper cites.
Algebraic aggregation theory
Rubinstein, A. and Fishburn, P. C · 1986
Earlier work this paper cites.
Independence of clones as a criterion for voting rules
Tideman, T. N · 1987
Earlier work this paper cites.
Social Choice Theory: An Introduction
Kelly, J. S · 1988
Earlier work this paper cites.
Classics of Social Choice
McLean, I. and Urken, A. (eds.) · 1995
Earlier work this paper cites.
Composition-consistent tournament solutions and social choice functions
Laffond, G., Lainé, J., and Laslier, J · 1996
Earlier work this paper cites.
Communication Complexity
Kushilevitz, E. and Nisan, N · 1997
Earlier work this paper cites.
Robust combinatorial auction protocol against false-name bids
Yokoo, M., Sakurai, Y., and Matsubara, S · 2001
Earlier work this paper cites.
Voting procedures
Brams, S. J. and Fishburn, P. C · 2002
Earlier work this paper cites.
Impossibility theorems in the Arrovian framework
Campbell, D. E. and Kelly, J. S · 2002
Earlier work this paper cites.
Vote elicitation: Complexity and strategy-proofness
Conitzer, V. and Sandholm, T · 2002
Earlier work this paper cites.
Social welfare functionals and interpersonal comparability
d’Aspremont, C. and Gevers, L · 2002
Earlier work this paper cites.
Aggregating sets of judgments: An impossibility result
List, C. and Pettit, P · 2002
Earlier work this paper cites.
The effect of false-name bids in combinatorial auctions: New fraud in Internet auctions
Yokoo, M., Sakurai, Y., and Matsubara, S · 2004
Earlier work this paper cites.
Communication complexity of common voting rules
Conitzer, V. and Sandholm, T · 2005
Earlier work this paper cites.
Social Choice and the Mathematics of Manipulation
Taylor, A. D · 2005
Earlier work this paper cites.
Preference elicitation in combinatorial auctions
Sandholm, T. and Boutilier, C · 2006
Earlier work this paper cites.
Vote and aggregation in combinatorial domains with structured preferences
Lang, J · 2007
Earlier work this paper cites.
Majority Judgement: Measuring, Ranking and Electing
Balinski, M. and Laraki, R · 2010
Earlier work this paper cites.
Using mechanism design to prevent false-name manipulations
Conitzer, V. and Yokoo, M · 2010
Earlier work this paper cites.
Handbook on Approval Voting
Laslier, J.-F. and Sanver, M. R. (eds.) · 2010
Earlier work this paper cites.
Computational Aspects of Cooperative Game Theory
Chalkiadakis, G., Elkind, E., and Wooldridge, M · 2011
Earlier work this paper cites.
Dominating manipulations in voting with partial information
Conitzer, V., Walsh, T., and Xia, L · 2011
Earlier work this paper cites.
Theory choice and social choice: Kuhn versus Arrow
Okasha, S · 2011
Earlier work this paper cites.
Social Choice and Individual Values
Arrow, K. J · 2012
Earlier work this paper cites.
Communication complexity of approximating voting rules
Service, T. C. and Adams, J. A · 2012
Earlier work this paper cites.
Scalable simple random sampling and stratified sampling
Meng, X · 2013
Earlier work this paper cites.
Handbook of Computational Social Choice
Brandt, F., Conitzer, V., Endriss, U., Lang, J., and Procaccia, A. D · 2015
Earlier work this paper cites.
Rationalizations of voting rules
Elkind, E. and Slinko, A · 2015
Earlier work this paper cites.
Judgment aggregation
Endriss, U · 2015
Earlier work this paper cites.
Probabilistic opinion pooling
Dietrich, F. and List, C · 2016
Cited alongside, same era.
Embedding ethical principles in collective decision support systems
Greene, J., Rossi, F., Tasioulas, J., Venable, K. B., and Williams, B. C · 2016
Cited alongside, same era.
Introduction to the theory of voting
Zwicker, W. S · 2016
Cited alongside, same era.
Justified representation in approval-based committee voting
Aziz, H., Brill, M., Conitzer, V., Elkind, E., Freeman, R., and Walsh, T · 2017
Cited alongside, same era.
Rolling the dice: Recent results in probabilistic social choice
Brandt, F · 2017
Cited alongside, same era.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Cited alongside, same era.
Moral machine or tyranny of the majority?
Feffer, M., Heidari, H., and Lipton, Z. C · 2023
Later among the works it cites.
Fish, S., Gölz, P., Parkes, D. C., Procaccia, A. D., Rusak, G., Shapira, I., and Wüthrich, M · 2023
Later among the works it cites.
Active teacher selection for reinforcement learning from human feedback
Freedman, R., Svegliato, J., Wray, K., and Russell, S · 2023
Later among the works it cites.
Representation with incomplete votes
Halpern, D., Kehne, G., Procaccia, A. D., Tucker-Foltz, J., and Wüthrich, M · 2023
Later among the works it cites.
Language agents as digital representatives in collective decision-making
Jarrett, D., Pislar, M., Bakker, M. A., Tessler, M. H., Koster, R., Balaguer, J., Elie, R., Summerfield, C., and Tacchetti, A · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Properties of multiwinner voting rules
Elkind, E., Faliszewski, P., Skowron, P., and Slinko, A · 2017
Cited alongside, same era.
Multiwinner voting: A new challenge for social choice theory
Faliszewski, P., Skowron, P., Slinko, A., and Talmon, N · 2017
Cited alongside, same era.
Epistemic democracy with correlated voters
Pivato, M · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
The Moral Machine experiment
Awad, E., Dsouza, S., Kim, R., Schulz, J., Henrich, J., Shariff, A., Bonnefon, J.-F., and Rahwan, I · 2018
Cited alongside, same era.
A voting-based system for ethical decision making
Noothigattu, R., Gaikwad, S., Awad, E., Dsouza, S., Rahwan, I., Ravikumar, P., and Procaccia, A · 2018
Cited alongside, same era.
Later among the works it cites.
Kirk, H. R., Vidgen, B., Röttger, P., and Hale, S. A · 2023
Later among the works it cites.
Multi-Winner Voting with Approval Preferences
Lackner, M. and Skowron, P · 2023
Later among the works it cites.
The alignment ceiling: Objective mismatch in reinforcement learning from human feedback
Lambert, N. and Calandra, R · 2023
Later among the works it cites.
The history and risks of reinforcement learning and human feedback
Lambert, N., Gilbert, T. K., and Zick, T · 2023
Later among the works it cites.
Rlaif: Scaling reinforcement learning from human feedback with ai feedback
Lee, H., Phatale, S., Mansoor, H., Lu, K., Mesnard, T., Bishop, C., Carbune, V., and Rastogi, A · 2023
Later among the works it cites.
Eliciting human preferences with language models
Li, B. Z., Tamkin, A., Goodman, N., and Andreas, J · 2023
Later among the works it cites.
AI alignment and social choice: Fundamental limitations and policy implications
Mishra, A · 2023
Later among the works it cites.
More human than human: measuring chatgpt political bias
Motoki, F., Neto, V. P., and Rodrigues, V · 2023
Later among the works it cites.
GPT-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
Do the rewards justify the means? measuring trade-offs between rewards and ethical behavior in the machiavelli benchmark
Pan, A., Chan, J. S., Zou, A., Li, N., Basart, S., Woodside, T., Zhang, H., Emmons, S., and Hendrycks, D · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2023
Later among the works it cites.
Distributional preference learning: Understanding and accounting for hidden context in rlhf
Siththaranjan, A., Laidlaw, C., and Hadfield-Menell, D · 2023
Later among the works it cites.
Survey on reinforcement learning for language processing
Uc-Cetina, V., Navarro-Guerrero, N., Martin-Gonzalez, A., Weber, C., and Wermter, S · 2023
Later among the works it cites.
Fine-grained human feedback gives better rewards for language model training
Wu, Z., Hu, Y., Shi, W., Dziri, N., Suhr, A., Ammanabrolu, P., Smith, N. A., Ostendorf, M., and Hajishirzi, H · 2023
Later among the works it cites.
RLHF and IIA: perverse incentives
Xu, W., Dong, S., Lu, X., Lam, G., Wen, Z., and Roy, B. V · 2023
Later among the works it cites.
Representation engineering: A top-down approach to AI transparency
Zou, A., Phan, L., Chen, S., Campbell, J., Guo, P., Ren, R., Pan, A., Yin, X., Mazeika, M., Dombrowski, A.-K., et al · 2023
Later among the works it cites.
Introducing claude, 2023
Anthropic · 2024
Closest in time.
Suppressing pink elephants with direct principle feedback, 2024
Castricato, L., Lile, N., Anand, S., Schoelkopf, H., Verma, S., and Biderman, S · 2024
Closest in time.
Maxmin-rlhf: Towards equitable alignment of large language models with diverse human preferences
Chakraborty, S., Qiu, J., Yuan, H., Koppel, A., Huang, F., Manocha, D., Bedi, A. S., and Wang, M · 2024
Closest in time.
Mapping social choice theory to RLHF
Dai, J. and Fleisig, E · 2024
Closest in time.
Human-centered loss functions (halos)
Ethayarajh, K., Xu, W., Jurafsky, D., and Kiela, D · 2024
Closest in time.
Collective constitutional AI: Aligning a language model with public input
Ganguli, D. et al · 2024
Closest in time.
Axioms for ai alignment from human feedback
Ge, L., Halpern, D., Micha, E., Procaccia, A. D., Shapira, I., Vorobeychik, Y., and Wu, J · 2024
Closest in time.
Bard, 2023
Google · 2024
Closest in time.
Meta and microsoft introduce the next generation of llama, 2023
Meta · 2024
Closest in time.
Democratic inputs to ai grant program: lessons learned and implementation plans, 2024
OpenAI · 2024
Closest in time.
Cultural bias in explainable ai research: A systematic analysis
Peters, U. and Carman, M · 2024
Closest in time.
The political preferences of llms
Rozado, D · 2024
Closest in time.
A critical evaluation of ai feedback for aligning large language models, 2024
Sharma, A., Keh, S., Mitchell, E., Finn, C., Arora, K., and Kollar, T · 2024
Closest in time.
Preference ranking optimization for human alignment
Song, F., Yu, B., Li, M., Yu, H., Huang, F., Li, Y., and Wang, H · 2024
Closest in time.
A minimaximalist approach to reinforcement learning from human feedback
Swamy, G., Dann, C., Kidambi, R., Wu, Z. S., and Agarwal, A · 2024
Closest in time.