Fetching the paper…
Reading the bibliography…
We argue that many general evaluation problems can be viewed through the lens of voting theory.
OpenSpiel: A framework for reinforcement learning in games
M. Lanctot, E. Lockhart, J.-B. Lespiau, V. Zambaldi, S. Upadhyay, J. Pérolat, S. Srinivasan, F. Timbers, K. Tuyls, S. Omidshafiei, D. Hennes, D. Morrill, P. Muller, T. Ewalds, R. Faulkner, J. Kramár, B. D. Vylder, B. Saeta, J. Bradbury, D. Ding, S. Borgeaud, M. Lai, J. Schrittwieser, T. Anthony, E. Hughes, I. Danihelka, and J. Ryan-Davis · 1908
Earlier work this paper cites.
A new measure of rank correlation
M. Kendall · 1938
Earlier work this paper cites.
A ’reasonable’ social welfare function, 1951
A. H. Copeland · 1951
Earlier work this paper cites.
Mathematics without numbers
J. Kemeny · 1959
Earlier work this paper cites.
Aggregation of preference orderings
G. Kreweras · 1965
Earlier work this paper cites.
Choice functions and revealed preference
A. K. Sen · 1971
Earlier work this paper cites.
Aggregation of preferences with variable electorate
J. Smith · 1973
Earlier work this paper cites.
Social choice scoring functions
H. P. Young · 1975
Earlier work this paper cites.
Social choice theory: A re-examination
A. K. Sen · 1977
Earlier work this paper cites.
The Ratings of Chess Players, Past and Present
A. E. Elo · 1978
Earlier work this paper cites.
A consistent extension of Condorcet’s election principle
H. P. Young and A. Levenglick · 1978
Earlier work this paper cites.
Probabilistic social choice based on simple voting comparisons
P. C. Fishburn · 1984
Earlier work this paper cites.
Independence of clones as a criterion for voting rules
N. Tideman · 1987
Earlier work this paper cites.
Measuring the non-transitivity in chess
R. Sanjaya, J. Wang, and Y. Yang · 1999
Earlier work this paper cites.
Discrete Mathematics and Its Applications
K. H. Rosen · 2003
Earlier work this paper cites.
Bayesian Elo rating, 2005
R. Coulom · 2005
Earlier work this paper cites.
Trueskill™: A Bayesian skill rating system
R. Herbrich, T. Minka, and T. Graepel · 2006
Earlier work this paper cites.
Human-level performance in no-press diplomacy via equilibrium search
J. Gray, A. Lerer, A. Bakhtin, and N. Brown · 2010
Earlier work this paper cites.
A new monotonic, clone-independent, reversal symmetric, and Condorcet-consistent single-winner election method
M. Schulze · 2011
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Cited alongside, same era.
General tiebreaking schemes for computational social choice
R. Freeman, M. Brill, and V. Conitzer · 2015
Cited alongside, same era.
Consistent probabilistic social choice
F. Brandl, F. Brandt, and H. G. Seedig · 2016
Cited alongside, same era.
Handbook of Computational Social Choice
F. Brandt, V. Conitzer, U. Endriss, J. Lang, and A. D. Procaccia, editors · 2016
Cited alongside, same era.
Rolling the dice: Recent results in probabilistic social choice
F. Brandt · 2017
Cited alongside, same era.
Arrovian aggregation of convex preferences
F. Brandl and F. Brandt · 2020
Later among the works it cites.
Real world games look like spinning tops
W. M. Czarnecki, G. Gidel, B. Tracey, K. Tuyls, S. Omidshafiei, D. Balduzzi, and M. Jaderberg · 2020
Later among the works it cites.
Evaluating the performance of reinforcement learning algorithms
S. Jordan, Y. Chandak, D. Cohen, M. Zhang, and P. Thomas · 2020
Later among the works it cites.
Explainable voting
D. Peters, A. D. Procaccia, A. Psomas, and Z. Zhou · 2020
Later among the works it cites.
Distortion in social choice problems: The first 15 years and beyond
E. Anshelevich, A. Filos-Ratsikas, N. Shah, and A. A. Voudouris · 2021
Later among the works it cites.
Open-ended learning leads to generally capable agents
O. Team, A. Stooke, A. Mahajan, C. Barros, C. Deck, J. Bauer, J. Sygnowski, M. Trebacz, M. Jaderberg, M. Mathieu, N. McAleese, N. Bradley-Schmieg, N. Wong, N. Porcel, R. Raileanu, S. Hughes-Fitt, V. Dalibard, and W. M. Czarnecki · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Trends in Computational Social Choice
U. Endriss · 2017
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
M. Hessel, J. Modayil, H. van Hasselt, T. Schaul, G. Ostrovski, W. Dabney, D. Horgan, B. Piot, M. G. Azar, and D. Silver · 2017
Cited alongside, same era.
Simple, robust and optimal ranking from pairwise comparisons
N. B. Shah and M. J. Wainwright · 2017
Cited alongside, same era.
Re-evaluating evaluation
D. Balduzzi, K. Tuyls, J. Perolat, and T. Graepel · 2018
Cited alongside, same era.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
M. C. Machado, M. G. Bellemare, E. Talvitie, J. Veness, M. J. Hausknecht, and M. Bowling · 2018
Cited alongside, same era.
A generalized method for empirical game theoretic analysis
K. Tuyls, J. Perolat, M. Lanctot, J. Z. Leibo, and T. Graepel · 2018
Cited alongside, same era.
Later among the works it cites.
A quantitative and qualitative analysis of the robustness of (real-world) election winners
N. Boehmer, R. Bredereck, P. Faliszewski, and R. Niedermeier · 2022
Later among the works it cites.
Human-level play in the game of Diplomacy by combining language models with strategic reasoning
M. F. A. R. D. T. (FAIR)†, A. Bakhtin, N. Brown, E. Dinan, G. Farina, C. Flaherty, D. Fried, A. Goff, J. Gray, H. Hu, A. P. Jacob, M. Komeili, K. Konath, M. Kwon, A. Lerer, M. Lewis, A. H. Miller, S. Mitts, A. Renduchintala, S. Roller, D. Rowe, W. Shi, J. Spisak, A. Wei, D. Wu, H. Zhang, and M. Zijlstra · 2022
Later among the works it cites.
Holistic evaluation of language models
P. Liang, R. Bommasani, T. Lee, D. Tsipras, D. Soylu, M. Yasunaga, Y. Zhang, D. Narayanan, Y. Wu, A. Kumar, B. Newman, B. Yuan, B. Yan, C. Zhang, C. Cosgrove, C. D. Manning, C. Ré, D. Acosta-Navas, D. A. Hudson, E. Zelikman, E. Durmus, F. Ladhak, F. Rong, H. Ren, H. Yao, J. Wang, K. Santhanam, L. Orr, L. Zheng, M. Yuksekgonul, M. Suzgun, N. Kim, N. Guha, N. Chatterji, O. Khattab, P. Henderson, Q. Huang, R. Chi, S. M. Xie, S. Santurkar, S. Ganguli, T. Hashimoto, T. Icard, T. Zhang, V. Chaudhary, W. Wang, X. Li, Y. Mai, Y. Zhang, and Y. Koreeda · 2022
Later among the works it cites.
Game theoretic rating in n n -player general-sum games with equilibria
L. Marris, M. Lanctot, I. Gemp, S. Omidshafiei, S. McAleer, J. Connor, K. Tuyls, and T. Graepel · 2022
Later among the works it cites.
Ghost ratings, 2023
T. Anthony · 2023
Closest in time.
On the limitations of the elo: Real-world games are transitive, not additive
Q. Bertand, W. M. Czarnecki, and G. Gidel · 2023
Closest in time.
Rethink reporting of evaluation results in AI
R. Burnell, W. Schellaert, J. Burden, T. D. Ullman, F. Martinez-Plumed, J. B. Tenenbaum, D. Rutar, L. G. Cheke, J. Sohl-Dickstein, M. Mitchell, et al · 2023
Closest in time.
Agentbench: Evaluating llms as agents, 2023
X. Liu, H. Yu, H. Zhang, Y. Xu, X. Lei, H. Lai, Y. Gu, H. Ding, K. Men, K. Yang, S. Zhang, X. Deng, A. Zeng, Z. Du, C. Zhang, S. Shen, T. Zhang, Y. Su, H. Sun, M. Huang, Y. Dong, and J. Tang · 2023
Closest in time.
Chatbot arena conversation dataset release, 2023
LMSysOrg · 2023
Closest in time.
Human-timescale adaptation in an open-ended task space
A. A. Team, J. Bauer, K. Baumli, S. Baveja, F. Behbahani, A. Bhoopchand, N. Bradley-Schmieg, M. Chang, N. Clay, A. Collister, V. Dasagi, L. Gonzalez, K. Gregor, E. Hughes, S. Kashem, M. Loks-Thompson, H. Openshaw, J. Parker-Holder, S. Pathak, N. Perez-Nieves, N. Rakicevic, T. Rocktäschel, Y. Schroecker, J. Sygnowski, K. Tuyls, S. York, A. Zacherl, and L. Zhang · 2023
Closest in time.
webdiplomacy
webDiplomacy Development Team · 2023
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena, 2023
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica · 2023
Closest in time.