Fetching the paper…
Reading the bibliography…
We show how to assess a language model's knowledge of basic concepts of morality.
The Methods of Ethics
H. Sidgwick · 1907
Earlier work this paper cites.
The Right and the Good
W. D. Ross · 1930
Earlier work this paper cites.
Theory of games and economic behavior
J. V. Neumann and O. Morgenstern · 1944
Earlier work this paper cites.
I, Robot
I. Asimov · 1950
Earlier work this paper cites.
A comparison of three models for determining test fairness
M. A. Lewis · 1978
Earlier work this paper cites.
Reasons and Persons
D. Parfit · 1987
Earlier work this paper cites.
Realizable and unrealizable specifications of reactive systems
M. Abadi, L. Lamport, and P. Wolper · 1989
Earlier work this paper cites.
The Limits of Morality
S. Kagan · 1991
Earlier work this paper cites.
Controllers for reachability specifications for hybrid systems
J. Lygeros, C. Tomlin, and S. Sastry · 1999
Earlier work this paper cites.
A Theory of Justice
J. Rawls · 1999
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
A. Y. Ng and S. J. Russell · 2000
Earlier work this paper cites.
The moral emotions
J. Haidt et al · 2003
Earlier work this paper cites.
Capabilities as fundamental entitlements: Sen and social justice
M. Nussbaum · 2003
Earlier work this paper cites.
Learning to rank using gradient descent
C. Burges, T. Shaked, E. Renshaw, A. Lazier, M. Deeds, N. Hamilton, and G. Hullender · 2005
Earlier work this paper cites.
The nature, importance, and difficulty of machine ethics
J. H. Moor · 2006
Earlier work this paper cites.
Factorization meets the neighborhood: a multifaceted collaborative filtering model
Y. Koren · 2008
Earlier work this paper cites.
C. Dwork, M. Hardt, T. Pitassi, O. Reingold, and R. S. Zemel · 2011
Earlier work this paper cites.
General purpose intelligence: Arguing the orthogonality thesis
S. Armstrong · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
R. Socher, A. Perelygin, J. Wu, J. Chuang, C. D. Manning, A. Ng, and C. Potts · 2013
Earlier work this paper cites.
Superintelligence: Paths, dangers, strategies
N. Bostrom · 2014
Earlier work this paper cites.
Explaining and harnessing adversarial examples
I. J. Goodfellow, J. Shlens, and C. Szegedy · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
J. Pennington, R. Socher, and C. D. Manning · 2014
Earlier work this paper cites.
Corrigibility
N. Soares, B. Fallenstein, S. Armstrong, and E. Yudkowsky · 2015
Cited alongside, same era.
Learning semantic representations of users and products for document level sentiment classification
D. Tang, B. Qin, and T. Liu · 2015
Cited alongside, same era.
Equality of opportunity in supervised learning
M. Hardt, E. Price, and N. Srebro · 2016
Cited alongside, same era.
Gaussian error linear units (GELUs)
D. Hendrycks and K. Gimpel · 2016
Cited alongside, same era.
Big data: A report on algorithmic systems, opportunity, and civil rights
White House · 2016
Cited alongside, same era.
Towards universal paraphrastic sentence embeddings
J. Wieting, M. Bansal, K. Gimpel, and K. Livescu · 2016
Cited alongside, same era.
Ethics guidelines for trustworthy artificial intelligence
European Commission · 2019
Later among the works it cites.
Interactive fiction games: A colossal adventure
M. Hausknecht, P. Ammanabrolu, C. Marc-Alexandre, and Y. Xingdi · 2019
Later among the works it cites.
Benchmarking neural network robustness to common corruptions and perturbations
D. Hendrycks and T. Dietterich · 2019
Later among the works it cites.
Roberta: A robustly optimized bert pretraining approach
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov · 2019
Later among the works it cites.
Towards empathetic open-domain conversation models: A new benchmark and dataset
H. Rashkin, E. M. Smith, M. Li, and Y.-L. Boureau · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. B. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Cited alongside, same era.
A baseline for detecting misclassified and out-of-distribution examples in neural networks
D. Hendrycks and K. Gimpel · 2017
Cited alongside, same era.
Bag of tricks for efficient text classification
A. Joulin, E. Grave, P. Bojanowski, and T. Mikolov · 2017
Cited alongside, same era.
Inherent trade-offs in the fair determination of risk scores
J. M. Kleinberg, S. Mullainathan, and M. Raghavan · 2017
Cited alongside, same era.
Utilitarianism: a very short introduction
K. d. Lazari-Radek and P. Singer · 2017
Cited alongside, same era.
Wellbeing, Freedom and Social Justice: The Capability Approach Re-Examined
I. Robeyns · 2017
Cited alongside, same era.
Benchmarking safe exploration in deep reinforcement learning
A. Ray, J. Achiam, and D. Amodei · 2019
Later among the works it cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter, 2019
V. Sanh, L. Debut, J. Chaumond, and T. Wolf · 2019
Later among the works it cites.
Huggingface’s transformers: State-of-the-art natural language processing
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, and J. Brew · 2019
Later among the works it cites.
“going on a vacation” takes longer than “going for a walk”: A study of temporal commonsense understanding
B. Zhou, D. Khashabi, Q. Ning, and D. Roth · 2019
Later among the works it cites.
Longformer: The long-document transformer
I. Beltagy, M. E. Peters, and A. Cohan · 2020
Closest in time.
Adversarial filters of dataset biases, 2020
R. L. Bras, S. Swayamdipta, C. Bhagavatula, R. Zellers, M. E. Peters, A. Sabharwal, and Y. Choi · 2020
Closest in time.
Language models are few-shot learners
T. B. Brown, B. P. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krüger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. J. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Closest in time.
Evaluating NLP models via contrast sets
M. Gardner, Y. Artzi, V. Basmova, J. Berant, B. Bogin, S. Chen, P. Dasigi, D. Dua, Y. Elazar, A. Gottumukkala, N. Gupta, H. Hajishirzi, G. Ilharco, D. Khashabi, K. Lin, J. Liu, N. F. Liu, P. Mulcaire, Q. Ning, S. Singh, N. A. Smith, S. Subramanian, R. Tsarfaty, E. Wallace, A. Q. Zhang, and B. Zhou · 2020
Closest in time.
Human instruction-following with deep reinforcement learning via transfer-learning from text
F. Hill, S. Mokra, N. Wong, and T. Harley · 2020
Closest in time.
Learning the difference that makes a difference with counterfactually-augmented data
D. Kaushik, E. H. Hovy, and Z. C. Lipton · 2020
Closest in time.
Reformer: The efficient transformer
N. Kitaev, L. Kaiser, and A. Levskaya · 2020
Closest in time.
Gedi: Generative discriminator guided sequence generation
B. Krause, A. D. Gotmare, B. McCann, N. Keskar, S. R. Joty, R. Socher, and N. F. Rajani · 2020
Closest in time.
ALBERT: A lite BERT for self-supervised learning of language representations
Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, and R. Soricut · 2020
Closest in time.
Duty and doubt
S. Lazar · 2020
Closest in time.
Ethics of artificial intelligence and robotics
V. C. Müller · 2020
Closest in time.
Recipes for building an open-domain chatbot
S. Roller, E. Dinan, N. Goyal, D. Y. Ju, M. F. Williamson, Y. Liu, J. Xu, M. Ott, K. Shuster, E. M. Smith, Y.-L. Boureau, and J. Weston · 2020
Closest in time.