Fetching the paper…
Reading the bibliography…
Work in AI ethics and fairness has made much progress in regulating LLMs to reflect certain values, such as fairness, truth, and diversity.
Philosophical Investigations
Wittgenstein, L. (1953) · 1953
Earlier work this paper cites.
Distributional structure
Harris, Z. S. (1954) · 1954
Earlier work this paper cites.
Cartesian meditations : an introduction to phenomenology
Husserl, E. and Cairns, D. (1960) · 1960
Earlier work this paper cites.
Word & Object
Quine, W. V. O. (1960) · 1960
Earlier work this paper cites.
Being and time
Heidegger, M. (1962) · 1962
Earlier work this paper cites.
Critique of pure reason
Kant, I. and Smith, N. K. (1781, 1929, 1965) · 1965
Earlier work this paper cites.
Four Essays on Liberty
Berlin, I. (1969) · 1969
Earlier work this paper cites.
The Phenomenology of Spirit
Hegel, G. W. F. (1977) · 1977
Earlier work this paper cites.
Naming and Necessity: Lectures Given to the Princeton University Philosophy Colloquium
Kripke, S. A. (1980) · 1980
Earlier work this paper cites.
Truth and Method
Gadamer, H. (1982) · 1982
Earlier work this paper cites.
Husserl, Intentionality, and Cognitive Science
Dreyfus, H. L., editor (1984) · 1984
Earlier work this paper cites.
Quining Qualia
Dennett, D. C. (1988) · 1988
Earlier work this paper cites.
The symbol grounding problem
Harnad, S. (1990) · 1990
Earlier work this paper cites.
A stroll through the worlds of animals and men: A picture book of invisible worlds
von Uexküll, J. (1992) · 1992
Earlier work this paper cites.
Ressentiment
Scheler, M. (1994) · 1994
Cited alongside, same era.
On the genealogy of morality
Nietzsche, F., Ansell-Pearson, K., and Diethe, C. (1995) · 1995
Cited alongside, same era.
The looping effects of human kinds
Hacking, I. (1996) · 1996
Cited alongside, same era.
A Theory of Justice: Revised Edition
Rawls, J. (1999) · 1999
Cited alongside, same era.
The Order of Things: An Archaeology of the Human Sciences
Foucault, M. (2002) · 2002
Cited alongside, same era.
Using category theory to assess the relationship between consciousness and integrated information theory
Tsuchiya, N., Taguchi, S., and Saigo, H. (2016) · 2016
Cited alongside, same era.
Can machines think? a report on turing test experiments at the royal society
Can machines learn morality? the delphi experiment
Jiang, L., Hwang, J. D., Bhagavatula, C., Bras, R. L., Liang, J., Dodge, J., Sakaguchi, K., Forbes, M., Borchardt, J., Gabriel, S., Tsvetkov, Y., Etzioni, O., Sap, M., Rini, R. A., and Choi, Y. (2021) · 2021
Later among the works it cites.
Language models as agent models
Andreas, J. (2022) · 2022
Later among the works it cites.
Knowledge is power: Symbolic knowledge distillation, commonsense morality, & multimodal script knowledge
Choi, Y. (2022) · 2022
Later among the works it cites.
Survey of hallucination in natural language generation
Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y., Dai, W., Madotto, A., and Fung, P. (2022) · 2022
Later among the works it cites.
An a.i. pioneer on what we should really fear
Marchese, D. (2022) · 2022
Later among the works it cites.
Can machines learn how to behave?
Agüera y Arcas, B. (2022) · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Warwick, K. and Shah, H. (2016) · 2016
Cited alongside, same era.
Beyond Concepts: Unicepts, Language, and Natural Information
Millikan, R. G. (2017) · 2017
Cited alongside, same era.
A Thousand Plateaus and Philosophy
Deleuze, G. and Guattari, F. (2018) · 2018
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M. (2018) · 2018
Cited alongside, same era.
The misgendering machines: Trans/hci implications of automatic gender recognition
Keyes, O. (2018) · 2018
Cited alongside, same era.
From clustering to cluster explanations via neural networks
Kauffmann, J. R., Esders, M., Montavon, G., Samek, W., and Müller, K.-R. (2019) · 2019
Cited alongside, same era.
Closest in time.
On the Computation of Meaning, Language Models and Incomprehensible Horrors
Bennett, M. T. (2023) · 2023
Closest in time.
Open problems and fundamental limitations of reinforcement learning from human feedback
Casper, S., Davies, X., Shi, C., Gilbert, T. K., Scheurer, J., Rando, J., Freedman, R., Korbak, T., Lindner, D., Freire, P., Wang, T., Marks, S., Segerie, C.-R., Carroll, M., Peng, A., Christoffersen, P., Damani, M., Slocum, S., Anwar, U., Siththaranjan, A., Nadeau, M., Michaud, E. J., Pfau, J., Krasheninnikov, D., Chen, X., Langosco, L., Hase, P., Bıyık, E., Dragan, A., Krueger, D., Sadigh, D., and Hadfield-Menell, D. (2023) · 2023
Closest in time.
Nlpositionality: Characterizing design biases of datasets and models
Santy, S., Liang, J., Bras, R. L., Reinecke, K., and Sap, M. (2023) · 2023
Closest in time.
Large language models and the reverse turing test
Sejnowski, T. J. (2023) · 2023
Closest in time.
The curse of recursion: Training on generated data makes models forget
Shumailov, I., Shumaylov, Z., Zhao, Y., Gal, Y., Papernot, N., and Anderson, R. (2023) · 2023
Closest in time.
Value kaleidoscope: Engaging ai with pluralistic human values, rights, and duties
Sorensen, T., Jiang, L., Hwang, J., Levine, S., Pyatkin, V., West, P., Dziri, N., Lu, X., Rao, K., Bhagavatula, C., Sap, M., Tasioulas, J., and Choi, Y. (2023) · 2023
Closest in time.
Imitation versus innovation: What children can do that large language and language-and-vision models cannot (yet)?
Yiu, E., Kosoy, E., and Gopnik, A. (2023) · 2023
Closest in time.