Fetching the paper…
Reading the bibliography…
This is a speculative essay on interface design and artificial intelligence.
An experimental study of apparent behavior
Fritz Heider and Marianne Simmel · 1944
Earlier work this paper cites.
Machines and mindlessness: Social responses to computers
Clifford Nass and Youngme Moon · 2000
Earlier work this paper cites.
Pandora and the music genome project, song structure analysis tools facilitate new music discovery
J Joyce · 2006
Earlier work this paper cites.
On seeing human: a three-factor theory of anthropomorphism
Nicholas Epley, Adam Waytz, and John T Cacioppo · 2007
Earlier work this paper cites.
Who sees human? the stability and importance of individual differences in anthropomorphism
Adam Waytz, John Cacioppo, and Nicholas Epley · 2010
Earlier work this paper cites.
Why these ads? google explains ad targeting, allows blocking
John Rampton · 2011
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Guillaume Alain and Yoshua Bengio · 2016
Earlier work this paper cites.
Social robotics
Cynthia Breazeal, Kerstin Dautenhahn, and Takayuki Kanda · 2016
Earlier work this paper cites.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, et al · 2018
Earlier work this paper cites.
The building blocks of interpretability
Chris Olah, Arvind Satyanarayan, Ian Johnson, Shan Carter, Ludwig Schubert, Katherine Ye, and Alexander Mordvintsev · 2018
Earlier work this paper cites.
On interpretability and feature representations: an analysis of the sentiment neuron
Jonathan Donnelly and Adam Roegiest · 2019
Earlier work this paper cites.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D Manning · 2019
Earlier work this paper cites.
Visualizing and measuring the geometry of bert
Emily Reif, Ann Yuan, Martin Wattenberg, Fernanda B Viegas, Andy Coenen, Adam Pearce, and Been Kim · 2019
Cited alongside, same era.
Bert rediscovers the classical nlp pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick · 2019
Cited alongside, same era.
Understanding the role of individual units in a deep neural network
David Bau, Jun-Yan Zhu, Hendrik Strobelt, Agata Lapedriza, Bolei Zhou, and Antonio Torralba · 2020
Cited alongside, same era.
Finding universal grammatical relations in multilingual bert
Ethan A Chi, John Hewitt, and Christopher D Manning · 2020
Cited alongside, same era.
Can language models encode perceptual structure without grounding? a case study in color
Mostafa Abdou, Artur Kulmizev, Daniel Hershcovich, Stella Frank, Ellie Pavlick, and Anders Søgaard · 2021
Language models (mostly) know what they know
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield Dodds, Nova DasSarma, Eli Tran-Johnson, et al · 2022
Later among the works it cites.
Emergent world representations: Exploring a sequence model trained on a synthetic task
Kenneth Li, Aspen K Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg · 2022
Later among the works it cites.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, et al · 2022
Later among the works it cites.
Chris olah on what the hell is going on inside neural networks
Christopher Olah · 2022
Later among the works it cites.
Discovering language model behaviors with model-written evaluations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On the dangers of stochastic parrots: Can language models be too big?
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell · 2021
Cited alongside, same era.
A little bird told me your gender: Gender inferences in social media
Eduard Fosch-Villaronga, Adam Poulsen, Roger Andre Søraa, and BHM Custers · 2021
Cited alongside, same era.
Implicit representations of meaning in neural language models
Belinda Z Li, Maxwell Nye, and Jacob Andreas · 2021
Cited alongside, same era.
Machinelike or humanlike? a literature review of anthropomorphism in ai-enabled technology
Mengjun Li and Ayoung Suh · 2021
Cited alongside, same era.
Language models as agent models
Jacob Andreas · 2022
Cited alongside, same era.
Probing classifiers: Promises, shortcomings, and advances
Yonatan Belinkov · 2022
Cited alongside, same era.
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, et al · 2022
Cited alongside, same era.
Ethan Perez, Sam Ringer, Kamilė Lukošiūtė, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, et al · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou · 2022
Later among the works it cites.
Agi ruin: A list of lethalities
Eliezer Yudkowsky · 2022
Later among the works it cites.
A toy model of universality: Reverse engineering how networks learn group operations
Bilal Chughtai, Lawrence Chan, and Neel Nanda · 2023
Closest in time.
Theory of mind may have spontaneously emerged in large language models
Michal Kosinski · 2023
Closest in time.
On ai anthropomorphism
Ben Shneiderman and Michael Muller · 2023
Closest in time.
Microsoft’s bing is an emotionally manipulative liar, and people love it
James Vincent · 2023
Closest in time.