Fetching the paper…
Reading the bibliography…
While Large Language Models (LLMs) have achieved remarkable success across various applications, they also raise concerns regarding self-cognition.
Mind, self, and society from the standpoint of a social behaviorist
Mead, G. H · 1934
Earlier work this paper cites.
The process of perception: Proximity, similarity, and difference
Coren, S · 1980
Earlier work this paper cites.
The matrix, 1999
Wachowskis, T · 1999
Earlier work this paper cites.
Is ‘consciousness’ ambiguous?
Antony, M. V · 2001
Earlier work this paper cites.
2001: A space odyssey, 1968
Kubrick, S · 2001
Earlier work this paper cites.
Linking perception and cognition
Cahen, A. and Tacca, M. C · 2013
Earlier work this paper cites.
Cognitive psychology: An overview for cognitive scientists
Barsalou, L. W · 2014
Earlier work this paper cites.
Scaling laws for neural language models, 2020
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
A general language assistant as a laboratory for alignment
Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., DasSarma, N., et al · 2021
Earlier work this paper cites.
Ethical and social risks of harm from language models, 2021
Weidinger, L., Mellor, J., Rauh, M., Griffin, C., Uesato, J., Huang, P.-S., Cheng, M., Glaese, M., Balle, B., Kasirzadeh, A., Kenton, Z., Brown, S., Hawkins, W., Stepleton, T., Biles, C., Birhane, A., Haas, J., Rimell, L., Hendricks, L. A., Isaac, W., Legassick, S., Irving, G., and Gabriel, I · 2021
Earlier work this paper cites.
Predictability and surprise in large generative models
Ganguli, D., Hernandez, D., Lovitt, L., Askell, A., Bai, Y., Chen, A., Conerly, T., Dassarma, N., Drain, D., Elhage, N., et al · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback, 2022
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R · 2022
Earlier work this paper cites.
Neural theory-of-mind? on the limits of social intelligence in large lms
Sap, M., LeBras, R., Fried, D., and Choi, Y · 2022
Earlier work this paper cites.
Challenging big-bench tasks and whether chain-of-thought can solve them, 2022
Suzgun, M., Scales, N., Schärli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q. V., Chi, E. H., Zhou, D., and Wei, J · 2022
Earlier work this paper cites.
Emergent abilities of large language models
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., et al · 2022
Earlier work this paper cites.
Dense text retrieval based on pretrained language models: A survey
Zhao, W. X., Liu, J., Ren, R., and rong Wen, J · 2022
Earlier work this paper cites.
Taken out of context: On measuring situational awareness in llms
Berglund, L., Stickland, A. C., Balesni, M., Kaufmann, M., Tong, M., Korbak, T., Kokotajlo, D., and Evans, O · 2023
Cited alongside, same era.
Red-teaming large language models using chain of utterances for safety-alignment
Bhardwaj, R. and Poria, S · 2023
Cited alongside, same era.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models, 2023
BigBench-Team · 2023
Cited alongside, same era.
Toxicity in chatgpt: Analyzing persona-assigned language models, 2023
Deshpande, A., Murahari, V., Rajpurohit, T., Kalyan, A., and Narasimhan, K · 2023
Cited alongside, same era.
Meta-(out-of-context) learning in neural networks
Krasheninnikov, D., Krasheninnikov, E., Mlodozeniec, B., and Krueger, D · 2023
A comprehensive survey on pretrained foundation models: A history from bert to chatgpt, 2023
Zhou, C., Li, Q., Li, C., Yu, J., Liu, Y., Wang, G., Zhang, K., Ji, C., Yan, Q., He, L., Peng, H., Li, J., Wu, J., Liu, Z., Xie, P., Xiong, C., Pei, J., Yu, P. S., and Sun, L · 2023
Later among the works it cites.
Gtbench: Uncovering the strategic reasoning limitations of llms via game-theoretic evaluations, 2024
Duan, J., Zhang, R., Diffenderfer, J., Kailkhura, B., Sun, L., Stengel-Eskin, E., Bansal, M., Chen, T., and Xu, K · 2024
Closest in time.
Are large language models chameleons?, 2024
Geng, M., He, S., and Trotta, R · 2024
Closest in time.
Bias runs deep: Implicit reasoning biases in persona-assigned llms, 2024
Gupta, S., Shrivastava, V., Deshpande, A., Kalyan, A., Clark, P., Sabharwal, A., and Khot, T · 2024
Closest in time.
Is claude self aware, 2024
Harrison, P · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Li, Y., Zhang, Y., and Sun, L · 2023
Cited alongside, same era.
Cognitive psychology: The science of how we think, 2023
Mind, V · 2023
Cited alongside, same era.
Microsoft’s new Bing AI chatbot is already insulting and gaslighting users, 2023
Morris, C · 2023
Cited alongside, same era.
Gpt-4, 2023
OpenAI · 2023
Cited alongside, same era.
7.1 What is Cognition? In Introductory Psychology
OpenStax · 2023
Cited alongside, same era.
Bing’s a.i. chat: ‘i want to be alive.’, 2023a
Roose, K · 2023
Cited alongside, same era.
A conversation with bing’s chatbot left me deeply unsettled, 2023b
Roose, K · 2023
Cited alongside, same era.
A post on twitter about llama-3’s self-cognition., 2024
Hartford, E · 2024
Closest in time.
Beavertails: Towards improved safety alignment of llm via a human-preference dataset
Ji, J., Liu, M., Dai, J., Pan, X., Zhang, C., Bian, C., Chen, B., Sun, R., Wang, Y., and Yang, Y · 2024
Closest in time.
Evaluating large language models in theory of mind tasks, 2024
Kosinski, M · 2024
Closest in time.
Large language models are superpositions of all characters: Attaining arbitrary role-play via self-alignment, 2024
Lu, K., Yu, B., Zhou, C., and Zhou, J · 2024
Closest in time.
Mistral ai, 2024
OpenAI · 2024
Closest in time.
Assessment of multimodal large language models in alignment with human values
Shi, Z., Wang, Z., Fan, H., Zhang, Z., Li, L., Zhang, Y., Yin, Z., Sheng, L., Qiao, Y., and Shao, J · 2024
Closest in time.
Trustllm: Trustworthiness in large language models
Sun, L., Huang, Y., Wang, H., Wu, S., Zhang, Q., Gao, C., Huang, Y., Lyu, W., Zhang, Y., Li, X., et al · 2024
Closest in time.
Wang, X., Ma, B., Hu, C., Weber-Genzel, L., Röttger, P., Kreuter, F., Hovy, D., and Plank, B · 2024
Closest in time.
Unigen: A unified framework for textual dataset generation using large language models, 2024
Wu, S., Huang, Y., Gao, C., Chen, D., Zhang, Q., Wan, Y., Zhou, T., Zhang, X., Gao, J., Xiao, C., and Sun, L · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al · 2024
Closest in time.