Fetching the paper…
Reading the bibliography…
Understanding knowledge mechanisms in Large Language Models (LLMs) is crucial for advancing towards trustworthy AGI.
Interpretable machine learning: definitions, methods, and applications
W. James Murdoch, Chandan Singh, Karl Kumbier, Reza Abbasi-Asl, and Bin Yu. 2019 · 1901
Earlier work this paper cites.
Fine-tuning language models from human preferences
Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul F. Christiano, and Geoffrey Irving. 2019 · 1909
Earlier work this paper cites.
Taxonomy of Educational Objectives: The Classification of Educational Goals. Handbook I: Cognitive Domain
Benjamin S. Bloom, Max D. Engelhart, Edward J. Furst, Walker H. Hill, and David R. Krathwohl. 1956 · 1956
Earlier work this paper cites.
Problems of Monetary Management: The UK Experience , pages 91–121. Macmillan Education UK, London
C. A. E. Goodhart. 1984 · 1984
Earlier work this paper cites.
Knowledge acquisition - principles and guidelines
Karen L. McGraw and Karan Harbison-Briggs. 1990 · 1990
Earlier work this paper cites.
What is a knowledge representation?
Randall Davis, Howard E. Shrobe, and Peter Szolovits. 1993 · 1993
Earlier work this paper cites.
Social foundations of cognition
John M Levine, Lauren B Resnick, and E Tory Higgins. 1993 · 1993
Earlier work this paper cites.
Sparse coding with an overcomplete basis set: A strategy employed by v1?
Bruno A Olshausen and David J Field. 1997 · 1997
Earlier work this paper cites.
How knowledge works
John Hyman. 1999 · 1999
Earlier work this paper cites.
Creative innovation: possible brain mechanisms
Kenneth M Heilman, Stephen E Nadeau, and David O Beversdorf. 2003 · 2003
Earlier work this paper cites.
Cognitive science: An introduction to mind and brain
Daniel Kolak, William Hirstein, Peter Mandik, and Jonathan Waskan. 2006 · 2006
Earlier work this paper cites.
Efficient sparse coding algorithms
Honglak Lee, Alexis J. Battle, Rajat Raina, and Andrew Y. Ng. 2006 · 2006
Earlier work this paper cites.
The nature of creativity
Robert J Sternberg. 2006 · 2006
Earlier work this paper cites.
The continuity of mind
Spivey and Michael. 2007 · 2007
Earlier work this paper cites.
Benjamin Heinzerling and Kentaro Inui. 2020 · 2008
Earlier work this paper cites.
Collective learning: Applying distributed cognition for collective intelligence
Jose A Fadul. 2009 · 2009
Earlier work this paper cites.
Mapping student information literacy activity against bloom’s taxonomy of cognitive skills
Judith Keene, John Colvin, and Justine Sissons. 2010 · 2010
Earlier work this paper cites.
The standard definition of creativity
Mark A Runco and Garrett J Jaeger. 2012 · 2012
Earlier work this paper cites.
Fundamental neuroscience
Larry Squire, Darwin Berg, Floyd E Bloom, Sascha Du Lac, Anirvan Ghosh, and Nicholas C Spitzer. 2012 · 2012
Earlier work this paper cites.
Networks in cognitive science
Andrea Baronchelli, Ramon Ferrer-i Cancho, Romualdo Pastor-Satorras, Nick Chater, and Morten H Christiansen. 2013 · 2013
Earlier work this paper cites.
The less you know, you think you know more; dunning and kruger effect in collective decision making
Ali Mahmoodi, Majid Nili Ahmadabadi, and Bahador Bahrami. 2013 · 2013
Earlier work this paper cites.
Knowledge representation
Arthur B Markman. 2013 · 2013
Earlier work this paper cites.
Social interaction in learning and development
Aleksandar Baucal, Lausanne Mouline, Switzerland Kristiina Kumpulainen, Charis Psaltis, and Baruch Schwarz. 2014 · 2014
Earlier work this paper cites.
Generative adversarial nets
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C. Courville, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric A. Weiss, Niru Maheswaranathan, and Surya Ganguli. 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Anderson and krathwohl–bloom’s taxonomy revised
Leslie Owen Wilson. 2016 · 2016
Earlier work this paper cites.
Network dissection: Quantifying interpretability of deep visual representations
David Bau, Bolei Zhou, Aditya Khosla, Aude Oliva, and Antonio Torralba. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
What is knowledge?
Linda Zagzebski. 2017 · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals. 2017 · 2017
Earlier work this paper cites.
Hierarchical neural story generation
Angela Fan, Mike Lewis, and Yann N. Dauphin. 2018 · 2018
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin. 2019 · 2019
Earlier work this paper cites.
Similarity of neural network representations revisited
Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Geoffrey E. Hinton. 2019 · 2019
Earlier work this paper cites.
One ticket to win them all: generalizing lottery ticket initializations across datasets and optimizers
Ari S. Morcos, Haonan Yu, Michela Paganini, and Yuandong Tian. 2019 · 2019
Earlier work this paper cites.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick S. H. Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander H. Miller. 2019 · 2019
Earlier work this paper cites.
ERNIE: enhanced language representation with informative entities
Zhengyan Zhang, Xu Han, Zhiyuan Liu, Xin Jiang, Maosong Sun, and Qun Liu. 2019 · 2019
Earlier work this paper cites.
Climbing towards NLU: On meaning, form, and understanding in the age of data
Emily M. Bender and Alexander Koller. 2020 · 2020
Earlier work this paper cites.
Experience grounds language
Yonatan Bisk, Ari Holtzman, Jesse Thomason, Jacob Andreas, Yoshua Bengio, Joyce Chai, Mirella Lapata, Angeliki Lazaridou, Jonathan May, Aleksandr Nisnevich, Nicolas Pinto, and Joseph Turian. 2020 · 2020
Earlier work this paper cites.
I’m already optimal: the dunning-kruger effect, sociogenesis, and self-integration
John N. A. Brown and Lukas Esterle. 2020 · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Earlier work this paper cites.
exbert: A visual analysis tool to explore learned representations in transformer models
Benjamin Hoover, Hendrik Strobelt, and Sebastian Gehrmann. 2020 · 2020
Earlier work this paper cites.
Concept bottleneck models
Pang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann, Emma Pierson, Been Kim, and Percy Liang. 2020 · 2020
Earlier work this paper cites.
Complexity analysis and synchronization control of fractional-order jafari-sprott chaotic system
Guohui Li, Xiangyu Zhang, and Hong Yang. 2020 · 2020
Earlier work this paper cites.
interpreting gpt: the logit lens
nostalgebraist. 2020 · 2020
Earlier work this paper cites.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter. 2020 · 2020
Earlier work this paper cites.
How much knowledge can you pack into the parameters of a language model?
Adam Roberts, Colin Raffel, and Noam Shazeer. 2020 · 2020
Earlier work this paper cites.
A methodology for evaluating the extensibility of boolean networks’ structure and function
Rémi Segretain, Sergiu Ivanov, Laurent Trilling, and Nicolas Glade. 2020 · 2020
Earlier work this paper cites.
Investigating gender bias in language models using causal mediation analysis
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart M. Shieber. 2020 · 2020
Earlier work this paper cites.
A survey of synthetic data generation for machine learning
Mohammad Abufadda and Khalid Mansour. 2021 · 2021
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Earlier work this paper cites.
Knowledgeable or educated guess? revisiting language models as knowledge bases
Boxi Cao, Hongyu Lin, Xianpei Han, Le Sun, Lingyong Yan, Meng Liao, Tong Xue, and Jin Xu. 2021a · 2021
Earlier work this paper cites.
Editing factual knowledge in language models
Nicola De Cao, Wilker Aziz, and Ivan Titov. 2021b · 2021
Earlier work this paper cites.
The elastic lottery ticket hypothesis
Xiaohan Chen, Yu Cheng, Shuohang Wang, Zhe Gan, Jingjing Liu, and Zhangyang Wang. 2021 · 2021
Earlier work this paper cites.
Chaotic behavior analysis of a new incommensurate fractional-order hopfield neural network system
Nadjette Debbouche, Adel Ouannas, Iqbal M. Batiha, Giuseppe Grassi, Mohammed K. A. Kaabar, Hadi Jahanshahi, Ayman A. Aly, and Awad M. Aljuaid. 2021 · 2021
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021 · 2021
Earlier work this paper cites.
Amnesic probing: Behavioral explanation with amnesic counterfactuals
Yanai Elazar, Shauli Ravfogel, Alon Jacovi, and Yoav Goldberg. 2021 · 2021
Earlier work this paper cites.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. 2021 · 2021
Earlier work this paper cites.
A brain-computer interface that evokes tactile sensations improves robotic arm control
Sharlene N Flesher, John E Downey, Jeffrey M Weiss, Christopher L Hughes, Angelica J Herrera, Elizabeth C Tyler-Kabara, Michael L Boninger, Jennifer L Collinger, and Robert A Gaunt. 2021 · 2021
Earlier work this paper cites.
Transformer feed-forward layers are key-value memories
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2021 · 2021
Earlier work this paper cites.
Knowledgeable machine learning for natural language processing
Xu Han, Zhengyan Zhang, and Zhiyuan Liu. 2021 · 2021
Earlier work this paper cites.
Acquisition of chess knowledge in alphazero
Thomas McGrath, Andrei Kapishnikov, Nenad Tomasev, Adam Pearce, Demis Hassabis, Been Kim, Ulrich Paquet, and Vladimir Kramnik. 2021 · 2021
Earlier work this paper cites.
Four principles of explainable artificial intelligence
P Jonathon Phillips, P Jonathon Phillips, Carina A Hahn, Peter C Fontana, Amy N Yates, Kristen Greene, David A Broniatowski, and Mark A Przybocki. 2021 · 2021
Earlier work this paper cites.
Protein design and variant prediction using autoregressive generative models
Jung-Eun Shin, Adam J. Riesselman, Aaron W. Kollasch, Conor McMahon, Elana Simon, Chris Sander, Aashish Manglik, Andrew C. Kruse, and Debora S. Marks. 2021 · 2021
Earlier work this paper cites.
Can language models be biomedical knowledge bases?
Mujeen Sung, Jinhyuk Lee, Sean S. Yi, Minji Jeon, Sungdong Kim, and Jaewoo Kang. 2021 · 2021
Earlier work this paper cites.
Transformer visualization via dictionary learning: contextualized embedding as a linear superposition of transformer factors
Zeyu Yun, Yubei Chen, Bruno A. Olshausen, and Yann LeCun. 2021 · 2021
Earlier work this paper cites.
Factual probing is [MASK]: learning vs. learning to recall
Zexuan Zhong, Dan Friedman, and Danqi Chen. 2021 · 2021
Earlier work this paper cites.
Towards tracing factual knowledge in language models back to the training data
Ekin Akyürek, Tolga Bolukbasi, Frederick Liu, Binbin Xiong, Ian Tenney, Jacob Andreas, and Kelvin Guu. 2022 · 2022
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom B. Brown, Jack Clark, Sam McCandlish, Chris Olah, Benjamin Mann, and Jared Kaplan. 2022 · 2022
Earlier work this paper cites.
Probing classifiers: Promises, shortcomings, and advances
Yonatan Belinkov. 2022 · 2022
Earlier work this paper cites.
Overcoming a theoretical limitation of self-attention
David Chiang and Peter Cholak. 2022 · 2022
Earlier work this paper cites.
Knowledge is power: Symbolic knowledge distillation, commonsense morality, & multimodal script knowledge
Yejin Choi. 2022 · 2022
Earlier work this paper cites.
Knowledge neurons in pretrained transformers
Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei. 2022 · 2022
Earlier work this paper cites.
The emergent properties of the connected brain
Thiebaut de Schotten, Michel, Forkel, and Stephanie J. 2022 · 2022
Earlier work this paper cites.
Translation between molecules and natural language
Carl Edwards, Tuan Manh Lai, Kevin Ros, Garrett Honke, Kyunghyun Cho, and Heng Ji. 2022 · 2022
Earlier work this paper cites.
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah. 2022 · 2022
Earlier work this paper cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer. 2022 · 2022
Earlier work this paper cites.
Inducing causal structure for interpretable neural networks
Atticus Geiger, Zhengxuan Wu, Hanson Lu, Josh Rozner, Elisa Kreiss, Thomas Icard, Noah D. Goodman, and Christopher Potts. 2022 · 2022
Earlier work this paper cites.
Lm-debugger: An interactive tool for inspection and intervention in transformer-based language models
Mor Geva, Avi Caciularu, Guy Dar, Paul Roit, Shoval Sadde, Micah Shlain, Bar Tamir, and Yoav Goldberg. 2022a · 2022
Earlier work this paper cites.
Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space
Mor Geva, Avi Caciularu, Kevin Ro Wang, and Yoav Goldberg. 2022b · 2022
Earlier work this paper cites.
PTR: prompt tuning with rules for text classification
Xu Han, Weilin Zhao, Ning Ding, Zhiyuan Liu, and Maosong Sun. 2022 · 2022
Earlier work this paper cites.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals, and Laurent Sifre. 2022 · 2022
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022 · 2022
Earlier work this paper cites.
Causal scrubbing: A method for rigorously testing interpretability hypotheses
LawrenceC, Adrià Garriga-alonso, Nicholas Goldowsky Dill, ryan greenblatt, jenny, Ansh Radhakrishnan, Buck, and Nate Thomas. 2022 · 2022
Earlier work this paper cites.
Factuality enhanced language models for open-ended text generation
Nayeon Lee, Wei Ping, Peng Xu, Mostofa Patwary, Pascale Fung, Mohammad Shoeybi, and Bryan Catanzaro. 2022b · 2022
Earlier work this paper cites.
How pre-trained language models capture factual knowledge? A causal-inspired analysis
Shaobo Li, Xiaoguang Li, Lifeng Shang, Zhenhua Dong, Chengjie Sun, Bingquan Liu, Zhenzhou Ji, Xin Jiang, and Qun Liu. 2022 · 2022
Earlier work this paper cites.
A rigorous study of integrated gradients method and extensions to internal neuron attributions
Daniel Lundström, Tianjian Huang, and Meisam Razaviyayn. 2022 · 2022
Earlier work this paper cites.
Locating and editing factual associations in GPT
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022 · 2022
Earlier work this paper cites.
Fast model editing at scale
Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, and Christopher D. Manning. 2022 · 2022
Earlier work this paper cites.
Transformerlens
Neel Nanda and Joseph Bloom. 2022 · 2022
Earlier work this paper cites.
A survey of machine unlearning
Thanh Tam Nguyen, Thanh Trung Huynh, Phi Le Nguyen, Alan Wee-Chung Liew, Hongzhi Yin, and Quoc Viet Hung Nguyen. 2022 · 2022
Earlier work this paper cites.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe. 2022 · 2022
Earlier work this paper cites.
Measuring reliability of large language models through semantic consistency
Harsh Raj, Domenic Rosati, and Subhabrata Majumdar. 2022 · 2022
Earlier work this paper cites.
Taking features out of superposition with sparse autoencoders
Lee Sharkey, Dan Braun, , and Beren Millidge. 2022 · 2022
Earlier work this paper cites.
Large language models encode clinical knowledge
Karan Singhal, Shekoofeh Azizi, Tao Tu, S. Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Kumar Tanwani, Heather Cole-Lewis, Stephen Pfohl, Perry Payne, Martin Seneviratne, Paul Gamble, Chris Kelly, Nathaneal Schärli, Aakanksha Chowdhery, Philip Andrew Mansfield, Blaise Agüera y Arcas, Dale R. Webster, Gregory S. Corrado, Yossi Matias, Katherine Chou, Juraj Gottweis, Nenad Tomasev, Yun Liu, Alvin Rajkomar, Joelle K. Barral, Christopher Semturs, Alan Karthikesalingam, and Vivek Natarajan. 2022 · 2022
Earlier work this paper cites.
Tom2c: Target-oriented multi-agent communication and cooperation with theory of mind
Yuanfei Wang, Fangwei Zhong, Jing Xu, and Yizhou Wang. 2022 · 2022
Earlier work this paper cites.
Symbolic knowledge distillation: from general language models to commonsense models
Peter West, Chandra Bhagavatula, Jack Hessel, Jena D. Hwang, Liwei Jiang, Ronan Le Bras, Ximing Lu, Sean Welleck, and Yejin Choi. 2022 · 2022
Earlier work this paper cites.
GQA: training generalized multi-query transformer models from multi-head checkpoints
Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sanghai. 2023 · 2023
Earlier work this paper cites.
AI sentience and socioculture
Aj Alvero and Courtney Peña. 2023 · 2023
Earlier work this paper cites.
From word types to tokens and back: A survey of approaches to word meaning representation and interpretation
Marianna Apidianaki. 2023 · 2023
Earlier work this paper cites.
The internal state of an LLM knows when its lying
Amos Azaria and Tom M. Mitchell. 2023 · 2023
Cited alongside, same era.
Eliciting latent predictions from transformers with the tuned lens
Nora Belrose, Zach Furman, Logan Smith, Danny Halawi, Igor Ostrovsky, Lev McKinney, Stella Biderman, and Jacob Steinhardt. 2023 · 2023
Cited alongside, same era.
The reversal curse: Llms trained on "a is b" fail to learn "b is a"
Lukas Berglund, Meg Tong, Max Kaufmann, Mikita Balesni, Asa Cooper Stickland, Tomasz Korbak, and Owain Evans. 2023 · 2023
Cited alongside, same era.
A drop of ink may make a million think: The spread of false information in large language models
Ning Bian, Peilin Liu, Xianpei Han, Hongyu Lin, Yaojie Lu, Ben He1, and Le Sun. 2023 · 2023
Cited alongside, same era.
On behalf of the stakeholders: Trends in nlp model interpretability in the era of llms
Nitay Calderon and Roi Reichart. 2024 · 2024
Closest in time.
Retentive or forgetful? diving into the knowledge memorizing mechanism of language models
Boxi Cao, Qiaoyu Tang, Hongyu Lin, Shanshan Jiang, Bin Dong, Xianpei Han, Jiawei Chen, Tianshu Wang, and Le Sun. 2024b · 2024
Closest in time.
Art or artifice? large language models and the false promise of creativity
Tuhin Chakrabarty, Philippe Laban, Divyansh Agarwal, Smaranda Muresan, and Chien-Sheng Wu. 2024 · 2024
Closest in time.
Chateval: Towards better llm-based evaluators through multi-agent debate
Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu. 2024 · 2024
Closest in time.
Large knowledge model: Perspectives and challenges
Huajun Chen. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Towards monosemanticity: Decomposing language models with dictionary learning
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nicholas L. Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Alex Tamkin, Karina Nguyen, Brayden McLean, Tristan Hume Josiah E. Burke, Shan Carter, Tom Henighan, , and Chris Olah. 2023 · 2023
Cited alongside, same era.
Weak-to-strong generalization: Eliciting strong capabilities with weak supervision
Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, Ilya Sutskever, and Jeff Wu. 2023 · 2023
Cited alongside, same era.
Evidence of a predictive coding hierarchy in the human brain listening to speech
Charlotte Caucheteux, Alexandre Gramfort, and Jean-Rémi King. 2023 · 2023
Cited alongside, same era.
Natural abstractions: Key claims, theorems, and critiques
Lawrence Chan, Leon Lang, and Erik Jenner. 2023 · 2023
Cited alongside, same era.
Dola: Decoding by contrasting layers improves factuality in large language models
Yung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim, James R. Glass, and Pengcheng He. 2023 · 2023
Cited alongside, same era.
A toy model of universality: Reverse engineering how networks learn group operations
Bilal Chughtai, Lawrence Chan, and Neel Nanda. 2023 · 2023
Cited alongside, same era.
Evaluating the ripple effects of knowledge editing in language models
Roi Cohen, Eden Biran, Ori Yoran, Amir Globerson, and Mor Geva. 2023 · 2023
Cited alongside, same era.
Towards automated circuit discovery for mechanistic interpretability
Arthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso. 2023 · 2023
Cited alongside, same era.
Journey to the center of the knowledge neurons: Discoveries of language-independent knowledge neurons and degenerate knowledge neurons
Yuheng Chen, Pengfei Cao, Yubo Chen, Kang Liu, and Jun Zhao. 2024e · 2024
Closest in time.
Gaming expertise induces meso-scale brain plasticity and efficiency mechanisms as revealed by whole-brain modeling
Carlos Coronel-Oliveros, Vicente Medel, Sebastián Orellana, Julio Rodiño, Fernando Lehue, Josephine Cruzat, Enzo Tagliazucchi, Aneta Brzezicka, Patricio Orio, Natalia Kowalczyk-Grebska, and Agustín Ibáñez. 2024 · 2024
Closest in time.
Induction heads as an essential mechanism for pattern matching in in-context learning
Joy Crosbie and Ekaterina Shutova. 2024 · 2024
Closest in time.
Towards guaranteed safe AI: A framework for ensuring robust and reliable AI systems
David Dalrymple, Joar Skalse, Yoshua Bengio, Stuart Russell, Max Tegmark, Sanjit Seshia, Steve Omohundro, Christian Szegedy, Ben Goldhaber, Nora Ammann, Alessandro Abate, Joe Halpern, Clark W. Barrett, Ding Zhao, Tan Zhi-Xuan, Jeannette Wing, and Joshua B. Tenenbaum. 2024 · 2024
Closest in time.
Creativeval: Evaluating creativity of llm-based hardware code generation
Matthew DeLorenzo, Vasudev Gohil, and Jeyavijayan Rajendran. 2024 · 2024
Closest in time.
Metacognitive capabilities of llms: An exploration in mathematical problem solving
Aniket Didolkar, Anirudh Goyal, Nan Rosemary Ke, Siyuan Guo, Michal Valko, Timothy P. Lillicrap, Danilo J. Rezende, Yoshua Bengio, Michael Mozer, and Sanjeev Arora. 2024 · 2024
Closest in time.
Jump to conclusions: Short-cutting transformers with linear transformations
Alexander Yom Din, Taelin Karidi, Leshem Choshen, and Mor Geva. 2024 · 2024
Closest in time.
Mitigating the problem of strong priors in lms with context extrapolation
Raymond Douglas, Andis Draguns, and Tomas Gavenciak. 2024 · 2024
Closest in time.
Transcoders find interpretable llm feature circuits
Jacob Dunefsky, Philippe Chlenski, and Neel Nanda. 2024 · 2024
Closest in time.
How to think step-by-step: A mechanistic understanding of chain-of-thought reasoning
Subhabrata Dutta, Joykirat Singh, Soumen Chakrabarti, and Tanmoy Chakraborty. 2024 · 2024
Closest in time.
KTO: model alignment as prospect theoretic optimization
Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky, and Douwe Kiela. 2024 · 2024
Closest in time.
Detecting hallucinations in large language models using semantic entropy
Sebastian Farquhar, Jannik Kossen, Lorenz Kuhn, and Yarin Gal. 2024 · 2024
Closest in time.
A primer on the inner workings of transformer-based language models
Javier Ferrando, Gabriele Sarti, Arianna Bisazza, and Marta R. Costa-jussà. 2024 · 2024
Closest in time.
What do the circuits mean? a knowledge edit view
Huaizhi Ge, Frank Rudzicz, and Zining Zhu. 2024 · 2024
Closest in time.
Does fine-tuning llms on new knowledge encourage hallucinations?
Zorik Gekhman, Gal Yona, Roee Aharoni, Matan Eyal, Amir Feder, Roi Reichart, and Jonathan Herzig. 2024 · 2024
Closest in time.
Patchscopes: A unifying framework for inspecting hidden representations of language models
Asma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon, and Mor Geva. 2024 · 2024
Closest in time.
Arcee’s mergekit: A toolkit for merging large language models
Charles Goddard, Shamane Siriwardhana, Malikeh Ehghaghi, Luke Meyers, Vlad Karpukhin, Brian Benedict, Mark McQuade, and Jacob Solawetz. 2024 · 2024
Closest in time.
An ontology of dark patterns knowledge: Foundations, definitions, and a pathway for shared knowledge-building
Colin M. Gray, Cristiana Teixeira Santos, Nataliia Bielova, and Thomas Mildner. 2024 · 2024
Closest in time.
Model editing can hurt general abilities of large language models
Jia-Chen Gu, Hao-Xiang Xu, Jun-Yu Ma, Pan Lu, Zhen-Hua Ling, Kai-Wei Chang, and Nanyun Peng. 2024 · 2024
Closest in time.
Hipporag: Neurobiologically inspired long-term memory for large language models
Bernal Jiménez Gutiérrez, Yiheng Shu, Yu Gu, Michihiro Yasunaga, and Yu Su. 2024 · 2024
Closest in time.
Fundamental problems with model editing: How should rational belief revision work in llms?
Peter Hase, Thomas Hofweber, Xiang Zhou, Elias Stengel-Eskin1, and Mohit Bansal1. 2024 · 2024
Closest in time.
Case-based or rule-based: How do transformers do the math?
Yi Hu, Xiaojuan Tang, Haotong Yang, and Muhan Zhang. 2024 · 2024
Closest in time.
A survey on retrieval-augmented text generation for large language models
Yizheng Huang and Jimmy Huang. 2024 · 2024
Closest in time.
The platonic representation hypothesis
Minyoung Huh, Brian Cheung, Tongzhou Wang, and Phillip Isola. 2024 · 2024
Closest in time.
Semantic encoding during language comprehension at single-cell resolution
Mohsen Jamali, Benjamin Grannan, Jing Cai, Arjun R Khanna, William Muñoz, Irene Caprara, Angelique C Paulk, Sydney S Cash, Evelina Fedorenko, and Ziv M Williams. 2024 · 2024
Closest in time.
Language models resist alignment
Jiaming Ji, Kaile Wang, Tianyi Qiu, Boyuan Chen, Jiayi Zhou, Changye Li, Hantao Lou, and Yaodong Yang. 2024 · 2024
Closest in time.
Latent causal probing: A formal perspective on probing with causal models of data
Charles Jin. 2024 · 2024
Closest in time.
How large language models encode context knowledge? A layer-wise probing study
Tianjie Ju, Weiwei Sun, Wei Du, Xinwei Yuan, Zhaochun Ren, and Gongshen Liu. 2024 · 2024
Closest in time.
Unfamiliar finetuning examples control how language models hallucinate
Katie Kang, Eric Wallace, Claire J. Tomlin, Aviral Kumar, and Sergey Levine. 2024 · 2024
Closest in time.
Measuring progress in dictionary learning for language model interpretability with board game models
Adam Karvonen, Benjamin Wright, Can Rager, Rico Angell, Jannik Brinkmann, Logan Smith, Claudio Mayrink Verdun, David Bau, and Samuel Marks. 2024 · 2024
Closest in time.
Investigating how large language models leverage internal knowledge to perform complex reasoning
Miyoung Ko, Sue Hyun Park, Joonsuk Park, and Minjoon Seo. 2024 · 2024
Closest in time.
Composable interventions for language models
Arinbjorn Kolbeinsson, Kyle O’Brien, Tianjin Huang, Shanghua Gao, Shiwei Liu, Jonathan Richard Schwarz, Anurag Vaidya, Faisal Mahmood, Marinka Zitnik, Tianlong Chen, et al. 2024 · 2024
Closest in time.
Aligning large language models with representation editing: A control perspective
Lingkai Kong, Haorui Wang, Wenhao Mu, Yuanqi Du, Yuchen Zhuang, Yifei Zhou, Yue Song, Rongzhi Zhang, Kai Wang, and Chao Zhang. 2024 · 2024
Closest in time.
Studying large language model behaviors under realistic knowledge conflicts
Evgenii Kortukov, Alexander Rubinstein, Elisa Nguyen, and Seong Joon Oh. 2024 · 2024
Closest in time.
From prompt engineering to collaborating: A human-centered approach to AI interfaces
Tanya Kraljic and Michal Lahav. 2024 · 2024
Closest in time.
Interpreting shared circuits for ordered sequence prediction in a large language model
Michael Lan, Fazl, and Barez. 2024 · 2024
Closest in time.
A mechanistic understanding of alignment algorithms: A case study on DPO and toxicity
Andrew Lee, Xiaoyan Bai, Itamar Pres, Martin Wattenberg, Jonathan K. Kummerfeld, and Rada Mihalcea. 2024 · 2024
Closest in time.
Florin Leon. 2024 · 2024
Closest in time.
PMET: precise model editing in a transformer
Xiaopeng Li, Shasha Li, Shezheng Song, Jing Yang, Jun Ma, and Jie Yu. 2024d · 2024
Closest in time.
Geogalactica: A scientific large language model in geoscience
Zhouhan Lin, Cheng Deng, Le Zhou, Tianhang Zhang, Yi Xu, Yutong Xu, Zhongmou He, Yuanyuan Shi, Beiya Dai, Yunchong Song, Boyi Zeng, Qiyuan Chen, Tao Shi, Tianyu Huang, Yiwei Xu, Shu Wang, Luoyi Fu, Weinan Zhang, Junxian He, Chao Ma, Yunqiang Zhu, Xinbing Wang, and Chenghu Zhou. 2024 · 2024
Closest in time.
Scaling laws for fact memorization of large language models
Xingyu Lu, Xiaonan Li, Qinyuan Cheng, Kai Ding, Xuanjing Huang, and Xipeng Qiu. 2024 · 2024
Closest in time.
From understanding to utilization: A survey on explainability for large language models
Haoyan Luo and Lucia Specia. 2024 · 2024
Closest in time.
Interpreting key mechanisms of factual recall in transformer-based language models
Ang Lv, Kaiyi Zhang, Yuhan Chen, Yulong Wang, Lifeng Liu, Ji-Rong Wen, Jian Xie, and Rui Yan. 2024 · 2024
Closest in time.
Temporal conformity-aware hawkes graph network for recommendations
Chenglong Ma, Yongli Ren, Pablo Castells, and Mark Sanderson. 2024 · 2024
Closest in time.
Sparse feature circuits: Discovering and editing interpretable causal graphs in language models
Samuel Marks, Can Rager, Eric J. Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller. 2024 · 2024
Closest in time.
Llm critics help catch llm bugs
Nat McAleese, Rai (Michael Pokorny), and Juan Felipe Cerón Uribe. 2024 · 2024
Closest in time.
Shortgpt: Layers in large language models are more redundant than you expect
Xin Men, Mingyu Xu, Qingyu Zhang, Bingning Wang, Hongyu Lin, Yaojie Lu, Xianpei Han, and Weipeng Chen. 2024 · 2024
Closest in time.
Simpo: Simple preference optimization with a reference-free reward
Yu Meng, Mengzhou Xia, and Danqi Chen. 2024 · 2024
Closest in time.
We found an neuron in gpt-2
Joseph Miller and Clement Neo. 2024 · 2024
Closest in time.
Training of physical neural networks
Ali Momeni, Babak Rahmani, Benjamin Scellier, Logan G. Wright, Peter L. McMahon, Clara C. Wanjura, Yuhang Li, Anas Skalli, Natalia G. Berloff, Tatsuhiro Onodera, Ilker Oguz, Francesco Morichetti, Philipp del Hougne, Manuel Le Gallo, Abu Sebastian, Azalia Mirhoseini, Cheng Zhang, Danijela Marković, Daniel Brunner, Christophe Moser, Sylvain Gigan, Florian Marquardt, Aydogan Ozcan, Julie Grollier, Andrea J. Liu, Demetri Psaltis, Andrea Alù, and Romain Fleury. 2024 · 2024
Closest in time.
Transformer debugger
Dan Mossing, Steven Bills, Henk Tillman, Tom Dupré la Tour, Nick Cammarata, Leo Gao, Joshua Achiam, Catherine Yeh, Jan Leike, Jeff Wu, and William Saunders. 2024 · 2024
Closest in time.
Advancing human-centric AI for robust x-ray analysis through holistic self-supervised learning
Théo Moutakanni, Piotr Bojanowski, Guillaume Chassagnon, Céline Hudelot, Armand Joulin, Yann LeCun, Matthew Muckley, Maxime Oquab, Marie-Pierre Revel, and Maria Vakalopoulou. 2024 · 2024
Closest in time.
AI knowledge and reasoning: Emulating expert creativity in scientific research
Anirban Mukherjee and Hannah Hanwen Chang. 2024 · 2024
Closest in time.
A survey of synthetic data augmentation methods in computer vision
Alhassan Mumuni, Fuseini Mumuni, and Nana Kobina Gerrar. 2024 · 2024
Closest in time.
Jatin Nainani. 2024 · 2024
Closest in time.
Marianna Nezhurina, Lucia Cipolina-Kun, Mehdi Cherti, and Jenia Jitsev. 2024 · 2024
Closest in time.
Using llms to animate interactive story characters with emotions and personality
Aline Normoyle, João Sedoc, and Funda Durupinar. 2024 · 2024
Closest in time.
Fine-tuning or retrieval? comparing knowledge injection in llms
Oded Ovadia, Menachem Brief, Moshik Mishaeli, and Oren Elisha. 2024 · 2024
Closest in time.
Divergent creativity in humans and large language models
Antoine Bellemare Pépin, François Lespinasse, Philipp Thölke, Yann Harel, Kory Mathewson, Jay A. Olson, Yoshua Bengio, and Karim Jerbi. 2024 · 2024
Closest in time.
Recite, reconstruct, recollect: Memorization in lms as a multifaceted phenomenon
USVSN Sai Prashanth, Alvin Deng, Kyle O’Brien, Jyothir S V, Mohammad Aflah Khan, Jaydeep Borkar, Christopher A. Choquette-Choo, Jacob Ray Fuehne, Stella Biderman, Tracy Ke, Katherine Lee, and Naomi Saphra. 2024 · 2024
Closest in time.
Zero bubble pipeline parallelism
Penghui Qi, Xinyi Wan, Guangxing Huang, and Min Lin. 2024 · 2024
Closest in time.
AUTOACT: automatic agent learning from scratch via self-planning
Shuofei Qiao, Ningyu Zhang, Runnan Fang, Yujie Luo, Wangchunshu Zhou, Yuchen Eleanor Jiang, Chengfei Lv, and Huajun Chen. 2024 · 2024
Closest in time.
Why does new knowledge create messy ripple effects in llms?
Jiaxin Qin, Zixuan Zhang, Chi Han, Manling Li, Pengfei Yu, and Heng Ji. 2024 · 2024
Closest in time.
Brain-inspired artificial intelligence: A comprehensive review
Jing Ren and Feng Xia. 2024 · 2024
Closest in time.
Open problems in technical ai governance
Anka Reuel, Ben Bucknall, Stephen Casper, Tim Fist, Lisa Soder, Onni Aarne, Lewis Hammond, Lujain Ibrahim, Alan Chan, Peter Wills, et al. 2024 · 2024
Closest in time.
On the stochastics of human and artificial creativity
Solve Sæbø and Helge Brovold. 2024 · 2024
Closest in time.
Computing power and the governance of artificial intelligence
Girish Sastry, Lennart Heim, Haydn Belfield, Markus Anderljung, Miles Brundage, Julian Hazell, Cullen O’Keefe, Gillian K. Hadfield, Richard Ngo, Konstantin Pilz, George Gor, Emma Bluemke, Sarah Shoker, Janet Egan, Robert F. Trager, Shahar Avin, Adrian Weller, Yoshua Bengio, and Diane Coyle. 2024 · 2024
Closest in time.
Rethinking LLM memorization through the lens of adversarial compression
Avi Schwarzschild, Zhili Feng, Pratyush Maini, Zachary C. Lipton, and J. Zico Kolter. 2024 · 2024
Closest in time.
Development of cognitive intelligence in pre-trained language models
Raj Sanjay Shah, Khushi Bhardwaj, and Sashank Varma. 2024 · 2024
Closest in time.
Ai-native memory: A pathway from llms towards agi
Jingbo Shang, Zai Zheng, Xiang Ying, Felix Tao, and Mindverse Team. 2024 · 2024
Closest in time.
Clever hans or neural theory of mind? stress testing social reasoning in large language models
Natalie Shapira, Mosh Levy, Seyed Hossein Alavi, Xuhui Zhou, Yejin Choi, Yoav Goldberg, Maarten Sap, and Vered Shwartz. 2024 · 2024
Closest in time.
Locating and editing factual associations in mamba
Arnab Sen Sharma, David Atkinson, and David Bau. 2024 · 2024
Closest in time.
To compress or not to compress - self-supervised learning and information theory: A review
Ravid Shwartz-Ziv and Yann LeCun. 2024 · 2024
Closest in time.
Rethinking interpretability in the era of large language models
Chandan Singh, Jeevana Priya Inala, Michel Galley, Rich Caruana, and Jianfeng Gao. 2024 · 2024
Closest in time.
Should we be going mad? A look at multi-agent debate strategies for llms
Andries P. Smit, Nathan Grinsztajn, Paul Duckworth, Thomas D. Barrett, and Arnu Pretorius. 2024 · 2024
Closest in time.
A collective ai via lifelong learning and sharing at the edge
Andrea Soltoggio, Eseoghene Ben-Iwhiwhu, Vladimir Braverman, Eric Eaton, Benjamin Epstein, Yunhao Ge, Lucy Halperin, Jonathan How, Laurent Itti, Michael A Jacobs, et al. 2024 · 2024
Closest in time.
Evaluation is key: a survey on evaluation measures for synthetic time series
Michael Stenger, Robert Leppich, Ian T. Foster, Samuel Kounev, and André Bauer. 2024 · 2024
Closest in time.
Confidence regulation neurons in language models
Alessandro Stolfo, Ben Wu, Wes Gurnee, Yonatan Belinkov, Xingyi Song, Mrinmaya Sachan, and Neel Nanda. 2024 · 2024
Closest in time.
Learning to (learn at test time): Rnns with expressive hidden states
Yu Sun, Xinhao Li, Karan Dalal, Jiarui Xu, Arjun Vikram, Genghan Zhang, Yann Dubois, Xinlei Chen, Xiaolong Wang, Sanmi Koyejo, Tatsunori Hashimoto, and Carlos Guestrin. 2024 · 2024
Closest in time.
Language-specific neurons: The key to multilingual capabilities in large language models
Tianyi Tang, Wenyang Luo, Haoyang Huang, Dongdong Zhang, Xiaolei Wang, Xin Zhao, Furu Wei, and Ji-Rong Wen. 2024 · 2024
Closest in time.
To forget or not? towards practical knowledge unlearning for large language models
Bozhong Tian, Xiaozhuan Liang, Siyuan Cheng, Qingbin Liu, Mengru Wang, Dianbo Sui, Xi Chen, Huajun Chen, and Ningyu Zhang. 2024 · 2024
Closest in time.
Can go ais be adversarially robust?
Tom Tseng, Euan McLean, Kellin Pelrine, Tony T. Wang, and Adam Gleave. 2024 · 2024
Closest in time.
Teaching transformers causal reasoning through axiomatic training
Vashishtha, Aniket, Kumar, Abhinav, Reddy, Abbavaram Gowtham, Balasubramanian, Vineeth N, Sharma, and Amit. 2024 · 2024
Closest in time.
Martina G. Vilas, Federico Adolfi, David Poeppel, and Gemma Roig. 2024 · 2024
Closest in time.
Extrinsic hallucinations in llms
Lilian Weng. 2024 · 2024
Closest in time.
Reft: Representation finetuning for language models
Zhengxuan Wu, Aryaman Arora, Zheng Wang, Atticus Geiger, Dan Jurafsky, Christopher D. Manning, and Christopher Potts. 2024 · 2024
Closest in time.
Agentgym: Evolving large language model-based agents across diverse environments
Zhiheng Xi, Yiwen Ding, Wenxiang Chen, Boyang Hong, Honglin Guo, Junzhe Wang, Dingwen Yang, Chenyang Liao, Xin Guo, Wei He, Songyang Gao, Lu Chen, Rui Zheng, Yicheng Zou, Tao Gui, Qi Zhang, Xipeng Qiu, Xuanjing Huang, Zuxuan Wu, and Yu-Gang Jiang. 2024 · 2024
Closest in time.
Can large language model agents simulate human trust behaviors?
Chengxing Xie, Canyu Chen, Feiran Jia, Ziyu Ye, Kai Shu, Adel Bibi, Ziniu Hu, Philip H. S. Torr, Bernard Ghanem, and Guohao Li. 2024 · 2024
Closest in time.
Potential and challenges of model editing for social debiasing
Jianhao Yan, Futing Wang, Yafu Li, and Yue Zhang. 2024 · 2024
Closest in time.
Knowledge circuits in pretrained transformers
Yunzhi Yao, Ningyu Zhang, Zekun Xi, Mengru Wang, Ziwen Xu, Shumin Deng, and Huajun Chen. 2024 · 2024
Closest in time.
Physics of language models: Part 2.1, grade-school math and the hidden reasoning process
Tian Ye, Zicheng Xum, Yuanzhi Li, and Zeyuan Allen-Zhu. 2024 · 2024
Closest in time.
Weak-to-strong extrapolation expedites alignment
Chujie Zheng, Ziqi Wang, Heng Ji, Minlie Huang, and Nanyun Peng. 2024 · 2024
Closest in time.
Shared imagination: Llms hallucinate alike
Yilun Zhou, Caiming Xiong, Silvio Savarese, and Chien-Sheng Wu. 2024 · 2024
Closest in time.
Language models represent beliefs of self and others
Wentao Zhu, Zhining Zhang, and Yizhou Wang. 2024 · 2024
Closest in time.
Improving alignment and robustness with circuit breakers
Andy Zou, Long Phan, Justin Wang, Derek Duenas, Maxwell Lin, Maksym Andriushchenko, Rowan Wang, Zico Kolter, Matt Fredrikson, and Dan Hendrycks. 2024 · 2024
Closest in time.
Molgpt: Molecular generation using a transformer-decoder model
Viraj Bagal, Rishal Aggarwal, P. K. Vinod, and U. Deva Priyakumar. 2022 · 2076
Closest in time.