Fetching the paper…
Reading the bibliography…
We introduce Concept Bottleneck Protein Language Models (CB-pLM), a generative masked language model with a layer where each neuron corresponds to an interpretable concept.
Prediction of protein antigenic determinants from amino acid sequences
TP Hopp and KR Woods · 1981
Earlier work this paper cites.
A simple method for displaying the hydropathic character of a protein
Jack Kyte and Russell F. Doolittle · 1982
Earlier work this paper cites.
Correlation between stability of a protein and its dipeptide composition: a novel approach for predicting in vivo stability of a protein from its primary sequence
Kunchur Guruprasad, B.V.Bhasker Reddy, and Madhusudan W. Pandit · 1990
Earlier work this paper cites.
The swiss-prot protein sequence database and its supplement trembl in 2000
Amos Bairoch and Rolf Apweiler · 2000
Earlier work this paper cites.
Hydrophobic, hydrophilic, and charged amino acid networks within protein
Md Aftabuddin and S Kundu · 2007
Earlier work this paper cites.
Biopython: freely available python tools for computational molecular biology and bioinformatics
Peter JA Cock, Tiago Antao, Jeffrey T Chang, Brad A Chapman, Cymon J Cox, Andrew Dalke, Iddo Friedberg, Thomas Hamelryck, Frank Kauff, Bartek Wilczynski, et al · 2009
Earlier work this paper cites.
How to explain individual classification decisions
David Baehrens, Timon Schroeter, Stefan Harmeling, Motoaki Kawanabe, Katja Hansen, and Klaus-Robert MÞller · 2010
Earlier work this paper cites.
Siltuximab, a novel anti–interleukin-6 monoclonal antibody, for castleman’s disease
Frits Van Rhee, Luis Fayad, Peter Voorhees, Richard Furman, Sagar Lonial, Hossein Borghaei, Lubomir Sokol, Julie Crawford, Mark Cornfeld, Ming Qi, et al · 2010
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Matthew D Zeiler and Rob Fergus · 2014
Earlier work this paper cites.
Aggregation risk prediction for antibodies and its application to biotherapeutic development
Olga Obrezanova, Andreas Arnell, Ramón Gómez De La Cuesta, Maud E Berthelot, Thomas RA Gallagher, Jesús Zurdo, and Yvette Stallwood · 2015
Earlier work this paper cites.
Uniref clusters: a comprehensive and scalable alternative for improving sequence similarity searches
Baris E Suzek, Yuqi Wang, Hongzhan Huang, Peter B McGarvey, Cathy H Wu, and UniProt Consortium · 2015
Earlier work this paper cites.
European union regulations on algorithmic decision-making and a “right to explanation”
Bryce Goodman and Seth Flaxman · 2017
Earlier work this paper cites.
Design and preparation of biomimetic and bioinspired materials
V Leiro, PM Moreno, B Sarmento, J Durão, L Gales, AP Pêgo, and CC Barrias · 2017
Earlier work this paper cites.
Not just a black box: Learning important features through propagating activation differences, 2017
Avanti Shrikumar, Peyton Greenside, Anna Shcherbina, and Anshul Kundaje · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan · 2017
Earlier work this paper cites.
Ashish Vaswani · 2017
Earlier work this paper cites.
Sanity checks for saliency maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Evaluating feature importance estimates
Sara Hooker, Dumitru Erhan, Pieter-Jan Kindermans, and Been Kim · 2018
Earlier work this paper cites.
Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences
Alexander Rives, Joshua Meier, Tom Sercu, Siddharth Goyal, Zeming Lin, Jason Liu, Demi Guo, Myle Ott, C. Lawrence Zitnick, Jerry Ma, and Rob Fergus · 2019
Earlier work this paper cites.
Solubility-Weighted Index: fast and accurate prediction of protein solubility
Bikash K Bhandari, Paul P Gardner, and Chun Shen Lim · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Artificial intelligence, values, and alignment
Iason Gabriel · 2020
Earlier work this paper cites.
Concept bottleneck models
Pang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann, Emma Pierson, Been Kim, and Percy Liang · 2020
Earlier work this paper cites.
Progen: Language modeling for protein generation
Ali Madani, Bryan McCann, Nikhil Naik, Nitish Shirish Keskar, Namrata Anand, Raphael R Eguchi, Po-Ssu Huang, and Richard Socher · 2020
Cited alongside, same era.
Promises and pitfalls of black-box concept learning models
Anita Mahinpei, Justin Clark, Isaac Lage, Finale Doshi-Velez, and Weiwei Pan · 2021
Cited alongside, same era.
Do concept bottleneck models learn as intended?
Andrei Margeloiu, Matthew Ashman, Umang Bhatt, Yanzhi Chen, Mateja Jamnik, and Adrian Weller · 2021
Cited alongside, same era.
Optimization of therapeutic antibodies by predicting antigen specificity from antibody sequence via deep learning
Derek M Mason, Simon Friedensohn, Cédric R Weber, Christian Jordi, Bastian Wagner, Simon M Meng, Roy A Ehling, Lucia Bonati, Jan Dahinden, Pablo Gainza, et al · 2021
Cited alongside, same era.
Orthogonal projection loss
xTrimoPGLM: Unified 100B-Scale Pre-trained Transformer for Deciphering the Language of Protein, 2023
Bo Chen, Xingyi Cheng, Yangli-ao Geng, Shen Li, Xin Zeng, Boyan Wang, Jing Gong, Chiming Liu, Aohan Zeng, Yuxiao Dong, Jie Tang, and Le Song · 2023
Later among the works it cites.
Sparse autoencoders find highly interpretable features in language models
Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey · 2023
Later among the works it cites.
Protein discovery with discrete walk-jump sampling
Nathan C Frey, Daniel Berenberg, Karina Zadorozhny, Joseph Kleinhenz, Julien Lafrance-Vanasse, Isidro Hotzel, Yan Wu, Stephen Ra, Richard Bonneau, Kyunghyun Cho, et al · 2023
Later among the works it cites.
Concept bottleneck generative models
Aya Abdelsalam Ismail, Julius Adebayo, Hector Corrada Bravo, Stephen Ra, and Kyunghyun Cho · 2023
Later among the works it cites.
Evolutionary-scale prediction of atomic-level protein structure with a language model
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kanchana Ranasinghe, Muzammal Naseer, Munawar Hayat, Salman Khan, and Fahad Shahbaz Khan · 2021
Cited alongside, same era.
Do input gradients highlight discriminative features?
Harshay Shah, Prateek Jain, and Praneeth Netrapalli · 2021
Cited alongside, same era.
Generative language modeling for antibody design
Richard W Shuai, Jeffrey A Ruffolo, and Jeffrey J Gray · 2021
Cited alongside, same era.
Antibody optimization enabled by artificial intelligence predictions of binding affinity and naturalness
Sharrol Bachas, Goran Rakocevic, David Spencer, Anand V Sastry, Robel Haile, John M Sutton, George Kasun, Andrew Stachyra, Jahir M Gutierrez, Edriss Yassine, et al · 2022
Cited alongside, same era.
ProtTrans: Toward Understanding the Language of Life Through Self-Supervised Learning
Ahmed Elnaggar, Michael Heinzinger, Christian Dallago, Ghalia Rehawi, Yu Wang, Llion Jones, Tom Gibbs, Tamas Feher, Christoph Angerer, Martin Steinegger, Debsindhu Bhowmik, and Burkhard Rost · 2022
Cited alongside, same era.
Addressing leakage in concept bottleneck models
Marton Havasi, Sonali Parbhoo, and Finale Doshi-Velez · 2022
Cited alongside, same era.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2022
Cited alongside, same era.
Language models of protein sequences at the scale of evolution enable accurate structure prediction
Zeming Lin, Halil Akin, Roshan Rao, Brian Hie, Zhongkai Zhu, Wenting Lu, Nikita Smetanin, Allan dos Santos Costa, Maryam Fazel-Zarandi, Tom Sercu, Sal Candido, et al · 2022
Cited alongside, same era.
Zeming Lin, Halil Akin, Roshan Rao, Brian Hie, Zhongkai Zhu, Wenting Lu, Nikita Smetanin, Robert Verkuil, Ori Kabeli, Yaniv Shmueli, et al · 2023
Later among the works it cites.
Label-free concept bottleneck models
Tuomas Oikarinen, Subhro Das, Lam M Nguyen, and Tsui-Wei Weng · 2023
Later among the works it cites.
Toward transparent ai: A survey on interpreting the inner structures of deep neural networks
Tilman Räuker, Anson Ho, Stephen Casper, and Dylan Hadfield-Menell · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
ProtST: Multi-Modality Learning of Protein Sequences and Biomedical Texts, 2023
Minghao Xu, Xinyu Yuan, Santiago Miret, and Jian Tang · 2023
Later among the works it cites.
Language in a bottle: Language model guided concept bottlenecks for interpretable image classification
Yue Yang, Artemis Panagopoulou, Shenghao Zhou, Daniel Jin, Chris Callison-Burch, and Mark Yatskar · 2023
Later among the works it cites.
scgpt: toward building a foundation model for single-cell multi-omics using generative ai
Haotian Cui, Chloe Wang, Hassaan Maan, Kuan Pang, Fengning Luo, Nan Duan, and Bo Wang · 2024
Closest in time.
Cramming protein language model training in 24 gpu hours
Nathan C Frey, Taylor Joren, Aya Ismail, Allen Goodman, Richard Bonneau, Kyunghyun Cho, and Vladimir Gligorijevic · 2024
Closest in time.
Protein design with guided discrete diffusion
Nate Gruver, Samuel Stanton, Nathan Frey, Tim GJ Rudner, Isidro Hotzel, Julien Lafrance-Vanasse, Arvind Rajpal, Kyunghyun Cho, and Andrew G Wilson · 2024
Closest in time.
Simulating 500 million years of evolution with a language model
Thomas Hayes, Roshan Rao, Halil Akin, Nicholas J. Sofroniew, Deniz Oktay, Zeming Lin, Robert Verkuil, Vincent Q. Tran, Jonathan Deaton, Marius Wiggert, Rohil Badkundri, Irhum Shafkat, Jun Gong, Alexander Derry, Raul S. Molina, Neil Thomas, Yousuf Khan, Chetan Mishra, Carolyn Kim, Liam J. Bartie, Matthew Nemeth, Patrick D. Hsu, Tom Sercu, Salvatore Candido, and Alexander Rives · 2024
Closest in time.
Efficient evolution of human antibodies from general protein language models
Brian L Hie, Varun R Shanker, Duo Xu, Theodora UJ Bruun, Payton A Weidenbacher, Shaogeng Tang, Wesley Wu, John E Pak, and Peter S Kim · 2024
Closest in time.
Discover-then-name: Task-agnostic concept bottlenecks via automated concept discovery
Sukrut Sridhar Rao, Sweta Mahajan, Moritz Böhle, and Bernt Schiele · 2024
Closest in time.
Adapting protein language models for structure-conditioned design
Jeffrey A Ruffolo, Aadyot Bhatnagar, Joel Beazer, Stephen Nayfach, Jordan Russ, Emily Hill, Riffat Hussain, Joseph Gallagher, and Ali Madani · 2024
Closest in time.
Crafting large language models for enhanced interpretability, 2024
Chung-En Sun, Tuomas Oikarinen, and Tsui-Wei Weng · 2024
Closest in time.
Implicitly guided design with propen: Match your data to follow the gradient
Nataša Tagasovska, Vladimir Gligorijević, Kyunghyun Cho, and Andreas Loukas · 2024
Closest in time.
Interpreting pretrained language models via concept bottlenecks
Zhen Tan, Lu Cheng, Song Wang, Bo Yuan, Jundong Li, and Huan Liu · 2024
Closest in time.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, Hoagy Cunningham, Nicholas L Turner, Callum McDougall, Monte MacDiarmid, C. Daniel Freeman, Theodore R. Sumers, Edward Rees, Joshua Batson, Adam Jermyn, Shan Carter, Chris Olah, and Tom Henighan · 2024
Closest in time.
Explainability for large language models: A survey
Haiyan Zhao, Hanjie Chen, Fan Yang, Ninghao Liu, Huiqi Deng, Hengyi Cai, Shuaiqiang Wang, Dawei Yin, and Mengnan Du · 2024
Closest in time.