Fetching the paper…
Reading the bibliography…
We describe HypotheSAEs, a general method to hypothesize interpretable relationships between text data (e.g., headlines) and a target variable (e.g., clicks).
Linguistic effects on news headline success: Evidence from thousands of online field experiments (Registered Report)
Gligorić, K., Lifchits, G., West, R., and Anderson, A · 1932
Earlier work this paper cites.
Fightin’ Words: Lexical Feature Selection and Evaluation for Identifying the Content of Political Conflict
Monroe, B. L., Colaresi, M. P., and Quinn, K. M · 1987
Earlier work this paper cites.
Regression Shrinkage and Selection Via the Lasso
Tibshirani, R · 1996
Earlier work this paper cites.
Latent dirichlet allocation
Blei, D. M., Ng, A. Y., and Jordan, M. I · 2003
Earlier work this paper cites.
Reading Tea Leaves: How Humans Interpret Topic Models
Chang, J., Gerrish, S., Wang, C., Boyd-graber, J., and Blei, D · 2009
Earlier work this paper cites.
What Drives Media Slant? Evidence From U.S. Daily Newspapers
Gentzkow, M. and Shapiro, J. M · 2010
Earlier work this paper cites.
A bayesian hierarchical topic model for political texts: Measuring expressed agendas in senate press releases
Grimmer, J · 2010
Earlier work this paper cites.
Twitter mood predicts the stock market
Bollen, J., Mao, H., and Zeng, X.-J · 2011
Earlier work this paper cites.
K-Sparse Autoencoders, March 2014
Makhzani, A. and Frey, B · 2014
Earlier work this paper cites.
Measuring Group Differences in High-Dimensional Choices: Method and Application to Congressional Speech, July 2016
Gentzkow, M., Shapiro, J. M., and Taddy, M · 2016
Earlier work this paper cites.
Pointer Sentinel Mixture Models, September 2016
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Earlier work this paper cites.
Yelp reviews of hospital care can supplement and inform traditional surveys of the patient experience of care
Ranard, B. L., Werner, R. M., Antanavicius, T., Schwartz, H. A., Smith, R. J., Meisel, Z. F., Asch, D. A., Ungar, L. H., and Merchant, R. M · 2016
Earlier work this paper cites.
Feature Visualization
Olah, C., Mordvintsev, A., and Schubert, L · 2017
Earlier work this paper cites.
Using big data and text analytics to understand how customer experiences posted on yelp. com impact the hospitality industry
Ting, P.-J. L., Chen, S.-L., Chen, H., and Fang, W.-C · 2017
Earlier work this paper cites.
Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV)
Kim, B., Wattenberg, M., Gilmer, J., Cai, C., Wexler, J., Viegas, F., and Sayres, R · 2018
Earlier work this paper cites.
Analyzing Polarization in Social Media: Method and Application to Tweets on 21 Mass Shootings, April 2019
Demszky, D., Garg, N., Voigt, R., Zou, J., Gentzkow, M., Shapiro, J., and Jurafsky, D · 2019
Earlier work this paper cites.
Concept Bottleneck Models, December 2020
Koh, P. W., Nguyen, T., Tang, Y. S., Mussmann, S., Pierson, E., Kim, B., and Liang, P · 2020
Earlier work this paper cites.
Computational grounded theory: A methodological framework
Nelson, L. K · 2020
Earlier work this paper cites.
Machine Learning for Social Science: An Agnostic Approach
Grimmer, J., Roberts, M. E., and Stewart, B. M · 2021
Cited alongside, same era.
Out-group animosity drives engagement on social media
Rathje, S., Van Bavel, J. J., and Van Der Linden, S · 2021
Cited alongside, same era.
BERTology Meets Biology: Interpreting Attention in Protein Language Models, March 2021
Vig, J., Madani, A., Varshney, L. R., Xiong, C., Socher, R., and Rajani, N. F · 2021
Cited alongside, same era.
Computational analysis of 140 years of US political speeches reveals more positive but increasingly polarized framing of immigration
Card, D., Chang, S., Becker, C., Mendelsohn, J., Voigt, R., Boustan, L., Abramitzky, R., and Jurafsky, D · 2022
Cited alongside, same era.
Toy Models of Superposition, September 2022
Elhage, N., Hume, T., Olsson, C., Schiefer, N., Henighan, T., Kravec, S., Hatfield-Dodds, Z., Lasenby, R., Drain, D., Chen, C., Grosse, R., McCandlish, S., Kaplan, J., Amodei, D., Wattenberg, M., and Olah, C · 2022
Cited alongside, same era.
Describing Differences in Image Sets with Natural Language, April 2024
Dunlap, L., Zhang, Y., Wang, X., Zhong, R., Darrell, T., Steinhardt, J., Gonzalez, J. E., and Yeung-Levy, S · 2024
Later among the works it cites.
Bayesian Concept Bottleneck Models with LLM Priors, October 2024
Feng, J., Kothari, A., Zier, L., Singh, C., and Tan, Y. S · 2024
Later among the works it cites.
Scaling and evaluating sparse autoencoders, June 2024
Gao, L., la Tour, T. D., Tillman, H., Goh, G., Troll, R., Radford, A., Sutskever, I., Leike, J., and Wu, J · 2024
Later among the works it cites.
LLM & Generative AI - BERTopic
Grootendorst, M · 2024
Later among the works it cites.
Concept induction: Analyzing unstructured text with high-level concepts using lloom, 2024
Lam, M. S., Teoh, J., Landay, J., Heer, J., and Bernstein, M. S · 2024
Later among the works it cites.
Interpretable-by-Design Text Understanding with Iteratively Generated Concept Bottleneck, April 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Text as Data: A New Framework for Machine Learning and the Social Sciences
Grimmer, J., Roberts, M. E., and Stewart, B. M · 2022
Cited alongside, same era.
BERTopic: Neural topic modeling with a class-based TF-IDF procedure, March 2022
Grootendorst, M · 2022
Cited alongside, same era.
Are Neural Topic Models Broken?
Hoyle, A. M., Sarkar, R., Goel, P., and Resnik, P · 2022
Cited alongside, same era.
Negative Patient Descriptors: Documenting Racial Bias In The Electronic Health Record
Sun, M., Oliwa, T., Peek, M. E., and Tung, E. L · 2022
Cited alongside, same era.
Language models can explain neurons in language models
Bills, S., Cammarata, N., Mossing, D., Tillman, H., Gao, L., Goh, G., Sutskever, I., Leike, J., Wu, J., and Saunders, W · 2023
Cited alongside, same era.
Towards monosemanticity: Decomposing language models with dictionary learning
Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N., Anil, C., Denison, C., Askell, A., et al · 2023
Cited alongside, same era.
Sparse autoencoders find highly interpretable features in language models
Cunningham, H., Ewart, A., Riggs, L., Huben, R., and Sharkey, L · 2023
Cited alongside, same era.
Ludan, J. M., Lyu, Q., Yang, Y., Dugan, L., Yatskar, M., and Callison-Burch, C · 2024
Later among the works it cites.
Disentangling Dense Embeddings with Sparse Autoencoders, August 2024
O’Neill, C., Ye, C., Iyer, K., and Wu, J. F · 2024
Later among the works it cites.
TopicGPT: A Prompt-based Topic Modeling Framework, April 2024
Pham, C. M., Hoyle, A., Sun, S., Resnik, P., and Iyyer, M · 2024
Later among the works it cites.
A large language model-based approach to quantifying the effects of social determinants in liver transplant decisions, January 2025
Robitschek, E., Bastani, A., Horwath, K., Sordean, S., Pletcher, M. J., Lai, J. C., Galletta, S., Ash, E., Ge, J., and Chen, I. Y · 2024
Later among the works it cites.
Concept Bottleneck Large Language Models, December 2024
Sun, C.-E., Oikarinen, T., Ustun, B., and Weng, T.-W · 2024
Later among the works it cites.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet. transformer circuits thread, 2024
Templeton, A., Conerly, T., Marcus, J., Lindsey, J., Bricken, T., Chen, B., Pearce, A., Citro, C., Ameisen, E., Jones, A., et al · 2024
Later among the works it cites.
Yelp Open Dataset
Yelp · 2024
Later among the works it cites.
Explaining Datasets in Words: Statistical Models with Natural Language Parameters, September 2024
Zhong, R., Wang, H., Klein, D., and Steinhardt, J · 2024
Later among the works it cites.
Hypothesis Generation with Large Language Models, August 2024
Zhou, Y., Liu, H., Srivastava, T., Mei, H., and Tan, C · 2024
Later among the works it cites.
Can Large Language Models Transform Computational Social Science?, February 2024
Ziems, C., Held, W., Shaikh, O., Chen, J., Zhang, Z., and Yang, D · 2024
Later among the works it cites.
Using AI to Summarize US Presidential Campaign TV Advertisement Videos, 1952-2012, March 2025
Breuer, A., Dietrich, B. J., Crespin, M. H., Butler, M., Pyrse, J. A., and Imai, K · 2025
Closest in time.
When curiosity gaps backfire: Effects of headline concreteness on information selection decisions
Aubin Le Quéré, M. and Matias, J. N · 2045
Closest in time.
The Upworthy Research Archive, a time series of 32,487 experiments in U.S. media
Matias, J. N., Munger, K., Le Quere, M. A., and Ebersole, C · 2052
Closest in time.