Fetching the paper…
Reading the bibliography…
To make sense of massive data, we often fit simplified models and then interpret the parameters; for example, we cluster the text embeddings and then interpret the mean parameters of each cluster.
Latent dirichlet allocation
David M Blei, Andrew Y Ng, and Michael I Jordan · 2003
Earlier work this paper cites.
Finding scientific topics
Thomas L Griffiths and Mark Steyvers · 2004
Earlier work this paper cites.
Stability-based validation of clustering solutions
Tilman Lange, Volker Roth, Mikio L Braun, and Joachim M Buhmann · 2004
Earlier work this paper cites.
Dynamic topic models
David M Blei and John D Lafferty · 2006
Earlier work this paper cites.
Automatically labeling hierarchical clusters
Pucktada Treeratpituk and Jamie Callan · 2006
Earlier work this paper cites.
The new york times annotated corpus
Evan Sandhaus · 2008
Earlier work this paper cites.
Enhancing cluster labeling using wikipedia
David Carmel, Haggai Roitman, and Naama Zwerdling · 2009
Earlier work this paper cites.
Reading tea leaves: How humans interpret topic models
Jonathan Chang, Sean Gerrish, Chong Wang, Jordan Boyd-Graber, and David Blei · 2009
Earlier work this paper cites.
Detecting influenza epidemics using search engine query data
Jeremy Ginsberg, Matthew H Mohebbi, Rajan S Patel, Lynnette Brammer, Mark S Smolinski, and Larry Brilliant · 2009
Earlier work this paper cites.
Understanding the intrinsic memorability of images
Phillip Isola, Devi Parikh, Antonio Torralba, and Aude Oliva · 2011
Earlier work this paper cites.
Baselines and bigrams: Simple, good sentiment and topic classification
Sida I Wang and Christopher D Manning · 2012
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun · 2015
Earlier work this paper cites.
Learning with latent language
Jacob Andreas, Dan Klein, and Sergey Levine · 2018
Earlier work this paper cites.
Fairness without demographics in repeated loss minimization
Tatsunori Hashimoto, Megha Srivastava, Hongseok Namkoong, and Percy Liang · 2018
Earlier work this paper cites.
Regularizing and optimizing LSTM language models
Stephen Merity, Nitish Shirish Keskar, and Richard Socher · 2018
Earlier work this paper cites.
Taxogen: Unsupervised topic taxonomy construction by adaptive term embedding and clustering
Chao Zhang, Fangbo Tao, Xiusi Chen, Jiaming Shen, Meng Jiang, Brian Sadler, Michelle Vanni, and Jiawei Han · 2018
Earlier work this paper cites.
Unsupervised domain clusters in pretrained language models
Roee Aharoni and Yoav Goldberg · 2020
Earlier work this paper cites.
Concept bottleneck models
Pang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann, Emma Pierson, Been Kim, and Percy Liang · 2020
Earlier work this paper cites.
Shaping visual representations with language for few-shot classification
Jesse Mu, Percy Liang, and Noah Goodman · 2020
Earlier work this paper cites.
How we do things with words: Analyzing text as social and cultural data
Dong Nguyen, Maria Liakata, Simon DeDeo, Jacob Eisenstein, David Mimno, Rebekah Tromble, and Jane Winters · 2020
Earlier work this paper cites.
Automatic detection of fake news
Pontus Nordberg, Joakim Kävrestad, and Marcus Nohlberg · 2020
Cited alongside, same era.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh · 2020
Cited alongside, same era.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Cited alongside, same era.
Measuring mathematical problem solving with the math dataset
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt · 2021
Cited alongside, same era.
What do users care about? detecting actionable insights from user feedback
Interpretable-by-design text classification with iteratively generated concept bottleneck
Josh Magnus Ludan, Qing Lyu, Yue Yang, Liam Dugan, Mark Yatskar, and Chris Callison-Burch · 2023
Later among the works it cites.
Demonstration of insightpilot: An llm-empowered automated data exploration system
Pingchuan Ma, Rui Ding, Shuai Wang, Shi Han, and Dongmei Zhang · 2023
Later among the works it cites.
Topicgpt: A prompt-based topic modeling framework
Chau Minh Pham, Alexander Hoyle, Simeng Sun, and Mohit Iyyer · 2023
Later among the works it cites.
Explaining black box text modules in natural language with language models
Chandan Singh, Aliyah R Hsu, Richard Antonello, Shailee Jain, Alexander G Huth, Bin Yu, and Jianfeng Gao · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kasturi Bhattacharjee, Rashmi Gangadharaiah, Kathleen McKeown, and Dan Roth · 2022
Cited alongside, same era.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al · 2022
Cited alongside, same era.
RLPrompt: Optimizing discrete text prompts with reinforcement learning
Mingkai Deng, Jianyu Wang, Cheng-Ping Hsieh, Yihan Wang, Han Guo, Tianmin Shu, Meng Song, Eric Xing, and Zhiting Hu · 2022
Cited alongside, same era.
Bertopic: Neural topic modeling with a class-based tf-idf procedure
Maarten Grootendorst · 2022
Cited alongside, same era.
Instruction induction: From few examples to natural language task descriptions
Or Honovich, Uri Shaham, Samuel R. Bowman, and Omer Levy · 2022
Cited alongside, same era.
Are neural topic models broken?
Alexander Miserlis Hoyle, Pranav Goel, Rupak Sarkar, and Philip Resnik · 2022
Cited alongside, same era.
Lila: Language-informed latent actions
Siddharth Karamcheti, Megha Srivastava, Percy Liang, and Dorsa Sadigh · 2022
Cited alongside, same era.
Coauthor: Designing a human-ai collaborative writing dataset for exploring language model capabilities
Mina Lee, Percy Liang, and Qian Yang · 2022
Cited alongside, same era.
Christopher T Small, Ivan Vendrov, Esin Durmus, Hadjar Homaei, Elizabeth Barry, Julien Cornebise, Ted Suzman, Deep Ganguli, and Colin Megill · 2023
Later among the works it cites.
Large language models enable few-shot clustering
Vijay Viswanathan, Kiril Gashteovski, Carolin Lawrence, Tongshuang Wu, and Graham Neubig · 2023
Later among the works it cites.
Goal-driven explainable clustering via language descriptions
Zihan Wang, Jingbo Shang, and Ruiqi Zhong · 2023
Later among the works it cites.
Learning adaptive planning representations with natural language guidance
Lionel Wong, Jiayuan Mao, Pratyusha Sharma, Zachary S Siegel, Jiahai Feng, Noa Korneev, Joshua B Tenenbaum, and Jacob Andreas · 2023
Later among the works it cites.
Goal driven discovery of distributional differences via language descriptions
Ruiqi Zhong, Peter Zhang, Steve Li, Jinwoo Ahn, Dan Klein, and Jacob Steinhardt · 2023
Later among the works it cites.
Chatbot arena: An open platform for evaluating llms by human preference
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E Gonzalez, et al · 2024
Closest in time.
Evolving interpretable visual classifiers with large language models
Mia Chiquier, Utkarsh Mall, and Carl Vondrick · 2024
Closest in time.
Concept induction: Analyzing unstructured text with high-level concepts using lloom
Michelle S Lam, Janice Teoh, James Landay, Jeffrey Heer, and Michael S Bernstein · 2024
Closest in time.
Gpt-3.5 turbo
OpenAI · 2024
Closest in time.
Phenomenal yet puzzling: Testing inductive reasoning capabilities of language models with hypothesis refinement
Linlu Qiu, Liwei Jiang, Ximing Lu, Melanie Sclar, Valentina Pyatkin, Chandra Bhagavatula, Bailin Wang, Yoon Kim, Yejin Choi, Nouha Dziri, and Xiang Ren · 2024
Closest in time.
Concept bottleneck models without predefined concepts
Simon Schrodi, Julian Schur, Max Argus, and Thomas Brox · 2024
Closest in time.
Rethinking interpretability in the era of large language models
Chandan Singh, Jeevana Priya Inala, Michel Galley, Rich Caruana, and Jianfeng Gao · 2024
Closest in time.
Wildchat: 1m chatGPT interaction logs in the wild
Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2024
Closest in time.
Goal driven discovery of distributional differences via language descriptions
Ruiqi Zhong, Peter Zhang, Steve Li, Jinwoo Ahn, Dan Klein, and Jacob Steinhardt · 2024
Closest in time.