Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have demonstrated remarkable capabilities in performing complex tasks.
Deep inside convolutional networks: Visualising image classification models and saliency maps
K. Simonyan, A. Vedaldi, and A. Zisserman · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
M. D. Zeiler and R. Fergus · 2014
Earlier work this paper cites.
“Why should I trust you?" Explaining the predictions of any classifier
M. T. Ribeiro, S. Singh, and C. Guestrin · 2016
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
F. Doshi-Velez and B. Kim · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
S. M. Lundberg and S.-I. Lee · 2017
Earlier work this paper cites.
Learning important features through propagating activation differences
A. Shrikumar, P. Greenside, and A. Kundaje · 2017
Earlier work this paper cites.
Smoothgrad: Removing noise by adding noise
D. Smilkov, N. Thorat, B. Kim, F. Viégas, and M. Wattenberg · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
M. Sundararajan, A. Taly, and Q. Yan · 2017
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
J. Devlin, M. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
CommonsenseQA: A question answering challenge targeting commonsense knowledge
A. Talmor, J. Herzig, N. Lourie, and J. Berant · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Robust and stable black box explanations
H. Lakkaraju, N. Arsov, and O. Bastani · 2020
Earlier work this paper cites.
On the opportunities and risks of foundation models
R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskill, et al · 2021
Cited alongside, same era.
Towards benchmarking the utility of explanations for model debugging
M. Idahl, L. Lyu, U. Gadiraju, and A. Anand · 2021
Cited alongside, same era.
How can i choose an explainer? an application-grounded evaluation of post-hoc explanations
S. Jesus, C. Belém, V. Balayan, J. Bento, P. Saleiro, P. Bizarro, and J. Gama · 2021
Cited alongside, same era.
Reliable post hoc explanations: Modeling uncertainty in explainability
D. Slack, A. Hilgard, S. Singh, and H. Lakkaraju · 2021
Cited alongside, same era.
https://platform.openai.com/docs/model-index-for-researchers , a
Gpt-3.5-turbo · 2022
Cited alongside, same era.
Post-hoc interpretability for neural nlp: A survey
A. Madsen, S. Reddy, and S. Chandar · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Later among the works it cites.
Talktomodel: Understanding machine learning models with open ended dialogues
D. Slack, S. Krishna, H. Lakkaraju, and S. Singh · 2022
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
A. Srivastava, A. Rastogi, A. Rao, A. A. M. Shoeb, A. Abid, A. Fisch, A. R. Brown, A. Santoro, A. Gupta, A. Garriga-Alonso, et al · 2022
Later among the works it cites.
Challenging big-bench tasks and whether chain-of-thought can solve them
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
https://cdn.openai.com/papers/gpt-4-system-card.pdf , b
Gpt-4 system card · 2022
Cited alongside, same era.
Interpretation of black box nlp models: A survey
S. Choudhary, N. Chatterjee, and S. K. Saha · 2022
Cited alongside, same era.
A survey for in-context learning
Q. Dong, L. Li, D. Dai, C. Zheng, Z. Wu, B. Chang, X. Sun, J. Xu, and Z. Sui · 2022
Cited alongside, same era.
The disagreement problem in explainable machine learning: A practitioner’s perspective
S. Krishna, T. Han, A. Gu, J. Pombra, S. Jabbari, S. Wu, and H. Lakkaraju · 2022
Cited alongside, same era.
Can language models learn from explanations in context?
A. K. Lampinen, I. Dasgupta, S. C. Chan, K. Matthewson, M. H. Tessler, A. Creswell, J. L. McClelland, J. X. Wang, and F. Hill · 2022
Cited alongside, same era.
Estimating the carbon footprint of bloom, a 176b parameter language model
A. S. Luccioni, S. Viguier, and A.-L. Ligozat · 2022
Cited alongside, same era.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever
Cited in the paper.
M. Suzgun, N. Scales, N. Schärli, S. Gehrmann, Y. Tay, H. W. Chung, A. Chowdhery, Q. V. Le, E. H. Chi, D. Zhou, et al · 2022
Later among the works it cites.
Interpreting language models with contrastive explanations
K. Yin and G. Neubig · 2022
Later among the works it cites.
Sparks of artificial general intelligence: Early experiments with gpt-4
S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. Lundberg, et al · 2023
Closest in time.
Survey of hallucination in natural language generation
Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y. Xu, E. Ishii, Y. J. Bang, A. Madotto, and P. Fung · 2023
Closest in time.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
P. Liu, W. Yuan, J. Fu, Z. Jiang, H. Hayashi, and G. Neubig · 2023
Closest in time.
M. Turpin, J. Michael, E. Perez, and S. R. Bowman · 2023
Closest in time.
Agieval: A human-centric benchmark for evaluating foundation models
W. Zhong, R. Cui, Y. Guo, Y. Liang, S. Lu, Y. Wang, A. Saied, W. Chen, and N. Duan · 2023
Closest in time.