Fetching the paper…
Reading the bibliography…
The text produced by language models (LMs) can exhibit specific `behaviors,' such as a failure to follow alignment training, that we hope to detect and react to during deployment.
The concept of exchangeability and its applications
J. M. Bernardo · 1996
Earlier work this paper cites.
Practical solutions to the problem of diagonal dominance in kernel document clustering
D. Greene and P. Cunningham · 2006
Earlier work this paper cites.
The new york times annotated corpus
E. Sandhaus · 2008
Earlier work this paper cites.
A tutorial on conformal prediction
G. Shafer and V. Vovk · 2008
Earlier work this paper cites.
Learning word vectors for sentiment analysis
A. L. Maas, R. E. Daly, P. T. Pham, D. Huang, A. Y. Ng, and C. Potts · 2011
Earlier work this paper cites.
A. Gammerman, V. Vovk, and V. Vapnik · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
R. Socher, A. Perelygin, J. Wu, J. Chuang, C. D. Manning, A. Y. Ng, and C. Potts · 2013
Earlier work this paper cites.
Good debt or bad debt: Detecting semantic orientations in economic texts
P. Malo, A. Sinha, P. Korhonen, J. Wallenius, and P. Takala · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Earlier work this paper cites.
Yelp dataset challenge: Review rating prediction
N. Asghar · 2016
Earlier work this paper cites.
MS MARCO: A human generated machine reading comprehension dataset
T. Nguyen, M. Rosenberg, X. Song, J. Gao, S. Tiwary, R. Majumder, and L. Deng · 2016
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes, 2017
G. Alain and Y. Bengio · 2017
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
M. Joshi, E. Choi, D. Weld, and L. Zettlemoyer · 2017
Earlier work this paper cites.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)
B. Kim, M. Wattenberg, J. Gilmer, C. J. Cai, J. Wexler, F. B. Viégas, and R. Sayres · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
T. Mihaylov, P. Clark, T. Khot, and A. Sabharwal · 2018
Earlier work this paper cites.
CARER: Contextualized affect representations for emotion recognition
E. Saravia, H.-C. T. Liu, Y.-H. Huang, J. Wu, and Y.-S. Chen · 2018
Earlier work this paper cites.
Fever: a large-scale dataset for fact extraction and verification
J. Thorne, A. Vlachos, C. Christodoulopoulos, and A. Mittal · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
Designing and interpreting probes with control tasks
J. Hewitt and P. Liang · 2019
Earlier work this paper cites.
Cosmos QA: Machine reading comprehension with contextual commonsense reasoning
L. Huang, R. Le Bras, C. Bhagavatula, and Y. Choi · 2019
Cited alongside, same era.
Natural questions: A benchmark for question answering research
T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, J. Devlin, K. Lee, K. Toutanova, L. Jones, M. Kelcey, M.-W. Chang, A. M. Dai, J. Uszkoreit, Q. Le, and S. Petrov · 2019
Cited alongside, same era.
Language models as knowledge bases?
F. Petroni, T. Rocktäschel, S. Riedel, P. Lewis, A. Bakhtin, Y. Wu, and A. Miller · 2019
Cited alongside, same era.
CommonsenseQA: A question answering challenge targeting commonsense knowledge
A. Talmor, J. Herzig, N. Lourie, and J. Berant · 2019
Cited alongside, same era.
Piqa: Reasoning about physical commonsense in natural language
Y. Bisk, R. Zellers, R. L. Bras, J. Gao, and Y. Choi · 2020
Cited alongside, same era.
Language models are few-shot learners
What learning algorithm is in-context learning? investigations with linear models
E. Akyürek, D. Schuurmans, J. Andreas, T. Ma, and D. Zhou · 2023
Later among the works it cites.
Knowledge of knowledge: Exploring known-unknowns uncertainty with large language models
A. Amayuelas, L. Pan, W. Chen, and W. Wang · 2023
Later among the works it cites.
Does it know?: Probing for uncertainty in language model latent beliefs
Anonymous · 2023
Later among the works it cites.
The internal state of an LLM knows when it’s lying
A. Azaria and T. Mitchell · 2023
Later among the works it cites.
Palm: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Cited alongside, same era.
Climate-fever: A dataset for verification of real-world climate claims, 2020
T. Diggelmann, J. Boyd-Graber, J. Bulian, M. Ciaramita, and M. Leippold · 2020
Cited alongside, same era.
Qasc: A dataset for question answering via sentence composition
T. Khot, P. Clark, M. Guerquin, P. Jansen, and A. Sabharwal · 2020
Cited alongside, same era.
DeeBERT: Dynamic early exiting for accelerating BERT inference
J. Xin, R. Tang, J. Lee, Y. Yu, and J. Lin · 2020
Cited alongside, same era.
Bert loses patience: Fast and robust inference with early exit
W. Zhou, C. Xu, T. Ge, J. McAuley, K. Xu, and F. Wei · 2020
Cited alongside, same era.
Newsmtsc: (multi-)target-dependent sentiment classification in news articles
F. Hamborg and K. Donnay · 2021
Cited alongside, same era.
How can we know when language models know? on the calibration of language models for question answering
Z. Jiang, J. Araki, H. Ding, and G. Neubig · 2021
Cited alongside, same era.
M. Jazbec, J. U. Allingham, D. Zhang, and E. Nalisnick · 2023
Later among the works it cites.
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. d. l. Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, et al · 2023
Later among the works it cites.
Conformal prediction with large language models for multi-choice question answering
B. Kumar, C. Lu, G. Gupta, A. Palepu, D. Bellamy, R. Raskar, and A. Beam · 2023
Later among the works it cites.
Inference-time intervention: Eliciting truthful answers from a language model
K. Li, O. Patel, F. Viégas, H. Pfister, and M. Wattenberg · 2023
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
S. Yao, D. Yu, J. Zhao, I. Shafran, T. L. Griffiths, Y. Cao, and K. R. Narasimhan · 2023
Later among the works it cites.
Do large language models know what they don’t know?
Z. Yin, Q. Sun, Q. Guo, J. Wu, X. Qiu, and X. Huang · 2023
Later among the works it cites.
Lookback lens: Detecting and mitigating contextual hallucinations in large language models using only attention maps
Y.-S. Chuang, L. Qiu, C.-Y. Hsieh, R. Krishna, Y. Kim, and J. R. Glass · 2024
Later among the works it cites.
A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Yang, A. Fan, et al · 2024
Later among the works it cites.
On large language models’ hallucination with regard to known facts
C. Jiang, B. Qi, X. Hong, D. Fu, Y. Cheng, F. Meng, M. Yu, B. Zhou, and J. Zhou · 2024
Later among the works it cites.
On the universal truthfulness hyperplane inside LLMs
J. Liu, S. Chen, Y. Cheng, and J. He · 2024
Later among the works it cites.
Probing language models for pre-training data detection
Z. Liu, T. Zhu, C. Tan, B. Liu, H. Lu, and W. Chen · 2024
Later among the works it cites.
Time is encoded in the weights of finetuned language models
K. Nylund, S. Gururangan, and N. Smith · 2024
Later among the works it cites.
Unsupervised real-time hallucination detection based on the internal states of large language models
W. Su, C. Wang, Q. Ai, Y. Hu, Z. Wu, Y. Zhou, and Y. Liu · 2024
Later among the works it cites.
Probing language models on their knowledge source
Z. Tighidet, J. Mei, B. Piwowarski, and P. Gallinari · 2024
Later among the works it cites.
Attention satisfies: A constraint-satisfaction lens on factual errors of language models
M. Yuksekgonul, V. Chandrasekaran, E. Jones, S. Gunasekar, R. Naik, H. Palangi, E. Kamar, and B. Nushi · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al · 2025
Closest in time.