Fetching the paper…
Reading the bibliography…
Self-evaluation using large language models (LLMs) has proven valuable not only in benchmarking but also methods like reward modeling, constitutional AI, and self-refinement.
A new readability yardstick
Flesch, R · 1939
Earlier work this paper cites.
Abstractive Text Summarization using Sequence-to-sequence RNNs and Beyond
Nallapati, R., Zhou, B., dos Santos, C., Gulcehre, C., and Xiang, B · 2016
Earlier work this paper cites.
Scalable agent alignment via reward modeling: a research direction
Leike, J., Krueger, D., Everitt, T., Martic, M., Maini, V., and Legg, S · 2018
Earlier work this paper cites.
Narayan, S., Cohen, S. B., and Lapata, M · 2018
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Automatic detection of machine generated text: A critical survey
Jawahar, G., Abdul-Mageed, M., and Lakshmanan, L. V · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F · 2020
Earlier work this paper cites.
Recursively summarizing books with human feedback
Wu, J., Ouyang, L., Ziegler, D. M., Stiennon, N., Lowe, R., Leike, J., and Christiano, P · 2021
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., et al · 2022
Earlier work this paper cites.
Language models (mostly) know what they know
Kadavath, S., Conerly, T., Askell, A., Henighan, T., Drain, D., Perez, E., Schiefer, N., Hatfield-Dodds, Z., DasSarma, N., Tran-Johnson, E., et al · 2022
Earlier work this paper cites.
Self-critiquing models for assisting human evaluators
Saunders, W., Yeh, C., Wu, J., Bills, S., Ouyang, L., Ward, J., and Leike, J · 2022
Earlier work this paper cites.
Knowledge of knowledge: Exploring known-unknowns uncertainty with large language models
Amayuelas, A., Pan, L., Chen, W., and Wang, W · 2023
Earlier work this paper cites.
Taken out of context: On measuring situational awareness in llms
Berglund, L., Stickland, A. C., Balesni, M., Kaufmann, M., Tong, M., Korbak, T., Kokotajlo, D., and Evans, O · 2023
Earlier work this paper cites.
Machine-generated text: A comprehensive survey of threat models and detection methods
Crothers, E., Japkowicz, N., and Viktor, H. L · 2023
Earlier work this paper cites.
GPTScore: Evaluate as You Desire, February 2023
Fu, J., Ng, S.-K., Jiang, Z., and Liu, P · 2023
Earlier work this paper cites.
Is GPT-4 a reliable rater? Evaluating Consistency in GPT-4 Text Ratings
Hackl, V., Müller, A. E., Granitzer, M., and Sailer, M · 2023
Cited alongside, same era.
TuringMirror: Evaluating the ability of LLMs to recognize LLM-generated text, August 2023
Hoelscher-Obermaier, J., Lutz, M. J., Feuillade-Montixi, and Modak, S · 2023
Cited alongside, same era.
Benchmarking cognitive biases in large language models as evaluators
Koo, R., Lee, M., Raheja, V., Park, J. I., Kim, Z. M., and Kang, D · 2023
Cited alongside, same era.
Towards a situational awareness benchmark for llms
Laine, R., Meinke, A., and Evans, O · 2023
Cited alongside, same era.
RLAIF: Scaling reinforcement learning from human feedback with ai feedback
Lee, H., Phatale, S., Mansoor, H., Lu, K., Mesnard, T., Bishop, C., Carbune, V., and Rastogi, A · 2023
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
A survey on llm-gernerated text detection: Necessity, methods, and future directions
Wu, J., Yang, S., Zhan, R., Yuan, Y., Wong, D. F., and Chao, L. S · 2023
Later among the works it cites.
A survey on detection of llms-generated content
Yang, X., Pan, L., Zhao, X., Chen, H., Petzold, L., Wang, W. Y., and Cheng, W · 2023
Later among the works it cites.
Do large language models know what they don’t know?
Yin, Z., Sun, Q., Guo, Q., Wu, J., Qiu, X., and Huang, X · 2023
Later among the works it cites.
Evaluating Instruction-Tuned Large Language Models on Code Comprehension and Generation, August 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
AlpacaEval: An Automatic Evaluator of Instruction-following Models, February 2024
Li, X., Zhang, T., Dubois, Y., Taori, R., Gulrajani, I., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Cited alongside, same era.
LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores, November 2023
Liu, Y., Moosavi, N. S., and Lin, C · 2023
Cited alongside, same era.
Self-Refine: Iterative Refinement with Self-Feedback, May 2023
Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., Gupta, S., Majumder, B. P., Hermann, K., Welleck, S., Yazdanbakhsh, A., and Clark, P · 2023
Cited alongside, same era.
Detectgpt: Zero-shot machine-generated text detection using probability curvature
Mitchell, E., Lee, Y., Khazatsky, A., Manning, C. D., and Finn, C · 2023
Cited alongside, same era.
OpenAI · 2023
Cited alongside, same era.
Towards Evaluating AI Systems for Moral Status Using Self-Reports, November 2023
Perez, E. and Long, R · 2023
Cited alongside, same era.
Large Language Models Sensitivity to The Order of Options in Multiple-Choice Questions, August 2023
Pezeshkpour, P. and Hruschka, E · 2023
Cited alongside, same era.
Yuan, Z., Liu, J., Zi, Q., Liu, M., Peng, X., and Lou, Y · 2023
Later among the works it cites.
Evaluating Large Language Models at Evaluating Instruction Following, October 2023
Zeng, Z., Yu, J., Gao, T., Meng, Y., Goyal, T., and Chen, D · 2023
Later among the works it cites.
Benchmarking foundation models with language-model-as-an-examiner
Bai, Y., Ying, J., Cao, Y., Lv, X., He, Y., Wang, X., Yu, J., Zeng, K., Xiao, Y., Lyu, H., et al · 2024
Closest in time.
Spotting llms with binoculars: Zero-shot detection of machine-generated text
Hans, A., Schwarzschild, A., Cherepanova, V., Kazemi, H., Saha, A., Goldblum, M., Geiping, J., and Goldstein, T · 2024
Closest in time.
A survey of ai-generated text forensic systems: Detection, attribution, and characterization
Kumarage, T., Agrawal, G., Sheth, P., Moraffah, R., Chadha, A., Garland, J., and Liu, H · 2024
Closest in time.
Feedback loops with language models drive in-context reward hacking
Pan, A., Jones, E., Jagadeesan, M., and Steinhardt, J · 2024
Closest in time.
Is llm-as-a-judge robust? investigating universal adversarial attacks on zero-shot llm assessment
Raina, V., Liusie, A., and Gales, M · 2024
Closest in time.
Wang, Y., Liao, Y., Liu, H., Liu, H., Wang, Y., and Wang, Y · 2024
Closest in time.
Perils of self-feedback: Self-bias amplifies in large language models
Xu, W., Zhu, G., Zhao, X., Pan, L., Li, L., and Wang, W. Y · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al · 2024
Closest in time.