Fetching the paper…
Reading the bibliography…
From grading papers to summarizing medical documents, large language models (LLMs) are evermore used for evaluation of text generated by humans and AI alike.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Nils Reimers and Iryna Gurevych. 2019 · 1908
Earlier work this paper cites.
Cognitive maps in rats and men
Edward C Tolman. 1948 · 1948
Earlier work this paper cites.
A Coefficient of Agreement for Nominal Scales
J. Cohen. 1960 · 1960
Earlier work this paper cites.
Learning and reasoning by analogy
Patrick H Winston. 1980 · 1980
Earlier work this paper cites.
Stochastic neighbor embedding
Geoffrey E Hinton and Sam Roweis. 2002 · 2002
Earlier work this paper cites.
A package for automatic evaluation of summaries. In Proceedings of Workshop on Text Summarization of ACL, Spain , Vol. 5
Lin CY ROUGE. 2004 · 2004
Earlier work this paper cites.
Hippocampal contributions to control: the third way
Máté Lengyel and Peter Dayan. 2007 · 2007
Earlier work this paper cites.
Neural representations of events arise from temporal community structure
Anna C Schapiro, Timothy T Rogers, Natalia I Cordova, Nicholas B Turk-Browne, and Matthew M Botvinick. 2013 · 2013
Earlier work this paper cites.
Universal and Transferable Adversarial Attacks on Aligned Language Models
Andy Zou, Zifan Wang, J. Zico Kolter, and Matt Fredrikson. 2023 · 2013
Earlier work this paper cites.
Logically-Constrained Neural Fitted Q-Iteration. In Proceedings of the 18th International Conference on Autonomous Agents and MultiAgent Systems . International Foundation for Autonomous Agents and Multiagent Systems, 2012–2014
Hosein Hasanbeig, Alessandro Abate, and Daniel Kroening. 2019a · 2014
Earlier work this paper cites.
Teaching machines to read and comprehend
Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015 · 2015
Earlier work this paper cites.
Charles Blundell, Benigno Uria, Alexander Pritzel, Yazhe Li, Avraham Ruderman, Joel Z Leibo, Jack Rae, Daan Wierstra, and Demis Hassabis. 2016 · 2016
Earlier work this paper cites.
Reinforcement Learning and Episodic Memory in Humans and Animals: An Integrative Framework
Samuel J Gershman and Nathaniel D Daw. 2017 · 2017
Earlier work this paper cites.
The successor representation in human reinforcement learning
Ida Momennejad, Evan M Russek, Jin H Cheong, Matthew M Botvinick, Nathaniel Douglass Daw, and Samuel J Gershman. 2017 · 2017
Earlier work this paper cites.
What is a cognitive map? Organizing knowledge for flexible behavior
Timothy EJ Behrens, Timothy H Muller, James CR Whittington, Shirley Mark, Alon B Baram, Kimberly L Stachenfeld, and Zeb Kurth-Nelson. 2018 · 2018
Earlier work this paper cites.
Logically-constrained reinforcement learning
Hosein Hasanbeig, Alessandro Abate, and Daniel Kroening. 2018 · 2018
Earlier work this paper cites.
Improving abstraction in text summarization
Wojciech Kryściński, Romain Paulus, Caiming Xiong, and Richard Socher. 2018 · 2018
Earlier work this paper cites.
Offline replay supports planning in human reinforcement learning
Ida Momennejad, A Ross Otto, Nathaniel D Daw, and Kenneth A Norman. 2018 · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto. 2018 · 2018
Earlier work this paper cites.
Bridge ties bind collective memories
Ida Momennejad, Ajua Duker, and Alin Coman. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019 · 2019
Cited alongside, same era.
SummEval: Re-evaluating Summarization Evaluation
Alexander R Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2020 · 2020
Cited alongside, same era.
Safe and certified reinforcement learning with logical constraints
Hosein Hasanbeig. 2020 · 2020
Cited alongside, same era.
Cautious reinforcement learning with logical constraints
Hosein Hasanbeig, Alessandro Abate, and Daniel Kroening. 2020 · 2020
OpenAI 2023. 2023 · 2023
Closest in time.
Introducing Claude
Anthropic. 2023 · 2023
Closest in time.
CogEval Prompts and Responses
CogMaps. 2023 · 2023
Closest in time.
Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers
Damai Dai, Yutao Sun, Li Dong, Yaru Hao, Shuming Ma, Zhifang Sui, and Furu Wei. 2023 · 2023
Closest in time.
How Ready are Pre-trained Abstractive Models and LLMs for Legal Case Judgement Summarization?
Aniket Deroy, Kripabandhu Ghosh, and Saptarshi Ghosh. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning structures: predictive representations, replay, and generalization
Ida Momennejad. 2020 · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Cited alongside, same era.
Pegasus: Pre-training with extracted gap-sentences for abstractive summarization. In International Conference on Machine Learning . PMLR, 11328–11339
Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter Liu. 2020 · 2020
Cited alongside, same era.
DeepSynth: Automata synthesis for automatic task segmentation in deep reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 35. 7647–7656
Hosein Hasanbeig, Natasha Yogananda Jeppu, Alessandro Abate, Tom Melham, and Daniel Kroening. 2021 · 2021
Cited alongside, same era.
What Makes Good In-Context Examples for GPT- 3 3 ?
Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, and Weizhu Chen. 2021 · 2021
Cited alongside, same era.
Bartscore: Evaluating generated text as text generation
Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021 · 2021
Cited alongside, same era.
Calibrate before use: Improving few-shot performance of language models. In International Conference on Machine Learning . PMLR, 12697–12706
Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021 · 2021
Cited alongside, same era.
Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Zhiyong Wu, Baobao Chang, Xu Sun, Jingjing Xu, Lei Li, and Zhifang Sui. 2023 · 2023
Closest in time.
Gptscore: Evaluate as you desire
Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu. 2023 · 2023
Closest in time.
Certified Reinforcement Learning with Logic Guidance
Hosein Hasanbeig, Daniel Kroening, and Alessandro Abate. 2023a · 2023
Closest in time.
Symbolic Task Inference in Deep Reinforcement Learning
Hosein Hasanbeig, Natasha Yogananda Jeppu, Alessandro Abate, Tom Melham, and Daniel Kroening. 2023b · 2023
Closest in time.
Causal Reasoning and Large Language Models: Opening a New Frontier for Causality
Emre Kiciman, Robert Ness, Amit Sharma, and Chenhao Tan. 2023 · 2023
Closest in time.
PRD: Peer Rank and Discussion Improve Large Language Model based Evaluations
Ruosen Li, Teerth Patel, and Xinya Du. 2023 · 2023
Closest in time.
Gpteval: Nlg evaluation using gpt-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023b · 2023
Closest in time.
Evaluating Cognitive Maps in Large Language Models: No Emergent Planning
Ida Momennejad, Hosein Hasanbeig, Felipe Frujeri, Hiteshi Sharma, Robert Ness, Nebojsa Jojic, Hamid Palangi, and Jonathan Larson. 2023 · 2023
Closest in time.
Large language models encode clinical knowledge
Karan Singhal, Shekoofeh Azizi, Tao Tu, S. Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, Perry Payne, Martin Seneviratne, Paul Gamble, Chris Kelly, Abubakr Babiker, Nathanael Schärli, Aakanksha Chowdhery, Philip Mansfield, Dina Demner-Fushman, Blaise Agüera y Arcas, Dale Webster, Greg S. Corrado, Yossi Matias, Katherine Chou, Juraj Gottweis, Nenad Tomasev, Yun Liu, Alvin Rajkomar, Joelle Barral, Christopher Semturs, Alan Karthikesalingam, and Vivek Natarajan. 2023 · 2023
Closest in time.
Stanford Alpaca: An Instruction-following LLaMA model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023 · 2023
Closest in time.
LLaMA: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.
Miles Turpin, Julian Michael, Ethan Perez, and Samuel R Bowman. 2023 · 2023
Closest in time.
Practical and ethical challenges of large language models in education: A systematic scoping review
Lixiang Yan, Lele Sha, Linxuan Zhao, Yuheng Li, Roberto Martinez-Maldonado, Guanliang Chen, Xinyu Li, Yueqiao Jin, and Dragan Gašević. 2023 · 2023
Closest in time.
A Survey of Large Language Models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie, and Ji-Rong Wen. 2023 · 2023
Closest in time.
Judging LLM-as-a-judge with MT-Bench and Chatbot Arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2023
Closest in time.
Red teaming ChatGPT via Jailbreaking: Bias, Robustness, Reliability and Toxicity
Terry Yue Zhuo, Yujin Huang, Chunyang Chen, and Zhenchang Xing. 2023 · 2023
Closest in time.