Fetching the paper…
Reading the bibliography…
This paper introduces ExpertLongBench, an expert-level benchmark containing 11 tasks from 9 domains that reflect realistic expert workflows and applications.
Cognitive processes in revision
John R Hayes, Linda Flower, Karen A Schriver, James Stratman, Linda Carey, et al · 1987
Earlier work this paper cites.
Smiles, a chemical language and information system. 1. introduction to methodology and encoding rules
David Weininger · 1988
Earlier work this paper cites.
An introduction to the bootstrap
Bradley Efron and Robert J Tibshirani · 1994
Earlier work this paper cites.
Critical thinking: How to prepare students for a rapidly changing world
Richard Paul · 1995
Earlier work this paper cites.
Scop: A structural classification of proteins database for the investigation of sequences and structures
Alexey G. Murzin, Steven E. Brenner, Tim Hubbard, and Cyrus Chothia · 1995
Earlier work this paper cites.
The effects of feedback interventions on performance: a historical review, a meta-analysis, and a preliminary feedback intervention theory
Avraham N Kluger and Angelo DeNisi · 1996
Earlier work this paper cites.
The influence of teacher commentary on student revision
Dana R Ferris · 1997
Earlier work this paper cites.
The metaphor of scaffolding: Its utility for the field of learning disabilities
C Addison Stone · 1998
Earlier work this paper cites.
The protein data bank
Helen M. Berman, John Westbrook, Zukang Feng, Gary Gilliland, T. N. Bhat, Helge Weissig, Ilya N. Shindyalov, and Philip E. Bourne · 2000
Earlier work this paper cites.
A revision of bloom’s taxonomy: An overview
David R Krathwohl · 2002
Earlier work this paper cites.
Teacher feedback, writing assignment quality, and third-grade students’ revision in lower-and higher-achieving urban schools
Lindsay Clare Matsumura, G Genevieve Patthey-Chavez, Rosa Valdés, and Helen Garnier · 2002
Earlier work this paper cites.
What we really value: Beyond rubrics in teaching and assessing writing
Bob Broad · 2003
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
The effect of different types of corrective feedback on esl student writing
John Bitchener, Stuart Young, and Denise Cameron · 2005
Earlier work this paper cites.
Standard deviations and standard errors
Douglas G Altman and J Martin Bland · 2005
Earlier work this paper cites.
The impact of teachers’ comment types on students’ revision
Yoshihito Sugita · 2006
Earlier work this paper cites.
The power of feedback
John Hattie and Helen Timperley · 2007
Earlier work this paper cites.
The cognitive process of decision making
Yingxu Wang and Guenther Ruhe · 2007
Earlier work this paper cites.
Active-constructive-interactive: A conceptual framework for differentiating learning activities
Michelene TH Chi · 2009
Earlier work this paper cites.
The rules of information aggregation and emergence of collective intelligent behavior
Luís M. A. Bettencourt · 2009
Earlier work this paper cites.
Mark my words: the role of assessment criteria in uk higher education grading practices
Sue Bloxham, Peter Boyd, and Susan Orr · 2011
Earlier work this paper cites.
The high-throughput highway to computational materials design
Stefano Curtarolo, Gus LW Hart, Marco Buongiorno Nardelli, Natalio Mingo, Stefano Sanvito, and Ohad Levy · 2013
Earlier work this paper cites.
Commentary: The materials project: A materials genome approach to accelerating materials innovation
Anubhav Jain, Shyue Ping Ong, Geoffroy Hautier, Wei Chen, William Davidson Richards, Stephen Dacek, Shreyas Cholia, Dan Gunter, David Skinner, Gerbrand Ceder, et al · 2013
Earlier work this paper cites.
Rubrics in the classroom: do teachers really follow them?
Heejeong Jeong · 2015
Earlier work this paper cites.
Direct vs. indirect written corrective feedback: Student perceptions
Anne Westmacott · 2017
Earlier work this paper cites.
Tethered to the ehr: Primary care physician workload assessment using ehr event log data and time-motion observations
Brian G Arndt, John W Beasley, Michael D Watkinson, Jonathan L Temte, Wen-Jan Tuan, Christine A Sinsky, and Valerie J Gilchrist · 2017
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi · 2019
Earlier work this paper cites.
Text-mined dataset of inorganic materials synthesis recipes
Olga Kononova, Haoyan Huo, Tanjin He, Ziqin Rong, Tiago Botari, Wenhao Sun, Vahe Tshitoyan, and Gerbrand Ceder · 2019
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2020
Earlier work this paper cites.
Hybrid ai for context understanding
Alessandro Oltramari, Cory Henson, Ruwan Wickramarachchi, Don Brutzman, and Richard Markeloff · 2020
Earlier work this paper cites.
Text2mol: Cross-modal molecule retrieval with natural language queries
Carl Edwards, ChengXiang Zhai, and Heng Ji · 2021
Cited alongside, same era.
Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences
Alexander Rives, Joshua Meier, Tom Sercu, Siddharth Goyal, Zeming Lin, Jason Liu, Dustin Guo, Myle Ott, C Lawrence Zitnick, Jerry Ma, et al · 2021
Cited alongside, same era.
The biogrid database: A comprehensive biomedical resource of curated protein, genetic, and chemical interactions
Rose Oughtred, Jenna Rust, Cecilia Chang, Bobby-Joe Breitkreutz, Chris Stark, Anjali Willems, Lorrie Boucher, Gabrielle Leung, Natalia Kolas, Frank Zhang, et al · 2021
Cited alongside, same era.
Multi-lexsum: Real-world summaries of civil rights lawsuits at multiple granularities
Zejiang Shen, Kyle Lo, Lauren Yu, Nathan Dahlberg, Margo Schlanger, and Doug Downey · 2022
Cited alongside, same era.
Self-consistency improves chain of thought reasoning in language models
Assessing gpt-4’s performance in delivering medical advice: comparative analysis with human experts
Eunbeen Jo, Sanghoun Song, Jong-Ho Kim, Subin Lim, Ju Hyeon Kim, Jung-Joon Cha, Young-Min Kim, Hyung Joon Joo, et al · 2024
Later among the works it cites.
Use of gpt-4 to diagnose complex clinical cases, 2024
Alexander V Eriksen, Sören Möller, and Jesper Ryg · 2024
Later among the works it cites.
Kola: Carefully benchmarking world knowledge of large language models
Jifan Yu, Xiaozhi Wang, Shangqing Tu, Shulin Cao, Daniel Zhang-Li, Xin Lv, Hao Peng, Zijun Yao, Xiaohan Zhang, Hanming Li, et al · 2024
Later among the works it cites.
Pedagogical alignment of large language models (llm) for personalized learning: a survey, trends and challenges
Mahefa Abel Razafinirina, William Germain Dimbisoa, and Thomas Mahatody · 2024
Later among the works it cites.
Inderjeet Nair, Jiaye Tan, Xiaotian Su, Anne Gere, Xu Wang, and Lu Wang · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou · 2022
Cited alongside, same era.
Translation between molecules and natural language, 2022
Carl Edwards, Tuan Lai, Kevin Ros, Garrett Honke, Kyunghyun Cho, and Heng Ji · 2022
Cited alongside, same era.
The art of problem solving and its translation into practice
J. Brooks · 2022
Cited alongside, same era.
Evaluating the feasibility of chatgpt in healthcare: an analysis of multiple clinical and research scenarios
Marco Cascella, Jonathan Montomoli, Valentina Bellini, and Elena Bignami · 2023
Cited alongside, same era.
Expertqa: Expert-curated questions and attributed answers
Chaitanya Malaviya, Subin Lee, Sihao Chen, Elizabeth Sieber, Mark Yatskar, and Dan Roth · 2023
Cited alongside, same era.
Agieval: A human-centric benchmark for evaluating foundation models
Wanjun Zhong, Ruixiang Cui, Yiduo Guo, Yaobo Liang, Shuai Lu, Yanlin Wang, Amin Saied, Weizhu Chen, and Nan Duan · 2023
Cited alongside, same era.
FActScore: Fine-grained atomic evaluation of factual precision in long form text generation
Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi · 2023
Cited alongside, same era.
WiCE: Real-world entailment for claims in Wikipedia
Ryo Kamoi, Tanya Goyal, Juan Diego Rodriguez, and Greg Durrett · 2023
Cited alongside, same era.
Later among the works it cites.
Sciknoweval: Evaluating multi-level scientific knowledge of large language models, 2024
Kehua Feng, Keyan Ding, Weijie Wang, Xiang Zhuang, Zeyuan Wang, Ming Qin, Yu Zhao, Jianhua Yao, Qiang Zhang, and Huajun Chen · 2024
Later among the works it cites.
Direct: Diagnostic reasoning for clinical notes via large language models
Bowen Wang, Jiuyang Chang, Yiming Qian, Guoxin Chen, Junhao Chen, Zhouqiang Jiang, Jiahao Zhang, Yuta Nakashima, and Hajime Nagahara · 2024
Later among the works it cites.
R-judge: Benchmarking safety risk awareness for llm agents
Tongxin Yuan, Zhiwei He, Lingzhong Dong, Yiming Wang, Ruijie Zhao, Tian Xia, Lizhen Xu, Binglin Zhou, Fangqi Li, Zhuosheng Zhang, Rui Wang, and Gongshen Liu · 2024
Later among the works it cites.
Agentharm: A benchmark for measuring harmfulness of llm agents
Maksym Andriushchenko, Alexandra Souly, Mateusz Dziemian, Derek Duenas, Maxwell Lin, Justin Wang, Dan Hendrycks, Andy Zou, Zico Kolter, Matt Fredrikson, Eric Winsor, Jerome Wynne, Yarin Gal, and Xander Davies · 2024
Later among the works it cites.
Factbench: A dynamic benchmark for in-the-wild language model factuality evaluation
Farima Fatahi Bayat, Lechen Zhang, Sheza Munir, and Lu Wang · 2024
Later among the works it cites.
An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al · 2024
Later among the works it cites.
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al · 2024
Later among the works it cites.
Workbench: a benchmark dataset for agents in a realistic workplace setting, 2024
Olly Styles, Sam Miller, Patricio Cerda-Mardini, Tanaya Guha, Victor Sanchez, and Bertie Vidgen · 2024
Later among the works it cites.
Evaluating large language models: Principles, approaches, and applications
Bo Li, Irina Sigler, and Yuan (Emily) Xue · 2024
Later among the works it cites.
Exploring llms applications in law: A literature review on current legal nlp approaches
Marco Siino, Mariana Falco, Daniele Croce, and Paolo Rosso · 2025
Closest in time.
Towards accurate differential diagnosis with large language models
Daniel McDuff, Mike Schaekermann, Tao Tu, Anil Palepu, Amy Wang, Jake Garrison, Karan Singhal, Yash Sharma, Shekoofeh Azizi, Kavita Kulkarni, Le Hou, Yong Cheng, Yun Liu, S. Sara Mahdavi, Sushant Prakash, Anupam Pathak, Christopher Semturs, Shwetak Patel, Dale R. Webster, Ewa Dominowska, Juraj Gottweis, Joelle Barral, Katherine Chou, Greg S. Corrado, Yossi Matias, Jake Sunshine, Alan Karthikesalingam, and Vivek Natarajan · 2025
Closest in time.
Can large language model analyze financial statements well?
Xinlin Wang and Mats Brorsson · 2025
Closest in time.
Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, Josephina Hu, Hugh Zhang, Chen Bo Calvin Zhang, Mohamed Shaaban, John Ling, Sean Shi, et al · 2025
Closest in time.
Dolomites: Domain-specific long-form methodical tasks
Chaitanya Malaviya, Priyanka Agrawal, Kuzman Ganchev, Pranesh Srinivasan, Fantine Huot, Jonathan Berant, Mark Yatskar, Dipanjan Das, Mirella Lapata, and Chris Alberti · 2025
Closest in time.
Li S Yifei, Allen Chang, Chaitanya Malaviya, and Mark Yatskar · 2025
Closest in time.
Decomposition dilemmas: Does claim decomposition boost or burden fact-checking performance?, 2025
Qisheng Hu, Quanyu Long, and Wenya Wang · 2025
Closest in time.
Rocketeval: Efficient automated llm evaluation via grading checklist
Tianjun Wei, Wei Wen, Ruizhi Qiao, Xing Sun, and Jianghong Ma · 2025
Closest in time.
Judge ai: Assessing large language models in judicial decision-making
Eric A Posner and Shivam Saran · 2025
Closest in time.
Can ai grade your essays? a comparative analysis of large language models and teacher ratings in multidimensional essay scoring
Kathrin Seßler, Maurice Fürstenberg, Babette Bühler, and Enkelejda Kasneci · 2025
Closest in time.
Statement of facts, n.d
LSD Law · 2025
Closest in time.
Infographic: The anatomy of a legal brief, 2021
TypeLaw · 2025
Closest in time.
A critical reflection on attempts to machine-learn materials synthesis insights from text-mined literature recipes
Wenhao Sun and Nicholas David · 2025
Closest in time.
Predicting long-term student outcomes from short-term edtech log data
Ge Gao, Amelia Leon, Andrea Jetten, Jasmine Turner, Husni Almoubayyed, Stephen Fancsali, and Emma Brunskill · 2025
Closest in time.
Soap notes, 2023
V. Podder, V. Lew, and S. Ghassemzadeh · 2025
Closest in time.
Commercial llm agents are already vulnerable to simple yet dangerous attacks, 2025
Ang Li, Yin Zhou, Vethavikashini Chithrra Raghuram, Tom Goldstein, and Micah Goldblum · 2025
Closest in time.