Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have revolutionized natural language processing and demonstrated impressive capabilities in various tasks.
Icicle plots: Better displays for hierarchical clustering
J. B. Kruskal and J. M. Landwehr · 1983
Earlier work this paper cites.
Educational assessment of students
A. J. Nitko · 1996
Earlier work this paper cites.
A taxonomy for learning, teaching, and assessing: A revision of Bloom’s taxonomy of educational objectives, abridged edition
L. W. Anderson, D. R. Krathwohl, P. W. Airasian, K. A. Cruikshank, R. E. Mayer, P. R. Pintrich, J. Raths, and M. C. Wittrock · 2000
Earlier work this paper cites.
Cumulated gain-based evaluation of ir techniques
K. Järvelin and J. Kekäläinen · 2002
Earlier work this paper cites.
Graphical knowledge display–mind mapping and concept mapping as efficient tools in mathematics education
A. Brinkmann · 2003
Earlier work this paper cites.
Learning from architects: the difference between knowledge visualization and information visualization
R. A. Burkhard · 2004
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
The effect of concept mapping on students’ learning achievements and interests
C.-C. Chiou · 2008
Earlier work this paper cites.
Identifying optimal prognostic parameters from data: a genetic algorithms approach
J. Coble and J. W. Hines · 2009
Earlier work this paper cites.
Does the mind map learning strategy facilitate information retrieval and critical thinking in medical students?
A. V. D’Antoni, G. P. Zipp, V. G. Olson, and T. F. Cahill · 2010
Earlier work this paper cites.
Mind mapping as a teaching resource
S. Edwards and N. Cooper · 2010
Earlier work this paper cites.
Treevis. net: A tree visualization reference
H.-J. Schulz · 2011
Earlier work this paper cites.
Visualizing and understanding neural models in NLP
J. Li, X. Chen, E. Hovy, and D. Jurafsky · 2015
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang · 2016
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
M. Joshi, E. Choi, D. S. Weld, and L. Zettlemoyer · 2017
Earlier work this paper cites.
Crowdsourcing multiple choice science questions
J. Welbl, N. F. Liu, and M. Gardner · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
T. Mihaylov, P. Clark, T. Khot, and A. Sabharwal · 2018
Earlier work this paper cites.
Face to face: Evaluating visual comparison
B. Ondov, N. Jardine, N. Elmqvist, and S. Franconeri · 2018
Earlier work this paper cites.
Case study on effective use of mind map in engineering education
R. T. Selvi and G. Chandramohan · 2018
Earlier work this paper cites.
Assessing student learning: A common sense guide
L. Suskie · 2018
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman · 2018
Cited alongside, same era.
Swag: A large-scale adversarial dataset for grounded commonsense inference
R. Zellers, Y. Bisk, R. Schwartz, and Y. Choi · 2018
Cited alongside, same era.
PubMedQA: A dataset for biomedical research question answering
Q. Jin, B. Dhingra, Z. Liu, W. Cohen, and X. Lu · 2019
Cited alongside, same era.
Natural questions: A benchmark for question answering research
T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, J. Devlin, K. Lee, K. Toutanova, L. Jones, M. Kelcey, M.-W. Chang, A. M. Dai, J. Uszkoreit, Q. Le, and S. Petrov · 2019
Cited alongside, same era.
Model cards for model reporting
Holistic evaluation of language models, 2022
P. Liang, R. Bommasani, T. Lee, D. Tsipras, D. Soylu, M. Yasunaga, Y. Zhang, D. Narayanan, Y. Wu, A. Kumar, B. Newman, B. Yuan, B. Yan, C. Zhang, C. Cosgrove, C. D. Manning, C. Ré, D. Acosta-Navas, D. A. Hudson, E. Zelikman, E. Durmus, F. Ladhak, F. Rong, H. Ren, H. Yao, J. Wang, K. Santhanam, L. Orr, L. Zheng, M. Yuksekgonul, M. Suzgun, N. Kim, N. Guha, N. Chatterji, O. Khattab, P. Henderson, Q. Huang, R. Chi, S. M. Xie, S. Santurkar, S. Ganguli, T. Hashimoto, T. Icard, T. Zhang, V. Chaudhary, W. Wang, X. Li, Y. Mai, Y. Zhang, and Y. Koreeda · 2022
Later among the works it cites.
TruthfulQA: Measuring how models mimic human falsehoods
S. Lin, J. Hilton, and O. Evans · 2022
Later among the works it cites.
Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering
A. Pal, L. K. Umapathi, and M. Sankarasubbu · 2022
Later among the works it cites.
BLOOM: A 176b-parameter open-access multilingual language model
T. L. Scao, A. Fan, C. Akiki, E. Pavlick, S. Ilić, D. Hesslow, R. Castagné, A. S. Luccioni, F. Yvon, M. Gallé, et al · 2022
Later among the works it cites.
Visual comparison of language model adaptation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Mitchell, S. Wu, A. Zaldivar, P. Barnes, L. Vasserman, B. Hutchinson, E. Spitzer, I. D. Raji, and T. Gebru · 2019
Cited alongside, same era.
Language models as knowledge bases?
F. Petroni, T. Rocktäschel, S. Riedel, P. Lewis, A. Bakhtin, Y. Wu, and A. Miller · 2019
Cited alongside, same era.
Hellaswag: Can a machine really finish your sentence?
R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi · 2019
Cited alongside, same era.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2020
Cited alongside, same era.
Closing the ai accountability gap: Defining an end-to-end framework for internal algorithmic auditing
I. D. Raji, A. Smart, R. N. White, M. Mitchell, T. Gebru, B. Hutchinson, J. Smith-Loud, D. Theron, and P. Barnes · 2020
Cited alongside, same era.
On the dangers of stochastic parrots: Can language models be too big?
E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell · 2021
Cited alongside, same era.
R. Sevastjanova, E. Cakmak, S. Ravfogel, R. Cotterell, and M. El-Assady · 2022
Later among the works it cites.
LMFingerprints: Visual explanations of language model embedding spaces through layerwise contextualization scores
R. Sevastjanova, A. Kalouli, C. Beck, H. Hauptmann, and M. El-Assady · 2022
Later among the works it cites.
Self-Instruct: Aligning language model with self generated instructions
Y. Wang, Y. Kordi, S. Mishra, A. Liu, N. A. Smith, D. Khashabi, and H. Hajishirzi · 2022
Later among the works it cites.
Taxonomy of risks posed by language models
L. Weidinger, J. Uesato, M. Rauh, C. Griffin, P.-S. Huang, J. Mellor, A. Glaese, M. Cheng, B. Balle, A. Kasirzadeh, et al · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V. Lin, et al · 2022
Later among the works it cites.
Sparks of artificial general intelligence: Early experiments with GPT-4, 2023
S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. Lundberg, H. Nori, H. Palangi, M. T. Ribeiro, and Y. Zhang · 2023
Closest in time.
ChatGPT goes to law school
J. H. Choi, K. E. Hickman, A. Monahan, and D. Schwarcz · 2023
Closest in time.
Survey of hallucination in natural language generation
Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y. Xu, E. Ishii, Y. J. Bang, A. Madotto, and P. Fung · 2023
Closest in time.
Performance of ChatGPT on USMLE: Potential for ai-assisted medical education using large language models
T. H. Kung, M. Cheatham, A. Medenilla, C. Sillos, L. De Leon, C. Elepaño, M. Madriaga, R. Aggabao, G. Diaz-Candido, J. Maningo, et al · 2023
Closest in time.
G-eval: Nlg evaluation using gpt-4 with better human alignment, may 2023
Y. Liu, D. Iter, Y. Xu, S. Wang, R. Xu, and C. Zhu · 2023
Closest in time.
Dissociating language and thought in large language models: a cognitive perspective
K. Mahowald, A. A. Ivanova, I. A. Blank, N. Kanwisher, J. B. Tenenbaum, and E. Fedorenko · 2023
Closest in time.
US Bar - MBE
National Conference of Bar Examiners · 2023
Closest in time.
Capabilities of GPT-4 on medical challenge problems
H. Nori, N. King, S. M. McKinney, D. Carignan, and E. Horvitz · 2023
Closest in time.
GPT-4 technical report, 2023
OpenAI · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto · 2023
Closest in time.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Closest in time.