Fetching the paper…
Reading the bibliography…
As the capabilities of Large Language Models (LLMs) in healthcare and medicine continue to advance, there is a growing need for competitive open-source models that can safeguard public interest.
Searching web data using minhash lsh
B. Rao and E. Zhu · 2016
Earlier work this paper cites.
CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models
N. N. et al · 2020
Earlier work this paper cites.
Distill and replay for continual language learning
J. Sun, S. Wang, J. Zhang, and C. Zong · 2020
Earlier work this paper cites.
Aligning ai with shared human values
D. H. et al · 2021
Earlier work this paper cites.
Discovering language model behaviors with model-written evaluations
E. P. et al · 2022
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
D. Ganguli, L. Lovitt, J. Kernion, A. Askell, Y. Bai, S. Kadavath, B. Mann, E. Perez, N. Schiefer, K. Ndousse, et al · 2022
Earlier work this paper cites.
ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection
T. Hartvigsen, S. Gabriel, H. Palangi, et al · 2022
Earlier work this paper cites.
TruthfulQA: Measuring How Models Mimic Human Falsehoods
S. Lin, J. Hilton, and O. Evans · 2022
Earlier work this paper cites.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
M. Wortsman, G. Ilharco, S. Y. Gadre, et al · 2022
Earlier work this paper cites.
Meditron-70b: Scaling medical pretraining for large language models
Z. Chen, A. H. Cano, A. Romanou, et al · 2023
Earlier work this paper cites.
Deep reinforcement learning from human preferences, 2023
P. Christiano, J. Leike, T. B. Brown, et al · 2023
Earlier work this paper cites.
Enhancing chat language models by scaling high-quality instructional conversations
N. Ding, Y. Chen, B. Xu, Y. Qin, Z. Zheng, S. Hu, Z. Liu, M. Sun, and B. Zhou · 2023
Earlier work this paper cites.
Neftune: Noisy embeddings improve instruction finetuning
N. J. et al · 2023
Earlier work this paper cites.
A framework for few-shot language model evaluation, 12 2023
L. Gao, J. Tow, B. Abbasi, et al · 2023
Earlier work this paper cites.
Chatgpt outperforms crowd workers for text-annotation tasks
F. Gilardi, M. Alizadeh, and M. Kubli · 2023
Earlier work this paper cites.
Large language model ai chatbots require approval as medical devices
S. Gilbert, H. Harvey, T. Melvin, et al · 2023
Earlier work this paper cites.
K. He, R. Mao, Q. Lin, et al · 2023
Earlier work this paper cites.
Medcpt: Contrastive pre-trained transformers with large-scale pubmed search logs for zero-shot biomedical information retrieval
Q. Jin, W. Kim, Q. Chen, D. C. Comeau, L. Yeganova, W. J. Wilbur, and Z. Lu · 2023
Cited alongside, same era.
Two directions for clinical data generation with large language models: Data-to-label and label-to-data
R. Li, X. Wang, and H. Yu · 2023
Cited alongside, same era.
Angle-optimized text embeddings
X. Li and J. Li · 2023
Cited alongside, same era.
W. Liu, W. Zeng, K. He, et al · 2023
Cited alongside, same era.
Embeddings for medical literature
D. Mezzetti · 2023
Cited alongside, same era.
Language models are super mario: Absorbing abilities from homologous models as a free lunch
L. Yu, B. Yu, H. Yu, F. Huang, and Y. Li · 2023
Later among the works it cites.
preprint arxiv , 2024
Benchmarking open healthcare llms at scale · 2024
Closest in time.
Z. Cai, M. Cao, H. Chen, K. Chen, K. Chen, X. Chen, X. Chen, et al · 2024
Closest in time.
Red teaming gpt-4v: Are gpt-4v safe against uni/multi-modal jailbreak attacks?
S. Chen, Z. Han, B. He, Z. Ding, W. Yu, P. Torr, V. Tresp, and J. Gu · 2024
Closest in time.
How abilities in large language models are affected by supervised fine-tuning data composition, 2024
G. Dong, H. Yuan, K. Lu, C. Li, M. Xue, D. Liu, W. Wang, Z. Yuan, C. Zhou, and J. Zhou · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Orca: Progressive learning from complex explanation traces of gpt-4, 2023
S. Mukherjee, A. Mitra, G. Jawahar, S. Agarwal, H. Palangi, and A. Awadallah · 2023
Cited alongside, same era.
Can generalist foundation models outcompete special-purpose tuning? case study in medicine, 2023
H. Nori, Y. T. Lee, S. Zhang, D. Carignan, R. Edgar, N. Fusi, N. King, J. Larson, Y. Li, W. Liu, R. Luo, S. M. McKinney, R. O. Ness, H. Poon, T. Qin, N. Usuyama, C. White, and E. Horvitz · 2023
Cited alongside, same era.
A study of generative large language model for medical research and healthcare
C. Peng, X. Yang, A. Chen, K. E. Smith, N. PourNejatian, A. B. Costa, C. Martin, M. G. Flores, Y. Zhang, T. Magoc, et al · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model, 2023
R. Rafailov, A. Sharma, E. Mitchell, et al · 2023
Cited alongside, same era.
Does synthetic data generation of llms help clinical text mining?
R. Tang, X. Han, X. Jiang, and X. Hu · 2023
Cited alongside, same era.
Zephyr: Direct distillation of lm alignment, 2023
L. Tunstall, E. Beeching, N. Lambert, et al · 2023
Cited alongside, same era.
Med-halt: Medical domain hallucination test for large language models
L. K. Umapathi, A. Pal, and M. Sankarasubbu · 2023
Cited alongside, same era.
Arcee’s mergekit: A toolkit for merging large language models
C. Goddard, S. Siriwardhana, M. Ehghaghi, et al · 2024
Closest in time.
Risks from language models for automated mental healthcare: Ethics and structure for implementation
D. Grabb, M. Lamparth, and N. Vasan · 2024
Closest in time.
A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. d. l. Casas, E. B. Hanna, F. Bressand, et al · 2024
Closest in time.
On the societal impact of open foundation models
S. Kapoor, R. Bommasani, K. Klyman, et al · 2024
Closest in time.
Biomistral: A collection of open-source pretrained large language models for medical domains
Y. Labrak, A. Bazoge, E. Morin, et al · 2024
Closest in time.
Best practices and lessons learned on synthetic data for language models
R. Liu, J. Wei, F. Liu, C. Si, Y. Zhang, J. Rao, S. Zheng, D. Peng, D. Yang, D. Zhou, et al · 2024
Closest in time.
Openmedlm: Prompt engineering can out-perform fine-tuning in medical question-answering with open-source large language models, 2024
J. Maharjan, A. Garikipati, N. P. Singh, L. Cyrus, M. Sharma, M. Ciobanu, G. Barnes, R. Thapa, Q. Mao, and R. Das · 2024
Closest in time.
Sfr-embedding-mistral:enhance text retrieval with transfer learning
R. Meng, Y. Liu, S. R. Joty, C. Xiong, Y. Zhou, and S. Yavuz · 2024
Closest in time.
Datatrove: large scale data processing, 2024
G. Penedo, A. Cappelli, T. Wolf, and M. Sasko · 2024
Closest in time.
A toolbox for surfacing health equity harms and biases in large language models
S. R. Pfohl, H. Cole-Lewis, R. Sayres, et al · 2024
Closest in time.
Rainbow teaming: Open-ended generation of diverse adversarial prompts
M. Samvelyan, S. C. Raparthy, A. Lupu, E. Hambro, A. H. Markosyan, M. Bhatt, Y. Mao, M. Jiang, J. Parker-Holder, J. Foerster, et al · 2024
Closest in time.
Openmathinstruct-1: A 1.8 million math instruction tuning dataset
S. Toshniwal, I. Moshkov, S. Narenthiran, D. Gitman, F. Jia, and I. Gitman · 2024
Closest in time.