Fetching the paper…
Reading the bibliography…
Recent works show that assembling multiple off-the-shelf large language models (LLMs) can harness their complementary abilities.
S. Kullback and R. A. Leibler, “On information and sufficiency,”
1951
Earlier work this paper cites.
J. MacQueen
1967
Earlier work this paper cites.
L. Van der Maaten and G. Hinton, “Visualizing data using t-SNE,”
2008
Earlier work this paper cites.
J. Van Gael, Y. Saatci, Y. W. Teh, and Z. Ghahramani, “Beam sampling for the infinite hidden markov model,” in
2008
Earlier work this paper cites.
2018
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in
2019
Earlier work this paper cites.
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of visual representations,” in
2020
Earlier work this paper cites.
K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick, “Momentum contrast for unsupervised visual representation learning,” in
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
P. He, X. Liu, J. Gao, and W. Chen, “DeBERTa: Decoding-enhanced BERT with disentangled attention,” in
2021
Earlier work this paper cites.
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt, “Measuring massive multitask language understanding,” in
2021
Earlier work this paper cites.
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt, “Measuring mathematical problem solving with the MATH dataset,” in
2021
Earlier work this paper cites.
J. Ni, G. H. Abrego, N. Constant, J. Ma, K. B. Hall, D. Cer, and Y. Yang, “Sentence-T5: Scalable sentence encoders from pre-trained text-to-text models,” in
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervision,” in
2021
Earlier work this paper cites.
G. Izacard, M. Caron, L. Hosseini, S. Riedel, P. Bojanowski, A. Joulin, and E. Grave, “Unsupervised dense information retrieval with contrastive learning,”
2022
Cited alongside, same era.
M. S. Matena and C. A. Raffel, “Merging models with fisher-weighted averaging,” in
2022
Cited alongside, same era.
2022
Cited alongside, same era.
K. Zhou, J. Yang, C. C. Loy, and Z. Liu, “Learning to prompt for vision-language models,”
2022
Cited alongside, same era.
2023
2023
Later among the works it cites.
X. Wang, J. Wei, D. Schuurmans, Q. V. Le, E. H. Chi, S. Narang, A. Chowdhery, and D. Zhou, “Self-consistency improves chain of thought reasoning in language models,” in
2023
Later among the works it cites.
Q. Zheng, X. Xia, X. Zou, Y. Dong, S. Wang, Y. Xue, L. Shen, Z. Wang, A. Wang, Y. Li
2023
Later among the works it cites.
Z. Chen, Y. Deng, H. Yuan, K. Ji, and Q. Gu, “Self-play fine-tuning converts weak language models to strong language models,” in
2024
Closest in time.
C. Computations, “cognitivecomputations/dolphin-2.6-mistral-7b,”
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
L. Gao, J. Tow, B. Abbasi, S. Biderman, S. Black, A. DiPofi, C. Foster, L. Golding, J. Hsu, A. Le Noac’h, H. Li, K. McDonell, N. Muennighoff, C. Ociepa, J. Phang, L. Reynolds, H. Schoelkopf, A. Skowron, L. Sutawika, E. Tang, A. Thite, B. Wang, K. Wang, and A. Zou, “A framework for few-shot language model evaluation,”
2023
Cited alongside, same era.
B. Huang, “Vigogne: French instruction-following and chat models,”
2023
Cited alongside, same era.
Y. Huang, Y. Bai, Z. Zhu, J. Zhang, J. Zhang, T. Su, J. Liu, C. Lv, Y. Zhang, J. Lei, Y. Fu, M. Sun, and J. He, “C-Eval: A multi-level multi-discipline chinese evaluation suite for foundation models,” in
2023
Cited alongside, same era.
G. Ilharco, M. T. Ribeiro, M. Wortsman, L. Schmidt, H. Hajishirzi, and A. Farhadi, “Editing models with task arithmetic,” in
2023
Cited alongside, same era.
2023
Cited alongside, same era.
D. Jiang, X. Ren, and B. Y. Lin, “LLM-Blender: Ensembling large language models with pairwise ranking and generative fusion,” in
2023
Cited alongside, same era.
X. Jin, X. Ren, D. Preotiuc-Pietro, and P. Cheng, “Dataless knowledge fusion by merging weights of language models,” in
2023
Cited alongside, same era.
C. Computations, “cognitivecomputations/dolphin-2.9-llama3-8b,”
2024
Closest in time.
N. Gupta, H. Narasimhan, W. Jitkrittum, A. S. Rawat, A. K. Menon, and S. Kumar, “Language model cascades: Token-level uncertainty and beyond,” in
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
K. Lu, H. Yuan, R. Lin, J. Lin, Z. Yuan, C. Zhou, and J. Zhou, “Routing to the expert: Efficient reward-guided ensemble of large language models,” in
2024
Closest in time.
L. Nie, Z. Ding, E. Hu, C. Jermaine, and S. Chaudhuri, “Online cascade learning for efficient inference over streams,” in
2024
Closest in time.
2024
Closest in time.
Llama Team. “The LLaMA 3 herd of models.” Preprint arXiv:2407.21783, 2024
2024
Closest in time.
H. Wang, F. M. Polo, Y. Sun, S. Kundu, E. Xing, and M. Yurochkin, “Fusing models with complementary expertise,” in
2024
Closest in time.
L. Yu, W. Jiang, H. Shi, J. Yu, Z. Liu, Y. Zhang, J. T. Kwok, Z. Li, A. Weller, and W. Liu, “MetaMath: Bootstrap your own mathematical questions for large language models,” in
2024
Closest in time.
M. Yue, J. Zhao, M. Zhang, L. Du, and Z. Yao, “Large language model cascades with mixture of thought representations for cost-efficient reasoning,” in
2024
Closest in time.
2024
Closest in time.
C. Zhou and B. Yuqi, “Chinese-Mistral: An efficient and effective chinese large language model,”
2024
Closest in time.