Fetching the paper…
Reading the bibliography…
Ensembling outputs from diverse sources is a straightforward yet effective approach to boost performance.
Introduction to the practice of statistics
H. J. Arnold · 1990
Earlier work this paper cites.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2020
Earlier work this paper cites.
Mathematical modelling and verbal abilities: How they determine students’ ability to solve mathematical word problems?
K. Sarjana, L. Hayati, and W. Wahidaturrahmi · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. D. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Competition-level code generation with alphacode
Y. Li, D. Choi, J. Chung, N. Kushman, J. Schrittwieser, R. Leblond, T. Eccles, J. Keeling, F. Gimeno, A. Dal Lago, et al · 2022
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, S. Narang, A. Chowdhery, and D. Zhou · 2022
Earlier work this paper cites.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Earlier work this paper cites.
Introducing claude, 2023
A. Anthropic · 2023
Earlier work this paper cites.
J. Bai, S. Bai, Y. Chu, Z. Cui, K. Dang, X. Deng, Y. Fan, W. Ge, Y. Han, F. Huang, et al · 2023
Earlier work this paper cites.
The vendi score: A diversity evaluation metric for machine learning
D. Dan Friedman and A. B. Dieng · 2023
Earlier work this paper cites.
Improving factuality and reasoning in language models through multiagent debate
Y. Du, S. Li, A. Torralba, J. B. Tenenbaum, and I. Mordatch · 2023
Earlier work this paper cites.
Encouraging divergent thinking in large language models through multi-agent debate
T. Liang, Z. He, W. Jiao, X. Wang, Y. Wang, R. Wang, Y. Yang, Z. Tu, and S. Shi · 2023
Earlier work this paper cites.
Routing to the expert: Efficient reward-guided ensemble of large language models, 2023
K. Lu, H. Yuan, R. Lin, J. Lin, Z. Yuan, C. Zhou, and J. Zhou · 2023
Cited alongside, same era.
Code llama: Open foundation models for code
B. Roziere, J. Gehring, F. Gloeckle, S. Sootla, I. Gat, X. E. Tan, Y. Adi, J. Liu, R. Sauvestre, T. Remez, et al · 2023
Cited alongside, same era.
Gpt-4 doesn’t know it’s wrong: An analysis of iterative prompting for reasoning problems
K. Stechly, M. Marquez, and S. Kambhampati · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
G. Team, R. Anil, S. Borgeaud, Y. Wu, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, et al · 2023
Cited alongside, same era.
Bonbon alignment for large language models and the sweetness of best-of-n sampling
L. Gui, C. Gârbacea, and V. Veitch · 2024
Later among the works it cites.
More agents is all you need, 2024
J. Li, Q. Zhang, Y. Yu, Q. Fu, and D. Ye · 2024
Later among the works it cites.
Mitigating the alignment tax of rlhf, 2024
Y. Lin, H. Lin, W. Xiong, S. Diao, J. Liu, J. Zhang, R. Pan, H. Wang, W. Hu, H. Zhang, H. Dong, R. Pi, H. Zhao, N. Jiang, H. Ji, Y. Yao, and T. Zhang · 2024
Later among the works it cites.
Self-refine: Iterative refinement with self-feedback
A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri, S. Prabhumoye, Y. Yang, et al · 2024
Later among the works it cites.
SimPO: Simple preference optimization with a reference-free reward
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al · 2023
Cited alongside, same era.
Can large language models really improve by self-critiquing their own plans?
K. Valmeekam, M. Marquez, and S. Kambhampati · 2023
Cited alongside, same era.
Wizardlm: Empowering large language models to follow complex instructions
C. Xu, Q. Sun, K. Zheng, X. Geng, P. Zhao, J. Feng, C. Tao, and D. Jiang · 2023
Cited alongside, same era.
Judging llm-as-a-judge with mt-bench and chatbot arena
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, et al · 2023
Cited alongside, same era.
Large language monkeys: Scaling inference compute with repeated sampling
B. Brown, J. Juravsky, R. Ehrlich, R. Clark, Q. V. Le, C. Ré, and A. Mirhoseini · 2024
Cited alongside, same era.
Moa is all you need: Building llm research team using mixture of agents
S. Chen, L. Zeng, A. Raghunathan, F. Huang, and T. C. Kim · 2024
Cited alongside, same era.
Length-controlled alpacaeval: A simple way to debias automatic evaluators
Y. Dubois, B. Galambosi, P. Liang, and T. B. Hashimoto · 2024
Cited alongside, same era.
A. P. Gema, J. O. J. Leang, G. Hong, A. Devoto, A. C. M. Mancino, R. Saxena, X. He, Y. Zhao, X. Du, M. R. G. Madani, et al · 2024
Cited alongside, same era.
Y. Meng, M. Xia, and D. Chen · 2024
Later among the works it cites.
Openpipe mixture of agents: Outperform gpt-4 at 1/25th the cost, 2024
OpenPipe · 2024
Later among the works it cites.
Warp: On the benefits of weight averaged rewarded policies, 2024
A. Ramé, J. Ferret, N. Vieillard, R. Dadashi, L. Hussenot, P.-L. Cedoz, P. G. Sessa, S. Girgin, A. Douillard, and O. Bachem · 2024
Later among the works it cites.
Scaling llm test-time compute optimally can be more effective than scaling model parameters, 2024
C. Snell, J. Lee, K. Xu, and A. Kumar · 2024
Later among the works it cites.
Introducing dbrx: A new state-of-the-art open llm, 2024
M. R. Team et al · 2024
Later among the works it cites.
An empirical analysis of compute-optimal inference for problem-solving with language models, 2024
Y. Wu, Z. Sun, S. Li, S. Welleck, and Y. Yang · 2024
Later among the works it cites.
A. Yang, B. Yang, B. Hui, B. Zheng, B. Yu, C. Zhou, C. Li, C. Li, D. Liu, F. Huang, G. Dong, H. Wei, H. Lin, J. Tang, J. Wang, J. Yang, J. Tu, J. Zhang, J. Ma, J. Yang, J. Xu, J. Zhou, J. Bai, J. He, J. Lin, K. Dang, K. Lu, K. Chen, K. Yang, M. Li, M. Xue, N. Ni, P. Zhang, P. Wang, R. Peng, R. Men, R. Gao, R. Lin, S. Wang, S. Bai, S. Tan, T. Zhu, T. Li, T. Liu, W. Ge, X. Deng, X. Zhou, X. Ren, X. Zhang, X. Wei, X. Ren, X. Liu, Y. Fan, Y. Yao, Y. Zhang, Y. Wan, Y. Chu, Y. Liu, Z. Cui, Z. Zhang, Z. Guo, and Z. Fan · 2024
Later among the works it cites.
Wpo: Enhancing rlhf with weighted preference optimization
W. Zhou, R. Agrawal, S. Zhang, S. R. Indurthi, S. Zhao, K. Song, S. Xu, and C. Zhu · 2024
Later among the works it cites.
Deepseek-coder-v2: Breaking the barrier of closed-source models in code intelligence
Q. Zhu, D. Guo, Z. Shao, D. Yang, P. Wang, R. Xu, Y. Wu, Y. Li, H. Gao, S. Ma, et al · 2024
Later among the works it cites.