Fetching the paper…
Reading the bibliography…
Inference-time techniques, such as repeated sampling or iterative revisions, are emerging as powerful ways to enhance large-language models (LLMs) at test time.
How neural networks learn from experience
Hinton, G. E. et al · 1992
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms, 2012
Snoek, J., Larochelle, H., and Adams, R. P · 2012
Earlier work this paper cites.
Neural architecture search with reinforcement learning, 2017
Zoph, B. and Le, Q. V · 2017
Earlier work this paper cites.
Hypermapper: a practical design space exploration framework
Nardi, L., Souza, A., Koeplinger, D., and Olukotun, K · 2019
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
A comprehensive survey of neural architecture search: Challenges and solutions
Ren, P., Xiao, Y., Chang, X., Huang, P.-Y., Li, Z., Chen, X., and Wang, X · 2021
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback, 2022
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., Chen, C., Olsson, C., Olah, C., Hernandez, D., Drain, D., Ganguli, D., Li, D., Tran-Johnson, E., Perez, E., Kerr, J., Mueller, J., Ladish, J., Landau, J., Ndousse, K., Lukosuite, K., Lovitt, L., Sellitto, M., Elhage, N., Schiefer, N., Mercado, N., DasSarma, N., Lasenby, R., Larson, R., Ringer, S., Johnston, S., Kravec, S., Showk, S. E., Fort, S., Lanham, T., Telleen-Lawton, T., Conerly, T., Henighan, T., Hume, T., Bowman, S. R., Hatfield-Dodds, Z., Mann, B., Amodei, D., Joseph, N., McCandlish, S., Brown, T., and Kaplan, J · 2022
Earlier work this paper cites.
Competition-level code generation with alphacode
Li, Y., Choi, D., Chung, J., Kushman, N., Schrittwieser, J., Leblond, R., Eccles, T., Keeling, J., Gimeno, F., Dal Lago, A., et al · 2022
Earlier work this paper cites.
Bai, J., Bai, S., Chu, Y., Cui, Z., Dang, K., Deng, X., Fan, Y., Ge, W., Han, Y., Huang, F., Hui, B., Ji, L., Li, M., Lin, J., Lin, R., Liu, D., Liu, G., Lu, C., Lu, K., Ma, J., Men, R., Ren, X., Ren, X., Tan, C., Tan, S., Tu, J., Wang, P., Wang, S., Wang, W., Wu, S., Xu, B., Xu, J., Yang, A., Yang, H., Yang, J., Yang, S., Yao, Y., Yu, B., Yuan, H., Yuan, Z., Zhang, J., Zhang, X., Zhang, Y., Zhang, Z., Zhou, C., Zhou, J., Zhou, X., and Zhu, T · 2023
Earlier work this paper cites.
Llm-blender: Ensembling large language models with pairwise comparison and generative fusion
Jiang, D., Ren, X., and Lin, B. Y · 2023
Earlier work this paper cites.
Dspy: Compiling declarative language model calls into self-improving pipelines
Khattab, O., Singhvi, A., Maheshwari, P., Zhang, Z., Santhanam, K., Vardhamanan, S., Haq, S., Sharma, A., Joshi, T. T., Moazam, H., Miller, H., Zaharia, M., and Potts, C · 2023
Earlier work this paper cites.
Alpacaeval: An automatic evaluator of instruction-following models
Li, X., Zhang, T., Dubois, Y., Taori, R., Gulrajani, I., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Earlier work this paper cites.
Gaia: a benchmark for general ai assistants, 2023
Mialon, G., Fourrier, C., Swift, C., Wolf, T., LeCun, Y., and Scialom, T · 2023
Earlier work this paper cites.
Gpqa: A graduate-level google-proof qa benchmark, 2023
Rein, D., Hou, B. L., Stickland, A. C., Petty, J., Pang, R. Y., Dirani, J., Michael, J., and Bowman, S. R · 2023
Earlier work this paper cites.
Iterative dpo alignment
Tran, H., Glaze, C., and Hancock, B · 2023
Cited alongside, same era.
Zephyr: Direct distillation of lm alignment
Tunstall, L., Beeching, E., Lambert, N., Rajani, N., Rasul, K., Belkada, Y., Huang, S., von Werra, L., Fourrier, C., Habib, N., et al · 2023
Cited alongside, same era.
Judging llm-as-a-judge with mt-bench and chatbot arena, 2023
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., and Stoica, I · 2023
Cited alongside, same era.
Phi-3 technical report: A highly capable language model locally on your phone
Abdin, M., Jacobs, S. A., Awan, A. A., Aneja, J., Awadallah, A., Awadalla, H., Bach, N., Bahree, A., Bakhtiari, A., Behl, H., et al · 2024
Cited alongside, same era.
The claude 3 model family: Opus, sonnet, haiku
Anthropic · 2024
Cited alongside, same era.
Mixtral of experts, 2024
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., de las Casas, D., Hanna, E. B., Bressand, F., Lengyel, G., Bour, G., Lample, G., Lavaud, L. R., Saulnier, L., Lachaux, M.-A., Stock, P., Subramanian, S., Yang, S., Antoniak, S., Scao, T. L., Gervet, T., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2024
Closest in time.
From live data to high-quality benchmarks: The arena-hard pipeline, April 2024b
Li, T., Chiang, W.-L., Frick, E., Dunlap, L., Banghua Zhu, J. E. G., and Stoica, I · 2024
Closest in time.
SimPO: Simple preference optimization with a reference-free reward
Meng, Y., Xia, M., and Chen, D · 2024
Closest in time.
Mixeval: Deriving wisdom of the crowd from llm benchmark mixtures, 2024
Ni, J., Xue, F., Yue, X., Deng, Y., Shah, M., Jain, K., Neubig, G., and You, Y · 2024
Closest in time.
Learning to reason with LLMs
OpenAI · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Llama 3 model card
at Meta, A · 2024
Cited alongside, same era.
Large language monkeys: Scaling inference compute with repeated sampling, 2024
Brown, B., Juravsky, J., Ehrlich, R., Clark, R., Le, Q. V., Ré, C., and Mirhoseini, A · 2024
Cited alongside, same era.
Are more llm calls all you need? towards scaling laws of compound inference systems, 2024
Chen, L., Davis, J. Q., Hanin, B., Bailis, P., Stoica, I., Zaharia, M., and Zou, J · 2024
Cited alongside, same era.
Chatbot arena: An open platform for evaluating llms by human preference, 2024
Chiang, W.-L., Zheng, L., Sheng, Y., Angelopoulos, A. N., Li, T., Li, D., Zhang, H., Zhu, B., Jordan, M., Gonzalez, J. E., and Stoica, I · 2024
Cited alongside, same era.
Dbrx technical report
Databricks · 2024
Cited alongside, same era.
Networks of networks: Complexity class principles applied to compound ai systems design, 2024
Davis, J. Q., Hanin, B., Chen, L., Bailis, P., Stoica, I., and Zaharia, M · 2024
Cited alongside, same era.
Guo, D., Zhu, Q., Yang, D., Xie, Z., Dong, K., Zhang, W., Chen, G., Bi, X., Wu, Y., Li, Y. K., Luo, F., Xiong, Y., and Liang, W · 2024
Cited alongside, same era.
Qwen · 2024
Closest in time.
GPQA: A graduate-level google-proof qa benchmark
Rein, D., Hou, B. L., Stickland, A. C., Petty, J., Pang, R. Y., Dirani, J., Michael, J., and Bowman, S. R · 2024
Closest in time.
Llm pruning and distillation in practice: The minitron approach, 2024
Sreenivas, S. T., Muralidharan, S., Joshi, R., Chochowski, M., Patwary, M., Shoeybi, M., Catanzaro, B., Kautz, J., and Molchanov, P · 2024
Closest in time.
Qwq: Reflect deeply on the boundaries of the unknown, November 2024
Team, Q · 2024
Closest in time.
WizardLM: Empowering large pre-trained language models to follow complex instructions
Xu, C., Sun, Q., Zheng, K., Geng, X., Zhao, P., Feng, J., Tao, C., Lin, Q., and Jiang, D · 2024
Closest in time.
Textgrad: Automatic "differentiation" via text
Yuksekgonul, M., Bianchi, F., Boen, J., Liu, S., Huang, Z., Guestrin, C., and Zou, J · 2024
Closest in time.
Aflow: Automating agentic workflow generation, 2024
Zhang, J., Xiang, J., Yu, Z., Teng, F., Chen, X., Chen, J., Zhuge, M., Cheng, X., Hong, S., Wang, J., Zheng, B., Liu, B., Luo, Y., and Wu, C · 2024
Closest in time.
Sky-t1: Train your own o1 preview model within 450
Team, N · 2025
Closest in time.