Fetching the paper…
Reading the bibliography…
Developing prompt-based methods with Large Language Models (LLMs) requires making numerous decisions, which give rise to a combinatorial search problem over hyper-parameters.
Note on the sampling error of the difference between correlated proportions or percentages
Q. McNemar · 1947
Earlier work this paper cites.
An introduction to the bootstrap
B. Efron and R. J. Tibshirani · 1994
Earlier work this paper cites.
Probability inequalities for sums of bounded random variables
W. Hoeffding · 1994
Earlier work this paper cites.
The bootstrap and its application in signal processing
A. M. Zoubir and B. Boashash · 1998
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
Pure exploration in multi-armed bandits problems
S. Bubeck, R. Munos, and G. Stoltz · 2009
Earlier work this paper cites.
Best arm identification in multi-armed bandits
J.-Y. Audibert and S. Bubeck · 2010
Earlier work this paper cites.
The power of convex relaxation: Near-optimal matrix completion
E. J. Candès and T. Tao · 2010
Earlier work this paper cites.
An asymptotically optimal policy for finite support models in the multiarmed bandit problem
J. Honda and A. Takemura · 2011
Earlier work this paper cites.
Exact matrix completion via convex optimization
E. Candes and B. Recht · 2012
Earlier work this paper cites.
Pac subset selection in stochastic multi-armed bandits
S. Kalyanakrishnan, A. Tewari, P. Auer, and P. Stone · 2012
Earlier work this paper cites.
On bayesian upper confidence bounds for bandit problems
E. Kaufmann, O. Cappé, and A. Garivier · 2012
Earlier work this paper cites.
Kullback-leibler upper confidence bounds for optimal sequential allocation
O. Cappé, A. Garivier, O.-A. Maillard, R. Munos, and G. Stoltz · 2013
Earlier work this paper cites.
Almost optimal exploration in multi-armed bandits
Z. Karnin, T. Koren, and O. Somekh · 2013
Earlier work this paper cites.
Information complexity in bandit subset selection
E. Kaufmann and S. Kalyanakrishnan · 2013
Earlier work this paper cites.
lil’ucb: An optimal exploration algorithm for multi-armed bandits
K. Jamieson, M. Malloy, R. Nowak, and S. Bubeck · 2014
Earlier work this paper cites.
Matrix completion and low-rank svd via fast alternating least squares
T. Hastie, R. Mazumder, J. D. Lee, and R. Zadeh · 2015
Earlier work this paper cites.
Sentiment analysis on twitter data
V. Sahayak, V. Shete, and A. Pathan · 2015
Earlier work this paper cites.
Wikiqa: A challenge dataset for open-domain question answering
Y. Yang, W.-t. Yih, and C. Meek · 2015
Earlier work this paper cites.
Findings of the 2016 conference on machine translation
O. r. Bojar, R. Chatterjee, C. Federmann, Y. Graham, B. Haddow, M. Huck, A. Jimeno Yepes, P. Koehn, V. Logacheva, C. Monz, M. Negri, A. Neveol, M. Neves, M. Popel, M. Post, R. Rubino, C. Scarton, L. Specia, M. Turchi, K. Verspoor, and M. Zampieri · 2016
Earlier work this paper cites.
Optimal best arm identification with fixed confidence
A. Garivier and E. Kaufmann · 2016
Earlier work this paper cites.
On the complexity of best-arm identification in multi-armed bandit models
E. Kaufmann, O. Cappé, and A. Garivier · 2016
Cited alongside, same era.
Abstractive text summarization using sequence-to-sequence rnns and beyond
R. Nallapati, B. Zhou, C. Gulcehre, B. Xiang, et al · 2016
Cited alongside, same era.
SQuAD: 100,000+ questions for machine comprehension of text
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang · 2016
Cited alongside, same era.
triviaqa: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
M. Joshi, E. Choi, D. Weld, and L. Zettlemoyer · 2017
Cited alongside, same era.
Think you have solved question answering? try arc, the ai2 reasoning challenge
P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord · 2018
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman · 2021
Later among the works it cites.
A dataset of information-seeking questions and answers anchored in research papers
P. Dasigi, K. Lo, I. Beltagy, A. Cohan, N. A. Smith, and M. Gardner · 2021
Later among the works it cites.
Bold: Dataset and metrics for measuring biases in open-ended language generation
J. Dhamala, T. Sun, V. Kumar, S. Krishna, Y. Pruksachatkun, K.-W. Chang, and R. Gupta · 2021
Later among the works it cites.
Bartscore: Evaluating generated text as text generation
W. Yuan, G. Neubig, and P. Liu · 2021
Later among the works it cites.
Large language models are zero-shot reasoners
T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Cited alongside, same era.
Newsroom: A dataset of 1.3 million summaries with diverse extractive strategies
M. Grusky, M. Naaman, and Y. Artzi · 2018
Cited alongside, same era.
S. Narayan, S. B. Cohen, and M. Lapata · 2018
Cited alongside, same era.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Z. Yang, P. Qi, S. Zhang, Y. Bengio, W. W. Cohen, R. Salakhutdinov, and C. D. Manning · 2018
Cited alongside, same era.
Nonconvex low-rank tensor completion from noisy data
C. Cai, G. Li, H. V. Poor, and Y. Chen · 2019
Cited alongside, same era.
Pac identification of many good arms in stochastic multi-armed bandits
A. R. Chaudhuri and S. Kalyanakrishnan · 2019
Cited alongside, same era.
Nonconvex optimization meets low-rank matrix factorization: An overview
Y. Chi, Y. M. Lu, and Y. Chen · 2019
Cited alongside, same era.
Solving quantitative reasoning problems with language models
A. Lewkowycz, A. Andreassen, D. Dohan, E. Dyer, H. Michalewski, V. Ramasesh, A. Slone, C. Anil, I. Schlag, T. Gutman-Solo, et al · 2022
Later among the works it cites.
Open problem: Optimal best arm identification with fixed-budget
C. Qin · 2022
Later among the works it cites.
Palm: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al · 2023
Later among the works it cites.
American stories: A large-scale structured text dataset of historical u.s. newspapers, 2023
M. Dell, J. Carlson, T. Bryan, E. Silcock, A. Arora, Z. Shen, L. D’Amico-Wong, Q. Le, P. Querubin, and L. Heldring · 2023
Later among the works it cites.
Mistral 7b, 2023
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. de las Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, L. R. Lavaud, M.-A. Lachaux, P. Stock, T. L. Scao, T. Lavril, T. Wang, T. Lacroix, and W. E. Sayed · 2023
Later among the works it cites.
Alpacaeval: An automatic evaluator of instruction-following models
X. Li, T. Zhang, Y. Dubois, R. Taori, I. Gulrajani, C. Guestrin, P. Liang, and T. B. Hashimoto · 2023
Later among the works it cites.
Gpqa: A graduate-level google-proof q&a benchmark, 2023
D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y. Pang, J. Dirani, J. Michael, and S. R. Bowman · 2023
Later among the works it cites.
How far can camels go? exploring the state of instruction tuning on open resources, 2023
Y. Wang, H. Ivison, P. Dasigi, J. Hessel, T. Khot, K. R. Chandu, D. Wadden, K. MacMillan, N. A. Smith, I. Beltagy, and H. Hajishirzi · 2023
Later among the works it cites.
Large language models as optimizers
C. Yang, X. Wang, Y. Lu, H. Liu, Q. V. Le, D. Zhou, and X. Chen · 2023
Later among the works it cites.
A survey on evaluation of large language models
Y. Chang, X. Wang, J. Wang, Y. Wu, L. Yang, K. Zhu, H. Chen, X. Yi, C. Wang, Y. Wang, et al · 2024
Closest in time.
Chatbot arena: An open platform for evaluating llms by human preference
W.-L. Chiang, L. Zheng, Y. Sheng, A. N. Angelopoulos, T. Li, D. Li, H. Zhang, B. Zhu, M. Jordan, J. E. Gonzalez, et al · 2024
Closest in time.
Rewardbench: Evaluating reward models for language modeling, 2024
N. Lambert, V. Pyatkin, J. Morrison, L. Miranda, B. Y. Lin, K. Chandu, N. Dziri, S. Kumar, T. Zick, Y. Choi, N. A. Smith, and H. Hajishirzi · 2024
Closest in time.
Mapping the increasing use of llms in scientific papers
W. Liang, Y. Zhang, Z. Wu, H. Lepp, W. Ji, X. Zhao, H. Cao, S. Liu, S. He, Z. Huang, et al · 2024
Closest in time.
Best arm identification for prompt learning under a limited budget
C. Shi, K. Yang, J. Yang, and C. Shen · 2024
Closest in time.
Aya dataset: An open-access collection for multilingual instruction tuning, 2024
S. Singh, F. Vargus, D. Dsouza, B. F. Karlsson, A. Mahendiran, W.-Y. Ko, H. Shandilya, J. Patel, D. Mataciunas, L. OMahony, M. Zhang, R. Hettiarachchi, J. Wilson, M. Machado, L. S. Moura, D. Krzemiński, H. Fadaei, I. Ergün, I. Okoh, A. Alaagib, O. Mudannayake, Z. Alyafeai, V. M. Chien, S. Ruder, S. Guthikonda, E. A. Alghamdi, S. Gehrmann, N. Muennighoff, M. Bartolo, J. Kreutzer, A. Üstün, M. Fadaee, and S. Hooker · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, et al · 2024
Closest in time.