Fetching the paper…
Reading the bibliography…
This paper studies ensembling in the era of Large Vision-Language Models (LVLMs).
Verification of forecasts expressed in terms of probability
G. W. Brier · 1950
Earlier work this paper cites.
Neural network ensembles
L. Hansen and P. Salamon · 1990
Earlier work this paper cites.
Cascade generalization
J. Gama and P. Brazdil · 2000
Earlier work this paper cites.
Ensemble learning
T. Dietterich · 2002
Earlier work this paper cites.
Exploration of classification confidence in ensemble learning
L. Li, Q. Hu, X. Wu, and D. Yu · 2014
Earlier work this paper cites.
Obtaining well calibrated probabilities using Bayesian binning
M. P. Naeini, G. F. Cooper, and M. Hauskrecht · 2015
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Earlier work this paper cites.
Revisiting the calibration of modern neural networks
M. Minderer, J. Djolonga, R. Romijnders, F. Hubis, X. Zhai, N. Houlsby, D. Tran, and M. Lucic · 2021
Earlier work this paper cites.
A survey on ensemble learning under the era of deep learning
Y. Yang, H. Lv, and N. Chen · 2021
Earlier work this paper cites.
Flamingo: A visual language model for few-shot learning
J. B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, R. Ring, E. Rutherford, S. Cabi, T. Han, Z. Gong, S. Samangooei, M. Monteiro, J. L. Menick, S. Borgeaud, A. Brock, A. Nematzadeh, S. Sharifzadeh, M. Bińkowski, R. Barreira, O. Vinyals, A. Zisserman, and K. Simonyan · 2022
Cited alongside, same era.
Tomayto, Tomahto. Beyond token-level answer equivalence for question answering evaluation
J. Bulian, C. Buck, W. Gajewski, B. Börschinger, and T. Schuster · 2022
Cited alongside, same era.
Language models are general-purpose interfaces
Y. Hao, H. Song, L. Dong, S. Huang, Z. Chi, W. Wang, S. Ma, and F. Wei · 2022
Cited alongside, same era.
An empirical study of GPT-3 for few-shot knowledge-based VQA
Z. Yang, Z. Gan, J. Wang, X. Hu, Y. Lu, Z. Liu, and L. Wang · 2022
Cited alongside, same era.
PaLI: A jointly-scaled multilingual language-image model
X. Chen, X. Wang, S. Changpinyo, A. Piergiovanni, P. Padlewski, D. Salz, S. Goodman, A. Grycner, B. Mustafa, L. Beyer, A. Kolesnikov, J. Puigcerver, N. Ding, K. Rong, H. Akbari, G. Mishra, L. Xue, A. Thapliyal, J. Bradbury, W. Kuo, M. Seyedhosseini, C. Jia, B. K. Ayan, C. Riquelme, A. Steiner, A. Angelova, X. Zhai, N. Houlsby, and R. Soricut · 2023
Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation
L. Kuhn, Y. Gal, and S. Farquhar · 2023
Closest in time.
Encyclopedic VQA: Visual questions about detailed properties of fine-grained categories
T. Mensink, J. Uijlings, L. Castrejon, A. Goel, F. Cadar, H. Zhou, F. Sha, A. Araujo, and V. Ferrari · 2023
Closest in time.
Large language models vote: Prompting for rare disease identification
D. Oniani, J. Hilsman, H. Dong, F. Gao, S. Verma, and Y. Wang · 2023
Closest in time.
Evaluation of confidence-based ensembling in deep learning image classification
R. Rosales, P. Popov, and M. Paulitsch · 2023
Closest in time.
Reflexion: Language agents with verbal reinforcement learning
N. Shinn, F. Cassano, B. Labash, A. Gopinath, K. Narasimhan, and S. Yao · 2023
Closest in time.
Prompting GPT-3 to be reliable
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Confidence-based ensembles of end-to-end speech recognition models
I. Gitman, V. Lavrukhin, A. Laptev, and B. Ginsburg · 2023
Cited alongside, same era.
Reasoning with language model is planning with world model
S. Hao, Y. Gu, H. Ma, J. J. Hong, Z. Wang, D. Z. Wang, and Z. Hu · 2023
Cited alongside, same era.
PromptCap: Prompt-guided task-aware image captioning
Y. Hu, H. Hua, Z. Yang, W. Shi, N. A. Smith, and J. Luo · 2023
Cited alongside, same era.
DiversiGATE: A comprehensive framework for reliable large language models
S. Imani, A. Beyram, and H. Shrivastava · 2023
Cited alongside, same era.
https://lens.google.com - Web interface available at
Google Lens
Cited in the paper.
PaLM: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, P. Schuh, K. Shi, S. Tsvyashchenko, J. Maynez, A. Rao, P. Barnes, Y. Tay, N. Shazeer, V. Prabhakaran, E. Reif, N. Du, B. Hutchinson, R. Pope, J. Bradbury, J. Austin, M. Isard, G. Gur-Ari, P. Yin, T. Duke, A. Levskaya, S. Ghemawat, S. Dev, H. Michalewski, X. Garcia, V. Misra, K. Robinson, L. Fedus, D. Zhou, D. Ippolito, D. Luan, H. Lim, B. Zoph, A. Spiridonov, R. Sepassi, D. Dohan, S. Agrawal, M. Omernick, A. M. Dai, T. S. Pillai, M. Pellat, A. Lewkowycz, E. Moreira, R. Child, O. Polozov, K. Lee, Z. Zhou, X. Wang, B. Saeta, M. Diaz, O. Firat, M. Catasta, J. Wei, K. Meier-Hellstern, D. Eck, J. Dean, S. Petrov, and N. Fiedel
Cited in the paper.
C. Si, Z. Gan, Z. Yang, S. Wang, J. Wang, J. Boyd-Graber, and L. Wang · 2023
Closest in time.
Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback
K. Tian, E. Mitchell, A. Zhou, A. Sharma, R. Rafailov, H. Yao, C. Finn, and C. D. Manning · 2023
Closest in time.
Can LLMs express their uncertainty? an empirical evaluation of confidence elicitation in LLMs
M. Xiong, Z. Hu, X. Lu, Y. Li, J. Fu, J. He, and B. Hooi · 2023
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
S. Yao, D. Yu, J. Zhao, I. Shafran, T. L. Griffiths, Y. Cao, and K. Narasimhan · 2023
Closest in time.