Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have shown impressive performance on reasoning benchmarks like math and logic.
A versatile stochastic model of a function of unknown and time varying form
H. J. Kushner · 1962
Earlier work this paper cites.
A new method of locating the maximum point of an arbitrary multipeak curve in the presence of noise
H. J. Kushner · 1964
Earlier work this paper cites.
On Bayesian methods for seeking the extremum
J. Moc̆kus · 1974
Earlier work this paper cites.
Learning concepts by asking questions
C. Sammut and R. B. Banerji · 1986
Earlier work this paper cites.
Queries and concept learning
D. Angluin · 1988
Earlier work this paper cites.
Active learning with statistical models
D. A. Cohn, Z. Ghahramani, and M. I. Jordan · 1996
Earlier work this paper cites.
Reinforcement learning: A survey
L. P. Kaelbling, M. L. Littman, and A. W. Moore · 1996
Earlier work this paper cites.
PDDL – the planning domain definition language
M. Ghallab, A. Howe, C. Knoblock, D. McDermott, A. Ram, M. Veloso, D. Weld, D. W. SRI, A. Barrett, D. Christianson, et al · 1998
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
L. P. Kaelbling, M. L. Littman, and A. R. Cassandra · 1998
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration tradeoffs
P. Auer · 2002
Earlier work this paper cites.
Children’s questions: A mechanism for cognitive development
M. M. Chouinard, P. L. Harris, and M. P. Maratsos · 2007
Earlier work this paper cites.
Active learning literature survey
B. Settles · 2009
Earlier work this paper cites.
Gaussian process optimization in the bandit setting: No regret and experimental design
N. Srinivas, A. Krause, S. M. Kakade, and M. Seeger · 2010
Earlier work this paper cites.
Bayesian active learning for classification and preference learning
N. Houlsby, F. Huszár, Z. Ghahramani, and M. Lengyel · 2011
Earlier work this paper cites.
Entropy search for information-efficient global optimization
P. Hennig and C. J. Schuler · 2012
Earlier work this paper cites.
Integrated task and motion planning in belief space
L. P. Kaelbling and T. Lozano-Pérez · 2013
Earlier work this paper cites.
Truth is a lie: Crowd truth and the seven myths of human annotation
L. Aroyo and C. Welty · 2015
Earlier work this paper cites.
Bayesian reinforcement learning: A survey
M. Ghavamzadeh, S. Mannor, J. Pineau, A. Tamar, et al · 2015
Earlier work this paper cites.
Deep Bayesian active learning with image data
Y. Gal, R. Islam, and Z. Ghahramani · 2017
Earlier work this paper cites.
Max-value entropy search for efficient Bayesian optimization
Z. Wang and S. Jegelka · 2017
Earlier work this paper cites.
Focused model-learning and planning for non-Gaussian continuous state-action systems
Z. Wang, S. Jegelka, L. P. Kaelbling, and T. Lozano-Pérez · 2017
Earlier work this paper cites.
MultiWOZ - a large-scale multi-domain wizard-of-oz dataset for task-oriented dialogue modelling
P. Budzianowski, T.-H. Wen, B.-H. Tseng, I. Casanueva, U. Stefan, R. Osman, and M. Gašić · 2018
Cited alongside, same era.
Practical transfer learning for Bayesian optimization
M. Feurer, B. Letham, F. Hutter, and E. Bakshy · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
R. S. Sutton · 2018
Cited alongside, same era.
Active model learning and diverse action sampling for task and motion planning
Z. Wang, C. R. Garrett, L. P. Kaelbling, and T. Lozano-Pérez · 2018
Cited alongside, same era.
Combined task and motion planning under partial observability: An optimization-based approach
C. Phiquepal and M. Toussaint · 2019
Cited alongside, same era.
Pyperplan
Eliciting human preferences with language models, 2023
B. Z. Li, A. Tamkin, N. Goodman, and J. Andreas · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
Why don’t you do it right? analysing annotators’ disagreement in subjective tasks
M. Sandri, E. Leonardelli, S. Tonelli, and E. Ježek · 2023
Later among the works it cites.
Everyone’s voice matters: Quantifying annotation disagreement using demographic information
R. Wan, J. Kim, and D. Kang · 2023
Later among the works it cites.
Large language models should ask clarifying questions to increase confidence in generated code
J. J. Wu · 2023
Later among the works it cites.
On the paradox of learning to reason from data
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Alkhazraji, M. Frorath, M. Grützner, M. Helmert, T. Liebetraut, R. Mattmüller, M. Ortlieb, J. Seipp, T. Springenberg, P. Stahl, and J. Wülfing · 2020
Cited alongside, same era.
AmbigQA: Answering ambiguous open-domain questions
S. Min, J. Michael, H. Hajishirzi, and L. Zettlemoyer · 2020
Cited alongside, same era.
Towards scalable multi-domain conversational agents: The schema-guided dialogue dataset
A. Rastogi, X. Zang, S. Sunkara, R. Gupta, and P. Khaitan · 2020
Cited alongside, same era.
Program synthesis with large language models
J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le, et al · 2021
Cited alongside, same era.
We need to consider disagreement in evaluation
V. Basile, M. Fell, T. Fornaciari, D. Hovy, S. Paun, B. Plank, M. Poesio, A. Uma, et al · 2021
Cited alongside, same era.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. D. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al · 2021
Cited alongside, same era.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman · 2021
Cited alongside, same era.
H. Zhang, L. H. Li, T. Meng, K.-W. Chang, and G. Van Den Broeck · 2023
Later among the works it cites.
Clarify when necessary: Resolving ambiguity through interaction with LMs
M. J. Zhang and E. Choi · 2023
Later among the works it cites.
STaR-GATE: Teaching language models to ask clarifying questions
C. Andukuri, J.-P. Fränken, T. Gerstenberg, and N. D. Goodman · 2024
Later among the works it cites.
Certainly uncertain: A benchmark and metric for multimodal epistemic and aleatoric awareness
K. R. Chandu, L. Li, A. Awadalla, X. Lu, J. S. Park, J. Hessel, L. Wang, and Y. Choi · 2024
Later among the works it cites.
Transfer learning for Bayesian optimization on heterogeneous search spaces
Z. Fan, X. Han, and Z. Wang · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team Google · 2024
Later among the works it cites.
Gemma: Open models based on Gemini research and technology
Gemma Team · 2024
Later among the works it cites.
Loose lips sink ships: Asking questions in battleship with language-informed program sampling, 2024
G. Grand, V. Pepe, J. Andreas, and J. B. Tenenbaum · 2024
Later among the works it cites.
Bayesian preference elicitation with language models, 2024
K. Handa, Y. Gal, E. Pavlick, N. Goodman, J. Andreas, A. Tamkin, and B. Z. Li · 2024
Later among the works it cites.
Z. Hu, C. Liu, X. Feng, Y. Zhao, S.-K. Ng, A. T. Luu, J. He, P. W. Koh, and B. Hooi · 2024
Later among the works it cites.
GSM-plus: A comprehensive benchmark for evaluating the robustness of LLMs as mathematical problem solvers
Q. Li, L. Cui, X. Zhao, L. Kong, and W. Bi · 2024
Later among the works it cites.
Empowering language models with active inquiry for deeper understanding
J.-C. Pang, H.-B. Fan, P. Wang, J.-H. Xiao, N. Tang, S.-H. Yang, C. Jia, S.-J. Huang, and Y. Yu · 2024
Later among the works it cites.
Active preference inference using language models and probabilistic reasoning, 2024
W. T. Piriyakulkij, V. Kuleshov, and K. Ellis · 2024
Later among the works it cites.
Generalized planning in PDDL domains with pretrained large language models
T. Silver, S. Dan, K. Srinivas, J. B. Tenenbaum, L. Kaelbling, and M. Katz · 2024
Later among the works it cites.
Proactive agents for multi-turn text-to-image generation under uncertainty
M. Hahn, W. Zeng, N. Kannen, R. Galt, K. Badola, B. Kim, and Z. Wang · 2025
Closest in time.