Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have transformed the natural language processing landscape and brought to life diverse applications.
R. A. Bradley and M. E. Terry, “Rank analysis of incomplete block designs: I. the method of paired comparisons,”
1952
Earlier work this paper cites.
R. Bellman, “A markovian decision process,”
1957
Earlier work this paper cites.
A. Newell, “On the analysis of human problem solving protocols,” 1966
1966
Earlier work this paper cites.
A. Newell, “Human problem solving,”
1972
Earlier work this paper cites.
R. L. Plackett, “The analysis of permutations,”
1975
Earlier work this paper cites.
B. P. Lowerre and B. R. Reddy, “Harpy, a connected speech recognition system,”
1976
Earlier work this paper cites.
G. Williams, “Substantive and adjectival law,” in
1982
Earlier work this paper cites.
A. G. Barto, R. S. Sutton, and C. W. Anderson, “Neuronlike adaptive elements that can solve difficult learning control problems,”
1983
Earlier work this paper cites.
V. Konda and J. Tsitsiklis, “Actor-critic algorithms,”
1999
Earlier work this paper cites.
S. M. Kakade, “A natural policy gradient,”
2001
Earlier work this paper cites.
E. Hovy, L. Gerber, U. Hermjakob, C.-Y. Lin, and D. Ravichandran, “Toward semantics-based answer pinpointing,” in
2001
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in
2002
Earlier work this paper cites.
X. Li and D. Roth, “Learning question classifiers,” in
2002
Earlier work this paper cites.
I. J. Myung, “Tutorial on maximum likelihood estimation,”
2003
Earlier work this paper cites.
R. Coulom, “Efficient selectivity and backup operators in monte-carlo tree search,” in
2006
Earlier work this paper cites.
L. Kocsis and C. Szepesvári, “Bandit based monte-carlo planning,” in
2006
Earlier work this paper cites.
S. Bhatnagar, M. Ghavamzadeh, M. Lee, and R. S. Sutton, “Incremental natural actor-critic algorithms,”
2007
Earlier work this paper cites.
A. Maas, R. E. Daly, P. T. Pham, D. Huang, A. Y. Ng, and C. Potts, “Learning word vectors for sentiment analysis,” in
2011
Earlier work this paper cites.
A. Graves, “Sequence transduction with recurrent neural networks,”
2012
Earlier work this paper cites.
C. B. Browne, E. Powley, D. Whitehouse, S. M. Lucas, P. I. Cowling, P. Rohlfshagen, S. Tavener, D. Perez, S. Samothrakis, and S. Colton, “A survey of monte carlo tree search methods,”
2012
Earlier work this paper cites.
J. Jiang, D. He, and J. Allan, “Searching, browsing, and clicking in a search session: Changes in user behavior by task and over time,” in
2014
Earlier work this paper cites.
R. Vedantam, C. L. Zitnick, and D. Parikh, “Cider: Consensus-based image description evaluation,” 2015
2015
Earlier work this paper cites.
M. Ranzato, S. Chopra, M. Auli, and W. Zaremba, “Sequence level training with recurrent neural networks,” 2016
2016
Earlier work this paper cites.
S. Shen, Y. Cheng, Z. He, W. He, H. Wu, M. Sun, and Y. Liu, “Minimum risk training for neural machine translation,” 2016
2016
Earlier work this paper cites.
V. Mnih, “Asynchronous methods for deep reinforcement learning,”
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
J. X. Chen, “The evolution of computing: Alphago,”
2016
Earlier work this paper cites.
S. Roy and D. Roth, “Solving general arithmetic word problems,” 2016
2016
Earlier work this paper cites.
A. Vaswani, “Attention is all you need,”
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
S. J. Rennie, E. Marcheret, Y. Mroueh, J. Ross, and V. Goel, “Self-critical sequence training for image captioning,” in
2017
Earlier work this paper cites.
L. Yu, W. Zhang, J. Wang, and Y. Yu, “Seqgan: Sequence generative adversarial nets with policy gradient,” 2017
2017
Earlier work this paper cites.
J. Schulman, S. Levine, P. Moritz, M. I. Jordan, and P. Abbeel, “Trust region policy optimization,” 2017
2017
Earlier work this paper cites.
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer, “Automatic differentiation in pytorch,” in
2017
Earlier work this paper cites.
J. Devlin, “Bert: Pre-training of deep bidirectional transformers for language understanding,”
2018
Earlier work this paper cites.
A. Fan, M. Lewis, and Y. Dauphin, “Hierarchical neural story generation,”
2018
Earlier work this paper cites.
A. Radford, “Improving language understanding by generative pre-training,” 2018
2018
Earlier work this paper cites.
J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel, “High-dimensional continuous control using generalized advantage estimation,” 2018
2018
Earlier work this paper cites.
A. Sergeev and M. D. Balso, “Horovod: fast and easy distributed deep learning in tensorflow,” 2018
2018
Earlier work this paper cites.
P. Moritz, R. Nishihara, S. Wang, A. Tumanov, R. Liaw, E. Liang, M. Elibol, Z. Yang, W. Paul, M. I. Jordan, and I. Stoica, “Ray: A distributed framework for emerging ai applications,” 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Z. Yang, “Xlnet: Generalized autoregressive pretraining for language understanding,”
2019
Earlier work this paper cites.
Z. Lan, “Albert: A lite bert for self-supervised learning of language representations,”
2019
Earlier work this paper cites.
P. Tillet, H.-T. Kung, and D. Cox, “Triton: an intermediate language and compiler for tiled neural network computations,” in
2019
Earlier work this paper cites.
M. Dukhan, “The indirect convolution algorithm,” 2019
2019
Earlier work this paper cites.
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. de Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly, “Parameter-efficient transfer learning for nlp,” 2019
2019
Earlier work this paper cites.
D. Dua, Y. Wang, P. Dasigi, G. Stanovsky, S. Singh, and M. Gardner, “DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs,” in
2019
Earlier work this paper cites.
W. H. Guss, B. Houghton, N. Topin, P. Wang, C. Codel, M. Veloso, and R. Salakhutdinov, “Minerl: A large-scale dataset of minecraft demonstrations,” 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
A. Amini, S. Gabriel, P. Lin, R. Koncel-Kedziorski, Y. Choi, and H. Hajishirzi, “Mathqa: Towards interpretable math word problem solving with operation-based formalisms,” 2019
2019
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell,
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,”
2020
Earlier work this paper cites.
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel,
2020
Earlier work this paper cites.
Y. Shao, J. Mao, Y. Liu, W. Ma, K. Satoh, M. Zhang, and S. Ma, “Bert-pli: Modeling paragraph-level interactions for legal case retrieval.,” in
2020
Earlier work this paper cites.
N. Vieillard, T. Kozuno, B. Scherrer, O. Pietquin, R. Munos, and M. Geist, “Leverage the average: an analysis of kl regularization in reinforcement learning,”
2020
Earlier work this paper cites.
M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro, “Megatron-lm: Training multi-billion parameter language models using model parallelism,” 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano, “Learning to summarize with human feedback,”
2020
Earlier work this paper cites.
Z. Xie, S. Thiem, J. Martin, E. Wainwright, S. Marmorstein, and P. Jansen, “WorldTree v2: A corpus of science-domain structured explanations and inference patterns supporting multi-hop inference,” in
2020
Earlier work this paper cites.
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine, “D4rl: Datasets for deep data-driven reinforcement learning,” 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
M. Eric, R. Goel, S. Paul, A. Sethi, S. Agarwal, S. Gao, and D. Hakkani-Tur, “MultiWOZ 2.1: A consolidated multi-domain dialogue dataset with state corrections and state tracking baselines,” in
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
R. Lee, O. J. Mengshoel, A. Saksena, R. Gardner, D. Genin, J. Silbermann, M. Owen, and M. J. Kochenderfer, “Adaptive stress testing: Finding likely failure events with reinforcement learning,” 2020
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
O. Lieber, O. Sharir, B. Lenz, and Y. Shoham, “Jurassic-1: Technical details and evaluation,”
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
Q. Lhoest, A. Villanova del Moral, Y. Jernite, A. Thakur, P. von Platen, S. Patil, J. Chaumond, M. Drame, J. Plu, L. Tunstall, J. Davison, M. Šaško, G. Chhablani, B. Malik, S. Brandeis, T. Le Scao, V. Sanh, C. Xu, N. Patry, A. McMillan-Major, P. Schmid, S. Gugger, C. Delangue, T. Matussière, L. Debut, S. Bekman, P. Cistac, T. Goehringer, V. Mustar, F. Lagunas, A. Rush, and T. Wolf, “Datasets: A community library for natural language processing,” in
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
X. L. Li and P. Liang, “Prefix-tuning: Optimizing continuous prompts for generation,”
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
F. Liu, E. Bugliarello, E. M. Ponti, S. Reddy, N. Collier, and D. Elliott, “Visually grounded reasoning across languages and cultures,” in
2021
Earlier work this paper cites.
S. Lin, J. Hilton, and O. Evans, “Truthfulqa: Measuring how models mimic human falsehoods,”
2021
Earlier work this paper cites.
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt, “Measuring massive multitask language understanding,”
2021
Earlier work this paper cites.
D. Hendrycks, C. Burns, S. Basart, A. Critch, J. Li, D. Song, and J. Steinhardt, “Aligning ai with shared human values,”
2021
Earlier work this paper cites.
N. Thakur, N. Reimers, A. Rücklé, A. Srivastava, and I. Gurevych, “BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models,” in
2021
Earlier work this paper cites.
M. Ma, P. D’Oro, Y. Bengio, and P.-L. Bacon, “Long-term credit assignment via model-based temporal shortcuts,” in
2021
Earlier work this paper cites.
H. Zhang and Y. Guo, “Generalization of reinforcement learning with policy-aware adversarial data augmentation,” 2021
2021
Earlier work this paper cites.
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou,
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray,
2022
Earlier work this paper cites.
Y. Wang, S. Wang, Y. Li, and D. Dou, “Recognizing medical search query intent by few-shot learning,” in
2022
Earlier work this paper cites.
R. Luo, L. Sun, Y. Xia, T. Qin, S. Zhang, H. Poon, and T.-Y. Liu, “Biogpt: generative pre-trained transformer for biomedical text generation and mining,”
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa, “Large language models are zero-shot reasoners,”
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, J. W. Rae, O. Vinyals, and L. Sifre, “Training compute-optimal large language models,” 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
S. Mangrulkar, S. Gugger, L. Debut, Y. Belkada, S. Paul, and B. Bossan, “Peft: State-of-the-art parameter-efficient fine-tuning methods,” 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
R. Y. Aminabadi, S. Rajbhandari, M. Zhang, A. A. Awan, C. Li, D. Li, E. Zheng, J. Rasley, S. Smith, O. Ruwase, and Y. He, “Deepspeed inference: Enabling efficient inference of transformer models at unprecedented scale,” 2022
2022
Earlier work this paper cites.
Y. Zhou and K. Yang, “Exploring tensorrt to improve real-time inference for deep learning,” in
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
J. Hilton and L. Gao, “Measuring goodhart’s law,”
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
S. et.al, “Beyond the imitation game: Quantifying and extrapolating the capabilities of language models,” 2022
2022
Earlier work this paper cites.
H. Chen, D. M. Vo, H. Takamura, Y. Miyao, and H. Nakayama, “Storyer: Automatic story evaluation via ranking, rating and reasoning,” 2022
2022
Earlier work this paper cites.
S. Mishra, D. Khashabi, C. Baral, and H. Hajishirzi, “Cross-task generalization via natural language crowdsourcing instructions,” in
2022
Earlier work this paper cites.
Y. Wang, S. Mishra, P. Alipoormolabashi, Y. Kordi, A. Mirzaei, A. Arunkumar, A. Ashok, A. S. Dhanasekaran, A. Naik, D. Stap,
2022
Earlier work this paper cites.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. F. Christiano, J. Leike, and R. Lowe, “Training language models to follow instructions with human feedback,” in
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
Y. Fu, H. Peng, A. Sabharwal, P. Clark, and T. Khot, “Complexity-based prompting for multi-step reasoning,” in
2022
Earlier work this paper cites.
B. Johnson, “Metacognition for artificial intelligence system safety–an approach to safe and desired behavior,”
2022
Earlier work this paper cites.
W. Yu, Z. Sun, J. Xu, Z. Dong, X. Chen, H. Xu, and J.-R. Wen, “Explainable legal case matching via inverse optimal transport-based rationale extraction,” in
2022
Earlier work this paper cites.
S. Yang and D. Song, “FPC: Fine-tuning with prompt curriculum for relation extraction,” in
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
A. G. Baydin, B. A. Pearlmutter, D. Syme, F. Wood, and P. Torr, “Gradients without backpropagation,” 2022
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
S. Ye, Y. Jo, D. Kim, S. Kim, H. Hwang, and M. Seo, “Selfee: Iterative self-revising llm empowered by self-feedback generation,”
2023
Earlier work this paper cites.
J. Ren, Y. Zhao, T. Vu, P. J. Liu, and B. Lakshminarayanan, “Self-evaluation improves selective generation in large language models,” in
2023
Cited alongside, same era.
Y. Lai, C. Li, Y. Wang, T. Zhang, R. Zhong, L. Zettlemoyer, W.-t. Yih, D. Fried, S. Wang, and T. Yu, “Ds-1000: A natural and reliable benchmark for data science code generation,” in
2023
Cited alongside, same era.
2023
Cited alongside, same era.
T. Hagendorff, S. Fabi, and M. Kosinski, “Human-like intuitive behavior and reasoning biases emerged in large language models but disappeared in chatgpt,”
2023
Cited alongside, same era.
R. OpenAI, “Gpt-4 technical report. arxiv 2303.08774,”
2023
2024
Later among the works it cites.
B. Hui, J. Yang, Z. Cui, J. Yang, D. Liu, L. Zhang, T. Liu, J. Zhang, B. Yu, K. Lu,
2024
Later among the works it cites.
A. Défossez, L. Mazaré, M. Orsini, A. Royer, P. Pérez, H. Jégou, E. Grave, and N. Zeghidour, “Moshi: a speech-text foundation model for real-time dialogue,” tech. rep., 2024
2024
Later among the works it cites.
R. Teknium, J. Quesnelle, and C. Guang, “Hermes 3 technical report,”
2024
Later among the works it cites.
Z. Cai, M. Cao, H. Chen, K. Chen, K. Chen, X. Chen, X. Chen, Z. Chen, Z. Chen, P. Chu,
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
2023
Cited alongside, same era.
S. Hao, Y. Gu, H. Ma, J. J. Hong, Z. Wang, D. Z. Wang, and Z. Hu, “Reasoning with language model is planning with world model,” in
2023
Cited alongside, same era.
Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y. Xu, E. Ishii, Y. J. Bang, A. Madotto, and P. Fung, “Survey of hallucination in natural language generation,”
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
S. Yue, W. Chen, S. Wang, B. Li, C. Shen, S. Liu, Y. Zhou, Y. Xiao, S. Yun, X. Huang,
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Later among the works it cites.
P. Agrawal, S. Antoniak,
2024
Later among the works it cites.
A. A. G. Intelligence, “The amazon nova family of models: Technical report and model card,”
2024
Later among the works it cites.
Fujitsu and T. I. of Technology, “Fugaku-llm: The largest cpu-only trained language model,”
2024
Later among the works it cites.
R. AI, “Nova: A family of ai models by rubik’s ai,”
2024
Later among the works it cites.
M. R. Team
2024
Later among the works it cites.
T. GLM, A. Zeng, B. Xu, B. Wang, C. Zhang, D. Yin, D. Zhang, D. Rojas, G. Feng, H. Zhao,
2024
Later among the works it cites.
A. Bartolome, J. Hong, N. Lee, K. Rasul, and L. Tunstall, “Zephyr 141b a39b,” 2024
2024
Later among the works it cites.
M. Abdin, J. Aneja, H. Behl, S. Bubeck, R. Eldan, S. Gunasekar, M. Harrison, R. J. Hewett, M. Javaheripi, P. Kauffmann, J. R. Lee, Y. T. Lee, Y. Li, W. Liu, C. C. T. Mendes, A. Nguyen, E. Price, G. de Rosa, O. Saarikivi, A. Salim, S. Shah, X. Wang, R. Ward, Y. Wu, D. Yu, C. Zhang, and Y. Zhang, “Phi-4 technical report,” 2024
2024
Later among the works it cites.
C. Team, “Chameleon: Mixed-modal early-fusion foundation models,”
2024
Later among the works it cites.
A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei,
2024
Later among the works it cites.
S. Guo, B. Zhang, T. Liu, T. Liu, M. Khalman, F. Llinares, A. Rame, T. Mesnard, Y. Zhao, B. Piot,
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
J. Hong, N. Lee, and J. Thorne, “Orpo: Monolithic preference optimization without reference model,” 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
H. Wang, S. Hao, H. Dong, S. Zhang, Y. Bao, Z. Yang, and Y. Wu, “Offline reinforcement learning for llm multi-step reasoning,” 2024
2024
Later among the works it cites.
C. Wang, Z. Zhao, C. Zhu, K. A. Sankararaman, M. Valko, X. Cao, Z. Chen, M. Khabsa, Y. Chen, H. Ma, and S. Wang, “Preference optimization with multi-sample comparisons,” 2024
2024
Later among the works it cites.
X. Zhang, C. Du, T. Pang, Q. Liu, W. Gao, and M. Lin, “Chain of preference optimization: Improving chain-of-thought reasoning in LLMs,” in
2024
Later among the works it cites.
G. Xu, P. Jin, H. Li, Y. Song, L. Sun, and L. Yuan, “Llava-cot: Let vision language models reason step-by-step,” 2024
2024
Later among the works it cites.
J. Castaño, S. Martínez-Fernández, X. Franch, and J. Bogner, “Analyzing the evolution and maintenance of ml models on hugging face,” 2024
2024
Later among the works it cites.
S. Hao, Y. Gu, H. Luo, T. Liu, X. Shao, X. Wang, S. Xie, H. Ma, A. Samavedhi, Q. Gao,
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
K. Kuckreja, M. S. Danish, M. Naseer, A. Das, S. Khan, and F. S. Khan, “Geochat: Grounded large vision-language model for remote sensing,” in
2024
Later among the works it cites.
2024
Later among the works it cites.
X. Zhu, J. Li, Y. Liu, C. Ma, and W. Wang, “A survey on model compression for large language models,”
2024
Later among the works it cites.
H. Sun, M. Haider, R. Zhang, H. Yang, J. Qiu, M. Yin, M. Wang, P. Bartlett, and A. Zanette, “Fast best-of-n decoding via speculative rejection,” 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
X. Chen, R. Aksitov, U. Alon, J. Ren, K. Xiao, P. Yin, S. Prakash, C. Sutton, X. Wang, and D. Zhou, “Universal self-consistency for large language models,” in
2024
Later among the works it cites.
2024
Later among the works it cites.
M. Besta, N. Blach, A. Kubicek, R. Gerstenberger, M. Podstawski, L. Gianinazzi, J. Gajda, T. Lehmann, H. Niewiadomski, P. Nyczyk, and T. Hoefler, “Graph of thoughts: Solving elaborate problems with large language models,”
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun, C. Wei, B. Yu, R. Yuan, R. Sun, M. Yin, B. Zheng, Z. Yang, Y. Liu, W. Huang, H. Sun, Y. Su, and W. Chen, “Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi,” in
2024
Later among the works it cites.
T. Schmied, M. Hofmarcher, F. Paischer, R. Pascanu, and S. Hochreiter, “Learning to modulate pre-trained models in rl,”
2024
Later among the works it cites.
T. Nguyen, C. V. Nguyen, V. D. Lai, H. Man, N. T. Ngo, F. Dernoncourt, R. A. Rossi, and T. H. Nguyen, “CulturaX: A cleaned, enormous, and multilingual dataset for large language models in 167 languages,” in
2024
Later among the works it cites.
2024
Later among the works it cites.
G. Son, D. Yoon, J. Suk, J. Aula-Blasco, M. Aslan, V. T. Kim, S. B. Islam, J. Prats-Cristià, L. Tormo-Bañuelos, and S. Kim, “Mm-eval: A multilingual meta-evaluation benchmark for llm-as-a-judge and reward models,” 2024
2024
Later among the works it cites.
C. Chhun, F. M. Suchanek, and C. Clavel, “Do language models enjoy their own stories? Prompting large language models for automatic story evaluation,”
2024
Later among the works it cites.
J. Ji, M. Liu, J. Dai, X. Pan, C. Zhang, C. Bian, B. Chen, R. Sun, Y. Wang, and Y. Yang, “Beavertails: Towards improved safety alignment of llm via a human-preference dataset,”
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri, S. Prabhumoye, Y. Yang,
2024
Later among the works it cites.
OpenAI, “Early access for safety testing,” 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Y. R. Dong, T. Hu, and N. Collier, “Can llm be a personalized judge?,”
2024
Later among the works it cites.
H. Du, S. Liu, L. Zheng, Y. Cao, A. Nakamura, and L. Chen, “Privacy in fine-tuning large language models: Attacks, defenses, and future directions,” 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
M. Fawi, “Curlora: Stable llm continual fine-tuning and catastrophic forgetting mitigation,” 2024
2024
Later among the works it cites.
Y. Sun, Z. Li, Y. Li, and B. Ding, “Improving lora in privacy-preserving federated learning,”
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
L. Zhang, A. Hosseini, H. Bansal, M. Kazemi, A. Kumar, and R. Agarwal, “Generative verifiers: Reward modeling as next-token prediction,” in
2024
Later among the works it cites.
T. Zhang, S. G. Patil, N. Jain, S. Shen, M. Zaharia, I. Stoica, and J. E. Gonzalez, “Raft: Adapting language model to domain specific rag,” in
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
L. Luo, Y. Liu, R. Liu, S. Phatale, H. Lara, Y. Li, L. Shu, Y. Zhu, L. Meng, J. Sun,
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
S. Lee, G. Lee, J. W. Kim, J. Shin, and M.-K. Lee, “Hetal: Efficient privacy-preserving transfer learning with homomorphic encryption,” 2024
2024
Later among the works it cites.
Y. Wei, J. Jia, Y. Wu, C. Hu, C. Dong, Z. Liu, X. Chen, Y. Peng, and S. Wang, “Distributed differential privacy via shuffling versus aggregation: A curious study,”
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
J. Zhang, J. Xiang, Z. Yu, F. Teng, X. Chen, J. Chen, M. Zhuge, X. Cheng, S. Hong, J. Wang,
2024
Later among the works it cites.
H. D. Le, X. Xia, and Z. Chen, “Multi-agent causal discovery using large language models,”
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Y. Liu, C. Cai, X. Zhang, X. Yuan, and C. Wang, “Arondight: Red teaming large vision language models with auto-generated multi-modal jailbreak prompts,” in
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
M. Kemmerling, D. Lütticke, and R. H. Schmitt, “Beyond games: a systematic review of neural monte carlo tree search applications,”
2024
Later among the works it cites.
2024
Later among the works it cites.
H. Li, L. Ding, M. Fang, and D. Tao, “Revisiting catastrophic forgetting in large language model tuning,” 2024
2024
Later among the works it cites.
N. Alzahrani, H. A. Alyahya, Y. Alnumay, S. Alrashed, S. Alsubaie, Y. Almushaykeh, F. Mirza, N. Alotaibi, N. Altwairesh, A. Alowisheq, M. S. Bari, and H. Khan, “When benchmarks are targets: Revealing the sensitivity of large language model leaderboards,” 2024
2024
Later among the works it cites.
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi,
2025
Closest in time.
Y. Ye, Z. Huang, Y. Xiao, E. Chern, S. Xia, and P. Liu, “Limo: Less is more for reasoning,” 2025
2025
Closest in time.
C. Li, W. Wu, H. Zhang, Y. Xia, S. Mao, L. Dong, I. Vulić, and F. Wei, “Imagine while reasoning in space: Multimodal visualization-of-thought,” 2025
2025
Closest in time.
(Accessed: 2025-12-6)
interconnects.ai, “blob reinforcement fine-tuning.” · 2025
Closest in time.
(Accessed: 2025-12-6)
OpenAI, “Reinforcement fine-tuning.” · 2025
Closest in time.
J. Wu, S. Yang, R. Zhan, Y. Yuan, L. S. Chao, and D. F. Wong, “A survey on llm-generated text detection: Necessity, methods, and future directions,”
2025
Closest in time.
J. Geiping, S. McLeish, N. Jain, J. Kirchenbauer, S. Singh, B. R. Bartoldson, B. Kailkhura, A. Bhatele, and T. Goldstein, “Scaling up test-time compute with latent reasoning: A recurrent depth approach,” 2025
2025
Closest in time.
Accessed: 2025-02-28
OpenAI, “Openai gpt-4.5 system card,” 2025 · 2025
Closest in time.
Accessed: 2025-02-26
Anthropic, “Claude 3.7 sonnet,” 2025 · 2025
Closest in time.
Accessed: 2025-02-26
Nexusflow, “Athene: An rlhf-enhanced language model,” 2024 · 2025
Closest in time.
500B, Zed AI, RLHF, Multi-modal, Open
Z. AI, “Zed,” 2025 · 2025
Closest in time.
220B, Supernova Labs, RLHF, Multi-modal, Open
S. Labs, “Supernova,” 2025 · 2025
Closest in time.
Accessed: 2025-02-24
xAI, “Grok-3: The next generation ai model by xai,” tech. rep., xAI, 2025 · 2025
Closest in time.
MiniMax, A. Li,
2025
Closest in time.
OpenAI, “Openai o3 system card,” technical report, OpenAI, 2025
2025
Closest in time.
K. Team, A. Du, B. Gao, B. Xing, C. Jiang, C. Chen, C. Li, C. Xiao, C. Du, C. Liao,
2025
Closest in time.
O. Thawakar, D. Dissanayake, K. More, R. Thawkar, A. Heakl, N. Ahsan, Y. Li, M. Zumri, J. Lahoud, R. M. Anwer, H. Cholakkal, I. Laptev, M. Shah, F. S. Khan, and S. Khan, “Llamav-o1: Rethinking step-by-step visual reasoning in llms,” 2025
2025
Closest in time.
R package version 19.0.0, https://arrow.apache.org/docs/r/
N. Richardson, I. Cook, N. Crane, D. Dunnington, R. François, J. Keane, D. Moldovan-Grünfeld, J. Ooms, J. Wujciak-Jens, and Apache Arrow, · 2025
Closest in time.
Runtime for Intel CPUs/iGPUs with pruning/quantization support
I. Corporation, “Openvino: Intel optimization toolkit,” 2025 · 2025
Closest in time.
Deterministic low-latency inference via custom tensor streaming processor
I. Groq, “Groq: Ai accelerator,” 2025 · 2025
Closest in time.
2025
Closest in time.
S. R. Motwani, C. Smith, R. J. Das, R. Rafailov, I. Laptev, P. H. S. Torr, F. Pizzati, R. Clark, and C. S. de Witt, “Malt: Improving reasoning with multi-agent llm training,” 2025
2025
Closest in time.
2025
Closest in time.
C. Wang, Z. Zhao, Y. Jiang, Z. Chen, C. Zhu, Y. Chen, J. Liu, L. Zhang, X. Fan, H. Ma, and S.-Y. Wang, “Beyond reward hacking: Causal rewards for large language model alignment,” 2025
2025
Closest in time.
Y. Yu, W. Ping, Z. Liu, B. Wang, J. You, C. Zhang, M. Shoeybi, and B. Catanzaro, “Rankrag: Unifying context ranking with retrieval-augmented generation in llms,”
2025
Closest in time.
H. Xia, Y. Li, C. T. Leong, W. Wang, and W. Li, “Tokenskip: Controllable chain-of-thought compression in llms,” 2025
2025
Closest in time.