Fetching the paper…
Reading the bibliography…
The rapid advancement in large language models (LLMs) has brought forth a diverse range of models with varying capabilities that excel in different tasks and domains.
Network randomization: A simple technique for generalization in deep reinforcement learning
Lee, K., Lee, K., Shin, J., and Lee, H · 1910
Earlier work this paper cites.
The proposed uscf rating system, its development, theory, and applications
Elo, A. E · 1967
Earlier work this paper cites.
The multi-armed bandit problem: decomposition and computation
Katehakis, M. N. and Veinott Jr, A. F · 1987
Earlier work this paper cites.
Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods
Platt, J. et al · 1999
Earlier work this paper cites.
Item response theory: Principles and applications
Hambleton, R. K. and Swaminathan, H · 2013
Earlier work this paper cites.
Hidden parameter markov decision processes: an emerging paradigm for modeling families of related tasks
Konidaris, G. and Doshi-Velez, F · 2014
Earlier work this paper cites.
Policy gradient approaches for multi-objective sequential decision making
Parisi, S., Pirotta, M., Smacchia, N., Bascetta, L., and Restelli, M · 2014
Earlier work this paper cites.
Multi-objective reinforcement learning using sets of pareto dominating policies
Van Moffaert, K. and Nowé, A · 2014
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M., and Abbeel, P · 2015
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
Zhang, H · 2017
Earlier work this paper cites.
Generalization and regularization in dqn
Farebrother, J., Machado, M. C., and Bowling, M · 2018
Earlier work this paper cites.
Learning invariances for policy generalization
Tachet, R., Bachman, P., and van Seijen, H · 2018
Earlier work this paper cites.
Dynamic weights in multi-objective deep reinforcement learning
Abels, A., Roijers, D., Lenaerts, T., Nowé, A., and Steckelmacher, D · 2019
Earlier work this paper cites.
A survey on practical applications of multi-armed and contextual bandits
Bouneffouf, D. and Rish, I · 2019
Earlier work this paper cites.
Quantifying generalization in reinforcement learning
Cobbe, K., Klimov, O., Hesse, C., Kim, T., and Schulman, J · 2019
Earlier work this paper cites.
A generalized algorithm for multi-objective reinforcement learning and policy adaptation
Yang, R., Sun, X., and Narasimhan, K · 2019
Earlier work this paper cites.
Generalization to new actions in reinforcement learning
Jain, A., Szot, A., and Lim, J. J · 2020
Earlier work this paper cites.
Reinforcement learning with augmented data
Laskin, M., Lee, K., Stooke, A., Pinto, L., Abbeel, P., and Srinivas, A · 2020
Earlier work this paper cites.
Improving generalization in reinforcement learning with mixture regularization
Wang, K., Kang, B., Shao, J., and Feng, J · 2020
Earlier work this paper cites.
Prediction-guided multi-objective reinforcement learning for continuous robot control
Xu, J., Tian, Y., Ma, P., Rus, D., Sueda, S., and Matusik, W · 2020
Earlier work this paper cites.
Learning invariant representations for reinforcement learning without reconstruction
Zhang, A., McAllister, R., Calandra, R., Gal, Y., and Levine, S · 2020
Cited alongside, same era.
Contrastive behavioral similarity embeddings for generalization in reinforcement learning
Agarwal, R., Machado, M. C., Castro, P. S., and Bellemare, M. G · 2021
Cited alongside, same era.
Learning one representation to optimize all rewards
Touati, A. and Ollivier, Y · 2021
Cited alongside, same era.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Yarats, D., Kostrikov, I., and Fergus, R · 2021
Cited alongside, same era.
Generalization of reinforcement learning with policy-aware adversarial data augmentation
Zhang, H. and Guo, Y · 2021
Alpacaeval: An automatic evaluator of instruction-following models, 2023
Li, X., Zhang, T., Dubois, Y., Taori, R., Gulrajani, I., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Later among the works it cites.
Routing to the expert: Efficient reward-guided ensemble of large language models
Lu, K., Yuan, H., Lin, R., Lin, J., Yuan, Z., Zhou, C., and Zhou, J · 2023
Later among the works it cites.
Automix: Automatically mixing language models
Madaan, A., Aggarwal, P., Anand, A., Potharaju, S. P., Mishra, S., Zhou, P., Gupta, A., Rajagopal, D., Kappaganthu, K., Yang, Y., et al · 2023
Later among the works it cites.
Code llama: Open foundation models for code
Roziere, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan, X. E., Adi, Y., Liu, J., Sauvestre, R., Remez, T., et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Pd-morl: Preference-driven multi-objective reinforcement learning algorithm
Basaklar, T., Gumussoy, S., and Ogras, U. Y · 2022
Cited alongside, same era.
Contextualize me–the case for context in reinforcement learning
Benjamins, C., Eimer, T., Schubert, F., Mohan, A., Döhler, S., Biedenkapp, A., Rosenhahn, B., Hutter, F., and Lindauer, M · 2022
Cited alongside, same era.
Rt-1: Robotics transformer for real-world control at scale
Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Dabis, J., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Hsu, J., et al · 2022
Cited alongside, same era.
Deep reinforcement learning policies learn shared adversarial features across mdps
Korkmaz, E · 2022
Cited alongside, same era.
Holistic evaluation of language models
Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., et al · 2022
Cited alongside, same era.
Transformers are meta-reinforcement learners
Melo, L. C · 2022
Cited alongside, same era.
Robust reinforcement learning: A review of foundations and recent advances
Moos, J., Hansel, K., Abdulsamad, H., Stark, S., Clever, D., and Peters, J · 2022
Cited alongside, same era.
Shnitzer, T., Ou, A., Silva, M., Soule, K., Sun, Y., Solomon, J., Thompson, N., and Yurochkin, M · 2023
Later among the works it cites.
Large language models encode clinical knowledge
Singhal, K., Azizi, S., Tu, T., Mahdavi, S. S., Wei, J., Chung, H. W., Scales, N., Tanwani, A., Cole-Lewis, H., Pfohl, S., et al · 2023
Later among the works it cites.
Fusing models with complementary expertise
Wang, H., Polo, F. M., Sun, Y., Kundu, S., Xing, E., and Yurochkin, M · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena, 2023
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., and Stoica, I · 2023
Later among the works it cites.
Starling-7b: Improving llm helpfulness & harmlessness with rlaif, November 2023
Zhu, B., Frick, E., Wu, T., Zhu, H., and Jiao, J · 2023
Later among the works it cites.
Hybrid llm: Cost-efficient and quality-aware query routing
Ding, D., Mallick, A., Wang, C., Sim, R., Mukherjee, S., Ruhle, V., Lakshmanan, L. V., and Awadallah, A. H · 2024
Later among the works it cites.
Open llm leaderboard v2
Fourrier, C., Habib, N., Lozovskaya, A., Szafer, K., and Wolf, T · 2024
Later among the works it cites.
Zero-shot reinforcement learning via function encoders
Ingebrand, T., Zhang, A., and Topcu, U · 2024
Later among the works it cites.
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al · 2024
Later among the works it cites.
A survey analyzing generalization in deep reinforcement learning
Korkmaz, E · 2024
Later among the works it cites.
Blending is all you need: Cheaper, better alternative to trillion-parameters llm
Lu, X., Liusie, A., Raina, V., Zhang, Y., and Beauchamp, W · 2024
Later among the works it cites.
Metallm: A high-performant and cost-efficient dynamic framework for wrapping llms
Nguyen, Q. H., Hoang, D. C., Decugis, J., Manchanda, S., Chawla, N. V., and Doan, K. D · 2024
Later among the works it cites.
Routellm: Learning to route llms with preference data
Ong, I., Almahairi, A., Wu, V., Chiang, W.-L., Wu, T., Gonzalez, J. E., Kadous, M. W., and Stoica, I · 2024
Later among the works it cites.
Optimising calls to large language models with uncertainty-based two-tier selection
Ramírez, G., Birch, A., and Titov, I · 2024
Later among the works it cites.
Fly-swat or cannon? cost-effective language model choice via meta-modeling
Šakota, M., Peyrard, M., and West, R · 2024
Later among the works it cites.
Learning pareto set for multi-objective continuous robot control
Shu, T., Shang, K., Gong, C., Nan, Y., and Ishibuchi, H · 2024
Later among the works it cites.