Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have made substantial strides in structured tasks through Reinforcement Learning (RL), demonstrating proficiency in mathematical reasoning and code generation.
Natural language processing: a historical review
Jones, K. S · 1994
Earlier work this paper cites.
Absolute identification by relative judgment
Stewart, N., Brown, G. D., and Chater, N · 2005
Earlier work this paper cites.
Introduction to modern information retrieval
Chowdhury, G. G · 2010
Earlier work this paper cites.
Do user preferences and evaluation measures line up?
Sanderson, M., Paramita, M. L., Clough, P., and Kanoulas, E · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Representation learning: A review and new perspectives
Bengio, Y., Courville, A., and Vincent, P · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Mikolov, T · 2013
Earlier work this paper cites.
Glove: Global vectors for word representation
Pennington, J., Socher, R., and Manning, C. D · 2014
Earlier work this paper cites.
Skip-thought vectors
Kiros, R., Zhu, Y., Salakhutdinov, R. R., Zemel, R., Urtasun, R., Torralba, A., and Fidler, S · 2015
Earlier work this paper cites.
Deep learning
LeCun, Y., Bengio, Y., and Hinton, G · 2015
Earlier work this paper cites.
Deep learning , volume 1
Goodfellow, I., Bengio, Y., Courville, A., and Bengio, Y · 2016
Earlier work this paper cites.
Relative judgement is relatively difficult: Evidence against the role of relative judgement in absolute identification
Guest, D., Adelman, J. S., and Kent, C · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Van Den Oord, A., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., Kavukcuoglu, K., et al · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Lightgbm: A highly efficient gradient boosting decision tree
Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., Ye, Q., and Liu, T.-Y · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A · 2017
Earlier work this paper cites.
Cer, D · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J · 2018
Earlier work this paper cites.
Universal language model fine-tuning for text classification
Howard, J. and Ruder, S · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A · 2018
Cited alongside, same era.
Deep learning for time series classification: a review
Ismail Fawaz, H., Forestier, G., Weber, J., Idoumghar, L., and Muller, P.-A · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
Hierarchical multi-scale gaussian transformer for stock movement prediction
Ding, Q., Wu, S., Sun, H., Guo, J., and Guo, J · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Cited alongside, same era.
Learning to summarize with human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F · 2020
Beyond one-preference-for-all: Multi-objective direct preference optimization
Zhou, Z., Liu, J., Yang, C., Shao, J., Liu, Y., Yue, X., Ouyang, W., and Qiao, Y · 2023
Later among the works it cites.
Scalable ensembling for mitigating reward overoptimisation
Ahmed, A. M., Rafailov, R., Sharkov, S., Li, X., and Koyejo, S · 2024
Later among the works it cites.
A survey on evaluation of large language models
Chang, Y., Wang, X., Wang, J., Wu, Y., Yang, L., Zhu, K., Chen, H., Yi, X., Wang, C., Wang, Y., et al · 2024
Later among the works it cites.
Rlhf workflow: From reward modeling to online rlhf
Dong, H., Xiong, W., Pang, B., Wang, H., Zhao, H., Zhou, Y., Jiang, N., Sahoo, D., Xiong, C., and Zhang, T · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Tabnet: Attentive interpretable tabular learning
Arik, S. Ö. and Pfister, T · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al · 2022
Cited alongside, same era.
Why do tree-based models still outperform deep learning on typical tabular data?
Grinsztajn, L., Oyallon, E., and Varoquaux, G · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Later among the works it cites.
Jaech, A., Kalai, A., Lerer, A., Richardson, A., El-Kishky, A., Low, A., Helyar, A., Madry, A., Beutel, A., Carney, A., et al · 2024
Later among the works it cites.
Rewardbench: Evaluating reward models for language modeling, 2024
Lambert, N., Pyatkin, V., Morrison, J., Miranda, L., Lin, B. Y., Chandu, K., Dziri, N., Kumar, S., Zick, T., Choi, Y., Smith, N. A., and Hajishirzi, H · 2024
Later among the works it cites.
Mitigating the alignment tax of rlhf
Lin, Y., Lin, H., Xiong, W., Diao, S., Liu, J., Zhang, J., Pan, R., Wang, H., Hu, W., Zhang, H., et al · 2024
Later among the works it cites.
Skywork-reward: Bag of tricks for reward modeling in llms
Liu, C. Y., Zeng, L., Liu, J., Yan, R., He, J., Wang, C., Yan, S., Liu, Y., and Zhou, Y · 2024
Later among the works it cites.
Uncertainty-aware reward model: Teaching reward models to know what is unknown
Lou, X., Yan, D., Shen, W., Yan, Y., Xie, J., and Zhang, J · 2024
Later among the works it cites.
Mahan, D., Van Phung, D., Rafailov, R., Blagden, C., Lile, N., Castricato, L., Fränken, J.-P., Finn, C., and Albalak, A · 2024
Later among the works it cites.
Active preference learning for large language models
Muldrew, W., Hayes, P., Zhang, M., and Barber, D · 2024
Later among the works it cites.
Offsetbias: Leveraging debiased data for tuning evaluators
Park, J., Jwa, S., Ren, M., Kim, D., and Choi, S · 2024
Later among the works it cites.
Gemma: Open models based on gemini research and technology
Team, G., Mesnard, T., Hardin, C., Dadashi, R., Bhupatiraju, S., Pathak, S., Sifre, L., Rivière, M., Kale, M. S., Love, J., et al · 2024
Later among the works it cites.
Solving olympiad geometry without human demonstrations
Trinh, T. H., Wu, Y., Le, Q. V., He, H., and Luong, T · 2024
Later among the works it cites.
Metametrics: Calibrating metrics for generation tasks using human preferences
Winata, G. I., Anugraha, D., Susanto, L., Kuwanto, G., and Wijaya, D. T · 2024
Later among the works it cites.
Yin, Y., Wang, Z., Gu, Y., Huang, H., Chen, W., and Zhou, M · 2024
Later among the works it cites.
Cheating automatic llm benchmarks: Null models achieve high win rates
Zheng, X., Pang, T., Du, C., Liu, Q., Jiang, J., and Lin, M · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al · 2025
Closest in time.