Fetching the paper…
Reading the bibliography…
In this paper, we propose R$^3$: Learning Reasoning through Reverse Curriculum Reinforcement Learning (RL), a novel method that employs only outcome supervision to achieve the benefits of process supervision for large language models.
Defeasible reasoning
Pollock, J. L · 1987
Earlier work this paper cites.
Introduction to reinforcement learning , volume 135
Sutton, R. S., Barto, A. G., et al · 1998
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. M. and Langford, J · 2002
Earlier work this paper cites.
Learning to summarize from human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D. M., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F · 2009
Earlier work this paper cites.
Dynamic programming and optimal control: Volume I , volume 4
Bertsekas, D · 2012
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Bowman, S. R., Angeli, G., Potts, C., and Manning, C. D · 2015
Earlier work this paper cites.
Playing atari games with deep reinforcement learning and human checkpoint replay, 2016
Hosu, I.-A. and Rebedea, T · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T. P., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Exploration from Demonstration for Interactive Reinforcement Learning
Subramanian, K., Isbell, C. L., and Thomaz, A. L · 2016
Earlier work this paper cites.
Exploration from demonstration for interactive reinforcement learning
Subramanian, K., Jr., C. L. I., and Thomaz, A. L · 2016
Earlier work this paper cites.
Reverse curriculum generation for reinforcement learning
Florensa, C., Held, D., Wulfmeier, M., Zhang, M., and Abbeel, P · 2017
Earlier work this paper cites.
RACE: large-scale reading comprehension dataset from examinations
Lai, G., Xie, Q., Liu, H., Yang, Y., and Hovy, E. H · 2017
Earlier work this paper cites.
Data-efficient Deep Reinforcement Learning for Dexterous Manipulation, April 2017
Popov, I., Heess, N., Lillicrap, T., Hafner, R., Barth-Maron, G., Vecerik, M., Lampe, T., Tassa, Y., Erez, T., and Riedmiller, M · 2017
Earlier work this paper cites.
An overview of multi-task learning in deep neural networks
Ruder, S · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Seq2sql: Generating structured queries from natural language using reinforcement learning
Zhong, V., Xiong, C., and Socher, R · 2017
Earlier work this paper cites.
On a flexible representation for defeasible reasoning variants
Hecham, A., Bisquert, P., and Croitoru, M · 2018
Earlier work this paper cites.
Overcoming Exploration in Reinforcement Learning with Demonstrations, February 2018
Nair, A., McGrew, B., Andrychowicz, M., Zaremba, W., and Abbeel, P · 2018
Earlier work this paper cites.
DeepMimic: Example-Guided Deep Reinforcement Learning of Physics-Based Character Skills
Peng, X. B., Abbeel, P., Levine, S., and van de Panne, M · 2018
Earlier work this paper cites.
Multi-goal reinforcement learning: Challenging robotics environments and request for research, 2018
Plappert, M., Andrychowicz, M., Ray, A., McGrew, B., Baker, B., Powell, G., Schneider, J., Tobin, J., Chociej, M., Welinder, P., Kumar, V., and Zaremba, W · 2018
Earlier work this paper cites.
Learning Montezuma’s Revenge from a single demonstration , 2018
Salimans, T. and Chen, R · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Williams, A., Nangia, N., and Bowman, S. R · 2018
Earlier work this paper cites.
SWAG: A large-scale adversarial dataset for grounded commonsense inference
Zellers, R., Bisk, Y., Schwartz, R., and Choi, Y · 2018
Earlier work this paper cites.
Rethinking defeasible reasoning: A scalable approach
Maher, M. J., Tachmazidis, I., Antoniou, G., Wade, S., and Cheng, L · 2020
Earlier work this paper cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Rasley, J., Rajbhandari, S., Ruwase, O., and He, Y · 2020
Cited alongside, same era.
A general language assistant as a laboratory for alignment
Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., DasSarma, N., Elhage, N., Hatfield-Dodds, Z., Hernandez, D., Kernion, J., Ndousse, K., Olsson, C., Amodei, D., Brown, T. B., Clark, J., McCandlish, S., Olah, C., and Kaplan, J · 2021
Cited alongside, same era.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R. B., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., Brynjolfsson, E., Buch, S., Card, D., Castellon, R., Chatterji, N. S., Chen, A. S., Creel, K., Davis, J. Q., Demszky, D., Donahue, C., Doumbouya, M., Durmus, E., Ermon, S., Etchemendy, J., Ethayarajh, K., Fei-Fei, L., Finn, C., Gale, T., Gillespie, L., Goel, K., Goodman, N. D., Grossman, S., Guha, N., Hashimoto, T., Henderson, P., Hewitt, J., Ho, D. E., Hong, J., Hsu, K., Huang, J., Icard, T., Jain, S., Jurafsky, D., Kalluri, P., Karamcheti, S., Keeling, G., Khani, F., Khattab, O., Koh, P. W., Krass, M. S., Krishna, R., Kuditipudi, R., and et al · 2021
Cited alongside, same era.
Tora: A tool-integrated reasoning agent for mathematical problem solving
Gou, Z., Shao, Z., Gong, Y., Shen, Y., Yang, Y., Huang, M., Duan, N., and Chen, W · 2023
Later among the works it cites.
Reinforced self-training (rest) for language modeling
Gülçehre, Ç., Paine, T. L., Srinivasan, S., Konyushkova, K., Weerts, L., Sharma, A., Siddhant, A., Ahern, A., Wang, M., Gu, C., Macherey, W., Doucet, A., Firat, O., and de Freitas, N · 2023
Later among the works it cites.
Design of chain-of-thought in math problem solving
Jie, Z., Luong, T. Q., Zhang, X., Jin, X., and Li, H · 2023
Later among the works it cites.
Boardgameqa: A dataset for natural language reasoning with contradictory information
Kazemi, M., Yuan, Q., Bhatia, D., Kim, N., Xu, X., Imbrasaite, V., and Ramachandran, D · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al · 2021
Cited alongside, same era.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Cited alongside, same era.
Webgpt: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., Jiang, X., Cobbe, K., Eloundou, T., Krueger, G., Button, K., Knight, M., Chess, B., and Schulman, J · 2021
Cited alongside, same era.
Are NLP models really able to solve simple math word problems?
Patel, A., Bhattamishra, S., and Goyal, N · 2021
Cited alongside, same era.
Scaling language models: Methods, analysis & insights from training gopher
Rae, J. W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, H. F., Aslanides, J., Henderson, S., Ring, R., Young, S., Rutherford, E., Hennigan, T., Menick, J., Cassirer, A., Powell, R., van den Driessche, G., Hendricks, L. A., Rauh, M., Huang, P., Glaese, A., Welbl, J., Dathathri, S., Huang, S., Uesato, J., Mellor, J., Higgins, I., Creswell, A., McAleese, N., Wu, A., Elsen, E., Jayakumar, S. M., Buchatskaya, E., Budden, D., Sutherland, E., Simonyan, K., Paganini, M., Sifre, L., Martens, L., Li, X. L., Kuncoro, A., Nematzadeh, A., Gribovskaya, E., Donato, D., Lazaridou, A., Mensch, A., Lespiau, J., Tsimpoukelli, M., Grigorev, N., Fritz, D., Sottiaux, T., Pajarskas, M., Pohlen, T., Gong, Z., Toyama, D., de Masson d’Autume, C., Li, Y., Terzi, T., Mikulik, V., Babuschkin, I., Clark, A., de Las Casas, D., Guy, A., Jones, C., Bradbury, J., Johnson, M. J., Hechtman, B. A., Weidinger, L., Gabriel, I., Isaac, W., Lockhart, E., Osindero, S., Rimell, L., Dyer, C., Vinyals, O., Ayoub, K., Stanway, J., Bennett, L., Hassabis, D., Kavukcuoglu, K., and Irving, G · 2021
Cited alongside, same era.
Generate & rank: A multi-task framework for math word problems
Shen, J., Yin, Y., Li, L., Shang, L., Jiang, X., Zhang, M., and Liu, Q · 2021
Cited alongside, same era.
Trajectory-based split hindsight reverse curriculum learning
Wu, J., Zhang, D., Zhong, S., and Qiao, H · 2021
Cited alongside, same era.
A survey on multi-task learning
Zhang, Y. and Yang, Q · 2021
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., Joseph, N., Kadavath, S., Kernion, J., Conerly, T., Showk, S. E., Elhage, N., Hatfield-Dodds, Z., Hernandez, D., Hume, T., Johnston, S., Kravec, S., Lovitt, L., Nanda, N., Olsson, C., Amodei, D., Brown, T. B., Clark, J., McCandlish, S., Olah, C., Mann, B., and Kaplan, J · 2022
Cited alongside, same era.
The cot collection: Improving zero-shot and few-shot learning of language models via chain-of-thought fine-tuning
Kim, S., Joo, S. J., Kim, D., Jang, J., Ye, S., Shin, J., and Seo, M · 2023
Later among the works it cites.
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K · 2023
Later among the works it cites.
Reinforcement learning, bit by bit
Lu, X., Roy, B. V., Dwaracherla, V., Ibrahimi, M., Osband, I., and Wen, Z · 2023
Later among the works it cites.
Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct
Luo, H., Sun, Q., Xu, C., Zhao, P., Lou, J., Tao, C., Geng, X., Lin, Q., Chen, S., and Zhang, D · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Failure modes of learning reward models for llms and other sequence models
Pitis, S · 2023
Later among the works it cites.
Webcpm: Interactive web search for chinese long-form question answering
Qin, Y., Cai, Z., Jin, D., Yan, L., Liang, S., Zhu, K., Lin, Y., Han, X., Ding, N., Wang, H., Xie, R., Qi, F., Liu, Z., Sun, M., and Zhou, J · 2023
Later among the works it cites.
Code llama: Open foundation models for code
Rozière, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan, X. E., Adi, Y., Liu, J., Remez, T., Rapin, J., Kozhevnikov, A., Evtimov, I., Bitton, J., Bhatt, M., Canton-Ferrer, C., Grattafiori, A., Xiong, W., Défossez, A., Copet, J., Azhar, F., Touvron, H., Martin, L., Usunier, N., Scialom, T., and Synnaeve, G · 2023
Later among the works it cites.
Large language models can be easily distracted by irrelevant context
Shi, F., Chen, X., Misra, K., Scales, N., Dohan, D., Chi, E. H., Schärli, N., and Zhou, D · 2023
Later among the works it cites.
Alpaca: A strong, replicable instruction-following model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Canton-Ferrer, C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T · 2023
Later among the works it cites.
Self-polish: Enhance reasoning in large language models via problem refinement
Xi, Z., Jin, S., Zhou, Y., Zheng, R., Gao, S., Liu, J., Gui, T., Zhang, Q., and Huang, X · 2023
Later among the works it cites.
Self-evaluation guided beam search for reasoning
Xie, Y., Kawaguchi, K., Zhao, Y., Zhao, X., Kan, M.-Y., He, J., and Xie, Q · 2023
Later among the works it cites.
Scaling relationship on learning mathematical reasoning with large language models
Yuan, Z., Yuan, H., Li, C., Dong, G., Tan, C., and Zhou, C · 2023
Later among the works it cites.
Mammoth: Building math generalist models through hybrid instruction tuning
Yue, X., Qu, X., Zhang, G., Fu, Y., Huang, W., Sun, H., Su, Y., and Chen, W · 2023
Later among the works it cites.
Siren’s song in the ai ocean: A survey on hallucination in large language models
Zhang, Y., Li, Y., Cui, L., Cai, D., Liu, L., Fu, T., Huang, X., Zhao, E., Zhang, Y., Chen, Y., et al · 2023
Later among the works it cites.
Secrets of RLHF in large language models part I: PPO
Zheng, R., Dou, S., Gao, S., Hua, Y., Shen, W., Wang, B., Liu, Y., Jin, S., Liu, Q., Zhou, Y., Xiong, L., Chen, L., Xi, Z., Xu, N., Lai, W., Zhu, M., Chang, C., Yin, Z., Weng, R., Cheng, W., Huang, H., Sun, T., Yan, H., Gui, T., Zhang, Q., Qiu, X., and Huang, X · 2023
Later among the works it cites.
Zhou, A., Wang, K., Lu, Z., Shi, W., Luo, S., Qin, Z., Lu, S., Jia, A., Song, L., Zhan, M., and Li, H · 2023
Later among the works it cites.
Reft: Reasoning with reinforced fine-tuning, 2024
Luong, T. Q., Zhang, X., Jie, Z., Sun, P., Jin, X., and Li, H · 2024
Closest in time.