Fetching the paper…
Reading the bibliography…
Eliciting reasoning capabilities from language models (LMs) is a critical direction on the path towards building intelligent systems.
Neural execution of graph algorithms
Veličković, P., Ying, R., Padovano, M., Hadsell, R., and Blundell, C · 2019
Earlier work this paper cites.
What can neural networks reason about?
Xu, K., Li, J., Zhang, M., Du, S. S., Kawarabayashi, K.-i., and Jegelka, S · 2019
Earlier work this paper cites.
Logiqa: A challenge dataset for machine reading comprehension with logical reasoning
Liu, J., Cui, L., Liu, H., Huang, D., Wang, Y., and Zhang, Y · 2020
Earlier work this paper cites.
Towards scale-invariant graph-related problem solving by iterative homogeneous gnns
Tang, H., Huang, Z., Gu, J., Lu, B.-L., and Su, H · 2020
Earlier work this paper cites.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Finetuned language models are zero-shot learners
Wei, J., Bosma, M., Zhao, V. Y., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V · 2021
Earlier work this paper cites.
Exploring length generalization in large language models
Anil, C., Wu, Y., Andreassen, A., Lewkowycz, A., Misra, V., Ramasesh, V., Slone, A., Gur-Ari, G., Dyer, E., and Neyshabur, B · 2022
Earlier work this paper cites.
Introduction to algorithms
Cormen, T. H., Leiserson, C. E., Rivest, R. L., and Stein, C · 2022
Earlier work this paper cites.
Neural networks and the chomsky hierarchy
Delétang, G., Ruoss, A., Grau-Moya, J., Genewein, T., Wenliang, L. K., Catt, E., Cundy, C., Hutter, M., Legg, S., Veness, J., et al · 2022
Earlier work this paper cites.
A generalist neural algorithmic learner
Ibarz, B., Kurin, V., Papamakarios, G., Nikiforou, K., Bennani, M., Csordás, R., Dudzik, A. J., Bošnjak, M., Vitvitskyi, A., Rubanova, Y., et al · 2022
Earlier work this paper cites.
The CLRS Algorithmic Reasoning Benchmark
Veličković, P., Badia, A. P., Budden, D., Pascanu, R., Banino, A., Dashevskiy, M., Hadsell, R., and Blundell, C · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Cited alongside, same era.
Teaching algorithmic reasoning via in-context learning
Zhou, H., Nova, A., Larochelle, H., Courville, A., Neyshabur, B., and Sedghi, H · 2022
Cited alongside, same era.
Transformers can achieve length generalization but not robustly
Zhou, Y., Alon, U., Chen, X., Wang, X., Agarwal, R., and Zhou, D · 2022
Cited alongside, same era.
The reversal curse: Llms trained on" a is b" fail to learn" b is a"
Berglund, L., Tong, M., Kaufmann, M., Balesni, M., Stickland, A. C., Korbak, T., and Evans, O · 2023
What algorithms can transformers learn? a study in length generalization
Zhou, H., Bradley, A., Littwin, E., Razin, N., Saremi, O., Susskind, J., Bengio, S., and Nakkiran, P · 2023
Later among the works it cites.
On the markov property of neural algorithmic reasoning: Analyses and methods
Bohde, M., Liu, M., Saxton, A., and Ji, S · 2024
Closest in time.
Faith and fate: Limits of transformers on compositionality
Dziri, N., Lu, X., Sclar, M., Li, X. L., Jiang, L., Lin, B. Y., Welleck, S., West, P., Bhagavatula, C., Le Bras, R., et al · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Gemma Team, Mesnard, T., Hardin, C., Dadashi, R., Bhupatiraju, S., Pathak, S., Sifre, L., Rivière, M., Kale, M. S., Love, J., et al · 2024
Closest in time.
Neural algorithmic reasoning for combinatorial optimisation
Georgiev, D. G., Numeroso, D., Bacciu, D., and Liò, P · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Neural algorithmic reasoning with causal regularisation
Bevilacqua, B., Nikiforou, K., Ibarz, B., Bica, I., Paganini, M., Blundell, C., Mitrovic, J., and Veličković, P · 2023
Cited alongside, same era.
The expresssive power of transformers with chain of thought
Merrill, W. and Sabharwal, A · 2023
Cited alongside, same era.
Salsa-clrs: A sparse and scalable benchmark for algorithmic reasoning
Minder, J., Grötschla, F., Mathys, J., and Wattenhofer, R · 2023
Cited alongside, same era.
OpenAI · 2023
Cited alongside, same era.
Gpqa: A graduate-level google-proof q&a benchmark
Rein, D., Hou, B. L., Stickland, A. C., Petty, J., Pang, R. Y., Dirani, J., Michael, J., and Bowman, S. R · 2023
Cited alongside, same era.
Randomized positional encodings boost length generalization of transformers
Ruoss, A., Delétang, G., Genewein, T., Grau-Moya, J., Csordás, R., Bennani, M., Legg, S., and Veness, J · 2023
Cited alongside, same era.
Positional description matters for transformers arithmetic
Shen, R., Bubeck, S., Eldan, R., Lee, Y. T., Li, Y., and Zhang, Y · 2023
Cited alongside, same era.
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Google · 2024
Closest in time.
Recursive algorithmic reasoning
Jürß, J., Jayalath, D. H., and Veličković, P · 2024
Closest in time.
Benchmarking chatgpt on algorithmic reasoning
McLeish, S., Schwarzschild, A., and Goldstein, T · 2024
Closest in time.
Beyond lines and circles: Unveiling the geometric reasoning gap in large language models
Mouselinos, S., Michalewski, H., and Malinowski, M · 2024
Closest in time.
Understanding transformer reasoning capabilities via graph algorithms
Sanford, C., Fatemi, B., Hall, E., Tsitsulin, A., Kazemi, M., Halcrow, J., Perozzi, B., and Mirrokni, V · 2024
Closest in time.
Toolformer: Language models can teach themselves to use tools
Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Hambro, E., Zettlemoyer, L., Cancedda, N., and Scialom, T · 2024
Closest in time.
Functional benchmarks for robust evaluation of reasoning performance, and the reasoning gap
Srivastava, S., PV, A., Menon, S., Sukumar, A., Philipose, A., Prince, S., Thomas, S., et al · 2024
Closest in time.