Fetching the paper…
Reading the bibliography…
While designing inductive bias in neural architectures has been widely studied, we hypothesize that transformer networks are flexible enough to learn inductive bias from suitable generic tasks.
Roberta: A robustly optimized BERT pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 1907
Earlier work this paper cites.
Enhancing the transformer with explicit relational encoding for math problem solving
Schlag, I., Smolensky, P., Fernandez, R., Jojic, N., Schmidhuber, J., and Gao, J · 1910
Earlier work this paper cites.
Reasoning and the logic of things: The Cambridge conferences lectures of 1898
Peirce, C. S · 1992
Earlier work this paper cites.
Long Short-Term Memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Object recognition with gradient-based learning
LeCun, Y., Haffner, P., Bottou, L., and Bengio, Y · 1999
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2005
Earlier work this paper cites.
The graph neural network model
Scarselli, F., Gori, M., Tsoi, A. C., Hagenbuchner, M., and Monfardini, G · 2008
Earlier work this paper cites.
Generative Language Modeling for Automated Theorem Proving
Polu, S. and Sutskever, I · 2009
Earlier work this paper cites.
Charles Sanders Peirce: Logic
Bellucci, F. and Pietarinen, A.-V · 2015
Earlier work this paper cites.
The lean theorem prover (system description)
de Moura, L. M., Kong, S., Avigad, J., van Doorn, F., and von Raumer, J · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Luong, T., Pham, H., and Manning, C. D · 2015
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z · 2016
Earlier work this paper cites.
Attention is All you Need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Premise selection for theorem proving by deep graph embedding
Wang, M., Tang, Y., Wang, J., and Deng, J · 2017
Earlier work this paper cites.
Can Neural Networks Understand Logical Entailment?
Evans, R., Saxton, D., Amos, D., Kohli, P., and Grefenstette, E · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2018
Earlier work this paper cites.
HOList: An Environment for Machine Learning of Higher Order Logic Theorem Proving
Bansal, K., Loos, S. M., Rabe, M. N., Szegedy, C., and Wilcox, S · 2019
Earlier work this paper cites.
Learning to Reason in Large Theories without Imitation
Bansal, K., Szegedy, C., Rabe, M. N., Loos, S. M., and Toman, V · 2019
Earlier work this paper cites.
Buzzard, K., Hughes, C., Lau, K., Livingston, A., Mir, R. F., and Morrison, S · 2019
Earlier work this paper cites.
Cross-lingual Language Model Pretraining
Conneau, A. and Lample, G · 2019
Cited alongside, same era.
Improving Graph Neural Network Representations of Logical Formulae with Subgraph Pooling
Crouse, M., Abdelaziz, I., Cornelio, C., Thost, V., Wu, L., Forbus, K., and Fokoue, A · 2019
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Unified Language Model Pre-training for Natural Language Understanding and Generation
Dong, L., Yang, N., Wang, W., Wei, F., Liu, X., Wang, Y., Gao, J., Zhou, M., and Hon, H · 2019
Cited alongside, same era.
GamePad: A learning environment for theorem proving
Huang, D., Dhariwal, P., Song, D., and Sutskever, I · 2019
Cited alongside, same era.
Deep learning for symbolic mathematics
Lample, G. and Charton, F · 2020
Later among the works it cites.
Learning heuristics for quantified boolean formulas through reinforcement learning
Lederman, G., Rabe, M., Seshia, S., and Lee, E. A · 2020
Later among the works it cites.
The lean mathematical library
mathlib · 2020
Later among the works it cites.
Universal linguistic inductive biases via meta-learning
McCoy, R. T., Grant, E., Smolensky, P., Griffiths, T., and Linzen, T · 2020
Later among the works it cites.
Graph representations for higher-order logic and theorem proving
Paliwal, A., Loos, S. M., Rabe, M. N., Bansal, K., and Szegedy, C · 2020
Later among the works it cites.
Learning Music Helps You Read: Using transfer to study linguistic structure in language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ott, M., Edunov, S., Baevski, A., Fan, A., Gross, S., Ng, N., Grangier, D., and Auli, M · 2019
Cited alongside, same era.
Analysing mathematical reasoning abilities of neural models
Saxton, D., Grefenstette, E., Hill, F., and Kohli, P · 2019
Cited alongside, same era.
Guiding High-Performance SAT solvers with Unsat-Core Predictions
Selsam, D. and Bjørner, N · 2019
Cited alongside, same era.
Learning a SAT solver from single-bit supervision
Selsam, D., Lamm, M., Bünz, B., Liang, P., de Moura, L., and Dill, D. L · 2019
Cited alongside, same era.
MASS: masked sequence to sequence pre-training for language generation
Song, K., Tan, X., Qin, T., Lu, J., and Liu, T · 2019
Cited alongside, same era.
Learning to Prove Theorems via Interacting with Proof Assistants
Yang, K. and Deng, J · 2019
Cited alongside, same era.
Xlnet: Generalized autoregressive pretraining for language understanding
Yang, Z., Dai, Z., Yang, Y., Carbonell, J., Salakhutdinov, R. R., and Le, Q. V · 2019
Cited alongside, same era.
Papadimitriou, I. and Jurafsky, D · 2020
Later among the works it cites.
Guiding Inferences in Connection Tableau by Recurrent Neural Networks
Piotrowski, B. and Urban, J · 2020
Later among the works it cites.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Later among the works it cites.
Liquid tensor experiment
Scholze, P · 2020
Later among the works it cites.
First Neural Conjecturing Datasets and Experiments
Urban, J. and Jakubův, J · 2020
Later among the works it cites.
Exploration of neural machine translation in autoformalization of mathematics in mizar
Wang, Q., Brown, C., Kaliszyk, C., and Urban, J · 2020
Later among the works it cites.
Can neural networks acquire a structural bias from raw linguistic data?
Warstadt, A. and Bowman, S. R · 2020
Later among the works it cites.
What can neural networks reason about?
Xu, K., Li, J., Zhang, M., Du, S. S., Kawarabayashi, K.-i., and Jegelka, S · 2020
Later among the works it cites.
PEGASUS: pre-training with extracted gap-sentences for abstractive summarization
Zhang, J., Zhao, Y., Saleh, M., and Liu, P. J · 2020
Later among the works it cites.
Proof artifact co-training for theorem proving with language models
Han, J. M., Rute, J., Wu, Y., Ayers, E. W., and Polu, S · 2021
Closest in time.
Isarstep: a benchmark for high-level mathematical reasoning
Li, W., Yu, L., Wu, Y., and Paulson, L. C · 2021
Closest in time.
Mathematical Reasoning via Self-supervised Skip-tree Training
Rabe, M. N., Lee, D., Bansal, K., and Szegedy, C · 2021
Closest in time.
Learning Branching Heuristics for Propositional Model Counting
Vaezipoor, P., Lederman, G., Wu, Y., Maddison, C. J., Grosse, R. B., Lee, E. A., Seshia, S. A., and Bacchus, F · 2021
Closest in time.
INT: An Inequality Benchmark for Evaluating Generalization in Theorem Proving
Wu, Y., Jiang, A., Ba, J., and Grosse, R · 2021
Closest in time.