Fetching the paper…
Reading the bibliography…
We introduce multiple physics pretraining (MPP), an autoregressive task-agnostic pretraining approach for physical surrogate modeling of spatiotemporal systems with transformers.
Stationary wave solutions of a system of reaction-diffusion equations derived from the fitzhugh–nagumo equations
Klaasen, G. A. and Troy, W. C · 1984
Earlier work this paper cites.
Numerical methods for conservation laws , volume 214
LeVeque, R. J. and Leveque, R. J · 1992
Earlier work this paper cites.
Partial differential equations for scientists and engineers
Farlow, S. J · 1993
Earlier work this paper cites.
Turbulence in two dimensions
Ouellette, N. T · 2012
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O., Fischer, P., and Brox, T · 2015
Earlier work this paper cites.
Gaussian error linear units (gelus), 2016
Hendrycks, D. and Gimpel, K · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Carreira, J. and Zisserman, A · 2017
Earlier work this paper cites.
The ”something something” video database for learning and evaluating visual common sense, 2017
Goyal, R., Kahou, S. E., Michalski, V., Materzyńska, J., Westphal, S., Kim, H., Haenel, V., Fruend, I., Yianilos, P., Mueller-Freitag, M., Hoppe, F., Thurau, C., Bax, I., and Memisevic, R · 2017
Earlier work this paper cites.
The kinetics human action video dataset, 2017
Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Viola, F., Green, T., Back, T., Natsev, P., Suleyman, M., and Zisserman, A · 2017
Earlier work this paper cites.
Instance normalization: The missing ingredient for fast stylization, 2017
Ulyanov, D., Vedaldi, A., and Lempitsky, V · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Solving high-dimensional partial differential equations using deep learning
Han, J., Jentzen, A., and E, W · 2018
Earlier work this paper cites.
Deep learning for universal linear embeddings of nonlinear dynamics
Lusch, B., Kutz, J. N., and Brunton, S. L · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al · 2018
Earlier work this paper cites.
Dgm: A deep learning algorithm for solving partial differential equations
Sirignano, J. and Spiliopoulos, K · 2018
Earlier work this paper cites.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification, 2018
Xie, S., Sun, C., Huang, J., Tu, Z., and Murphy, K · 2018
Earlier work this paper cites.
The deep ritz method: a deep learning-based numerical algorithm for solving variational problems
Yu, B. et al · 2018
Earlier work this paper cites.
Unsupervised deep learning algorithm for pde-based forward and inverse problems
Bar, L. and Sochen, N · 2019
Earlier work this paper cites.
Turbulence modeling in the age of data
Duraisamy, K., Iaccarino, G., and Xiao, H · 2019
Earlier work this paper cites.
Learning to predict the cosmological structure formation
He, S., Li, Y., Feng, Y., Ho, S., Ravanbakhsh, S., Chen, W., and Póczos, B · 2019
Earlier work this paper cites.
Axial attention in multidimensional transformers, 2019
Ho, J., Kalchbrenner, N., Weissenborn, D., and Salimans, T · 2019
Earlier work this paper cites.
Ccnet: Criss-cross attention for semantic segmentation
Huang, Z., Wang, X., Huang, L., Huang, C., Wei, Y., and Liu, W · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Earlier work this paper cites.
Lu, L., Jin, P., and Karniadakis, G. E · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations
Raissi, M., Perdikaris, P., and Karniadakis, G. E · 2019
Earlier work this paper cites.
Experiment tracking with weights and biases, 2020
Biewald, L · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations, 2020
Chen, T., Kornblith, S., Norouzi, M., and Hinton, G · 2020
Earlier work this paper cites.
Chemberta: Large-scale self-supervised pretraining for molecular property prediction, 2020
Chithrananda, S., Grand, G., and Ramsundar, B · 2020
Cited alongside, same era.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Cited alongside, same era.
Fourier neural operator for parametric partial differential equations
Li, Z., Kovachki, N., Azizzadenesheli, K., Liu, B., Bhattacharya, K., Stuart, A., and Anandkumar, A · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer, 2020
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Cited alongside, same era.
On layer normalization in the transformer architecture, 2020
Xiong, R., Yang, Y., He, D., Zheng, K., Zheng, S., Xing, C., Zhang, H., Lan, Y., Wang, L., and Liu, T.-Y · 2020
Pathak, J., Subramanian, S., Harrington, P., Raja, S., Chattopadhyay, A., Mardani, M., Kurth, T., Hall, D., Li, Z., Azizzadenesheli, K., Hassanzadeh, P., Kashinath, K., and Anandkumar, A · 2022
Later among the works it cites.
Learned coarse models for efficient turbulence simulation, 2022
Stachenfeld, K., Fielding, D. B., Kochkov, D., Cranmer, M., Pfaff, T., Godwin, J., Cui, C., Ho, S., Battaglia, P., and Sanchez-Gonzalez, A · 2022
Later among the works it cites.
PDEBench: An Extensive Benchmark for Scientific Machine Learning
Takamoto, M., Praditia, T., Leiteritz, R., MacKinlay, D., Alesiani, F., Pflüger, D., and Niepert, M · 2022
Later among the works it cites.
VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Tong, Z., Song, Y., Wang, J., and Wang, L · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Gradient surgery for multi-task learning, 2020
Yu, T., Kumar, S., Gupta, A., Levine, S., Hausman, K., and Finn, C · 2020
Cited alongside, same era.
Weak adversarial networks for high-dimensional partial differential equations
Zang, Y., Bao, G., Ye, X., and Zhou, H · 2020
Cited alongside, same era.
Vivit: A video vision transformer, 2021
Arnab, A., Dehghani, M., Heigold, G., Sun, C., Lučić, M., and Schmid, C · 2021
Cited alongside, same era.
Is space-time attention all you need for video understanding?
Bertasius, G., Wang, H., and Torresani, L · 2021
Cited alongside, same era.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al · 2021
Cited alongside, same era.
Choose a transformer: Fourier or galerkin, 2021
Cao, S · 2021
Cited alongside, same era.
A bayesian neural network predicts the dissolution of compact planetary systems
Cranmer, M., Tamayo, D., Rein, H., Battaglia, P., Hadden, S., Armitage, P. J., Ho, S., and Spergel, D. N · 2021
Cited alongside, same era.
Touvron, H., Cord, M., El-Nouby, A., Verbeek, J., and Jegou, H · 2022
Later among the works it cites.
Emergent abilities of large language models
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., et al · 2022
Later among the works it cites.
Scaling vision transformers
Zhai, X., Kolesnikov, A., Houlsby, N., and Beyer, L · 2022
Later among the works it cites.
A cookbook of self-supervised learning, 2023
Balestriero, R., Ibrahim, M., Sobal, V., Morcos, A., Shekhar, S., Goldstein, T., Bordes, F., Bardes, A., Mialon, G., Tian, Y., Schwarzschild, A., Wilson, A. G., Geiping, J., Garrido, Q., Fernandez, P., Bar, A., Pirsiavash, H., LeCun, Y., and Goldblum, M · 2023
Closest in time.
The rise of data-driven weather forecasting, 2023
Ben-Bouallegue, Z., Clare, M. C. A., Magnusson, L., Gascon, E., Maier-Gerber, M., Janousek, M., Rodwell, M., Pinault, F., Dramsch, J. S., Lang, S. T. K., Raoult, B., Rabier, F., Chevallier, M., Sandu, I., Dueben, P., Chantry, M., and Pappenberger, F · 2023
Closest in time.
Accurate medium-range global weather forecasting with 3d neural networks
Bi, K., Xie, L., Zhang, H., Chen, X., Gu, X., and Tian, Q · 2023
Closest in time.
Chemcrow: Augmenting large-language models with chemistry tools, 2023
Bran, A. M., Cox, S., White, A. D., and Schwaller, P · 2023
Closest in time.
Learning-rate-free learning by d-adaptation, 2023
Defazio, A. and Mishchenko, K · 2023
Closest in time.
Scaling vision transformers to 22 billion parameters, 2023
Dehghani, M., Djolonga, J., Mustafa, B., Padlewski, P., Heek, J., Gilmer, J., Steiner, A., Caron, M., Geirhos, R., Alabdulmohsin, I., Jenatton, R., Beyer, L., Tschannen, M., Arnab, A., Wang, X., Riquelme, C., Minderer, M., Puigcerver, J., Evci, U., Kumar, M., van Steenkiste, S., Elsayed, G. F., Mahendran, A., Yu, F., Oliver, A., Huot, F., Bastings, J., Collier, M. P., Gritsenko, A., Birodkar, V., Vasconcelos, C., Tay, Y., Mensink, T., Kolesnikov, A., Pavetić, F., Tran, D., Kipf, T., Lučić, M., Zhai, X., Keysers, D., Harmsen, J., and Houlsby, N · 2023
Closest in time.
Learning to correct spectral methods for simulating turbulent flows
Dresdner, G., Kochkov, D., Norgaard, P. C., Zepeda-Nunez, L., Smith, J., Brenner, M., and Hoyer, S · 2023
Closest in time.
Can physics-informed neural networks beat the finite element method?, 2023
Grossmann, T. G., Komorowska, U. J., Latz, J., and Schönlieb, C.-B · 2023
Closest in time.
Field-level neural network emulator for cosmological n-body simulations
Jamieson, D., Li, Y., de Oliveira, R. A., Villaescusa-Navarro, F., Ho, S., and Spergel, D. N · 2023
Closest in time.
Health system-scale language models are all-purpose prediction engines
Jiang, L. Y., Liu, X. C., Nejatian, N. P., Nasir-Moin, M., Wang, D., Abidin, A., Eaton, K., Riina, H. A., Laufer, I., Punjabi, P., et al · 2023
Closest in time.
Neural operator: Learning maps between function spaces, 2023
Kovachki, N., Li, Z., Liu, B., Azizzadenesheli, K., Bhattacharya, K., Stuart, A., and Anandkumar, A · 2023
Closest in time.
Learning skillful medium-range global weather forecasting
Lam, R., Sanchez-Gonzalez, A., Willson, M., Wirnsberger, P., Fortunato, M., Alet, F., Ravuri, S., Ewalds, T., Eaton-Rosen, Z., Hu, W., Merose, A., Hoyer, S., Holland, G., Vinyals, O., Stott, J., Pritzel, A., Mohamed, S., and Battaglia, P · 2023
Closest in time.
Towards an astronomical foundation model for stars with a transformer-based model, 2023
Leung, H. W. and Bovy, J · 2023
Closest in time.
Self-supervised learning with lie symmetries for partial differential equations, 2023
Mialon, G., Garrido, Q., Lawrence, H., Rehman, D., LeCun, Y., and Kiani, B. T · 2023
Closest in time.
Cross-modal fine-tuning: Align then refine, 2023
Shen, J., Li, L., Dery, L. M., Staten, C., Khodak, M., Neubig, G., and Talwalkar, A · 2023
Closest in time.
Deep learning closure models for large-eddy simulation of flows around bluff bodies
Sirignano, J. and MacArt, J. F · 2023
Closest in time.
Explaining the physics of transfer learning in data-driven turbulence modeling
Subel, A., Guan, Y., Chattopadhyay, A., and Hassanzadeh, P · 2023
Closest in time.
Towards foundation models for scientific machine learning: Characterizing scaling and transfer behavior, 2023
Subramanian, S., Harrington, P., Keutzer, K., Bhimji, W., Morozov, D., Mahoney, M., and Gholami, A · 2023
Closest in time.
Learning neural pde solvers with parameter-guided channel attention, 2023
Takamoto, M., Alesiani, F., and Niepert, M · 2023
Closest in time.
Towards generalist biomedical ai, 2023
Tu, T., Azizi, S., Driess, D., Schaekermann, M., Amin, M., Chang, P.-C., Carroll, A., Lau, C., Tanno, R., Ktena, I., Mustafa, B., Chowdhery, A., Liu, Y., Kornblith, S., Fleet, D., Mansfield, P., Prakash, S., Wong, R., Virmani, S., Semturs, C., Mahdavi, S. S., Green, B., Dominowska, E., y Arcas, B. A., Barral, J., Webster, D., Corrado, G. S., Matias, Y., Singhal, K., Florence, P., Karthikesalingam, A., and Natarajan, V · 2023
Closest in time.
Adan: Adaptive nesterov momentum algorithm for faster optimizing deep models, 2023
Xie, X., Zhou, P., Li, H., Lin, Z., and Yan, S · 2023
Closest in time.
Transfer learning enhanced deeponet for long-time prediction of evolution equations
Xu, W., Lu, Y., and Wang, L · 2023
Closest in time.
In-context operator learning for differential equation problems
Yang, L., Liu, S., Meng, T., and Osher, S. J · 2023
Closest in time.