Fetching the paper…
Reading the bibliography…
Foundation models, such as large language models, have demonstrated success in addressing various language and image processing tasks.
Review: Jan lukasiewicz, jerzy slupecki, panstwowe wydawnictwo, remarks on nicod’s axiom and on ”generalizing deduction”
H. A. Pogorzelski · 1965
Earlier work this paper cites.
Interaction of ”solitons” in a collisionless plasma and the recurrence of initial states
N. J. Zabusky and M. D. Kruskal · 1965
Earlier work this paper cites.
Approximations of continuous functionals by neural networks with application to dynamic systems
Tianping Chen and Hong Chen · 1993
Earlier work this paper cites.
Universal approximation to nonlinear operators by neural networks with arbitrary activation functions and its application to dynamical systems
Tianping Chen and Hong Chen · 1995
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Alex Graves · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Improved semantic representations from tree-structured long short-term memory networks
Kai Sheng Tai, Richard Socher, and Christopher D. Manning · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio · 2015
Earlier work this paper cites.
Recurrent neural network grammars
Chris Dyer, Adhiguna Kuncoro, Miguel Ballesteros, and Noah A. Smith · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation, 2016
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Lukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean · 2016
Earlier work this paper cites.
Learning partial differential equations via data discovery and sparse optimization
Hayden Schaeffer · 2017
Earlier work this paper cites.
Sparse model selection via integral terms
Hayden Schaeffer and Scott G McCalla · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Extracting sparse high-dimensional dynamics from limited data
Hayden Schaeffer, Giang Tran, and Rachel Ward · 2018
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov · 2019
Earlier work this paper cites.
Physically consistent numerical solver for time-dependent fokker-planck equations
Viktor Holubec, Klaus Kroy, and Stefano Steffenoni · 2019
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang · 2019
Earlier work this paper cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Videobert: A joint model for video and language representation learning
Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid · 2019
Earlier work this paper cites.
Lxmert: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal · 2019
Earlier work this paper cites.
Multimodal transformer for unaligned multimodal language sequences
Yao-Hung Hubert Tsai, Shaojie Bai, Paul Pu Liang, J Zico Kolter, Louis-Philippe Morency, and Ruslan Salakhutdinov · 2019
Earlier work this paper cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E Peters, and Arman Cohan · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Gpt-3: Its nature, scope, limits, and consequences
Luciano Floridi and Massimo Chiriatti · 2020
Earlier work this paper cites.
Fourier neural operator for parametric partial differential equations
Zongyi Li, Nikola Kovachki, Kamyar Azizzadenesheli, Burigede Liu, Kaushik Bhattacharya, Andrew Stuart, and Anima Anandkumar · 2020
Cited alongside, same era.
Neural operator: Graph kernel network for partial differential equations
Zongyi Li, Nikola Kovachki, Kamyar Azizzadenesheli, Burigede Liu, Kaushik Bhattacharya, Andrew Stuart, and Anima Anandkumar · 2020
Cited alongside, same era.
Ai feynman: A physics-inspired method for symbolic regression
Silviu-Marian Udrescu and Max Tegmark · 2020
Cited alongside, same era.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Cited alongside, same era.
Vilt: Vision-and-language transformer without convolution or region supervision
Finite expression methods for discovering physical laws from data
Zhongyi Jiang, Chunmei Wang, and Haizhao Yang · 2023
Later among the works it cites.
Zhongyi Jiang, Min Zhu, Dongzhuo Li, Qiuzi Li, Yanhua O Yuan, and Lu Lu · 2023
Later among the works it cites.
Fourier neural operator with learned deformations for pdes on general geometries
Zongyi Li, Daniel Zhengyu Huang, Burigede Liu, and Anima Anandkumar · 2023
Later among the works it cites.
B-deeponet: An enhanced bayesian deeponet for solving noisy parametric pdes using accelerated replica exchange sgld
Guang Lin, Christian Moya, and Zecheng Zhang · 2023
Later among the works it cites.
Learning the dynamical response of nonlinear non-autonomous dynamical systems with deep operator neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wonjae Kim, Bokyung Son, and Ildoo Kim · 2021
Cited alongside, same era.
Ai choreographer: Music conditioned 3d dance generation with aist++
Ruilong Li, Shan Yang, David A Ross, and Angjoo Kanazawa · 2021
Cited alongside, same era.
Learning nonlinear operators via deeponet based on the universal approximation theorem of operators
Lu Lu, Pengzhan Jin, Guofei Pang, Zhongqiang Zhang, and George Em Karniadakis · 2021
Cited alongside, same era.
Deepm&mnet for hypersonics: Predicting the coupled flow and finite-rate chemistry behind a normal shock using neural-network approximation of operators
Zhiping Mao, Lu Lu, Olaf Marxen, Tamer A Zaki, and George Em Karniadakis · 2021
Cited alongside, same era.
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever · 2021
Cited alongside, same era.
Linear algebra with transformers
Francois Charton · 2022
Cited alongside, same era.
Deep symbolic regression for recurrent sequences
Stephane d’Ascoli, Pierre-Alexandre Kamienny, Guillaume Lample, and Francois Charton · 2022
Cited alongside, same era.
Multifidelity deep operator networks
Amanda A Howard, Mauro Perego, George E Karniadakis, and Panos Stinis · 2022
Cited alongside, same era.
Guang Lin, Christian Moya, and Zecheng Zhang · 2023
Later among the works it cites.
PROSE: Predicting operators and symbolic expressions using multimodal transformers
Yuxuan Liu, Zecheng Zhang, and Hayden Schaeffer · 2023
Later among the works it cites.
Ppdonet: Deep operator networks for fast prediction of steady-state solutions in disk–planet systems
Shunyuan Mao, Ruobing Dong, Lu Lu, Kwang Moo Yi, Sifan Wang, and Paris Perdikaris · 2023
Later among the works it cites.
Multiple physics pretraining for physical surrogate models
Michael McCabe, Bruno Régaldo-Saint Blancard, Liam Holden Parker, Ruben Ohana, Miles Cranmer, Alberto Bietti, Michael Eickenberg, Siavash Golkar, Geraud Krawezik, Francois Lanusse, et al · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
Multimodal learning with transformers: A survey
Peng Xu, Xiatian Zhu, and David A Clifton · 2023
Later among the works it cites.
In-context operator learning with data prompts for differential equation problems
Liu Yang, Siting Liu, Tingwei Meng, and Stanley J Osher · 2023
Later among the works it cites.
Prompting in-context operator learning with sensor data, equations, and natural language
Liu Yang, Tingwei Meng, Siting Liu, and Stanley J Osher · 2023
Later among the works it cites.
Belnet: basis enhanced learning, a mesh-free neural operator
Zecheng Zhang, Wing Tat Leung, and Hayden Schaeffer · 2023
Later among the works it cites.
A discretization-invariant extension and analysis of some deep operator networks
Zecheng Zhang, Wing Tat Leung, and Hayden Schaeffer · 2023
Later among the works it cites.
Bayesian deep operator learning for homogenized1 to fine-scale maps for multiscale pde
Zecheng Zhang, Christian Moya, Wing Tat Leung, Guang Lin, and Hayden Schaeffer · 2023
Later among the works it cites.
Zecheng Zhang, Christian Moya, Lu Lu, Guang Lin, and Hayden Schaeffer · 2023
Later among the works it cites.
Min Zhu, Shihang Feng, Youzuo Lin, and Lu Lu · 2023
Later among the works it cites.
Christian Moya, Amirhossein Mollaali, Zecheng Zhang, Lu Lu, and Guang Lin · 2024
Closest in time.
Ups: Towards foundation models for pde solving via cross-modal adaptation
Junhong Shen, Tanya Marwah, and Ameet Talwalkar · 2024
Closest in time.
Language models don’t always say what they think: unfaithful explanations in chain-of-thought prompting
Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman · 2024
Closest in time.
Pde generalization of in-context operator networks: A study on 1d scalar nonlinear conservation laws
Liu Yang and Stanley J Osher · 2024
Closest in time.
Pdeformer: Towards a foundation model for one-dimensional partial differential equations
Zhanhong Ye, Xiang Huang, Leheng Chen, Hongsheng Liu, Zidong Wang, and Bin Dong · 2024
Closest in time.
Minglang Yin, Nicolas Charon, Ryan Brody, Lu Lu, Natalia Trayanova, and Mauro Maggioni · 2024
Closest in time.
Modno: Multi operator learning with distributed neural operators
Zecheng Zhang · 2024
Closest in time.
Toolqa: A dataset for llm question answering with external tools
Yuchen Zhuang, Yue Yu, Kuan Wang, Haotian Sun, and Chao Zhang · 2024
Closest in time.