Fetching the paper…
Reading the bibliography…
Pretrained large language models (LLMs) are surprisingly effective at performing zero-shot tasks, including time-series forecasting.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Über die von der molekularkinetischen Theorie der Wärme geforderte Bewegung von in ruhenden Flüssigkeiten suspendierten Teilchen
Albert Einstein. 1905 · 1905
Earlier work this paper cites.
Mouvement brownien et réalité moléculaire
Jean Perrin. 1909 · 1909
Earlier work this paper cites.
On a measure of divergence between two statistical populations defined by their probability distribution
Anil Bhattacharyya. 1943 · 1943
Earlier work this paper cites.
On a measure of divergence between two multinomial populations
Anil Bhattacharyya. 1946 · 1946
Earlier work this paper cites.
Deterministic nonperiodic flow
Edward N Lorenz. 1963 · 1963
Earlier work this paper cites.
The divergence and Bhattacharyya distance measures in signal selection
Thomas Kailath. 1967 · 1967
Earlier work this paper cites.
Detecting strange attractors in turbulence
Floris Takens. 2006 · 1980
Earlier work this paper cites.
Lognormal distributions
Edwin L Crow and Kunio Shimizu. 1987 · 1987
Earlier work this paper cites.
Comparing measures of sample skewness and kurtosis
Derrick N Joanes and Christine A Gill. 1998 · 1998
Earlier work this paper cites.
An introduction to numerical methods for stochastic differential equations
Eckhard Platen. 1999 · 1999
Earlier work this paper cites.
Comparing simulated and measured values using mean squared deviation and its components
Kazuhiko Kobayashi and Moin Us Salam. 2000 · 2000
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2001
Earlier work this paper cites.
Feature extraction based on the Bhattacharyya distance
Euisun Choi and Chulhee Lee. 2003 · 2003
Earlier work this paper cites.
A scalable hierarchical distributed language model
Andriy Mnih and Geoffrey E Hinton. 2008 · 2008
Cited alongside, same era.
Stochastic differential equations: an introduction with applications
Bernt Oksendal. 2013 · 2013
Cited alongside, same era.
Continuous martingales and Brownian motion , volume 293
Daniel Revuz and Marc Yor. 2013 · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Cited alongside, same era.
Nonlinear dynamics and chaos: With applications to physics, biology, chemistry, and engineering , 2nd edition
Steven H Strogatz. 2015 · 2015
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Yakun Chen, Xianzhi Wang, and Guandong Xu. 2023 · 2023
Later among the works it cites.
ForecastPFN: Synthetically-Trained Zero-Shot Forecasting
Samuel Dooley, Gurnoor Singh Khurana, Chirag Mohapatra, Siddartha Naidu, and Colin White. 2023 · 2023
Later among the works it cites.
Large language models are zero-shot time series forecasters
Nate Gruver, Marc Finzi, Shikai Qiu, and Andrew Gordon Wilson. 2023 · 2023
Later among the works it cites.
Language models represent space and time
Wes Gurnee and Max Tegmark. 2023 · 2023
Later among the works it cites.
Large language model prediction capabilities: Evidence from a real-world forecasting tournament
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Options, futures, and other derivatives , 11th edition
John C Hull. 2021 · 2021
Cited alongside, same era.
Show your work: Scratchpads for intermediate computation with language models
Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, et al. 2021 · 2021
Cited alongside, same era.
Statistical mechanics: Entropy, order parameters, and complexity
James P. Sethna. 2021 · 2021
Cited alongside, same era.
GraB: Finding provably better data permutations than random reshuffling
Yucheng Lu, Wentao Guo, and Christopher M De Sa. 2022 · 2022
Cited alongside, same era.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. 2022 · 2022
Cited alongside, same era.
Challenging big-bench tasks and whether chain-of-thought can solve them
Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V Le, Ed H Chi, Denny Zhou, et al. 2022 · 2022
Cited alongside, same era.
Philipp Schoenegger and Peter S Park. 2023 · 2023
Later among the works it cites.
Do pretrained transformers really learn in-context by gradient descent?
Lingfeng Shen, Aayush Mishra, and Daniel Khashabi. 2023 · 2023
Later among the works it cites.
Beyond neural scaling laws: beating power law scaling via data pruning
Ben Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli, and Ari S. Morcos. 2023 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023 · 2023
Later among the works it cites.
Transformers learn in-context by gradient descent
Johannes Von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov. 2023 · 2023
Later among the works it cites.
Prompt-based domain discrimination for multi-source time series domain adaptation
Junxiang Wang, Guangji Bai, Wei Cheng, Zhengzhang Chen, Liang Zhao, and Haifeng Chen. 2023 · 2023
Later among the works it cites.
Transformer multivariate forecasting: Less is more?
Jingjing Xu, Caesar Wu, Yuan-Fang Li, and Pascal Bouvry. 2023 · 2023
Later among the works it cites.
Unveiling transformers with lego: a synthetic reasoning task
Yi Zhang, Arturs Backurs, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, and Tal Wagner. 2023 · 2023
Later among the works it cites.
In-context language learning: Architectures and algorithms
Ekin Akyürek, Bailin Wang, Yoon Kim, and Jacob Andreas. 2024 · 2024
Closest in time.