Fetching the paper…
Reading the bibliography…
We describe a family of architectures to support transductive inference by allowing memory to grow to a finite but a-priori unknown bound while making efficient use of finite resources for inference.
Extrapolation, interpolation, and smoothing of stationary time series: with engineering applications
Norbert Wiener · 1949
Earlier work this paper cites.
A new approach to linear filtering and prediction problems, 1960
Rudolph Emil Kalman · 1960
Earlier work this paper cites.
A formal theory of inductive inference. i
RJ Solmonoff · 1964
Earlier work this paper cites.
On a matrix riccati equation of stochastic control
William M Wonham · 1968
Earlier work this paper cites.
On the algebraic structure of bilinear systems. theory and applica-tions of variable structure systems, 1972
R Brockett · 1972
Earlier work this paper cites.
Realization and structure theory of bilinear dynamical systems
Paolo D’Alessandro, Alberto Isidori, and Antonio Ruberti · 1974
Earlier work this paper cites.
Bilinear and nonlinear realizations of input-output maps
Arthur J Krener · 1975
Earlier work this paper cites.
On a measure of lack of fit in time series models
Greta M Ljung and George EP Box · 1978
Earlier work this paper cites.
Compression of individual sequences via variable-rate coding
Jacob Ziv and Abraham Lempel · 1978
Earlier work this paper cites.
On the stochastic realization problem
Anders Lindquist and Giorgio Picci · 1979
Earlier work this paper cites.
Linear systems
Thomas Kailath · 1980
Earlier work this paper cites.
Results on the non existence of finite dimensional filters
Mireille Chaleyat Maurel and Dominique Michel · 1984
Earlier work this paper cites.
A technique for high-performance data compression
Terry A. Welch · 1984
Earlier work this paper cites.
Nonlinear control systems: an introduction
Alberto Isidori · 1985
Earlier work this paper cites.
System identification toolbox: User’s guide
Lennart Ljung · 1995
Earlier work this paper cites.
Transductive inference for estimating values of functions
Olivier Chapelle, Vladimir Vapnik, and Jason Weston · 1999
Earlier work this paper cites.
Transductive inference and semi-supervised learning, 2006
Vladimir Vapnik · 2006
Earlier work this paper cites.
Stochastic processes and filtering theory
Andrew H Jazwinski · 2007
Earlier work this paper cites.
Bilinear control systems: matrices in action
David LeRoy Elliott · 2009
Cited alongside, same era.
Subspace identification for linear systems: Theory—Implementation—Applications
Peter Van Overschee and Bart De Moor · 2012
Cited alongside, same era.
The lambada dataset: Word prediction requiring a broad discourse context
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández · 2016
Cited alongside, same era.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord · 2018
Cited alongside, same era.
Hellaswag: Can a machine really finish your sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
SCROLLS: Standardized CompaRison over long language sequences
Uri Shaham, Elad Segal, Maor Ivgi, Avia Efrat, Ori Yoran, Adi Haviv, Ankit Gupta, Wenhan Xiong, Mor Geva, Jonathan Berant, and Omer Levy · 2022
Later among the works it cites.
Stacked residuals of dynamic layers for time series anomaly detection, 2022
L. Zancato, A. Achille, G. Paolini, A. Chiuso, and S. Soatto · 2022
Later among the works it cites.
A framework for few-shot language model evaluation, 12 2023
Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou · 2023
Later among the works it cites.
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Piqa: Reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Cited alongside, same era.
Test-time training with self-supervision for generalization under distribution shifts
Yu Sun, Xiaolong Wang, Zhuang Liu, John Miller, Alexei Efros, and Moritz Hardt · 2020
Cited alongside, same era.
Inductive learning of transductive inference:
B. Bowman et al · 2021
Cited alongside, same era.
Winogrande: An adversarial winograd schema challenge at scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi · 2021
Cited alongside, same era.
A novel deep neural network architecture for non-linear system identification
Luca Zancato and Alessandro Chiuso · 2021
Cited alongside, same era.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al · 2023
Later among the works it cites.
Landmark attention: Random-access infinite context length for transformers
Amirkeivan Mohtashami and Martin Jaggi · 2023
Later among the works it cites.
Efficient streaming language models with attention sinks
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis · 2023
Later among the works it cites.
Gated linear attention transformers with hardware-efficient training
Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda, and Yoon Kim · 2023
Later among the works it cites.
Train/test-time adaptation with retrieval
Luca Zancato, Alessandro Achille, Tian Yu Liu, Matthew Trager, Pramuditha Perera, and Stefano Soatto · 2023
Later among the works it cites.
Simple linear attention language models balance the recall-throughput tradeoff
Simran Arora, Sabri Eyuboglu, Michael Zhang, Aman Timalsina, Silas Alberti, Dylan Zinsley, James Zou, Atri Rudra, and Christopher Ré · 2024
Closest in time.
xlstm: Extended long short-term memory
Maximilian Beck, Korbinian Pöppel, Markus Spanring, Andreas Auer, Oleksandra Prudnikova, Michael Kopp, Günter Klambauer, Johannes Brandstetter, and Sepp Hochreiter · 2024
Closest in time.
Griffin: Mixing gated linear recurrences with local attention for efficient language models
Soham De, Samuel L Smith, Anushan Fernando, Aleksandar Botev, George Cristian-Muraru, Albert Gu, Ruba Haroun, Leonard Berrada, Yutian Chen, Srivatsan Srinivasan, et al · 2024
Closest in time.
Multi-modal hallucination control by visual information grounding
Alessandro Favero, Luca Zancato, Matthew Trager, Siddharth Choudhary, Pramuditha Perera, Alessandro Achille, Ashwin Swaminathan, and Stefano Soatto · 2024
Closest in time.
Jamba: A hybrid transformer-mamba language model
Opher Lieber, Barak Lenz, Hofit Bata, Gal Cohen, Jhonathan Osin, Itay Dalmedigos, Erez Safahi, Shaked Meirom, Yonatan Belinkov, Shai Shalev-Shwartz, et al · 2024
Closest in time.
Megalodon: Efficient llm pretraining and inference with unlimited context length
Xuezhe Ma, Xiaomeng Yang, Wenhan Xiong, Beidi Chen, Lili Yu, Hao Zhang, Jonathan May, Luke Zettlemoyer, Omer Levy, and Chunting Zhou · 2024
Closest in time.
Mechanistic design and scaling of hybrid architectures
Michael Poli, Armin W Thomas, Eric Nguyen, Pragaash Ponnusamy, Björn Deiseroth, Kristian Kersting, Taiji Suzuki, Brian Hie, Stefano Ermon, Christopher Ré, et al · 2024
Closest in time.
Megabyte: Predicting million-byte sequences with multiscale transformers
Lili Yu, Dániel Simig, Colin Flaherty, Armen Aghajanyan, Luke Zettlemoyer, and Mike Lewis · 2024
Closest in time.