Fetching the paper…
Reading the bibliography…
Recent State Space Models (SSMs) such as S4, S5, and Mamba have shown remarkable computational benefits in long-range temporal dependency modeling.
Discovering important people and objects for egocentric video summarization
Yong Jae Lee, Joydeep Ghosh, and Kristen Grauman · 2012
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2013
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Jimmy Ba, Jamie Ryan Kiros, and Geoffrey Hinton · 2016
Earlier work this paper cites.
Attend, infer, repeat: Fast scene understanding with generative models
SM Ali Eslami, Nicolas Heess, Theophane Weber, Yuval Tassa, David Szepesvari, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Neural expectation maximization
Klaus Greff, Sjoerd van Steenkiste, and Jürgen Schmidhuber · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Sequential attend, infer, repeat: Generative modelling of moving objects
Adam Kosiorek, Hyunjik Kim, Yee Whye Teh, and Ingmar Posner · 2018
Earlier work this paper cites.
Relational neural expectation maximization: Unsupervised discovery of objects and their interactions
Sjoerd Van Steenkiste, Michael Chang, Klaus Greff, and Jürgen Schmidhuber · 2018
Earlier work this paper cites.
Monet: Unsupervised scene decomposition and representation
Christopher P Burgess, Loic Matthey, Nicholas Watters, Rishabh Kabra, Irina Higgins, Matt Botvinick, and Alexander Lerchner · 2019
Earlier work this paper cites.
Exploiting spatial invariance for scalable unsupervised object tracking
Eric Crawford and Joelle Pineau · 2019
Earlier work this paper cites.
Spatially invariant unsupervised object detection with convolutional neural networks
Eric Crawford and Joelle Pineau · 2019
Earlier work this paper cites.
Multi-object representation learning with iterative variational inference
Klaus Greff, Raphaël Lopez Kaufmann, Rishab Kabra, Nick Watters, Chris Burgess, Daniel Zoran, Loic Matthey, Matthew Botvinick, and Alexander Lerchner · 2019
Earlier work this paper cites.
Scalor: Generative world models with scalable object representations
Jindong Jiang, Sepehr Janghorbani, Gerard De Melo, and Sungjin Ahn · 2019
Earlier work this paper cites.
Spatial broadcast decoder: A simple architecture for learning disentangled representations in vaes
Nicholas Watters, Loic Matthey, Christopher P. Burgess, and Alexander Lerchner · 2019
Earlier work this paper cites.
Object-centric image generation with factored depths, locations, and appearances
Titas Anciukevicius, Christoph H Lampert, and Paul Henderson · 2020
Earlier work this paper cites.
GENESIS: Generative scene inference and sampling with object-centric latent representations
Martin Engelcke, Adam R. Kosiorek, Oiwi Parker Jones, and Ingmar Posner · 2020
Earlier work this paper cites.
CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning
Rohit Girdhar and Deva Ramanan · 2020
Earlier work this paper cites.
HiPPO: Recurrent memory with optimal polynomial projections
Albert Gu, Tri Dao, Stefano Ermon, Atri Rudra, and Christopher Ré · 2020
Earlier work this paper cites.
Generative neurosymbolic machines
Jindong Jiang and Sungjin Ahn · 2020
Earlier work this paper cites.
Improving generative imagination in object-centric world models
Zhixuan Lin, Yi-Fu Wu, Skand Vishwanath Peri, Jindong Jiang, and Sungjin Ahn · 2020
Earlier work this paper cites.
Space: Unsupervised object-oriented scene representation via spatial attention and decomposition
Zhixuan Lin, Yi-Fu Wu, Skand Vishwanath Peri, Weihao Sun, Gautam Singh, Fei Deng, Jindong Jiang, and Sungjin Ahn · 2020
Earlier work this paper cites.
Object-centric learning with slot attention, 2020
Francesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran, Georg Heigold, Jakob Uszkoreit, Alexey Dosovitskiy, and Thomas Kipf · 2020
Earlier work this paper cites.
Scalability in perception for autonomous driving: Waymo open dataset
Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, Vijay Vasudevan, Wei Han, Jiquan Ngiam, Hang Zhao, Aleksei Timofeev, Scott Ettinger, Maxim Krivokon, Amy Gao, Aditya Joshi, Yu Zhang, Jonathon Shlens, Zhifeng Chen, and Dragomir Anguelov · 2020
Earlier work this paper cites.
Capsules with inverted dot-product attention routing
Yao-Hung Hubert Tsai, Nitish Srivastava, Hanlin Goh, and Ruslan Salakhutdinov · 2020
Earlier work this paper cites.
Towards causal generative scene models via competition of experts
Julius von Kügelgen, Ivan Ustyuzhaninov, Peter V. Gehler, Matthias Bethge, and Bernhard Schölkopf · 2020
Cited alongside, same era.
TransDreamer: Reinforcement learning with Transformer world models
Chang Chen, Yi-Fu Wu, Jaesik Yoon, and Sungjin Ahn · 2021
Cited alongside, same era.
Generative scene graph networks
Fei Deng, Zhuo Zhi, Donghun Lee, and Sungjin Ahn · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Cited alongside, same era.
GENESIS-V2: Inferring unordered object representations without iterative refinement
Martin Engelcke, Oiwi Parker Jones, and Ingmar Posner · 2021
Cited alongside, same era.
Slotformer: Unsupervised visual dynamics simulation with object-centric models
Ziyi Wu, Nikita Dvornik, Klaus Greff, Thomas Kipf, and Animesh Garg · 2022
Later among the works it cites.
Efficient long sequence modeling via state space augmented Transformer
Simiao Zuo, Xiaodong Liu, Jian Jiao, Denis Charles, Eren Manavoglu, Tuo Zhao, and Jianfeng Gao · 2022
Later among the works it cites.
Decision S4: Efficient sequence-based RL via state spaces layers
Shmuel Bar David, Itamar Zimerman, Eliya Nachmani, and Lior Wolf · 2023
Later among the works it cites.
Hungry Hungry Hippos: Towards language modeling with state space models
Daniel Y Fu, Tri Dao, Khaled Kamal Saab, Armin W Thomas, Atri Rudra, and Christopher Ré · 2023
Later among the works it cites.
Mamba: Linear-time sequence modeling with selective state spaces
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Recurrent independent mechanisms
Anirudh Goyal, Alex Lamb, Jordan Hoffmann, Shagun Sodhani, Sergey Levine, Yoshua Bengio, and Bernhard Schölkopf · 2021
Cited alongside, same era.
Combining recurrent, convolutional, and continuous-time models with linear state space layers
Albert Gu, Isys Johnson, Karan Goel, Khaled Kamal Saab, Tri Dao, Atri Rudra, and Christopher Ré · 2021
Cited alongside, same era.
Learning high fidelity depths of dressed humans by watching social media dance videos
Yasamin Jafarian and Hyun Soo Park · 2021
Cited alongside, same era.
Rishabh Kabra, Daniel Zoran, Goker Erdogan, Loic Matthey, Antonia Creswell, Matthew Botvinick, Alexander Lerchner, and Christopher P. Burgess · 2021
Cited alongside, same era.
Conditional Object-Centric Learning from Video
Thomas Kipf, Gamaleldin F. Elsayed, Aravindh Mahendran, Austin Stone, Sara Sabour, Georg Heigold, Rico Jonschkowski, Alexey Dosovitskiy, and Klaus Greff · 2021
Cited alongside, same era.
Long Range Arena: A benchmark for efficient Transformers
Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler · 2021
Cited alongside, same era.
Generative video transformer: Can objects be the words?
Yi-Fu Wu, Jaesik Yoon, and Sungjin Ahn · 2021
Cited alongside, same era.
Albert Gu and Tri Dao · 2023
Later among the works it cites.
Improving Object-centric Learning with Query Optimization
Baoxiong Jia, Yu Liu, and Siyuan Huang · 2023
Later among the works it cites.
Modelling long range dependencies in N N D: From task-specific to a general purpose CNN
David M Knigge, David W Romero, Albert Gu, Efstratios Gavves, Erik J Bekkers, Jakub Mikolaj Tomczak, Mark Hoogendoorn, and Jan-jakob Sonke · 2023
Later among the works it cites.
Structured state space models for in-context reinforcement learning
Chris Lu, Yannick Schroecker, Albert Gu, Emilio Parisotto, Jakob Foerster, Satinder Singh, and Feryal Behbahani · 2023
Later among the works it cites.
Long range language modeling via gated state spaces
Harsh Mehta, Ankit Gupta, Ashok Cutkosky, and Behnam Neyshabur · 2023
Later among the works it cites.
Resurrecting recurrent neural networks for long sequences
Antonio Orvieto, Samuel L Smith, Albert Gu, Anushan Fernando, Caglar Gulcehre, Razvan Pascanu, and Soham De · 2023
Later among the works it cites.
Neural Systematic Binder
Gautam Singh, Yeongbin Kim, and Sungjin Ahn · 2023
Later among the works it cites.
Simplified state space layers for sequence modeling
Jimmy T.H. Smith, Andrew Warrington, and Scott Linderman · 2023
Later among the works it cites.
Selective structured state-spaces for long-form video understanding
Jue Wang, Wentao Zhu, Pichao Wang, Xiang Yu, Linda Liu, Mohamed Omar, and Raffay Hamid · 2023
Later among the works it cites.
Inverted-attention transformers can learn object representations: Insights from slot attention
Yi-Fu Wu, Klaus Greff, Gamaleldin Fathy Elsayed, Michael Curtis Mozer, Thomas Kipf, and Sjoerd van Steenkiste · 2023
Later among the works it cites.
Slotdiffusion: Object-centric generative modeling with diffusion models
Ziyi Wu, Jingyu Hu, Wuyue Lu, Igor Gilitschenski, and Animesh Garg · 2023
Later among the works it cites.
Diffusion models without attention, 2023
Jing Nathan Yan, Jiatao Gu, and Alexander M. Rush · 2023
Later among the works it cites.
Deep latent state space models for time-series generation
Linqi Zhou, Michael Poli, Winnie Xu, Stefano Massaroli, and Stefano Ermon · 2023
Later among the works it cites.
Griffin: Mixing gated linear recurrences with local attention for efficient language models
Soham De, Samuel L Smith, Anushan Fernando, Aleksandar Botev, George Cristian-Muraru, Albert Gu, Ruba Haroun, Leonard Berrada, Yutian Chen, Srivatsan Srinivasan, et al · 2024
Closest in time.
Facing off world model backbones: RNNs, Transformers, and S4
Fei Deng, Junyeong Park, and Sungjin Ahn · 2024
Closest in time.
Animate anyone: Consistent and controllable image-to-video synthesis for character animation
Li Hu · 2024
Closest in time.
Object-centric slot diffusion
Jindong Jiang, Fei Deng, Gautam Singh, and Sungjin Ahn · 2024
Closest in time.
Mastering memory tasks with world models
Mohammad Reza Samsami, Artem Zholus, Janarthanan Rajendran, and Sarath Chandar · 2024
Closest in time.
Parallelized spatiotemporal binding, 2024
Gautam Singh, Yue Wang, Jiawei Yang, Boris Ivanovic, Sungjin Ahn, Marco Pavone, and Tong Che · 2024
Closest in time.
Mm-ego: Towards building egocentric multimodal llms
Hanrong Ye, Haotian Zhang, Erik Daxberger, Lin Chen, Zongyu Lin, Yanghao Li, Bowen Zhang, Haoxuan You, Dan Xu, Zhe Gan, et al · 2024
Closest in time.
Vision mamba: Efficient visual representation learning with bidirectional state space model
Lianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang, Wenyu Liu, and Xinggang Wang · 2024
Closest in time.