Fetching the paper…
Reading the bibliography…
This paper presents a fascinating find: By training an auto-regressive LLM model on text tokens, the text model inherently develops internally an ability to understand images and audio, thereby developing the ability to see and hear just by reading.
“Automatic musical genre classification of audio signals,” 2001
George Tzanetakis et al., · 2001
Earlier work this paper cites.
“Frequency estimation from waveforms using multi-layered neural networks,”
Prateek Verma and Ronald W. Schafer, · 2016
Earlier work this paper cites.
“Cnn architectures for large-scale audio classification,”
Shawn Hershey et al., · 2017
Earlier work this paper cites.
“Neural discrete representation learning,”
Aaron Van Den Oord, Oriol Vinyals, et al., · 2017
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani et al., · 2017
Earlier work this paper cites.
“Language models are few-shot learners,” 2020
Tom B. Brown, Benjamin Mann, et al., · 2020
Earlier work this paper cites.
“A framework for contrastive and generative learning of audio representations,”
Prateek Verma and Julius Smith, · 2020
Earlier work this paper cites.
“Progen: Language modeling for protein generation,”
Ali Madani et al., · 2020
Earlier work this paper cites.
“An image is worth 16x16 words: Transformers for image recognition at scale,”
Alexey Dosovitskiy et. al, · 2020
Earlier work this paper cites.
“FSD50K: an open dataset of human-labeled sound events,”
Eduardo Fonseca et al., · 2020
Earlier work this paper cites.
“Jukebox: A generative model for music,”
Prafulla et. al Dhariwal, · 2020
Earlier work this paper cites.
“Scaling laws for neural language models,” 2020
Jared Kaplan et al., · 2020
Earlier work this paper cites.
“Videogpt: Video generation using vq-vae and transformers,” 2021
Wilson Yan, Yunzhi Zhang, Pieter Abbeel, and Aravind Srinivas, · 2021
Earlier work this paper cites.
“A generative model for raw audio using transformer architectures,”
Prateek Verma et al., · 2021
Cited alongside, same era.
“AST: Audio Spectrogram Transformer,”
Yuan Gong, Yu-An Chung, and James Glass, · 2021
Cited alongside, same era.
“Lora: Low-rank adaptation of large language models,” 2021
Edward J. Hu et. al, · 2021
Cited alongside, same era.
Prateek Verma et al., · 2021
Cited alongside, same era.
“Leaf: A learnable frontend for audio classification,”
Neil Zeghidour et al., · 2021
Cited alongside, same era.
“Pretrained transformers as universal computation engines,” 2021
Kevin Lu et al., · 2021
Cited alongside, same era.
“Audiopalm: A large language model that can speak and listen,”
Paul K Rubenstein et al., · 2023
Later among the works it cites.
“Neural codec language models are zero-shot text to speech synthesizers,”
Chengyi et. al Wang, · 2023
Later among the works it cites.
“Visual instruction tuning,”
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee, · 2023
Later among the works it cites.
“Joint audio and speech understanding,”
Yuan Gong et al., · 2023
Later among the works it cites.
“Qwen-audio: Advancing universal audio understanding,”
Yunfei Chu et al., · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Frozen pretrained transformers as universal computation engines,”
Kevin Lu et al., · 2022
Cited alongside, same era.
“High fidelity neural audio compression,”
Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi, · 2022
Cited alongside, same era.
“Flamingo: a visual language model for few-shot learning,”
Jean-Baptiste et. al Alayrac, · 2022
Cited alongside, same era.
“Blip: Bootstrapping language-image pre-training for unified vision-language understanding,”
Junnan Li et al., · 2022
Cited alongside, same era.
“Robust speech recognition via large-scale weak supervision,”
Alec Radford et al., · 2022
Cited alongside, same era.
“Audiolm: a language modeling approach to audio generation,” 2023
Zalán Borsos, Raphaël Marinier, et al., · 2023
Cited alongside, same era.
“A content adaptive learnable” time-frequency” representation for audio signal processing,”
Prateek Verma et al., · 2023
Later among the works it cites.
“One-peace: Exploring one general representation model toward unlimited modalities,”
Peng Wang et al., · 2023
Later among the works it cites.
“Solving olympiad geometry without human demonstrations,”
Trieu H et. al Trinh, · 2024
Later among the works it cites.
“Fine-tuning enhances existing mechanisms: A case study on entity tracking,” 2024
Nikhil Prakash et al., · 2024
Later among the works it cites.
“Audio flamingo: A novel audio language model with few-shot learning and dialogue abilities,”
Zhifeng et. al Kong, · 2024
Later among the works it cites.
Sonia Bbouzidi, Ghazala Hcini, Imen Jdey, and Fadoua Drira, · 2024
Later among the works it cites.
“s1: Simple test-time scaling,”
Niklas et. al Muennighoff, · 2025
Closest in time.
“PyTorch from Scratch: Vision Transformer (ViT),” https://github.com/s-chh/PyTorch-Scratch-Vision-Transformer-ViT , 2021,
S. Chh, · 2025
Closest in time.