Fetching the paper…
Reading the bibliography…
Foundation models (FMs) are increasingly spearheading recent advances on a variety of tasks that fall under the purview of computer audition -- the use of machines to understand sounds.
“Language Models are Few-Shot Learners”
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever and Dario Amodei · 1901
Earlier work this paper cites.
“Signature Verification using a "Siamese" Time Delay Neural Network”
Jane Bromley, Isabelle Guyon, Yann LeCun, Eduard Säckinger and Roopak Shah · 1993
Earlier work this paper cites.
“Computational auditory scene analysis”
Guy Brown and Martin Cooke · 1994
Earlier work this paper cites.
“Advances in residual vector quantization: A review”
Christopher Barnes, Syed Rizvi and Nasser Nasrabadi · 1996
Earlier work this paper cites.
“Weight averaging for neural networks and local resampling schemes”
Joachim Utans · 1996
Earlier work this paper cites.
“Acoustical sound database in real environments for sound scene understanding and hands-free speech recognition.”
Satoshi Nakamura, Kazuo Hiyane, Futoshi Asano, Takanobu Nishiura and Takeshi Yamada · 2000
Earlier work this paper cites.
“Computational auditory scene analysis: Principles, algorithms, and applications”
DeLiang Wang and Guy Brown · 2006
Earlier work this paper cites.
“Imagenet: A large-scale hierarchical image database”
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li and Li Fei-Fei · 2009
Earlier work this paper cites.
“Challenges of using bioacoustics to globally monitor bats”
Charlotte Walters, Alanna Collen, Tim Lucas, Kim Mroz, Catherine Sayer and Kate Jones · 2013
Earlier work this paper cites.
“Acoustic scene classification: Classifying environments from the sounds they produce”
Daniele Barchiesi, Dimitrios Giannoulis, Dan Stowell and Mark Plumbley · 2015
Earlier work this paper cites.
“ESC: Dataset for environmental sound classification”
Karol Piczak · 2015
Earlier work this paper cites.
“Modelling human factors in perceptual multimedia quality: On the role of personality and culture”
Michael Scott, Sharath Guntuku, Yang Huan, Weisi Lin and Gheorghita Ghinea · 2015
Earlier work this paper cites.
“Very Deep Convolutional Networks for Large-Scale Image Recognition”
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
“Chime-home: A dataset for sound source recognition in a domestic environment”
Peter Foster, Siddharth Sigtia, Sacha Krstulovic, Jon Barker and Mark Plumbley · 2015
Earlier work this paper cites.
“TUT database for acoustic scene classification and sound event detection”
Annamaria Mesaros, Toni Heittola and Tuomas Virtanen · 2016
Earlier work this paper cites.
“Hierarchical learning for DNN-based acoustic scene classification”
Yong Xu, Qiang Huang, Wenwu Wang and Mark Plumbley · 2016
Earlier work this paper cites.
“Adieu features? end-to-end speech emotion recognition using a deep convolutional recurrent network”
George Trigeorgis, Fabien Ringeval, Raymond Brueckner, Erik Marchi, Mihalis Nicolaou, Björn Schuller and Stefanos Zafeiriou · 2016
Earlier work this paper cites.
“Deep Residual Learning for Image Recognition”
Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun · 2016
Earlier work this paper cites.
“Human and machine hearing: extracting meaning from sound”
Richard Lyon · 2017
Earlier work this paper cites.
“Audio Set: An ontology and human-labeled dataset for audio events”
Jort. Gemmeke, Daniel.. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Moore, Manoj Plakal and Marvin Ritter · 2017
Earlier work this paper cites.
“Prototypical Networks for Few-shot Learning”
Jake Snell, Kevin Swersky and Richard Zemel · 2017
Earlier work this paper cites.
“Attention is All you Need”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, Łukasz Kaiser and Illia Polosukhin · 2017
Earlier work this paper cites.
“Freesound datasets: a platform for the creation of open audio datasets”
Eduardo Fonseca, Jordi Pons, Xavier Favory, Frederic Font, Dmitry Bogdanov, Andres Ferraro, Sergio Oramas, Alastair Porter and Xavier Serra · 2017
Earlier work this paper cites.
“Ecoacoustics: The ecological role of sounds”
Almo Farina and Stuart Gage · 2017
Earlier work this paper cites.
“A multi-device dataset for urban acoustic scene classification”
Annamaria Mesaros, Toni Heittola and Tuomas Virtanen · 2018
Earlier work this paper cites.
“Acoustic scene classification: an overview of DCASE 2017 challenge entries”
Annamaria Mesaros, Toni Heittola and Tuomas Virtanen · 2018
Earlier work this paper cites.
“On the influence of cultural differences on the perception of audio coding artifacts in music”
Sascha Dick, Jiandong Zhang, Yili Qin, Nadja Schinkel-Bielefeld, Anna Leschanowsky and Frederik Nagel · 2018
Earlier work this paper cites.
“Acoustics and psychoacoustics of sound scenes and events”
Guillaume Lemaitre, Nicolas Grimault and Clara Suied · 2018
Earlier work this paper cites.
“Sound event localization and detection of overlapping sources using convolutional recurrent neural networks”
Sharath Adavanne, Archontis Politis, Joonas Nikunen and Tuomas Virtanen · 2018
Earlier work this paper cites.
“Pyroomacoustics: A python package for audio room simulation and array processing algorithms”
Robin Scheibler, Eric Bezzam and Ivan Dokmanić · 2018
Earlier work this paper cites.
“Bert: Pre-training of deep bidirectional transformers for language understanding”
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 2018
Earlier work this paper cites.
“MAVD: a dataset for sound event detection in urban environments”, 2019
Pablo Zinemanas, Pablo Cancela and Martı́n Rocamora · 2019
Earlier work this paper cites.
“Sound event detection in domestic environments with weakly labeled data and soundscape synthesis”
Nicolas Turpault, Romain Serizel, Ankit Shah and Justin Salamon · 2019
Earlier work this paper cites.
“Learning sound event classifiers from web audio with noisy labels”
Eduardo Fonseca, Manoj Plakal, Daniel Ellis, Frederic Font, Xavier Favory and Xavier Serra · 2019
Earlier work this paper cites.
“Hierarchy-aware loss function on a tree structured label space for audio event detection”
Arindam Jati, Naveen Kumar, Ruxin Chen and Panayiotis Georgiou · 2019
Earlier work this paper cites.
“AudioCaps: Generating Captions for Audios in The Wild”
Chris Kim, Byeongchang Kim, Hyunmin Lee and Gunhee Kim · 2019
Earlier work this paper cites.
“Exploring deep spectrum representations via attention-based recurrent and convolutional neural networks for speech emotion recognition”
Ziping Zhao, Zhongtian Bao, Yiqin Zhao, Zixing Zhang, Nicholas Cummins, Zhao Ren and Björn Schuller · 2019
Earlier work this paper cites.
“Language models are unsupervised multitask learners”
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei and Ilya Sutskever · 2019
Earlier work this paper cites.
“SONYC Urban Sound Tagging (SONYC-UST): A Multilabel Dataset from an Urban Acoustic Sensor Network”
Mark Cartwright, Ana Mendez, Jason Cramer, Vincent Lostanlen, Graham Dove, Ho-Hsiang Wu, Justin Salamon, Oded Nov and Juan Bello · 2019
Earlier work this paper cites.
“Model cards for model reporting”
Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Raji and Timnit Gebru · 2019
Earlier work this paper cites.
“Clotho: an Audio Captioning Dataset”
Konstantinos Drossos, Samuel Lipping and Tuomas Virtanen · 2020
Earlier work this paper cites.
“Temporal Reasoning via Audio Question Answering”
Haytham. Fayek and Justin Johnson · 2020
Earlier work this paper cites.
“The LOCATA challenge: Acoustic source localization and tracking”
Christine Evers, Heinrich Löllmann, Heinrich Mellmann, Alexander Schmidt, Hendrik Barfuss, Patrick Naylor and Walter Kellermann · 2020
Earlier work this paper cites.
“A Dataset of Reverberant Spatial Sound Scenes with Moving Sources for Sound Event Localization and Detection”
Archontis Politis, Sharath Adavanne and Tuomas Virtanen · 2020
Earlier work this paper cites.
“Learning representations from audio-visual spatial alignment”
Pedro Morgado, Yi Li and Nuno Nvasconcelos · 2020
Earlier work this paper cites.
“Sound source localization and reconstruction using a wearable microphone array and inertial sensors”
Clas Veibäck, Martin Skoglund, Fredrik Gustafsson and Gustaf Hendeby · 2020
Earlier work this paper cites.
“Panns: Large-scale pretrained audio neural networks for audio pattern recognition”
Qiuqiang Kong, Yin Cao, Turab Iqbal, Yuxuan Wang, Wenwu Wang and Mark Plumbley · 2020
Earlier work this paper cites.
“A comprehensive survey on transfer learning”
Fuzhen Zhuang, Zhiyuan Qi, Keyu Duan, Dongbo Xi, Yongchun Zhu, Hengshu Zhu, Hui Xiong and Qing He · 2020
Earlier work this paper cites.
“What is being transferred in transfer learning?”
Behnam Neyshabur, Hanie Sedghi and Chiyuan Zhang · 2020
Earlier work this paper cites.
“wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations”
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed and Michael Auli · 2020
Earlier work this paper cites.
“Retrieval-augmented generation for knowledge-intensive nlp tasks”
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih and Tim Rocktäschel · 2020
Earlier work this paper cites.
“Optimizing mode connectivity via neuron alignment”
Norman Tatro, Pin-Yu Chen, Payel Das, Igor Melnyk, Prasanna Sattigeri and Rongjie Lai · 2020
Earlier work this paper cites.
“Addressing missing labels in large-scale sound event recognition using a teacher-student framework with loss masking”
Eduardo Fonseca, Shawn Hershey, Manoj Plakal, Daniel Ellis, Aren Jansen and R Moore · 2020
Cited alongside, same era.
“What is the state of neural network pruning?”
Davis Blalock, Jose Gonzalez, Jonathan Frankle and John Guttag · 2020
Cited alongside, same era.
“DDSP: Differentiable Digital Signal Processing”
Jesse Engel, Chenjie Gu and Adam Roberts · 2020
Cited alongside, same era.
“On the opportunities and risks of foundation models”
Rishi Bommasani, Drew Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael Bernstein, Jeannette Bohg, Antoine Bosselut and Emma Brunskill · 2021
Cited alongside, same era.
“Sound event detection: A tutorial”
Annamaria Mesaros, Toni Heittola, Tuomas Virtanen and Mark Plumbley · 2021
Cited alongside, same era.
“BEATs: Audio Pre-Training with Acoustic Tokenizers”
Sanyuan Chen, Yu Wu, Chengyi Wang, Shujie Liu, Daniel Tompkins, Zhuo Chen, Wanxiang Che, Xiangzhan Yu and Furu Wei · 2023
Later among the works it cites.
“Clap learning audio concepts from natural language supervision”
Benjamin Elizalde, Soham Deshmukh, Mahmoud Al and Huaming Wang · 2023
Later among the works it cites.
“Dawn of the transformer era in speech emotion recognition: closing the valence gap”
Johannes Wagner, Andreas Triantafyllopoulos, Hagen Wierstorf, Maximilian Schmitt, Felix Burkhardt, Florian Eyben and Björn Schuller · 2023
Later among the works it cites.
“Robust speech recognition via large-scale weak supervision”
Alec Radford, Jong Kim, Tao Xu, Greg Brockman, Christine McLeavey and Ilya Sutskever · 2023
Later among the works it cites.
“Whisper-at: Noise-robust automatic speech recognizers are also strong general audio event taggers”
Yuan Gong, Sameer Khurana, Leonid Karlinsky and James Glass · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“The benefit of temporally-strong labels in audio event classification”
Shawn Hershey, Daniel Ellis, Eduardo Fonseca, Aren Jansen, Caroline Liu, R Moore and Manoj Plakal · 2021
Cited alongside, same era.
“FSD50K: an open dataset of human-labeled sound events”
Eduardo Fonseca, Xavier Favory, Jordi Pons, Frederic Font and Xavier Serra · 2021
Cited alongside, same era.
“A curated dataset of urban scenes for audio-visual scene analysis”
Shanshan Wang, Annamaria Mesaros, Toni Heittola and Tuomas Virtanen · 2021
Cited alongside, same era.
“New avenues in audio intelligence: Towards holistic real-life audio understanding”
Björn Schuller, Alice Baird, Alexander Gebhard, Shahin Amiriparian, Gil Keren, Maximilian Schmitt and Nicholas Cummins · 2021
Cited alongside, same era.
“Description and discussion on DCASE 2021 challenge task 2: Unsupervised anomalous sound detection for machine condition monitoring under domain shifted conditions”
Yohei Kawaguchi, Keisuke Imoto, Yuma Koizumi, Noboru Harada, Daisuke Niizumi, Kota Dohi, Ryo Tanabe, Harsh Purohit and Takashi Endo · 2021
Cited alongside, same era.
“Blind room parameter estimation using multiple multichannel speech recordings”
Prerak Srivastava, Antoine Deleforge and Emmanuel Vincent · 2021
Cited alongside, same era.
“gpuRIR: A python library for room impulse response simulation with GPU acceleration”
David Diaz-Guerra, Antonio Miguel and Jose Beltran · 2021
Cited alongside, same era.
“High Fidelity Neural Audio Compression”
Alexandre Défossez, Jade Copet, Gabriel Synnaeve and Yossi Adi · 2023
Later among the works it cites.
“Self-Instruct: Aligning Language Models with Self-Generated Instructions”
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah. Smith, Daniel Khashabi and Hannaneh Hajishirzi · 2023
Later among the works it cites.
“QLoRA: Efficient Finetuning of Quantized LLMs”
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman and Luke Zettlemoyer · 2023
Later among the works it cites.
“Pengi: An audio language model for audio tasks”
Soham Deshmukh, Benjamin Elizalde, Rita Singh and Huaming Wang · 2023
Later among the works it cites.
“Acoustic Prompt Tuning: Empowering Large Language Models with Audition Capabilities”, 2023
Jinhua Liang, Xubo Liu, Wenwu Wang, Mark Plumbley, Huy Phan and Emmanouil Benetos · 2023
Later among the works it cites.
“Joint audio and speech understanding”
Yuan Gong, Alexander Liu, Hongyin Luo, Leonid Karlinsky and James Glass · 2023
Later among the works it cites.
“Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models”
Yunfei Chu, Jin Xu, Xiaohuan Zhou, Qian Yang, Shiliang Zhang, Zhijie Yan, Chang Zhou and Jingren Zhou · 2023
Later among the works it cites.
“Uniaudio: An audio foundation model toward universal audio generation”
Dongchao Yang, Jinchuan Tian, Xu Tan, Rongjie Huang, Songxiang Liu, Xuankai Chang, Jiatong Shi, Sheng Zhao, Jiang Bian and Xixin Wu · 2023
Later among the works it cites.
“Lauragpt: Listen, attend, understand, and regenerate audio with gpt”
Qian Chen, Yunfei Chu, Zhifu Gao, Zerui Li, Kai Hu, Xiaohuan Zhou, Jin Xu, Ziyang Ma, Wen Wang and Siqi Zheng · 2023
Later among the works it cites.
“Audio Retrieval with WavText5K and CLAP Training”
Soham Deshmukh, Benjamin Elizalde and Huaming Wang · 2023
Later among the works it cites.
“Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation”
Yusong Wu, Ke Chen, Tianyu Zhang, Yuchen Hui, Taylor Berg-Kirkpatrick and Shlomo Dubnov · 2023
Later among the works it cites.
“Nonspeech7k dataset: Classification and analysis of human non-speech sound”
Muhammad Rashid, Guiqing Li and Chengrui Du · 2023
Later among the works it cites.
“Llama: Open and efficient foundation language models”
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro and Faisal Azhar · 2023
Later among the works it cites.
“Audiobox: Unified audio generation with natural language prompts”
Apoorv Vyas, Bowen Shi, Matthew Le, Andros Tjandra, Yi-Chiao Wu, Baishan Guo, Jiemin Zhang, Xinyue Zhang, Robert Adkins and William Ngan · 2023
Later among the works it cites.
“Merging Vision Transformers from Different Tasks and Domains”
Peng Ye, Chenyu Huang, Mingzhu Shen, Tao Chen, Yongqi Huang, Yuning Zhang and Wanli Ouyang · 2023
Later among the works it cites.
“Ecology & computer audition: Applications of audio technology to monitor organisms and environment”
Björn Schuller, Alican Akman, Yi Chang, Harry Coppock, Alexander Gebhard, Alexander Kathan, Esther Rituerto-González, Andreas Triantafyllopoulos and Florian Pokorny · 2023
Later among the works it cites.
“Can synthetic data boost the training of deep acoustic vehicle counting networks?”
Stefano Damiano, Luca Bondi, Shabnam Ghaffarzadegan, Andre Guntoro and Toon van Waterschoot · 2024
Closest in time.
“Knowledge fusion of large language models”
Fanqi Wan, Xinting Huang, Deng Cai, Xiaojun Quan, Wei Bi and Shuming Shi · 2024
Closest in time.
“The Revolution of Multimodal Large Language Models: A Survey”
Davide Caffagni, Federico Cocchi, Luca Barsellotti, Nicholas Moratelli, Sara Sarto, Lorenzo Baraldi, Marcella Cornia and Rita Cucchiara · 2024
Closest in time.
“A survey on multimodal large language models”
Shukang Yin, Chaoyou Fu, Sirui Zhao, Ke Li, Xing Sun, Tong Xu and Enhong Chen · 2024
Closest in time.
“Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research”
Xinhao Mei, Chutong Meng, Haohe Liu, Qiuqiang Kong, Tom Ko, Chengqi Zhao, Mark Plumbley, Yuexian Zou and Wenwu Wang · 2024
Closest in time.
“Listen, Think, and Understand”
Yuan Gong, Hongyin Luo, Alexander. Liu, Leonid Karlinsky and James. Glass · 2024
Closest in time.
“Speaker Distance Estimation in Enclosures from Single-Channel Audio”
Michael Neri, Archontis Politis, Daniel Krause, Marco Carli and Tuomas Virtanen · 2024
Closest in time.
“STARSS23: An audio-visual dataset of spatial recordings of real scenes with spatiotemporal annotations of sound events”
Kazuki Shimada, Archontis Politis, Parthasaarathy Sudarsanam, Daniel Krause, Kengo Uchida, Sharath Adavanne, Aapo Hakala, Yuichiro Koyama, Naoya Takahashi and Shusuke Takahashi · 2024
Closest in time.
“Mind the Domain Gap: a Systematic Analysis on Bioacoustic Sound Event Detection”
Jinhua Liang, Ines Nolasco, Burooj Ghani, Huy Phan, Emmanouil Benetos and Dan Stowell · 2024
Closest in time.
“Exploring Meta Information for Audio-based Zero-Shot Bird Classification”
Alexander Gebhard, Andreas Triantafyllopoulos, Teresa Bez, Lukas Christ, Alexander Kathan and Björn. Schuller · 2024
Closest in time.
“Semanticodec: An ultra low bitrate semantic audio codec for general sound”
Haohe Liu, Xuenan Xu, Yi Yuan, Mengyue Wu, Wenwu Wang and Mark Plumbley · 2024
Closest in time.
“RULER: What’s the Real Context Size of Your Long-Context Language Models?”
Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia and Boris Ginsburg · 2024
Closest in time.
“The Evolution of Multimodal Model Architectures”
Shakti Wadekar, Abhishek Chaurasia, Aman Chadha and Eugenio Culurciello · 2024
Closest in time.
“BAT: Learning to Reason about Spatial Sounds with Large Language Models”
Zhisheng Zheng, Puyuan Peng, Ziyang Ma, Xie Chen, Eunsol Choi and David Harwath · 2024
Closest in time.
“Scaling Instruction-Finetuned Language Models”
Hyung Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinson, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Zhao, Yanping Huang, Andrew Dai, Hongkun Yu, Slav Petrov, Ed. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc. Le and Jason Wei · 2024
Closest in time.
“Recap: retrieval-augmented audio captioning”
Sreyan Ghosh, Sonal Kumar, Chandra Evuru, Ramani Duraiswami and Dinesh Manocha · 2024
Closest in time.
“Arcee’s MergeKit: A Toolkit for Merging Large Language Models”
Charles Goddard, Shamane Siriwardhana, Malikeh Ehghaghi, Luke Meyers, Vlad Karpukhin, Brian Benedict, Mark McQuade and Jacob Solawetz · 2024
Closest in time.
“Llm augmented llms: Expanding capabilities through composition”
Rachit Bansal, Bidisha Samanta, Siddharth Dalmia, Nitish Gupta, Shikhar Vashishth, Sriram Ganapathy, Abhishek Bapna, Prateek Jain and Partha Talukdar · 2024
Closest in time.
“SALMONN: Towards Generic Hearing Abilities for Large Language Models”
Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun MA and Chao Zhang · 2024
Closest in time.
“Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities”
Zhifeng Kong, Arushi Goel, Rohan Badlani, Wei Ping, Rafael Valle and Bryan Catanzaro · 2024
Closest in time.
“Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context”
Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean-baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat and Julian Schrittwieser · 2024
Closest in time.
“Efficient Streaming Language Models with Attention Sinks”
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han and Mike Lewis · 2024
Closest in time.
“AI models collapse when trained on recursively generated data”
Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Nicolas Papernot, Ross Anderson and Yarin Gal · 2024
Closest in time.
“Listenable Maps for Zero-Shot Audio Classifiers”
Francesco Paissan, Luca Della, Mirco Ravanelli and Cem Subakan · 2024
Closest in time.
“EMR-Merging: Tuning-Free High-Performance Model Merging”
Chenyu Huang, Peng Ye, Tao Chen, Tong He, Xiangyu Yue and Wanli Ouyang · 2024
Closest in time.
“Expressivity and Speech Synthesis”
Andreas Triantafyllopoulos and Björn Schuller · 2024
Closest in time.
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten and Alex Vaughan · 2024
Closest in time.
“Data-efficient low-complexity acoustic scene classification in the dcase 2024 challenge”
Florian Schmid, Paul Primus, Toni Heittola, Annamaria Mesaros, Irene Martin-Morató, Khaled Koutini and Gerhard Widmer · 2024
Closest in time.
“Early-exit deep neural network-a comprehensive survey”
Haseena Rahmath, Vishal Srivastava, Kuldeep Chaurasia, Roberto Pacheco and Rodrigo Couto · 2024
Closest in time.
“Foundation Models Defining a New Era in Vision: a Survey and Outlook”
Muhammad Awais, Muzammal Naseer, Salman Khan, Rao Anwer, Hisham Cholakkal, Mubarak Shah, Ming-Hsuan Yang and Fahad Khan · 2025
Closest in time.