Fetching the paper…
Reading the bibliography…
In recent years, foundation models (FMs) such as large language models (LLMs) and latent diffusion models (LDMs) have profoundly impacted diverse sectors, including music.
The probable error of a mean
William Sealy Gosset · 1908
Earlier work this paper cites.
The annals of mathematical statistics
Harvy Clyde Carver, AL O’TOOLE, and TE RAIFORD · 1930
Earlier work this paper cites.
A scale for the measurement of the psychological magnitude pitch
Stanley Smith Stevens, John Volkmann, and Edwin Broomell Newman · 1937
Earlier work this paper cites.
Report on a general problem-solving program
A. Newell, J.C. Shaw, and H.A. Simon · 1959
Earlier work this paper cites.
Circularity in judgments of relative pitch
Roger N Shepard · 1964
Earlier work this paper cites.
Aspects of Tone Perception
R Plomp · 1976
Earlier work this paper cites.
The Need for a New Medical Model: A Challenge for Biomedicine
George L. Engel · 1977
Earlier work this paper cites.
Updating quasi-newton matrices with limited storage
Jorge Nocedal · 1980
Earlier work this paper cites.
The fréchet distance between multivariate normal distributions
DC Dowson and BV Landau · 1982
Earlier work this paper cites.
Multiple stage vector quantization for speech coding
Biing-Hwang Juang and A. Gray · 1982
Earlier work this paper cites.
Geometrical approximations to the structure of musical pitch
Roger N Shepard · 1982
Earlier work this paper cites.
Signal estimation from modified short-time fourier transform
Daniel Griffin and Jae Lim · 1984
Earlier work this paper cites.
Learning internal representations by error propagation
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1986
Earlier work this paper cites.
Pengi: An implementation of a theory of activity
Philip E Agre and David Chapman · 1987
Earlier work this paper cites.
Situated knowledges: The science question in feminism and the privilege of partial perspective
Donna Haraway · 1988
Earlier work this paper cites.
A tutorial on hidden Markov models and selected applications in speech recognition
L.R. Rabiner · 1989
Earlier work this paper cites.
Nonlinear principal component analysis using autoassociative neural networks
Mark A. Kramer · 1991
Earlier work this paper cites.
Value-free science?: Purity and power in modern knowledge
Robert Proctor · 1991
Earlier work this paper cites.
Individual comparisons by ranking methods
Frank Wilcoxon · 1992
Earlier work this paper cites.
Ain’t Nothin’ Like the Real Thing, Baby : The Right of Publicity and the Singing Voice
Russel A. Stamets · 1994
Earlier work this paper cites.
Strong objectivity: A response to the new objectivity question
Sandra Harding · 1995
Earlier work this paper cites.
Stanford encyclopedia of philosophy, 1995
Edward N Zalta, Uri Nodelman, Colin Allen, and John Perry · 1995
Earlier work this paper cites.
Bias in computer systems
Batya Friedman and Helen Nissenbaum · 1996
Earlier work this paper cites.
Toward a Critical Technical Practice: Lessons Learned in Trying to Reform AI
Philip E Agre · 1998
Earlier work this paper cites.
The GUIDO notation format: A novel approach for adequately representing score-level music
Holger H. Hoos, Keith Hamel, Kai Renz, and Jürgen Kilian · 1998
Earlier work this paper cites.
Applications of machine learning to music research: Empirical investigations into the phenomenon of musical expression
Gerhard Widmer · 1998
Earlier work this paper cites.
Realtime chord recognition of musical sound: Asystem using common lisp music
Takuya Fujishima · 1999
Earlier work this paper cites.
Musicxml for notation and analysis
Michael Good · 2001
Earlier work this paper cites.
Representing score-level music using the guido music-notation format, 2001
Holger Hoos, Keith Hamel, Kai Renz, and Jürgen Killian · 2001
Earlier work this paper cites.
The Future of Ideas: The Fate of the Commons in a Connected World
Lawrence Lessig · 2001
Earlier work this paper cites.
Music information processing using the humdrum toolkit: Concepts, examples, and lessons
David Huron · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
The music encoding initiative (mei)
Perry Roland · 2002
Earlier work this paper cites.
Music information retrieval
J Stephen Downie · 2003
Earlier work this paper cites.
When creators, corporations and consumers collide: Napster and the development of on-line music distribution
Tom McCourt and Patrick Burkart · 2003
Earlier work this paper cites.
Lilypond, a system for automated music engraving
Han-Wen Nienhuys and Jan Nieuwenhuizen · 2003
Earlier work this paper cites.
Score following: State of the art and new developments
Nicola Orio, Serge Lemouton, and Diemo Schwarz · 2003
Earlier work this paper cites.
Automatic singer identification
Tong Zhang · 2003
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Improvisation: Methods and techniques for music therapy clinicians, educators, and students
Tony Wigram · 2004
Earlier work this paper cites.
“the way it sounds‘’: Timbre models for analysis and retrieval of music signals
J-J Aucouturier, François Pachet, and Mark Sandler · 2005
Earlier work this paper cites.
A tutorial on onset detection in music signals
Juan Pablo Bello, Laurent Daudet, Samer Abdallah, Chris Duxbury, Mike Davies, and Mark B Sandler · 2005
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
A review of audio fingerprinting
Pedro Cano, Eloi Batlle, Ton Kalker, and Jaap Haitsma · 2005
Earlier work this paper cites.
Dimensionality reduction by learning an invariant mapping
Raia Hadsell, Sumit Chopra, and Yann LeCun · 2006
Earlier work this paper cites.
Meeting the universe halfway: Quantum physics and the entanglement of matter and meaning
Karen Barad · 2007
Earlier work this paper cites.
Classifying music audio with timbral and chroma features
Daniel P. W. Ellis · 2007
Earlier work this paper cites.
Hard bargaining on the hard drive: gender bias in the music technology classroom
Victoria Armstrong · 2008
Earlier work this paper cites.
Evaluation methods for musical audio beat tracking algorithms
Matthew EP Davies, Norberto Degara, and Mark D Plumbley · 2009
Earlier work this paper cites.
Causal inference in statistics: An overview
Judea Pearl · 2009
Earlier work this paper cites.
For a Relational Musicology: Music and Interdisciplinarity, Beyond the Practice Turn: The 2007 Dent Medal Address
Georgina Born · 2010
Earlier work this paper cites.
A survey of audio-based music classification and annotation
Zhouyu Fu, Guojun Lu, Kai Ming Ting, and Dengsheng Zhang · 2010
Earlier work this paper cites.
Sonification Report: Status of the Field and Research Agenda
Gregory Kramer, Bruce Walker, Terri Bonebright, Perry Cook, John H. Flowers, Nadine Miner, and John Neuhoff · 2010
Earlier work this paper cites.
Audio cover song identification and similarity: background, approaches, evaluation, and beyond
Joan Serra, Emilia Gómez, and Perfecto Herrera · 2010
Earlier work this paper cites.
Automated sleep quality measurement using eeg signal: first step towards a domain specific music recommendation system
Wei Zhao, Xinxi Wang, and Ye Wang · 2010
Earlier work this paper cites.
Amazon mechanical turk: Gold mine or coal mine
G Adda and Kevin Bretonnel Cohen · 2011
Earlier work this paper cites.
The million song dataset
Thierry Bertin-Mahieux, Daniel PW Ellis, Brian Whitman, and Paul Lamere · 2011
Earlier work this paper cites.
Automatic singer identification based on auditory features
Wei Cai, Qiang Li, and Xin Guan · 2011
Earlier work this paper cites.
Musicglove: Motivating and quantifying hand movement rehabilitation by using functional grips to play music
Nizan Friedman, Vicky Chan, Danny Zondervan, Mark Bachman, and David J Reinkensmeyer · 2011
Earlier work this paper cites.
A multicultural approach in music information research
Xavier Serra · 2011
Earlier work this paper cites.
Musicxml
Michael D Good · 2012
Earlier work this paper cites.
Moving beyond feature design: Deep architectures and automatic feature learning in music informatics
Eric J. Humphrey, Juan Pablo Bello, and Yann LeCun · 2012
Earlier work this paper cites.
A scape plot representation for visualizing repetitive structures of music recordings
Meinard Müller and Nanzhu Jiang · 2012
Earlier work this paper cites.
Secure binary embeddings of front-end factor analysis for privacy preserving speaker verification
José Portelo, Alberto Abad, Bhiksha Raj, and Isabel Trancoso · 2013
Earlier work this paper cites.
Exploiting domain knowledge in music information research
Xavier Serra · 2013
Earlier work this paper cites.
Roadmap for music information research
Xavier Serra, Michela Magas, Emmanouil Benetos, Magdalena Chudy, Simon Dixon, Arthur Flexer, Emilia Gómez Gutiérrez, Fabien Gouyon, Herrera Boyer, Sergi Jordà Puig, et al · 2013
Earlier work this paper cites.
Improving singing language identification through i-vector extraction
Anna M. Kruspe · 2014
Earlier work this paper cites.
Auto-Encoding Variational Bayes
Diederik P. Kingma and Max Welling · 2014
Earlier work this paper cites.
Automatic chord estimation from audio: A review of the state of the art
Matt McVicar, Raúl Santos-Rodríguez, Yizhao Ni, and Tijl De Bie · 2014
Earlier work this paper cites.
Melody extraction from polyphonic music signals: Approaches, applications, and challenges
Justin Salamon, Emilia Gómez, Daniel PW Ellis, and Gaël Richard · 2014
Earlier work this paper cites.
Music information retrieval: Recent developments and applications
Markus Schedl, Emilia Gómez, Julián Urbano, et al · 2014
Earlier work this paper cites.
On the changing regulations of privacy and personal information in music information retrieval
Pierre Saurel, Francis Rousseaux, and Marc Danger · 2014
Earlier work this paper cites.
Deep metric learning using triplet network
Elad Hoffer and Nir Ailon · 2015
Earlier work this paper cites.
Objectivity and diversity: Another logic of scientific research
Sandra Harding · 2015
Earlier work this paper cites.
Fundamentals of Music Processing: Audio, Analysis, Algorithms, Applications
Meinard Müller · 2015
Earlier work this paper cites.
Sipth: Singing transcription based on hysteresis defined on the pitch-time curve
Emilio Molina, Lorenzo J. Tardón, Ana M. Barbancho, and Isabel Barbancho · 2015
Earlier work this paper cites.
Fundamentals of music processing: Audio, analysis, algorithms, applications
Meinard Müller · 2015
Earlier work this paper cites.
Hmm-based expressive singing voice synthesis with singing style control and robust pitch modeling
Takashi Nose, Misa Kanemoto, Tomoki Koriyama, and Takao Kobayashi · 2015
Earlier work this paper cites.
Acousticbrainz: a community platform for gathering music information obtained from audio
Alastair Porter, Dmitry Bogdanov, Robert Kaye, Roman Tsukanov, and Xavier Serra · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
An embarrassingly simple approach to zero-shot learning
Bernardino Romera-Paredes and Philip Torr · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli · 2015
Earlier work this paper cites.
Current Emotion Research in Music Psychology
Swathi Swaminathan and E. Glenn Schellenberg · 2015
Earlier work this paper cites.
Advances and perspectives in web technologies for music representation
Adriano Baratè, Goffredo Haus, Luca Andrea Ludovico, and Giorgio Presti · 2016
Earlier work this paper cites.
Madmom: A new python audio and music signal processing library
Sebastian Böck, Filip Korzeniowski, Jan Schlüter, Florian Krebs, and Gerhard Widmer · 2016
Earlier work this paper cites.
Posthuman Critical Theory
Rosi Braidotti · 2016
Earlier work this paper cites.
Big data’s disparate impact
Solon Barocas and Andrew D Selbst · 2016
Earlier work this paper cites.
Automatic tagging using deep convolutional neural networks
Keunwoo Choi, George Fazekas, and Mark Sandler · 2016
Earlier work this paper cites.
Fma: A dataset for music analysis
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, and Xavier Bresson · 2016
Earlier work this paper cites.
Automatic transcription of flamenco singing from polyphonic music recordings
Nadine Kroher and Emilia Gómez · 2016
Earlier work this paper cites.
The Effects of Music on Pain: A Meta-Analysis
Jin Hyung Lee · 2016
Earlier work this paper cites.
Samplernn: An unconditional end-to-end neural audio generation model
Soroush Mehri, Kundan Kumar, Ishaan Gulrajani, Rithesh Kumar, Shubham Jain, Jose Sotelo, Aaron Courville, and Yoshua Bengio · 2016
Earlier work this paper cites.
C-rnn-gan: Continuous recurrent neural networks with adversarial training
Olof Mogren · 2016
Earlier work this paper cites.
WORLD: A vocoder-based high-quality speech synthesis system for real-time applications
Masanori Morise, Fumiya Yokomori, and Kenji Ozawa · 2016
Earlier work this paper cites.
Musically Informed Sonification for Chronic Pain Rehabilitation: Facilitating Progress & Avoiding Over-Doing
Joseph W. Newbold, Nadia Bianchi-Berthouze, Nicolas E. Gold, Ana Tajadura-Jiménez, and Amanda CdC Williams · 2016
Earlier work this paper cites.
Exploring customer reviews for music genre classification and evolutionary studies
Sergio Oramas, Luis Espinosa-Anke, Aonghus Lawlor, et al · 2016
Earlier work this paper cites.
Amazon’s mechanical turk a digital sweatshop? transparency and accountability in crowdsourced online research
Matthew Pittman and Kim Sheehan · 2016
Earlier work this paper cites.
Learning-based methods for comparing sequences, with applications to audio-to-midi alignment and matching
Colin Raffel · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Earlier work this paper cites.
Improved deep metric learning with multi-class n-pair loss objective
Kihyuk Sohn · 2016
Earlier work this paper cites.
Music transcription modelling and composition using deep learning
Bob L Sturm, Joao Felipe Santos, Oded Ben-Tal, and Iryna Korshunova · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Aaron Van Den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, Koray Kavukcuoglu, et al · 2016
Earlier work this paper cites.
Project magenta: Generating long-term structure in songs and stories, 2016
Elliot Waite, Douglas Eck, Adam Roberts, and Daniel Abolafia · 2016
Earlier work this paper cites.
Towards deep interpretability (mus-rover ii): Learning hierarchical representations of tonal music
Haizi Yu and Lav R Varshney · 2016
Earlier work this paper cites.
Transfer learning for music classification and regression tasks
Keunwoo Choi, George Fazekas, Mark Sandler, and Kyunghyun Cho · 2017
Earlier work this paper cites.
Monoaural audio source separation using deep convolutional neural networks
Pritish Chandna, Marius Miron, Jordi Janer, and Emilia Gómez · 2017
Earlier work this paper cites.
Towards end-to-end polyphonic music transcription: Transforming music audio directly to a score
Ralf Gunter Correa Carvalho and Paris Smaragdis · 2017
Earlier work this paper cites.
Domain adaptation for visual applications: A comprehensive survey
Gabriela Csurka · 2017
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
Finale Doshi-Velez and Been Kim · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Earlier work this paper cites.
Cnn architectures for large-scale audio classification
Shawn Hershey, Sourish Chaudhuri, Daniel PW Ellis, Jort F Gemmeke, Aren Jansen, R Channing Moore, Manoj Plakal, Devin Platt, Rif A Saurous, Bryan Seybold, et al · 2017
Earlier work this paper cites.
Voice conversion from unaligned corpora using variational autoencoding wasserstein generative adversarial networks
Chin-Cheng Hsu, Hsin-Te Hwang, Yi-Chiao Wu, Yu Tsao, and Hsin-Min Wang · 2017
Earlier work this paper cites.
Interactive music generation with positional constraints using anticipation-rnns
Gaëtan Hadjeres and Frank Nielsen · 2017
Earlier work this paper cites.
DeepBach: A steerable model for bach chorales generation
Gaëtan Hadjeres, François Pachet, and Frank Nielsen · 2017
Earlier work this paper cites.
Musicowl: The music score ontology
Jim Jones, Diego de Siqueira Braga, Kleber Tertuliano, and Tomi Kauppinen · 2017
Earlier work this paper cites.
Automatic Stylistic Composition of Bach Chorales with Deep LSTM
Feynman Liang, Mark Gotham, Matthew Johnson, and Jamie Shotton · 2017
Earlier work this paper cites.
Sample-level deep convolutional neural networks for music auto-tagging using raw waveforms
Jongpil Lee, Jiyoung Park, Keunhyoung Luke Kim, and Juhan Nam · 2017
Earlier work this paper cites.
Chord generation from symbolic melody using blstm networks
Hyungui Lim, Seungyeon Rhyu, and Kyogu Lee · 2017
Earlier work this paper cites.
Program induction by rationale generation: Learning to solve and explain algebraic word problems
Wang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom · 2017
Earlier work this paper cites.
Human and Machine Hearing: Extracting Meaning from Sound
Richard F Lyon · 2017
Earlier work this paper cites.
Neural Discrete Representation Learning
Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aaron Van Den Oord, Oriol Vinyals, et al · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aaron Van Den Oord, Oriol Vinyals, et al · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
The biopsychosocial model of illness: A model whose time has come
Derick T Wade and Peter W Halligan · 2017
Earlier work this paper cites.
Tacotron: Towards end-to-end speech synthesis
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al · 2017
Earlier work this paper cites.
Improvised duet interaction: learning improvisation techniques for automatic accompaniment
Guangyu Xia and Roger B Dannenberg · 2017
Earlier work this paper cites.
Zero-shot learning-the good, the bad and the ugly
Yongqin Xian, Bernt Schiele, and Zeynep Akata · 2017
Earlier work this paper cites.
Midinet: A convolutional generative adversarial network for symbolic-domain music generation
Li-Chia Yang, Szu-Yu Chou, and Yi-Hsuan Yang · 2017
Earlier work this paper cites.
Seqgan: Sequence generative adversarial nets with policy gradient
Lantao Yu, Weinan Zhang, Jun Wang, and Yong Yu · 2017
Earlier work this paper cites.
Interpretable classification models for recidivism prediction
Jiaming Zeng, Berk Ustun, and Cynthia Rudin · 2017
Earlier work this paper cites.
Exploring tonal-dramatic relationships in richard wagner’s ring cycle
Frank Zalkow, Christof Weiß, and Meinard Müller · 2017
Earlier work this paper cites.
Platforms, promotion, and product discovery: Evidence from spotify playlists
Luis Aguiar and Joel Waldfogel · 2018
Earlier work this paper cites.
The relationship between music, culture, and society: meaning in music
Georgina Barton and Georgina Barton · 2018
Earlier work this paper cites.
Rhythmic entrainment for hand rehabilitation using the leap motion controller
Scott Beveridge, Estefanıa Cano, and Kat Agres · 2018
Earlier work this paper cites.
Midi-Vae: Modeling Dynamics and Instrumentation of Music with Applications to Style Transfer
Gino Brunner, Andres Konrad, Yuyi Wang, and Roger Wattenhofer · 2018
Earlier work this paper cites.
Musical source separation: An introduction
Estefania Cano, Derry FitzGerald, Antoine Liutkus, Mark D Plumbley, and Fabian-Robert Stöter · 2018
Earlier work this paper cites.
Bachprop: Learning to compose music in multiple styles
Florian Colombo and Wulfram Gerstner · 2018
Earlier work this paper cites.
Modeling temporal tonal relations in polyphonic music through deep networks with a novel image-based representation
Ching Hua Chuan and Dorien Herremans · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
A broader view on bias in automated decision-making: Reflecting on epistemology and dynamics
Roel Dobbe, Sarah Dean, Thomas Gilbert, and Nitin Kohli · 2018
Earlier work this paper cites.
Musegan: Multi-track sequential generative adversarial networks for symbolic music generation and accompaniment
Hao-Wen Dong, Wen-Yi Hsiao, Li-Chia Yang, and Yi-Hsuan Yang · 2018
Earlier work this paper cites.
The nes music database: A multi-instrumental dataset with expressive performance attributes
Chris Donahue, Huanru Henry Mao, and Julian McAuley · 2018
Earlier work this paper cites.
The challenge of realistic music generation: modelling raw audio at scale
Sander Dieleman, Aaron van den Oord, and Karen Simonyan · 2018
Earlier work this paper cites.
SVSGAN: singing voice separation via generative adversarial network
Zhe-Cheng Fan, Yen-Lin Lai, and Jyh-Shing Roger Jang · 2018
Earlier work this paper cites.
A data-driven analysis of workers’ earnings on amazon mechanical turk
Kotaro Hara, Abigail Adams, Kristy Milland, Saiph Savage, Chris Callison-Burch, and Jeffrey P Bigham · 2018
Earlier work this paper cites.
Aspects of data ethics in a changing world: Where are we now?
David J Hand · 2018
Earlier work this paper cites.
Transformer-nade for piano performances
Curtis Hawthorne, Anna Huang, Daphne Ippolito, and Douglas Eck · 2018
Earlier work this paper cites.
CBVMR: content-based video-music retrieval using soft intra-modal structure constraint
Sungeun Hong, Woobin Im, and Hyun Seung Yang · 2018
Earlier work this paper cites.
Ethical dimensions of music information retrieval technology
Andre Holzapfel, Bob Sturm, and Mark Coeckelbergh · 2018
Earlier work this paper cites.
Ethical Dimensions of Music Information Retrieval Technology
Andre Holzapfel, Bob L. Sturm, and Mark Coeckelbergh · 2018
Earlier work this paper cites.
Enabling factorized no music modeling and generation with the maestro dataset
Curtis Hawthorne, Andriy Stasyuk, Adam Roberts, Ian Simon, Cheng-Zhi Anna Huang, Sander Dieleman, Erich Elsen, Jesse Engel, and Douglas Eck · 2018
Earlier work this paper cites.
Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit, Noam Shazeer, Ian Simon, Curtis Hawthorne, Andrew M Dai, Matthew D Hoffman, Monica Dinculescu, and Douglas Eck · 2018
Earlier work this paper cites.
Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit, Noam Shazeer, Ian Simon, Curtis Hawthorne, Andrew M Dai, Matthew D Hoffman, Monica Dinculescu, and Douglas Eck · 2018
Earlier work this paper cites.
Imposing higher-level structure in polyphonic music generation using convolutional restricted boltzmann machines and constraints
Stefan Lattner, Maarten Grachten, and Gerhard Widmer · 2018
Earlier work this paper cites.
Creating a multitrack classical music performance dataset for multimodal music analysis: Challenges, insights, and applications
Bochen Li, Xinzhao Liu, Karthik Dinesh, Zhiyao Duan, and Gaurav Sharma · 2018
Earlier work this paper cites.
The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english
Steven R Livingstone and Frank A Russo · 2018
Earlier work this paper cites.
A music theory ontology
Sabbir M Rashid, David De Roure, and Deborah L McGuinness · 2018
Earlier work this paper cites.
A hierarchical latent vector model for learning long-term structure in music
Adam Roberts, Jesse Engel, Colin Raffel, Curtis Hawthorne, and Douglas Eck · 2018
Earlier work this paper cites.
A Hierarchical Latent Vector Model for Learning Long-Term Structure in Music
Adam Roberts, Jesse Engel, Colin Raffel, Curtis Hawthorne, and Douglas Eck · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training, 2018
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Adversarial semi-supervised audio source separation applied to singing voice extraction
Daniel Stoller, Sebastian Ewert, and Simon Dixon · 2018
Earlier work this paper cites.
Jointly detecting and separating singing voice: A multi-task approach
Daniel Stoller, Sebastian Ewert, and Simon Dixon · 2018
Earlier work this paper cites.
X-vectors: Robust DNN embeddings for speaker recognition
David Snyder, Daniel Garcia-Romero, Gregory Sell, Daniel Povey, and Sanjeev Khudanpur · 2018
Earlier work this paper cites.
Music lyrics summarization method using textrank algorithm
Jiyoung Son and Yongtae Shin · 2018
Earlier work this paper cites.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Earlier work this paper cites.
The sound of pixels
Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh McDermott, and Antonio Torralba · 2018
Earlier work this paper cites.
Effectiveness of self-supervised pre-training for speech recognition
Alexei Baevski, Michael Auli, and Abdelrahman Mohamed · 2019
Earlier work this paper cites.
Automatic music transcription: An overview
Emmanouil Benetos, Simon Dixon, Zhiyao Duan, and Sebastian Ewert · 2019
Earlier work this paper cites.
The mtg-jamendo dataset for automatic music tagging
Dmitry Bogdanov, Minz Won, Philip Tovstogan, Alastair Porter, and Xavier Serra · 2019
Earlier work this paper cites.
Wgansing: A multi-voice singing voice synthesizer based on the wasserstein-gan
Pritish Chandna, Merlijn Blaauw, Jordi Bonada, and Emilia Gómez · 2019
Earlier work this paper cites.
Generating long sequences with sparse transformers, 2019
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 2019
Earlier work this paper cites.
The ieee global initiative on ethics of autonomous and intelligent systems
Raja Chatila and John C Havens · 2019
Earlier work this paper cites.
Data Usage in MIR: History & Future Recommendations
Wenqin Chen, Jessica Keast, Jordan Moody, Corinne Moriarty, Felicia Villalobos, Virtue Winter, Xueqi Zhang, Xuanqi Lyu, Elizabeth Freeman, Jessie Wang, Sherry Cai, and Katherine Kinnaird · 2019
Earlier work this paper cites.
Zero-shot learning and knowledge transfer in music classification and tagging
Jeong Choi, Jongpil Lee, Jiyoung Park, and Juhan Nam · 2019
Earlier work this paper cites.
Audio content descriptors of timbre
Marcelo Caetano, Charalampos Saitis, and Kai Siedenburg · 2019
Earlier work this paper cites.
Automatic lyric transcription from karaoke vocal tracks: Resources and a baseline system
Gerardo Roa Dabike and Jon Barker · 2019
Earlier work this paper cites.
Location attention for extrapolation to longer sequences
Yann Dubois, Gautier Dagan, Dieuwke Hupkes, and Elia Bruni · 2019
Earlier work this paper cites.
Decomposed: the political ecology of music
Kyle Devine · 2019
Earlier work this paper cites.
Music source separation in the waveform domain
Alexandre Défossez, Nicolas Usunier, Léon Bottou, and Francis Bach · 2019
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G. Carbonell, Quoc Viet Le, and Ruslan Salakhutdinov · 2019
Earlier work this paper cites.
Spotify teardown: Inside the black box of streaming music
Maria Eriksson, Rasmus Fleischer, Anna Johansson, Pelle Snickars, and Patrick Vonderau · 2019
Earlier work this paper cites.
Song lyrics summarization inspired by audio thumbnailing
Michael Fell, Elena Cabrio, Fabien Gandon, and Alain Giboin · 2019
Earlier work this paper cites.
Value sensitive design: shaping technology with moral imagination
Batya Friedman and David F. Hendry · 2019
Earlier work this paper cites.
“telling me not to worry…” hyperscanning and neural dynamics of emotion processing during guided imagery and music
Jörg C Fachner, Clemens Maidhof, Denise Grocke, Inge Nygaard Pedersen, Gro Trondalen, Gerhard Tucek, and Lars O Bonde · 2019
Earlier work this paper cites.
Fairness, accountability and transparency in music information research (fat-mir)
Emilia Gomez, Andre Holzapfel, Marius Miron, and Bob L Sturm · 2019
Earlier work this paper cites.
Star-transformer
Qipeng Guo, Xipeng Qiu, Pengfei Liu, Yunfan Shao, Xiangyang Xue, and Zheng Zhang · 2019
Earlier work this paper cites.
A gan model with self-attention mechanism to generate multi-instruments symbolic music
Faqian Guan, Chunyan Yu, and Suqiong Yang · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for nlp
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly · 2019
Earlier work this paper cites.
Timbretron: Awavenet(yclegan(Cqt(Audio))) Pipeline for Musical Timbre Transfer
Sicong Huang, Qiyang Li, Cem Anil, Xuchan Bao, Sageev Oore, and Roger B. Grosse · 2019
Earlier work this paper cites.
Joint singing voice separation and F0 estimation with deep u-net architectures
Andreas Jansson, Rachel M. Bittner, Sebastian Ewert, and Tillman Weyde · 2019
Earlier work this paper cites.
Virtuosonet: A hierarchical rnn-based system for modeling expressive piano performance
Dasaem Jeong, Taegyun Kwon, Yoojin Kim, Kyogu Lee, and Juhan Nam · 2019
Earlier work this paper cites.
Graph neural network for music score data and modeling expressive piano performance
Dasaem Jeong, Taegyun Kwon, Yoojin Kim, and Juhan Nam · 2019
Earlier work this paper cites.
Melgan: Generative adversarial networks for conditional waveform synthesis
Kundan Kumar, Rithesh Kumar, Thibault De Boissiere, Lucas Gestin, Wei Zhen Teoh, Jose Sotelo, Alexandre De Brebisson, Yoshua Bengio, and Aaron C Courville · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova · 2019
Earlier work this paper cites.
Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms
Kevin Kilgour, Mauricio Zuluaga, Dominik Roblek, and Matthew Sharifi · 2019
Earlier work this paper cites.
Learning Disentangled Representations of Timbre and Pitch for Musical Instrument Sounds Using Gaussian Mixture Variational Autoencoders
Yin-Jyun Luo, Kat Agres, and Dorien Herremans · 2019
Earlier work this paper cites.
Query by video: Cross-modal music retrieval
Bochen Li and Aparna Kumar · 2019
Earlier work this paper cites.
Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation
Yi Luo and Nima Mesgarani · 2019
Earlier work this paper cites.
Learning a joint embedding space of monophonic and mixed music signals for singing voice
Kyungyun Lee and Juhan Nam · 2019
Earlier work this paper cites.
Cross-modal music retrieval and applications: An overview of key methodologies
Meinard Müller, Andreas Arzt, Stefan Balke, Matthias Dorfer, and Gerhard Widmer · 2019
Earlier work this paper cites.
This thing called fairness: Disciplinary confusion realizing a value in technology
Deirdre K Mulligan, Joshua A Kroll, Nitin Kohli, and Richmond Y Wong · 2019
Earlier work this paper cites.
Virtual adversarial training: A regularization method for supervised and semi-supervised learning
Takeru Miyato, Shin-ichi Maeda, Masanori Koyama, and Shin Ishii · 2019
Earlier work this paper cites.
Explaining explanations in ai
Brent Mittelstadt, Chris Russell, and Sandra Wachter · 2019
Earlier work this paper cites.
Automatic singing transcription based on encoder-decoder recurrent neural networks with a weakly-supervised attention mechanism
Ryo Nishikimi, Eita Nakamura, Satoru Fukayama, Masataka Goto, and Kazuyoshi Yoshii · 2019
Earlier work this paper cites.
End-to-end melody note transcription based on a beat-synchronous attention mechanism
Ryo Nishikimi, Eita Nakamura, Masataka Goto, and Kazuyoshi Yoshii · 2019
Earlier work this paper cites.
Hierarchical transformers for long document classification
Raghavendra Pappagari, Piotr Zelasko, Jesús Villalba, Yishay Carmiel, and Najim Dehak · 2019
Earlier work this paper cites.
Blockwise self-attention for long document understanding
Jiezhong Qiu, Hao Ma, Omer Levy, Scott Wen-tau Yih, Sinong Wang, and Jie Tang · 2019
Earlier work this paper cites.
Autovc: Zero-shot voice style transfer with only autoencoder loss
Kaizhi Qian, Yang Zhang, Shiyu Chang, Xuesong Yang, and Mark Hasegawa-Johnson · 2019
Earlier work this paper cites.
Fastspeech: Fast, robust and controllable text to speech
Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu · 2019
Earlier work this paper cites.
Generating diverse high-fidelity images with vq-vae-2
Ali Razavi, Aaron Van den Oord, and Oriol Vinyals · 2019
Earlier work this paper cites.
wav2vec: Unsupervised pre-training for speech recognition
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli · 2019
Earlier work this paper cites.
End-to-end lyrics alignment for polyphonic music using an audio-to-character recognition model
Daniel Stoller, Simon Durand, and Sebastian Ewert · 2019
Earlier work this paper cites.
On the importance of audio-source separation for singer identification in polyphonic music
Bidisha Sharma, Rohan Kumar Das, and Haizhou Li · 2019
Earlier work this paper cites.
Artificial Intelligence and Music: Open Questions of Copyright Law and Engineering Praxis
Bob L. T. Sturm, Maria Iglesias, Oded Ben-Tal, Marius Miron, and Emilia Gómez · 2019
Earlier work this paper cites.
Artificial intelligence and music: open questions of copyright law and engineering praxis
Bob LT Sturm, Maria Iglesias, Oded Ben-Tal, Marius Miron, and Emilia Gómez · 2019
Earlier work this paper cites.
Timbre: Acoustics, Perception, and Cognition
Kai Siedenburg, Charalampos Saitis, Stephen McAdams, Arthur N Popper, and Richard R Fay · 2019
Earlier work this paper cites.
Sex and gender analysis improves science and engineering
Cara Tannenbaum, Robert P Ellis, Friederike Eyssel, James Zou, and Londa Schiebinger · 2019
Earlier work this paper cites.
Learning affective correspondence between music and image
Gaurav Verma, Eeshan Gunesh Dhekane, and Tanaya Guha · 2019
Earlier work this paper cites.
Gender differences in the global music industry: Evidence from musicbrainz and the echo nest
Yixue Wang and Emőke-Ágnes Horvát · 2019
Earlier work this paper cites.
Dermoscopy diagnosis of cancerous lesions utilizing dual deep learning algorithms via visual and audio (sonification) outputs: Laboratory and prospective observational studies
B. N. Walker, J. M. Rehg, A. Kalra, R. M. Winters, P. Drews, J. Dascalu, E. O. David, and A. Dascalu · 2019
Earlier work this paper cites.
Neural source-filter waveform models for statistical parametric speech synthesis
Xin Wang, Shinji Takaki, and Junichi Yamagishi · 2019
Earlier work this paper cites.
Bp-transformer: Modelling long-range context via binary partitioning, 2019
Zihao Ye, Qipeng Guo, Quan Gan, Xipeng Qiu, and Zheng Zhang · 2019
Earlier work this paper cites.
A Self-Consistent Sonification Method to Translate Amino Acid Sequences into Musical Compositions and Application in Protein Design Using Artificial Intelligence
Chi-Hua Yu, Zhao Qin, Francisco J. Martin-Martinez, and Markus J. Buehler · 2019
Earlier work this paper cites.
Deep Music Analogy Via Latent Representation Disentanglement
Ruihan Yang, Dingsu Wang, Ziyu Wang, Tianyao Chen, Junyan Jiang, and Gus Xia · 2019
Earlier work this paper cites.
Bandnet: A neural network-based, multi-instrument beatles-style midi music composition machine
Yichao Zhou, Wei Chu, Sam Young, and Xin Chen · 2019
Earlier work this paper cites.
Explainable artificial intelligence (xai): Concepts, taxonomies, opportunities and challenges toward responsible ai
Alejandro Barredo Arrieta, Natalia Díaz-Rodríguez, Javier Del Ser, Adrien Bennetot, Siham Tabik, Alberto Barbado, Salvador García, Sergio Gil-López, Daniel Molina, Richard Benjamins, et al · 2020
Earlier work this paper cites.
Effectiveness of self-supervised pre-training for speech recognition, May 2020
Alexei Baevski, Michael Auli, and Abdelrahman Mohamed · 2020
Earlier work this paper cites.
From ethics washing to ethics bashing: a view on tech ethics from within moral philosophy
Elettra Bietti · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Diversifying MIR: Knowledge and Real-World Challenges, and New Interdisciplinary Futures
Georgina Born · 2020
Earlier work this paper cites.
Longformer: The long-document transformer, 2020
Iz Beltagy, Matthew E. Peters, and Arman Cohan · 2020
Earlier work this paper cites.
wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations
Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli · 2020
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli · 2020
Earlier work this paper cites.
Intersectionality, 2nd Edition
Patricia Hill Collins and Sirma Bilge · 2020
Earlier work this paper cites.
Artificial creative intelligence: Breaking the imitation barrier
Rowland Chen, Roger B. Dannenberg, Bhiksha Raj, and Rita Singh · 2020
Earlier work this paper cites.
Improved baselines with momentum contrastive learning
Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He · 2020
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Earlier work this paper cites.
An Automatic Method to Develop Music With Music Segment and Long Short Term Memory for Tinnitus Music Therapy
Jiemei Chen, Fan Pan, Ping Zhong, Tiantian He, Leiyu Qi, Jingzhe Lu, Peiyu He, and Yun Zheng · 2020
Earlier work this paper cites.
Hifisinger: Towards high-fidelity neural singing voice synthesis
Jiawei Chen, Xu Tan, Jian Luan, Tao Qin, and Tie-Yan Liu · 2020
Earlier work this paper cites.
Music Sketchnet: Controllable Music Generation Via Factorized Representations of Pitch and Rhythm
Ke Chen, Cheng I. Wang, Taylor Berg-Kirkpatrick, and Shlomo Dubnov · 2020
Earlier work this paper cites.
Understanding optical music recognition
Jorge Calvo-Zaragoza, Jan Hajič Jr, and Alexander Pacha · 2020
Earlier work this paper cites.
Automatic lyrics transcription using dilated convolutional neural networks with self-attention
Emir Demirel, Sven Ahlbäck, and Simon Dixon · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Earlier work this paper cites.
Jukebox: A generative model for music
Prafulla Dhariwal, Heewoo Jun, Christine Payne, Jong Wook Kim, Alec Radford, and Ilya Sutskever · 2020
Earlier work this paper cites.
Jukebox: A generative model for music
Prafulla Dhariwal, Heewoo Jun, Christine Payne, Jong Wook Kim, Alec Radford, and Ilya Sutskever · 2020
Earlier work this paper cites.
Rhythm, Chord and Melody Generation for Lead Sheets Using Recurrent Neural Networks
Cedric De Boom, Stephanie Van Laere, Tim Verbelen, and Bart Dhoedt · 2020
Earlier work this paper cites.
Artist gender representation in music streaming
Avriel Epps-Darling, Henriette Cramer, and Romain Takeo Bouyer · 2020
Earlier work this paper cites.
Mmm : Exploring conditional multi-track music generation with the transformer, 2020
Jeff Ens and Philippe Pasquier · 2020
Earlier work this paper cites.
Coala: Co-aligned autoencoders for learning semantically enriched audio representations
Xavier Favory, Konstantinos Drossos, Tuomas Virtanen, and Xavier Serra · 2020
Earlier work this paper cites.
Computer-generated music for tabletop role-playing games
Lucas N. Ferreira, Levi H. S. Lelis, and Jim Whitehead · 2020
Earlier work this paper cites.
Foley music: Learning to generate music from videos
Chuang Gan, Deng Huang, Peihao Chen, Joshua B. Tenenbaum, and Antonio Torralba · 2020
Earlier work this paper cites.
Foley music: Learning to generate music from videos
Chuang Gan, Deng Huang, Peihao Chen, Joshua B Tenenbaum, and Antonio Torralba · 2020
Earlier work this paper cites.
Interacting with gpt-2 to generate controlled and believable musical sequences in abc notation
Carina Geerlings and Albert Merono-Penuela · 2020
Earlier work this paper cites.
Generative adversarial networks
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2020
Earlier work this paper cites.
Bootstrap your own latent-a new approach to self-supervised learning
Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, et al · 2020
Earlier work this paper cites.
Addressing the confounds of accompaniments in singer identification
Tsung-Han Hsieh, Kai-Hsiang Cheng, Zhe-Cheng Fan, Yu-Ching Yang, and Yi-Hsuan Yang · 2020
Earlier work this paper cites.
Towards a critical race methodology in algorithmic fairness
Alex Hanna, Emily Denton, Andrew Smart, and Jamila Smith-Loud · 2020
Earlier work this paper cites.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
Scaling laws for autoregressive generative modeling
Tom Henighan, Jared Kaplan, Mor Katz, Mark Chen, Christopher Hesse, Jacob Jackson, Heewoo Jun, Tom B Brown, Prafulla Dhariwal, Scott Gray, et al · 2020
Earlier work this paper cites.
Pop music transformer: Beat-based modeling and generation of expressive pop piano compositions
Yu Siang Huang and Yi Hsuan Yang · 2020
Earlier work this paper cites.
Pop music transformer: Beat-based modeling and generation of expressive pop piano compositions
Yu-Siang Huang and Yi-Hsuan Yang · 2020
Earlier work this paper cites.
Revisiting representation learning for singing voice separation with sinkhorn distances
Stylianos Ioannis Mimilakis, Konstantinos Drossos, and Gerald Schuller · 2020
Earlier work this paper cites.
Exploring aligned lyrics-informed singing voice separation
Chang-Bin Jeon, Hyeong-Seok Choi, and Kyogu Lee · 2020
Earlier work this paper cites.
RL-Duet: Online music accompaniment generation using deep reinforcement learning
Nan Jiang, Sheng Jin, Zhiyao Duan, and Changshui Zhang · 2020
Earlier work this paper cites.
Shulei Ji, Jing Luo, and Xinyu Yang · 2020
Earlier work this paper cites.
Transformer VAE: A Hierarchical Model for Structure-Aware and Interpretable Music Representation Learning
Junyan Jiang, Gus G. Xia, Dave B. Carlton, Chris N. Anderson, and Ryan H. Miyakawa · 2020
Earlier work this paper cites.
Panns: Large-scale pretrained audio neural networks for audio pattern recognition
Qiuqiang Kong, Yin Cao, Turab Iqbal, Yuxuan Wang, Wenwu Wang, and Mark D Plumbley · 2020
Earlier work this paper cites.
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae · 2020
Earlier work this paper cites.
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae · 2020
Earlier work this paper cites.
Reformer: The efficient transformer
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya · 2020
Earlier work this paper cites.
Giantmidi-piano: A large-scale midi dataset for classical piano music
Qiuqiang Kong, Bochen Li, Jitong Chen, and Yuxuan Wang · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
Sight to sound: An end-to-end approach for visual piano transcription
A. Sophia Koepke, Olivia Wiles, Yael Moses, and Andrew Zisserman · 2020
Earlier work this paper cites.
Disentangling timbre and singing style with multi-singer singing synthesis system
Juheon Lee, Hyeong-Seok Choi, Junghyun Koo, and Kyogu Lee · 2020
Earlier work this paper cites.
Singing voice conversion with disentangled representations of singer and vocal technique using variational autoencoders
Yin-Jyun Luo, Chin-Cheng Hsu, Kat Agres, and Dorien Herremans · 2020
Earlier work this paper cites.
Contrastive representation learning: A framework and review
Phuc H Le-Khac, Graham Healy, and Alan F Smeaton · 2020
Earlier work this paper cites.
Octopus: Comprehensive and elastic user representation for the generation of recommendation candidates
Zheng Liu, Jianxun Lian, Junhan Yang, Defu Lian, and Xing Xie · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive NLP tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela · 2020
Earlier work this paper cites.
Xiaoicesing: A high-quality and integrated singing voice synthesis system
Peiling Lu, Jie Wu, Jian Luan, Xu Tan, and Li Zhou · 2020
Earlier work this paper cites.
VAW-GAN for singing voice conversion with non-parallel training data
Junchen Lu, Kun Zhou, Berrak Sisman, and Haizhou Li · 2020
Earlier work this paper cites.
Feedback loop and bias amplification in recommender systems
Masoud Mansoury, Himan Abdollahpouri, Mykola Pechenizkiy, Bamshad Mobasher, and Robin Burke · 2020
Earlier work this paper cites.
Diversity and inclusion metrics in subset selection
Margaret Mitchell, Dylan Baker, Nyalleng Moorosi, Emily Denton, Ben Hutchinson, Alex Hanna, Timnit Gebru, and Jamie Morgenstern · 2020
Earlier work this paper cites.
Music Therapy in the Treatment of Dementia: A Systematic Review and Meta-Analysis
Celia Moreno-Morales, Raul Calero, Pedro Moreno-Morales, and Cristina Pintado · 2020
Earlier work this paper cites.
Unsupervised interpretable representation learning for singing voice separation
Stylianos I. Mimilakis, Konstantinos Drossos, and Gerald Schuller · 2020
Earlier work this paper cites.
A Novel Method for Sleep-Stage Classification Based on Sonification of Sleep Electroencephalogram Signals Using Wavelet Transform and Recurrent Neural Network
Foad Moradi, Hiwa Mohammadi, Mohammad Rezaei, Payam Sariaslani, Nazanin Razazian, Habibolah Khazaie, and Hojjat Adeli · 2020
Cited alongside, same era.
Content based singing voice source separation via strong conditioning using aligned phonemes
Gabriel Meseguer-Brocal and Geoffroy Peeters · 2020
Cited alongside, same era.
Zero-shot singing voice conversion
Shahan Nercessian · 2020
Cited alongside, same era.
DRUMGAN: synthesis of drum sounds with timbral feature conditioning using generative adversarial networks
Javier Nistal, Stefan Lattner, and Gaël Richard · 2020
Cited alongside, same era.
Deep learning based source separation applied to choir ensembles
Darius Petermann, Pritish Chandna, Helena Cuesta, Jordi Bonada, and Emilia Gómez · 2020
Cited alongside, same era.
Lp-musiccaps: Llm-based pseudo music captioning
Seungheon Doh, Keunwoo Choi, Jongpil Lee, and Juhan Nam · 2023
Later among the works it cites.
Lp-musiccaps: Llm-based pseudo music captioning
SeungHeon Doh, Keunwoo Choi, Jongpil Lee, and Juhan Nam · 2023
Later among the works it cites.
Singsong: Generating musical accompaniments from singing
Chris Donahue, Antoine Caillon, Adam Roberts, Ethan Manilow, Philippe Esling, Andrea Agostinelli, Mauro Verzetti, Ian Simon, Olivier Pietquin, Neil Zeghidour, et al · 2023
Later among the works it cites.
Pengi: An audio language model for audio tasks
Soham Deshmukh, Benjamin Elizalde, Rita Singh, and Huaming Wang · 2023
Later among the works it cites.
Fairness in recommender systems: research landscape and future directions
Yashar Deldjoo, Dietmar Jannach, Alejandro Bellogin, Alessandro Difonzo, and Dario Zanzonelli · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Blockwise self-attention for long document understanding
Jiezhong Qiu, Hao Ma, Omer Levy, Wen-tau Yih, Sinong Wang, and Jie Tang · 2020
Cited alongside, same era.
Popmag: Pop music accompaniment generation
Yi Ren, Jinzheng He, Xu Tan, Tao Qin, Zhou Zhao, and Tie-Yan Liu · 2020
Cited alongside, same era.
Fastspeech 2: Fast and high-quality end-to-end text to speech
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu · 2020
Cited alongside, same era.
Compressive transformers for long-range sequence modelling
Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier, and Timothy P. Lillicrap · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Cited alongside, same era.
Multi-task self-supervised learning for robust speech recognition
Mirco Ravanelli, Jianyuan Zhong, Santiago Pascual, Pawel Swietojanski, Joao Monteiro, Jan Trmal, and Yoshua Bengio · 2020
Cited alongside, same era.
Sound of Care: Towards a Co-Operative AI Digital Pain Companion to Support People with Chronic Primary Pain
Bleiz Macsen Del Sette, Dawn Carnes, and Charalampos Saitis · 2023
Later among the works it cites.
Contrastive learning-based audio to lyrics alignment for multiple languages
Simon Durand, Daniel Stoller, and Sebastian Ewert · 2023
Later among the works it cites.
High quality audio coding with mdctnet
Grant Davidson, Mark Vinton, Per Ekstrand, Cong Zhou, Lars Villemoes, and Lie Lu · 2023
Later among the works it cites.
Toward universal text-to-music retrieval
SeungHeon Doh, Minz Won, Keunwoo Choi, and Juhan Nam · 2023
Later among the works it cites.
CLAP learning audio concepts from natural language supervision
Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Ismail, and Huaming Wang · 2023
Later among the works it cites.
Byte pair encoding for symbolic music
Nathan Fradet, Nicolas Gutowski, Fabien Chhel, and Jean-Pierre Briot · 2023
Later among the works it cites.
How much context does my attention-based asr system need?, 2023
Robert Flynn and Anton Ragni · 2023
Later among the works it cites.
CTCBERT: Advancing hidden-unit bert with ctc objectives
Ruchao Fan, Yiming Wang, Yashesh Gaur, and Jinyu Li · 2023
Later among the works it cites.
Mamba: Linear-time sequence modeling with selective state spaces, 2023
Albert Gu and Tri Dao · 2023
Later among the works it cites.
Llark: A multimodal foundation model for music
Josh Gardner, Simon Durand, Daniel Stoller, and Rachel M. Bittner · 2023
Later among the works it cites.
Llark: A multimodal instruction-following language model for music
Joshua P Gardner, Simon Durand, Daniel Stoller, and Rachel M Bittner · 2023
Later among the works it cites.
Audiovisual masked autoencoders
Mariana-Iuliana Georgescu, Eduardo Fonseca, Radu Tudor Ionescu, Mario Lucic, Cordelia Schmid, and Anurag Arnab · 2023
Later among the works it cites.
Polyscriber: Integrated fine-tuning of extractor and lyrics transcriber for polyphonic music
Xiaoxue Gao, Chitralekha Gupta, and Haizhou Li · 2023
Later among the works it cites.
Planting a SEED of vision in large language model
Yuying Ge, Yixiao Ge, Ziyun Zeng, Xintao Wang, and Ying Shan · 2023
Later among the works it cites.
A domain-knowledge-inspired music embedding space and a novel attention mechanism for symbolic music modeling
Zixun Guo, Jaeyong Kang, and Dorien Herremans · 2023
Later among the works it cites.
Yuan Gong, Hongyin Luo, Alexander H Liu, Leonid Karlinsky, and James Glass · 2023
Later among the works it cites.
Text-to-audio generation using instruction guided latent diffusion model
Deepanway Ghosal, Navonil Majumder, Ambuj Mehrish, and Soujanya Poria · 2023
Later among the works it cites.
Text-to-audio generation using instruction guided latent diffusion model
Deepanway Ghosal, Navonil Majumder, Ambuj Mehrish, and Soujanya Poria · 2023
Later among the works it cites.
Contrastive Audio-Visual Masked Autoencoder
Yuan Gong, Andrew Rouditchenko, Alexander H. Liu, David Harwath, Leonid Karlinsky, Hilde Kuehne, and James R. Glass · 2023
Later among the works it cites.
Ultimate negative sampling for contrastive learning
Huijie Guo and Lei Shi · 2023
Later among the works it cites.
Vampnet: Music generation via masked acoustic token modeling
Hugo F Flores Garcia, Prem Seetharaman, Rithesh Kumar, and Bryan Pardo · 2023
Later among the works it cites.
A dual-path cross-modal network for video-music retrieval
Xin Gu, Yinghua Shen, and Chaohui Lv · 2023
Later among the works it cites.
Self-transriber: Few-shot lyrics transcription with self-training
Xiaoxue Gao, Xianghu Yue, and Haizhou Li · 2023
Later among the works it cites.
Multi-Source Contrastive Learning from Musical Audio, May 2023
Christos Garoufis, Athanasia Zlatintsi, and Petros Maragos · 2023
Later among the works it cites.
Towards reasoning in large language models: A survey
Jie Huang and Kevin Chen-Chuan Chang · 2023
Later among the works it cites.
Unisinger: Unified end-to-end singing voice synthesis with cross-modality information matching
Zhiqing Hong, Chenye Cui, Rongjie Huang, Lichao Zhang, Jinglin Liu, Jinzheng He, and Zhou Zhao · 2023
Later among the works it cites.
The impact of algorithmically driven recommendation systems on music consumption and production: A literature review
David Hesmondhalgh, Raquel Campos Valverde, D Kaye, and Zhongwei Li · 2023
Later among the works it cites.
Instructme: An instruction guided music edit and remix framework with latent diffusion models
Bing Han, Junyu Dai, Xuchen Song, Weituo Hao, Xinyan He, Dong Guo, Jitong Chen, Yuxuan Wang, and Yanmin Qian · 2023
Later among the works it cites.
InstructME: An instruction guided music edit and remix framework with latent diffusion models
Bing Han, Junyu Dai, Xuchen Song, Weituo Hao, Xinyan He, Dong Guo, Jitong Chen, Yuxuan Wang, and Yanmin Qian · 2023
Later among the works it cites.
Reasoning with language model is planning with world model
Shibo Hao, Yi Gu, Haodi Ma, Joshua Hong, Zhen Wang, Daisy Wang, and Zhiting Hu · 2023
Later among the works it cites.
ALCAP: alignment-augmented music captioner
Zihao He, Weituo Hao, Wei Tsung Lu, Changyou Chen, Kristina Lerman, and Xuchen Song · 2023
Later among the works it cites.
Beyond Diverse Datasets: Responsible MIR, Interdisciplinarity, and the Fractured Worlds of Music
Rujing Stacy Huang, Andre Holzapfel, Bob L. T. Sturm, and Anna-Kaisa Kaila · 2023
Later among the works it cites.
Make-An-Audio: Text-to-audio generation with latent diffusion models
Rongjie Huang, Jiawei Huang, Dongchao Yang, Yi Ren, Luping Liu, Mingze Li, Zhenhui Ye, Jinglin Liu, Xiang Yin, and Zhou Zhao · 2023
Later among the works it cites.
Contrastive masked autoencoders are stronger vision learners
Zhicheng Huang, Xiaojie Jin, Chengze Lu, Qibin Hou, Ming-Ming Cheng, Dongmei Fu, Xiaohui Shen, and Jiashi Feng · 2023
Later among the works it cites.
Improving sample quality of diffusion models using self-attention guidance
Susung Hong, Gyuseong Lee, Wooseok Jang, and Seungryong Kim · 2023
Later among the works it cites.
Ji-Sang Hwang, Sang-Hoon Lee, and Seong-Whan Lee · 2023
Later among the works it cites.
Atin Sakkeer Hussain, Shansong Liu, Chenshuo Sun, and Ying Shan · 2023
Later among the works it cites.
Audiogpt: Understanding and generating speech, music, sound, and talking head
Rongjie Huang, Mingze Li, Dongchao Yang, Jiatong Shi, Xuankai Chang, Zhenhui Ye, Yuning Wu, Zhiqing Hong, Jiawei Huang, Jinglin Liu, et al · 2023
Later among the works it cites.
Music composition with deep learning: A review
Carlos Hernandez-Olivan and José R. Beltrán · 2023
Later among the works it cites.
Noise2music: Text-conditioned music generation with diffusion models
Qingqing Huang, Daniel S Park, Tao Wang, Timo I Denk, Andy Ly, Nanxin Chen, Zhengdong Zhang, Zhishuai Zhang, Jiahui Yu, Christian Frank, et al · 2023
Later among the works it cites.
Noise2music: Text-conditioned music generation with diffusion models
Qingqing Huang, Daniel S. Park, Tao Wang, Timo I. Denk, Andy Ly, Nanxin Chen, Zhengdong Zhang, Zhishuai Zhang, Jiahui Yu, Christian Havnø Frank, Jesse H. Engel, Quoc V. Le, William Chan, and Wei Han · 2023
Later among the works it cites.
Make-An-Audio 2: Temporal-enhanced text-to-audio generation
Jiawei Huang, Yi Ren, Rongjie Huang, Dongchao Yang, Zhenhui Ye, Chen Zhang, Jinglin Liu, Xiang Yin, Zejun Ma, and Zhou Zhao · 2023
Later among the works it cites.
W2VC: wavlm representation based one-shot voice conversion with gradient reversal distillation and CTC supervision
Hao Huang, Lin Wang, Jichen Yang, Ying Hu, and Liang He · 2023
Later among the works it cites.
Emotion based Music Recommendation System using Deep Learning Model
Jaladi Sam Joel, B. Ernest Thompson, Steve Renny Thomas, T. Revanth Kumar, Shajin Prince, and D Bini · 2023
Later among the works it cites.
A Systematic Review of Scientific Studies on the Effects of Music Therapy on Individuals with Autism Spectrum Disorder
Udeme Samuel Jacob and Jace Pillay · 2023
Later among the works it cites.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al · 2023
Later among the works it cites.
Self-supervised representations for singing voice conversion
Tejas Jayashankar, Jilong Wu, Leda Sari, David Kant, Vimal Manohar, and Qing He · 2023
Later among the works it cites.
A survey on deep learning for symbolic music generation: Representations, algorithms, evaluations, and challenges
Shulei Ji, Xinyu Yang, and Jing Luo · 2023
Later among the works it cites.
Musical voice separation as link prediction: Modeling a musical perception task as a multi-trajectory tracking problem
Emmanouil Karystinaios, Francesco Foscarin, and Gerhard Widmer · 2023
Later among the works it cites.
Latent diffusion models: Is the generative ai revolution happening in latent space?, 2023
Arash Vahdat Karsten Kreis, Ruiqi Gao · 2023
Later among the works it cites.
Muse-svs: Multi-singer emotional singing voice synthesizer that controls emotional intensity
Sungjae Kim, Yewon Kim, Jewoo Jun, and Injung Kim · 2023
Later among the works it cites.
Video2music: Suitable music generation from videos using an affective multimodal transformer model
Jaeyong Kang, Soujanya Poria, and Dorien Herremans · 2023
Later among the works it cites.
High-fidelity audio compression with improved rvqgan
Rithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar, and Kundan Kumar · 2023
Later among the works it cites.
Audiogen: Textually guided audio generation
Felix Kreuk, Gabriel Synnaeve, Adam Polyak, Uriel Singer, Alexandre Défossez, Jade Copet, Devi Parikh, Yaniv Taigman, and Yossi Adi · 2023
Later among the works it cites.
Speak, read and prompt: High-fidelity text-to-speech with minimal supervision
Eugene Kharitonov, Damien Vincent, Zalán Borsos, Raphaël Marinier, Sertan Girgin, Olivier Pietquin, Matt Sharifi, Marco Tagliasacchi, and Neil Zeghidour · 2023
Later among the works it cites.
A survey on deep reinforcement learning for audio-based applications
Siddique Latif, Heriberto Cuayáhuitl, Farrukh Pervez, Fahad Shamshad, Hafiz Shehbaz Ali, and Erik Cambria · 2023
Later among the works it cites.
JEN-1: text-guided universal music generation with omnidirectional diffusion models
Peike Li, Boyu Chen, Yao Yao, Yikai Wang, Allen Wang, and Alex Wang · 2023
Later among the works it cites.
Jen-1: Text-guided universal music generation with omnidirectional diffusion models
Peike Li, Boyu Chen, Yao Yao, Yikai Wang, Allen Wang, and Alex Wang · 2023
Later among the works it cites.
AudioLDM: Text-to-audio generation with latent diffusion models
Haohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei, Xubo Liu, Danilo Mandic, Wenwu Wang, and Mark D Plumbley · 2023
Later among the works it cites.
AudioLDM: Text-to-audio generation with latent diffusion models
Haohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei, Xubo Liu, Danilo Mandic, Wenwu Wang, and Mark D Plumbley · 2023
Later among the works it cites.
Annotator subjectivity in the musiccaps dataset
Minhee Lee, Seungheon Doh, and Dasaem Jeong · 2023
Later among the works it cites.
Music understanding llama: Advancing text-to-music generation with question answering and captioning
Shansong Liu, Atin Sakkeer Hussain, Chenshuo Sun, and Ying Shan · 2023
Later among the works it cites.
Music understanding llama: Advancing text-to-music generation with question answering and captioning
Shansong Liu, Atin Sakkeer Hussain, Chenshuo Sun, and Ying Shan · 2023
Later among the works it cites.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi · 2023
Later among the works it cites.
A System Out of Balance: A Critical Analysis of Philosophical Justififications for Copyright Law Through the Lenz of Users’ Rights
Mitchell Longen · 2023
Later among the works it cites.
Inference-time intervention: Eliciting truthful answers from a language model
Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg · 2023
Later among the works it cites.
Melodydiffusion: chord-conditioned melody generation using a transformer-based diffusion model
Shuyu Li and Yunsick Sung · 2023
Later among the works it cites.
MRBERT: Pre-Training of Melody and Rhythm for Automatic Music Generation
Shuyu Li and Yunsick Sung · 2023
Later among the works it cites.
Sparks of large audio models: A survey and outlook
Siddique Latif, Moazzam Shoukat, Fahad Shamshad, Muhammad Usama, Heriberto Cuayáhuitl, and Björn W Schuller · 2023
Later among the works it cites.
Efficient neural music generation
Max WY Lam, Qiao Tian, Tang Li, Zongyu Yin, Siyuan Feng, Ming Tu, Yuliang Ji, Rui Xia, Mingbo Ma, Xuchen Song, et al · 2023
Later among the works it cites.
Getmusic: Generating any music tracks with a unified representation and diffusion framework
Ang Lv, Xu Tan, Peiling Lu, Wei Ye, Shikun Zhang, Jiang Bian, and Rui Yan · 2023
Later among the works it cites.
AudioLDM 2: Learning holistic audio generation with self-supervised pretraining
Haohe Liu, Qiao Tian, Yi Yuan, Xubo Liu, Xinhao Mei, Qiuqiang Kong, Yuping Wang, Wenwu Wang, Yuxuan Wang, and Mark D. Plumbley · 2023
Later among the works it cites.
AudioLDM 2: Learning holistic audio generation with self-supervised pretraining
Haohe Liu, Qiao Tian, Yi Yuan, Xubo Liu, Xinhao Mei, Qiuqiang Kong, Yuping Wang, Wenwu Wang, Yuxuan Wang, and Mark D Plumbley · 2023
Later among the works it cites.
Content-based controls for music large language modeling
Liwei Lin, Gus Xia, Junyan Jiang, and Yixiao Zhang · 2023
Later among the works it cites.
Musecoco: Generating symbolic music from text
Peiling Lu, Xin Xu, Chenfei Kang, Botao Yu, Chengyi Xing, Xu Tan, and Jiang Bian · 2023
Later among the works it cites.
MuseCoco: Generating Symbolic Music from Text, May 2023
Peiling Lu, Xin Xu, Chenfei Kang, Botao Yu, Chengyi Xing, Xu Tan, and Jiang Bian · 2023
Later among the works it cites.
Musecoco: Generating symbolic music from text
Peiling Lu, Xin Xu, Chenfei Kang, Botao Yu, Chengyi Xing, Xu Tan, and Jiang Bian · 2023
Later among the works it cites.
Wavjourney: Compositional audio creation with large language models
Xubo Liu, Zhongkai Zhu, Haohe Liu, Yi Yuan, Meng Cui, Qiushi Huang, Jinhua Liang, Yin Cao, Qiuqiang Kong, Mark D Plumbley, et al · 2023
Later among the works it cites.
Serge modular archive instrument (smai): Bridging skeuomorphic & machine learning enabled interfaces
Ted Moore and Jean Brazeau · 2023
Later among the works it cites.
The unwitting labourer: extracting humanness in AI training
Fabio Morreale, Elham Bahmanteymouri, Brent Burmester, Andrew Chen, and Michelle Thorp · 2023
Later among the works it cites.
MusTango: Toward controllable text-to-music generation
Jan Melechovsky, Zixun Guo, Deepanway Ghosal, Navonil Majumder, Dorien Herremans, and Soujanya Poria · 2023
Later among the works it cites.
Polyffusion: A Diffusion Model for Polyphonic Score Generation with Internal and External Controls
Lejun Min, Junyan Jiang, Gus Xia, and Jingwei Zhao · 2023
Later among the works it cites.
A review of deep learning techniques for speech processing
Ambuj Mehrish, Navonil Majumder, Rishabh Bharadwaj, Rada Mihalcea, and Soujanya Poria · 2023
Later among the works it cites.
Language-guided music recommendation for video via prompt analogies
Daniel McKee, Justin Salamon, Josef Sivic, and Bryan C. Russell · 2023
Later among the works it cites.
Data collection in music generation training sets: A critical analysis
Fabio Morreale, Megha Sharma, and I-Chieh Wei · 2023
Later among the works it cites.
The song describer dataset: a corpus of audio captions for music-and-language evaluation, 2023
I Manco, B Weck, S Doh, M Won, Y Zhang, D Bodganov, Y Wu, K Chen, P Tovstogan, E Benetos, et al · 2023
Later among the works it cites.
The song describer dataset: A corpus of audio captions for music-and-language evaluation
Ilaria Manco, Benno Weck, Seungheon Doh, Minz Won, Yixiao Zhang, Dmitry Bogdanov, Yusong Wu, Ke Chen, Philip Tovstogan, Emmanouil Benetos, Elio Quinton, György Fazekas, and Juhan Nam · 2023
Later among the works it cites.
On the effectiveness of speech self-supervised learning for music
Yinghao Ma, Ruibin Yuan, Yizhi Li, Ge Zhang, Chenghua Lin, Xingran Chen, Anton Ragni, Hanzhi Yin, Emmanouil Benetos, Norbert Gyenge, Ruibo Liu, Gus Xia, Roger B. Dannenberg, Yike Guo, and Jie Fu · 2023
Later among the works it cites.
MT4SSL: Boosting self-supervised speech representation learning by integrating multiple targets
Ziyang Ma, Zhisheng Zheng, Changli Tang, Yujin Wang, and Xie Chen · 2023
Later among the works it cites.
Pushing the limits of unsupervised unit discovery for ssl speech representation
Ziyang Ma, Zhisheng Zheng, Guanrou Yang, Yu Wang, Chao Zhang, and Xie Chen · 2023
Later among the works it cites.
Diffusion models as masked audio-video learners
Elvis Nunez, Yanzi Jin, Mohammad Rastegari, Sachin Mehta, and Maxwell Horton · 2023
Later among the works it cites.
Vits-based singing voice conversion leveraging whisper and multi-scale F0 modeling
Ziqian Ning, Yuepeng Jiang, Zhichao Wang, Bin Zhang, and Lei Xie · 2023
Later among the works it cites.
Generative spoken dialogue language modeling
Tu Anh Nguyen, Eugene Kharitonov, Jade Copet, Yossi Adi, Wei-Ning Hsu, Ali Elkahky, Paden Tomasello, Robin Algayres, Benoit Sagot, Abdelrahman Mohamed, et al · 2023
Later among the works it cites.
Human-ai music creation: Understanding the perceptions and experiences of music creators for ethical and productive collaboration
Michele Newman, Lidia Morris, and Jin Ha Lee · 2023
Later among the works it cites.
A comparative analysis of latent regressor losses for singing voice conversion
Brendan O’Connor and Simon Dixon · 2023
Later among the works it cites.
Loaf-m2l: Joint learning of wording and formatting for singable melody-to-lyric generation
Longshen Ou, Xichu Ma, and Ye Wang · 2023
Later among the works it cites.
mir_ref: A representation evaluation framework for music information retrieval tasks
Christos Plachouras, Pablo Alonso-Jiménez, and Dmitry Bogdanov · 2023
Later among the works it cites.
Automatic note-level score-to-performance alignments in the asap dataset
Silvan David Peter, Carlos Eduardo Cancino-Chacón, Francesco Foscarin, Andrew Philip McLeod, Florian Henkel, Emmanouil Karystinaios, and Gerhard Widmer · 2023
Later among the works it cites.
Sounding out reconstruction error-based evaluation of generative models of expressive performance
Silvan David Peter, Carlos Eduardo Cancino-Chacón, Emmanouil Karystinaios, and Gerhard Widmer · 2023
Later among the works it cites.
Fairness and diversity in information access systems
Lorenzo Porcaro, Carlos Castillo, Emilia Gómez, and João Vinagre · 2023
Later among the works it cites.
Carnatic singing voice separation using cold diffusion on training data with bleeding
Genís Plaja-Roglans, Marius Miron, Adithi Shankar, and Xavier Serra · 2023
Later among the works it cites.
Scalable diffusion models with transformers
William Peebles and Saining Xie · 2023
Later among the works it cites.
Audiopalm: A large language model that can speak and listen
Paul K Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen, Ankur Bapna, Zalán Borsos, Félix de Chaumont Quitry, Peter Chen, Dalia El Badawy, Wei Han, Eugene Kharitonov, et al · 2023
Later among the works it cites.
Robust speech recognition via large-scale weak supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever · 2023
Later among the works it cites.
Robust speech recognition via large-scale weak supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever · 2023
Later among the works it cites.
Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation
Ludan Ruan, Yiyang Ma, Huan Yang, Huiguo He, Bei Liu, Jianlong Fu, Nicholas Jing Yuan, Qin Jin, and Baining Guo · 2023
Later among the works it cites.
“legal challenges to generative ai, part ii: Deliberating on inconclusive ai-generated policy questions
Pamela Samuelson · 2023
Later among the works it cites.
Yang Song, Prafulla Dhariwal, Mark Chen, and Ilya Sutskever · 2023
Later among the works it cites.
A length-extrapolatable transformer
Yutao Sun, Li Dong, Barun Patra, Shuming Ma, Shaohan Huang, Alon Benhaim, Vishrav Chaudhary, Xia Song, and Furu Wei · 2023
Later among the works it cites.
Toolformer: Language models can teach themselves to use tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom · 2023
Later among the works it cites.
Moûsai: Text-to-music generation with long-context latent diffusion
Flavio Schneider, Zhijing Jin, and Bernhard Schölkopf · 2023
Later among the works it cites.
Moûsai: Text-to-music generation with long-context latent diffusion
Flavio Schneider, Zhijing Jin, and Bernhard Schölkopf · 2023
Later among the works it cites.
V2meow: Meowing to the visual beat via music generation
Kun Su, Judith Yue Li, Qingqing Huang, Dima Kuzmin, Joonseok Lee, Chris Donahue, Fei Sha, Aren Jansen, Yu Wang, Mauro Verzetti, and Timo I. Denk · 2023
Later among the works it cites.
Unsupervised music source separation using differentiable parametric source models
Kilian Schulze-Forster, Gaël Richard, Liam Kelley, Clement S. J. Doire, and Roland Badeau · 2023
Later among the works it cites.
Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface
Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang · 2023
Later among the works it cites.
Taskbench: Benchmarking large language models for task automation
Yongliang Shen, Kaitao Song, Xu Tan, Wenqi Zhang, Kan Ren, Siyu Yuan, Weiming Lu, Dongsheng Li, and Yueting Zhuang · 2023
Later among the works it cites.
Introducing mpt-7b: A new standard for open-source, commercially usable llms, 2023
MosaicML NLP Team · 2023
Later among the works it cites.
Anticipatory music transformer
John Thickstun, David Hall, Chris Donahue, and Percy Liang · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
Singer identity representation learning using self-supervised techniques
Bernardo Torres, Stefan Lattner, and Gaël Richard · 2023
Later among the works it cites.
Emomv: Affective music-video correspondence learning datasets for classification and retrieval
Ha Thi Phuong Thao, Gemma Roig, and Dorien Herremans · 2023
Later among the works it cites.
Salmonn: Towards generic hearing abilities for large language models
Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun Ma, and Chao Zhang · 2023
Later among the works it cites.
FIGARO: Controllable music generation using learned and expert features
Dimitri von Rütte, Luca Biggio, Yannic Kilcher, and Thomas Hofmann · 2023
Later among the works it cites.
Audiobox: Unified audio generation with natural language prompts
Apoorv Vyas, Bowen Shi, Matthew Le, Andros Tjandra, Yi-Chiao Wu, Baishan Guo, Jiemin Zhang, Xinyue Zhang, Robert Adkins, William Ngan, et al · 2023
Later among the works it cites.
Neural codec language models are zero-shot text to speech synthesizers
Chengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, et al · 2023
Later among the works it cites.
Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Yusong Wu, Ke Chen, Tianyu Zhang, Yuchen Hui, Taylor Berg-Kirkpatrick, and Shlomo Dubnov · 2023
Later among the works it cites.
Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Yusong Wu, Ke Chen, Tianyu Zhang, Yuchen Hui, Taylor Berg-Kirkpatrick, and Shlomo Dubnov · 2023
Later among the works it cites.
Next-gpt: Any-to-any multimodal llm
Shengqiong Wu, Hao Fei, Leigang Qu, Wei Ji, and Tat-Seng Chua · 2023
Later among the works it cites.
Origin stories: Plantations, computers, and industrial control
Meredith Whittaker · 2023
Later among the works it cites.
Training a singing transcription model using connectionist temporal classification loss and cross-entropy loss
Jun-You Wang and Jyh-Shing Roger Jang · 2023
Later among the works it cites.
Audit: Audio editing by following instructions with latent diffusion models
Yuancheng Wang, Zeqian Ju, Xu Tan, Lei He, Zhizheng Wu, Jiang Bian, and Sheng Zhao · 2023
Later among the works it cites.
Adapting pretrained speech model for mandarin lyrics transcription and alignment
Jun-You Wang, Chon-In Leong, Yu-Chen Lin, Li Su, and Jyh-Shing Roger Jang · 2023
Later among the works it cites.
Chord-conditioned melody harmonization with controllable harmonicity
Shangda Wu, Xiaobing Li, and Maosong Sun · 2023
Later among the works it cites.
Tunesformer: Forming irish tunes with control codes by bar patching
Shangda Wu, Xiaobing Li, Feng Yu, and Maosong Sun · 2023
Later among the works it cites.
SongDriver2: Real-time Emotion-based Music Arrangement with Soft Transition, May 2023
Zihao Wang, Le Ma, Chen Zhang, Bo Han, Yikai Wang, Xinyi Chen, HaoRong Hong, Wenbo Liu, Xinda Wu, and Kejun Zhang · 2023
Later among the works it cites.
Musicyolo: A vision-based framework for automatic singing transcription
Xianke Wang, Bowen Tian, Weiming Yang, Wei Xu, and Wenqing Cheng · 2023
Later among the works it cites.
Compose & Embellish: Well-structured piano performance generation via a two-stage approach
Shih-Lun Wu and Yi-Hsuan Yang · 2023
Later among the works it cites.
MuseMorphose: Full-Song and Fine-Grained Piano Music Style Transfer With One Transformer VAE
Shih-Lun Wu and Yi-Hsuan Yang · 2023
Later among the works it cites.
MuseMorphose: Full-song and fine-grained piano music style transfer with one transformer vae
Shih-Lun Wu and Yi-Hsuan Yang · 2023
Later among the works it cites.
Visual chatgpt: Talking, drawing and editing with visual foundation models
Chenfei Wu, Shengming Yin, Weizhen Qi, Xiaodong Wang, Zecheng Tang, and Nan Duan · 2023
Later among the works it cites.
Clamp: Contrastive language-music pre-training for cross-modal symbolic music information retrieval
Shangda Wu, Dingyao Yu, Xu Tan, and Maosong Sun · 2023
Later among the works it cites.
Clamp: Contrastive language-music pre-training for cross-modal symbolic music information retrieval
Shangda Wu, Dingyao Yu, Xu Tan, and Maosong Sun · 2023
Later among the works it cites.
The rise and potential of large language model based agents: A survey
Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, Junzhe Wang, Senjie Jin, Enyu Zhou, et al · 2023
Later among the works it cites.
Language model beats diffusion–tokenizer is key to visual generation
Lijun Yu, José Lezama, Nitesh B Gundavarapu, Luca Versari, Kihyuk Sohn, David Minnen, Yong Cheng, Agrim Gupta, Xiuye Gu, Alexander G Hauptmann, et al · 2023
Later among the works it cites.
Hifi-codec: Group-residual vector quantization for high fidelity audio codec
Dongchao Yang, Songxiang Liu, Rongjie Huang, Jinchuan Tian, Chao Weng, and Yuexian Zou · 2023
Later among the works it cites.
MARBLE: music audio representation benchmark for universal evaluation
Ruibin Yuan, Yinghao Ma, Yizhi Li, Ge Zhang, Xingran Chen, Hanzhi Yin, Le Zhuo, Yiqi Liu, Jiawen Huang, Zeyue Tian, Binyue Deng, Ningzhi Wang, Chenghua Lin, Emmanouil Benetos, Anton Ragni, Norbert Gyenge, Roger B. Dannenberg, Wenhu Chen, Gus Xia, Wei Xue, Si Liu, Shi Wang, Ruibo Liu, Yike Guo, and Jie Fu · 2023
Later among the works it cites.
Fast-hubert: an efficient training framework for self-supervised speech representation learning
Guanrou Yang, Ziyang Ma, Zhisheng Zheng, Yakun Song, Zhikang Niu, and Xie Chen · 2023
Later among the works it cites.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, et al · 2023
Later among the works it cites.
Zero-shot duet singing voices separation with diffusion models
Chin-Yun Yu, Emilian Postolache, Emanuele Rodolà, and György Fazekas · 2023
Later among the works it cites.
Musicagent: An ai agent for music understanding and generation with large language models
Dingyao Yu, Kaitao Song, Peiling Lu, Tianyu He, Xu Tan, Wei Ye, Shikun Zhang, and Jiang Bian · 2023
Later among the works it cites.
A phoneme-informed neural network model for note-level singing transcription
Sangeon Yong, Li Su, and Juhan Nam · 2023
Later among the works it cites.
Uniaudio: An audio foundation model toward universal audio generation
Dongchao Yang, Jinchuan Tian, Xu Tan, Rongjie Huang, Songxiang Liu, Xuankai Chang, Jiatong Shi, Sheng Zhao, Jiang Bian, Xixin Wu, et al · 2023
Later among the works it cites.
Source-filter hifi-gan: Fast and pitch controllable high-fidelity neural vocoder
Reo Yoneyama, Yi-Chiao Wu, and Tomoki Toda · 2023
Later among the works it cites.
Unsupervised deep unfolded representation learning for singing voice separation
Weitao Yuan, Shengbei Wang, Jianming Wang, Masashi Unoki, and Wenwu Wang · 2023
Later among the works it cites.
Comospeech: One-step speech and singing voice synthesis via consistency model
Zhen Ye, Wei Xue, Xu Tan, Jie Chen, Qifeng Liu, and Yike Guo · 2023
Later among the works it cites.
Diffsound: Discrete diffusion model for text-to-sound generation
Dongchao Yang, Jianwei Yu, Helin Wang, Wen Wang, Chao Weng, Yuexian Zou, and Dong Yu · 2023
Later among the works it cites.
A survey of ai music generation tools and models
Yueyue Zhu, Jared Baca, Banafsheh Rekabdar, and Reza Rawassizadeh · 2023
Later among the works it cites.
Judging LLM-as-a-judge with MT-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica · 2023
Later among the works it cites.
Llama-adapter: Efficient fine-tuning of language models with zero-init attention
Renrui Zhang, Jiaming Han, Chris Liu, Peng Gao, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hongsheng Li, and Yu Qiao · 2023
Later among the works it cites.
Llama-adapter: Efficient fine-tuning of language models with zero-init attention
Renrui Zhang, Jiaming Han, Chris Liu, Peng Gao, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hongsheng Li, and Yu Qiao · 2023
Later among the works it cites.
Symbolic music representations for classification tasks: A systematic evaluation
Huan Zhang, Emmanouil Karystinaios, Simon Dixon, Gerhard Widmer, and Carlos Eduardo Cancino-Chacón · 2023
Later among the works it cites.
A comprehensive survey on pretrained foundation models: A history from bert to chatgpt
Ce Zhou, Qian Li, Chen Li, Jun Yu, Yixin Liu, Guangjing Wang, Kai Zhang, Cheng Ji, Qiben Yan, Lifang He, et al · 2023
Later among the works it cites.
Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities
Dong Zhang, Shimin Li, Xin Zhang, Jun Zhan, Pengyu Wang, Yaqian Zhou, and Xipeng Qiu · 2023
Later among the works it cites.
Loop copilot: Conducting ai ensembles for music generation and iterative editing
Yixiao Zhang, Akira Maezawa, Gus Xia, Kazuhiko Yamamoto, and Simon Dixon · 2023
Later among the works it cites.
Ernie-music: Text-to-waveform music generation with diffusion models
Pengfei Zhu, Chao Pang, Shuohuan Wang, Yekun Chai, Yu Sun, Hao Tian, and Hua Wu · 2023
Later among the works it cites.
Adding conditional control to text-to-image diffusion models
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala · 2023
Later among the works it cites.
SDMuse: Stochastic differential music editing and generation via hybrid representation
Chen Zhang, Yi Ren, Kejun Zhang, and Shuicheng Yan · 2023
Later among the works it cites.
Discrete contrastive diffusion for cross-modal music and image generation
Ye Zhu, Yu Wu, Kyle Olszewski, Jian Ren, Sergey Tulyakov, and Yan Yan · 2023
Later among the works it cites.
Video background music generation: Dataset, method and evaluation
Le Zhuo, Zhaokai Wang, Baisen Wang, Yue Liao, Chenxi Bao, Stanley Peng, Songhao Han, Aixi Zhang, Fei Fang, and Si Liu · 2023
Later among the works it cites.
Video background music generation: Dataset, method and evaluation
Le Zhuo, Zhaokai Wang, Baisen Wang, Yue Liao, Chenxi Bao, Stanley Peng, Songhao Han, Aixi Zhang, Fei Fang, and Si Liu · 2023
Later among the works it cites.
Jingwei Zhao, Gus Xia, and Ye Wang · 2023
Later among the works it cites.
Lyricwhiz: Robust multilingual zero-shot lyrics transcription by whispering to chatgpt
Le Zhuo, Ruibin Yuan, Jiahao Pan, Yinghao Ma, Yizhi Li, Ge Zhang, Si Liu, Roger B. Dannenberg, Jie Fu, Chenghua Lin, Emmanouil Benetos, Wenhu Chen, Wei Xue, and Yike Guo · 2023
Later among the works it cites.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al · 2023
Later among the works it cites.
Speak foreign languages with your own voice: Cross-lingual neural codec language modeling
Ziqiang Zhang, Long Zhou, Chengyi Wang, Sanyuan Chen, Yu Wu, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, et al · 2023
Later among the works it cites.
VatLM: Visual-Audio-Text Pre-Training with Unified Masked Prediction for Speech Representation Learning
Qiushi Zhu, Long Zhou, Ziqiang Zhang, Shujie Liu, Binxing Jiao, Jie Zhang, Lirong Dai, Daxin Jiang, Jinyu Li, and Furu Wei · 2023
Later among the works it cites.
ASiT: Local-Global Audio Spectrogram vIsion Transformer for Event Classification, March 2024
Sara Atito, Muhammad Awais, Wenwu Wang, Mark D. Plumbley, and Josef Kittler · 2024
Closest in time.
Automating Society – Taking Stock of Automated Decision-Making in the EU, 2019
AlgorithmWatch, Bertelsmann Stiftung, Open Society Foundations · 2024
Closest in time.
Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers
Anonymous · 2024
Closest in time.
Artificial Intelligence and Musicking
Adam Eric Berkowitz · 2024
Closest in time.
Scaling transformer to 1m tokens and beyond with rmt, 2024
Aydar Bulatov, Yuri Kuratov, Yermek Kapushev, and Mikhail S. Burtsev · 2024
Closest in time.
BERT-like Pre-training for Symbolic Piano Music Classification Tasks, April 2024
Yi-Hui Chou, I.-Chun Chen, Chin-Jui Chang, Joann Ching, and Yi-Hsuan Yang · 2024
Closest in time.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al · 2024
Closest in time.
Simple and controllable music generation
Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, and Alexandre Défossez · 2024
Closest in time.
Eat: Self-supervised pre-training with efficient audio transformer
Wenxi Chen, Yuzhe Liang, Ziyang Ma, Zhisheng Zheng, and Xie Chen · 2024
Closest in time.
Scaling properties of speech language models
Santiago Cuervo and Ricard Marxer · 2024
Closest in time.
Melfusion: Synthesizing music from image and language cues using diffusion models
Sanjoy Chowdhury, Sayan Nag, KJ Joseph, Balaji Vasan Srinivasan, and Dinesh Manocha · 2024
Closest in time.
Cocola: Coherence-oriented contrastive learning of musical audio representations
Ruben Ciranni, Emilian Postolache, Giorgio Mariani, Michele Mancusi, Luca Cosmo, and Emanuele Rodolà · 2024
Closest in time.
Generative AI’s environmental costs are soaring — and mostly secret
Kate Crawford · 2024
Closest in time.
Mmtrail: A multimodal trailer video dataset with language and music descriptions
Xiaowei Chi, Yatian Wang, Aosong Cheng, Pengjun Fang, Zeyue Tian, Yingqing He, Zhaoyang Liu, Xingqun Qi, Jiahao Pan, Rongyu Zhang, et al · 2024
Closest in time.
Yunfei Chu, Jin Xu, Qian Yang, Haojie Wei, Xipin Wei, Zhifang Guo, Yichong Leng, Yuanjun Lv, Jinzheng He, Junyang Lin, et al · 2024
Closest in time.
Songcomposer: A large language model for lyric and melody composition in song generation
Shuangrui Ding, Zihan Liu, Xiaoyi Dong, Pan Zhang, Rui Qian, Conghui He, Dahua Lin, and Jiaqi Wang · 2024
Closest in time.
Songcomposer: A large language model for lyric and melody composition in song generation
Shuangrui Ding, Zihan Liu, Xiaoyi Dong, Pan Zhang, Rui Qian, Conghui He, Dahua Lin, and Jiaqi Wang · 2024
Closest in time.
Pablo González de la Torre, Marta Pérez-Verdugo, and Xabier E Barandiaran · 2024
Closest in time.
From real to cloned singer identification, 2024
Dorian Desblancs, Gabriel Meseguer-Brocal, Romain Hennequin, and Manuel Moussallam · 2024
Closest in time.
Musilingo: Bridging music and text with pre-trained language models for music captioning and query response
Zihao Deng, Yinghao Ma, Yudong Liu, Rongchen Guo, Ge Zhang, Wenhu Chen, Wenhao Huang, and Emmanouil Benetos · 2024
Closest in time.
Joint music and language attention models for zero-shot music tagging
Xingjian Du, Zhesong Yu, Jiaju Lin, Bilei Zhu, and Qiuqiang Kong · 2024
Closest in time.
Scaling up masked audio encoder learning for general audio classification, June 2024
Heinrich Dinkel, Zhiyong Yan, Yongqing Wang, Junbo Zhang, Yujun Wang, and Bin Wang · 2024
Closest in time.
Composerx: Multi-agent symbolic music composition with llms
Qixin Deng, Qikai Yang, Ruibin Yuan, Yipeng Huang, Yi Wang, Xubo Liu, Zeyue Tian, Jiahao Pan, Ge Zhang, Hanfeng Lin, et al · 2024
Closest in time.
Fast timing-conditioned latent audio diffusion
Zach Evans, CJ Carr, Josiah Taylor, Scott H Hawley, and Jordi Pons · 2024
Closest in time.
Long-form music generation with latent diffusion
Zach Evans, Julian D Parker, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons · 2024
Closest in time.
Audio mamba: Bidirectional state space model for audio representation learning
Mehmet Hamza Erol, Arda Senocak, Jiu Feng, and Joon Son Chung · 2024
Closest in time.
A-JEPA: Joint-Embedding Predictive Architecture Can Listen, January 2024
Zhengcong Fei, Mingyuan Fan, and Junshi Huang · 2024
Closest in time.
Automatic lyric transcription and automatic music transcription from multimodal singing
Xiangming Gu, Longshen Ou, Wei Zeng, Jianan Zhang, Nicholas Wong, and Ye Wang · 2024
Closest in time.
EMO-Music: Emotion Recognition Based Music Therapy with Deep Learning on Physiological Signals
Hanzhe Guo, Jiawen Zhang, Yueyao Jiang, Yifei Qi, Simeng Chen, Zhen Chen, Weiran Lin, Junwei Cao, and Shuangs hou Li · 2024
Closest in time.
Towards building an end-to-end multilingual automatic lyrics transcription model
Jiawen Huang and Emmanouil Benetos · 2024
Closest in time.
Llms meet multimodal generation and editing: A survey
Yingqing He, Zhaoyang Liu, Jingye Chen, Zeyue Tian, Hongyu Liu, Xiaowei Chi, Runtao Liu, Ruibin Yuan, Yazhou Xing, Wenhai Wang, et al · 2024
Closest in time.
A review of differentiable digital signal processing for music and speech synthesis
B Hayes, J Shier, G Fazekas, A McPherson, and C Saitis · 2024
Closest in time.
Mavil: Masked audio-video learners
Po-Yao Huang, Vasu Sharma, Hu Xu, Chaitanya Ryali, Yanghao Li, Shang-Wen Li, Gargi Ghosh, Jitendra Malik, Christoph Feichtenhofer, et al · 2024
Closest in time.
Knowledge diffusion for distillation
Tao Huang, Yuan Zhang, Mingkai Zheng, Shan You, Fei Wang, Chen Qian, and Chang Xu · 2024
Closest in time.
Glossary for discussion of ethics of autonomous and intelligent systems, version 1, 2017
IEEE Global Initiative on Ethics of Autonomous and Intelligent Systems · 2024
Closest in time.
The Modernization of Oriental Music Therapy: Five-Element Music Therapy Combined with Artificial Intelligence
Chan-Young Kwon, Hyunsu Kim, and Sung-Hee Kim · 2024
Closest in time.
High-fidelity audio compression with improved rvqgan
Rithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar, and Kundan Kumar · 2024
Closest in time.
Disco-10m: A large-scale music dataset
Luca Lanzendörfer, Florian Grötschla, Emil Funke, and Roger Wattenhofer · 2024
Closest in time.
Music understanding llama: Advancing text-to-music generation with question answering and captioning
Shansong Liu, Atin Sakkeer Hussain, Chenshuo Sun, and Ying Shan · 2024
Closest in time.
Bytecomposer: a human-like melody composition method based on language model agent
Xia Liang, Jiaju Lin, and Xinjian Du · 2024
Closest in time.
Mertech: Instrument playing technique detection using self-supervised pretrained model with multi-task finetuning
Dichucheng Li, Yinghao Ma, Weixing Wei, Qiuqiang Kong, Yulun Wu, Mingjin Che, Fan Xia, Emmanouil Benetos, and Wei Li · 2024
Closest in time.
Diff-BGM: A Diffusion Model for Video Background Music Generation
Sizhe Li, Yiming Qin, Minghang Zheng, Xin Jin, and Yang Liu · 2024
Closest in time.
Diff-bgm: A diffusion model for video background music generation
Sizhe Li, Yiming Qin, Minghang Zheng, Xin Jin, and Yang Liu · 2024
Closest in time.
Music source separation with band-split rope transformer
Wei-Tsung Lu, Ju-Chiang Wang, Qiuqiang Kong, and Yun-Ning Hung · 2024
Closest in time.
Semanticodec: An ultra low bitrate semantic audio codec for general sound
Haohe Liu, Xuenan Xu, Yi Yuan, Mengyue Wu, Wenwu Wang, and Mark D Plumbley · 2024
Closest in time.
Liwei Lin, Gus Xia, Yixiao Zhang, and Junyan Jiang · 2024
Closest in time.
Jiajia Li, Lu Yang, Mingni Tang, Cong Chen, Zuchao Li, Ping Wang, and Hai Zhao · 2024
Closest in time.
Comosvc: Consistency model-based singing voice conversion
Yiwen Lu, Zhen Ye, Wei Xue, Xu Tan, Qifeng Liu, and Yike Guo · 2024
Closest in time.
Mert: Acoustic music understanding model with large-scale self-supervised training
Yizhi Li, Ruibin Yuan, Ge Zhang, Yinghao Ma, Xingran Chen, Hanzhi Yin, Chenghua Lin, Anton Ragni, Emmanouil Benetos, Norbert Gyenge, et al · 2024
Closest in time.
Wavcraft: Audio editing and generation with large language models
Jinhua Liang, Huan Zhang, Haohe Liu, Yin Cao, Qiuqiang Kong, Xubo Liu, Wenwu Wang, Mark D. Plumbley, Huy Phan, and Emmanouil Benetos · 2024
Closest in time.
Models All The Way Down, 2024
Knowing Machines · 2024
Closest in time.
“Music therapy is the very definition of white privilege”: Music therapists’ perspectives on race and class in UK music therapy
Tamsin Mains, Victoria Clarke, and Luke Annesley · 2024
Closest in time.
Matthew C McCallum, Matthew EP Davies, Florian Henkel, Jaehun Kim, and Samuel E Sandberg · 2024
Closest in time.
Scaling data-constrained language models
Niklas Muennighoff, Alexander Rush, Boaz Barak, Teven Le Scao, Nouamane Tazi, Aleksandra Piktus, Sampo Pyysalo, Thomas Wolf, and Colin A Raffel · 2024
Closest in time.
Automatic identification of preferred music genres: an exploratory machine learning approach to support personalized music therapy
Ingrid Bruno Nunes, Maíra Araújo de Santana, Nicole Charron, Hyngrid Souza e Silva, Caylane Mayssa de Lima Simões, Camila Lins, Ana Beatriz de Souza Sampaio, Arthur Moreira Nogueira de Melo, Thailson Caetano Valdeci da Silva, Camila Tiodista, et al · 2024
Closest in time.
Diff-a-riff: Musical accompaniment co-creation via latent diffusion models
Javier Nistal, Marco Pasini, Cyran Aouameur, Maarten Grachten, and Stefan Lattner · 2024
Closest in time.
Masked Modeling Duo: Towards a Universal Audio Pre-Training Framework
Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi, Noboru Harada, and Kunio Kashino · 2024
Closest in time.
Harmonic Healing and Neural Networks: Enhancing Music Therapy Through AI Integration
Yogesh Prabhakar Pingle and Lakshmappa K. Ragha · 2024
Closest in time.
EnCodecMAE: Leveraging neural codecs for universal audio representation learning, May 2024
Leonardo Pepino, Pablo Riera, and Luciana Ferrer · 2024
Closest in time.
Mupt: A generative symbolic music pretrained transformer
Xingwei Qu, Yuelin Bai, Yinghao Ma, Ziya Zhou, Ka Man Lo, Jiaheng Liu, Ruibin Yuan, Lejun Min, Xueling Liu, Tianyu Zhang, et al · 2024
Closest in time.
Exploring the relationship between music, medicine and physics. Why pluralism is necessary in music therapy?
Madalina Dana Rucsanda, Alexandra Belibou, and Alexandra Ioana Rucsanda · 2024
Closest in time.
A fully differentiable model for unsupervised singing voice separation
Gael Richard, Pierre Chouteau, and Bernardo Torres · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean-baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, et al · 2024
Closest in time.
Emotion-aligned contrastive learning between images and music
Shanti Stewart, Kleanthis Avramidis, Tiantian Feng, and Shrikanth Narayanan · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu · 2024
Closest in time.
Ssamba: Self-supervised audio representation learning with mamba state space model
Siavash Shams, Sukru Samet Dindar, Xilin Jiang, and Nima Mesgarani · 2024
Closest in time.
Dolma: An open corpus of three trillion tokens for language model pretraining research
Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, et al · 2024
Closest in time.
Gendered innovations in science, health & medicine, engineering, and environment, 2011
L. Schiebinger, I. Klinge, H. Paik, I. Sanchez, M. Schraudner, and M. Stefanick · 2024
Closest in time.
Fish speech v1
Tianyu Li Shijia Liao · 2024
Closest in time.
A feeling for the algorithm: Diversity, expertise, and artificial intelligence
Catherine Stinson and Sofie Vlaad · 2024
Closest in time.
Understanding Human-AI Collaboration in Music Therapy Through Co-Design with Therapists, February 2024
Jingjing Sun, Jingyi Yang, Guyue Zhou, Yucheng Jin, and Jiangtao Gong · 2024
Closest in time.
Agentgpt: Assemble, configure, and deploy autonomous ai agents in your browser
AgentGPT Team · 2024
Closest in time.
Autogpt: build & use ai agents
AutoGPT Team · 2024
Closest in time.
so-vits-svc
SVC Develop Team · 2024
Closest in time.
Vidmuse: A simple video-to-music generation framework with long-short-term modeling
Zeyue Tian, Zhaoyang Liu, Ruibin Yuan, Jiahao Pan, Xiaoqiang Huang, Qifeng Liu, Xu Tan, Qifeng Chen, Wei Xue, and Yike Guo · 2024
Closest in time.
Computer audition: From task-specific machine learning to foundation models
Andreas Triantafyllopoulos, Iosif Tsangko, Alexander Gebhard, Annamaria Mesaros, Tuomas Virtanen, and Björn Schuller · 2024
Closest in time.
SALMONN: Towards generic hearing abilities for large language models
Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun MA, and Chao Zhang · 2024
Closest in time.
Salmonn: Towards generic hearing abilities for large language models
Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, MA Zejun, and Chao Zhang · 2024
Closest in time.
Towards audio language modeling-an overview
Haibin Wu, Xuanjun Chen, Yi-Cheng Lin, Kai-wei Chang, Ho-Lam Chung, Alexander H Liu, and Hung-yi Lee · 2024
Closest in time.
Music controlnet: Multiple time-varying controls for music generation
Shih-Lun Wu, Chris Donahue, Shinji Watanabe, and Nicholas J Bryan · 2024
Closest in time.
A foundation model for music informatics
Minz Won, Yun-Ning Hung, and Duc Le · 2024
Closest in time.
A Foundation Model for Music Informatics
Minz Won, Yun-Ning Hung, and Duc Le · 2024
Closest in time.
Muchomusic: Evaluating music understanding in multimodal audio-language models
Benno Weck, Ilaria Manco, Emmanouil Benetos, Elio Quinton, György Fazekas, and Dmitry Bogdanov · 2024
Closest in time.
A survey on large language model based autonomous agents
Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, et al · 2024
Closest in time.
Whole-song hierarchical generation of symbolic music using cascaded diffusion models
Ziyu Wang, Lejun Min, and Gus Xia · 2024
Closest in time.
Toksing: Singing voice synthesis based on discrete tokens
Yuning Wu, Jiatong Shi, Yuxun Tang, Shan Yang, Qin Jin, et al · 2024
Closest in time.
Generating chord progression from melody with flexible harmonic rhythm and controllable harmonic density
Shangda Wu, Yue Yang, Zhaowen Wang, Xiaobing Li, and Maosong Sun · 2024
Closest in time.
Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners
Yazhou Xing, Yingqing He, Zeyue Tian, Xintao Wang, and Qifeng Chen · 2024
Closest in time.
Doremi: Optimizing data mixtures speeds up language model pretraining
Sang Michael Xie, Hieu Pham, Xuanyi Dong, Nan Du, Hanxiao Liu, Yifeng Lu, Percy S Liang, Quoc V Le, Tengyu Ma, and Adams Wei Yu · 2024
Closest in time.
Research on the Improvement of Children’s Attention Through Binaural Beats Music Therapy in the Context of AI Music Generation
Weijia Yang, Chih-Fang Huang, Hsun-Yi Huang, Zixue Zhang, Wenjun Li, and Chunmei Wang · 2024
Closest in time.
Towards unified alignment between agents, humans, and environment
Zonghan Yang, An Liu, Zijun Liu, Kaiming Liu, Fangzhou Xiong, Yile Wang, Zeyuan Yang, Qingyuan Hu, Xinrui Chen, Zhenhe Zhang, Fuwen Luo, Zhicheng Guo, Peng Li, and Yang Liu · 2024
Closest in time.
Chatmusician: Understanding and generating music intrinsically with llm
Ruibin Yuan, Hanfeng Lin, Yi Wang, Zeyue Tian, Shangda Wu, Tianhao Shen, Ge Zhang, Yuhang Wu, Cong Liu, Ziya Zhou, et al · 2024
Closest in time.
Chatmusician: Understanding and generating music intrinsically with LLM
Ruibin Yuan, Hanfeng Lin, Yi Wang, Zeyue Tian, Shangda Wu, Tianhao Shen, Ge Zhang, Yuhang Wu, Cong Liu, Ziya Zhou, Ziyang Ma, Liumeng Xue, Ziyu Wang, Qin Liu, Tianyu Zheng, Yizhi Li, Yinghao Ma, Yiming Liang, Xiaowei Chi, Ruibo Liu, Zili Wang, Pengfei Li, Jingcheng Wu, Chenghua Lin, Qifeng Liu, Tao Jiang, Wenhao Huang, Wenhu Chen, Emmanouil Benetos, Jie Fu, Gus Xia, Roger B. Dannenberg, Wei Xue, Shiyin Kang, and Yike Guo · 2024
Closest in time.
React meets actre: Autonomous annotations of agent trajectories for contrastive self-training
Zonghan Yang, Peng Li, Ming Yan, Ji Zhang, Fei Huang, and Yang Liu · 2024
Closest in time.
Marble: Music audio representation benchmark for universal evaluation
Ruibin Yuan, Yinghao Ma, Yizhi Li, Ge Zhang, Xingran Chen, Hanzhi Yin, Yiqi Liu, Jiawen Huang, Zeyue Tian, Binyue Deng, et al · 2024
Closest in time.
Easytool: Enhancing llm-based agents with concise tool instruction
Siyu Yuan, Kaitao Song, Jiangjie Chen, Xu Tan, Yongliang Shen, Ren Kan, Dongsheng Li, and Deqing Yang · 2024
Closest in time.
Masked Autoencoders with Multi-Window Local-Global Attention Are Better Audio Learners
Sarthak Yadav, Sergios Theodoridis, Lars Kai Hansen, and Zheng-Hua Tan · 2024
Closest in time.
Beatdance: A beat-based model-agnostic contrastive learning framework for music-dance retrieval
Kaixing Yang, Xukun Zhou, Xulong Tang, Ran Diao, Hongyan Liu, Jun He, and Zhaoxin Fan · 2024
Closest in time.
Measuring diversity in datasets
Dorothy Zhao, Jerone TA Andrews, Orestis Papakyriakopoulos, and Alice Xiang · 2024
Closest in time.
Dexter: Learning and controlling performance expression with diffusion models
Huan Zhang, Shreyan Chowdhury, Carlos Eduardo Cancino-Chacón, Jinhua Liang, Simon Dixon, and Gerhard Widmer · 2024
Closest in time.
Anygpt: Unified multimodal llm with discrete sequence modeling
Jun Zhan, Junqi Dai, Jiasheng Ye, Yunhua Zhou, Dong Zhang, Zhigeng Liu, Xin Zhang, Ruibin Yuan, Ge Zhang, Linyang Li, et al · 2024
Closest in time.
Vision-language models for vision tasks: A survey
Jingyi Zhang, Jiaxing Huang, Sheng Jin, and Shijian Lu · 2024
Closest in time.
Stylesinger: Style transfer for out-of-domain singing voice synthesis
Yu Zhang, Rongjie Huang, Ruiqi Li, Jinzheng He, Yan Xia, Feiyang Chen, Xinyu Duan, Baoxing Huai, and Zhou Zhao · 2024
Closest in time.
From audio encoders to piano judges: Benchmarking performance understanding for solo piano
Huan Zhang, Jinhua Liang, and Simon Dixon · 2024
Closest in time.
Easygen: Easing multimodal generation with a bidirectional conditional diffusion model and llms, 2024
Xiangyu Zhao, Bo Liu, Qijiong Liu, Guangyuan Shi, and Xiao-Ming Wu · 2024
Closest in time.
Speechagents: Human-communication simulation with multi-modal multi-agent systems
Dong Zhang, Zhaowei Li, Pengyu Wang, Xin Zhang, Yaqian Zhou, and Xipeng Qiu · 2024
Closest in time.
Map-neo: Highly capable and transparent bilingual large language model series
Ge Zhang, Scott Qu, Jiaheng Liu, Chenchen Zhang, Chenghua Lin, Chou Leuang Yu, Danny Pan, Esther Cheng, Jie Liu, Qunshu Lin, et al · 2024
Closest in time.