Fetching the paper…
Reading the bibliography…
Audio inpainting aims to reconstruct missing segments in corrupted recordings.
A. Gray and J. Markel, “Distance measures for speech processing,” IEEE Trans. Acoustics, Speech, Signal Process. , vol. 24, no. 5, pp. 380–391 (1976 Oct.), 10.1109/TASSP.1976.1162849
1976
Earlier work this paper cites.
A. J. E. M. Janssen, R. N. J. Veldhuis, and L. B. Vries, “Adaptive Interpolation of Discrete-Time Signals that Can Be Modeled as Autoregressive Processes,” IEEE Trans. Acoust. Speech Signal Process. , vol. 34, no. 2, pp. 317–330 (1986 Apr.), 10.1109/TASSP.1986.1164824
1986
Earlier work this paper cites.
D. Goodman, G. Lockhart, O. Wasem, and W.-C. Wong, “Waveform Substitution Techniques for Recovering Missing Speech Segments in Packet Voice Communications,” IEEE Trans. Acoust. Speech Signal Process. , vol. 34, no. 6, pp. 1440–1448 (1986 Dec.), 10.1109/TASSP.1986.1164984
1986
Earlier work this paper cites.
W. Etter, “Restoration of a Discrete-Time Signal Segment by Interpolation based on the Left-Sided and Right-Sided Autoregressive Parameters,” IEEE Trans. Signal Processing , vol. 44, no. 5, pp. 1124–1135 (1996), 10.1109/78.502326
1996
Earlier work this paper cites.
S. J. Godsill and P. J. Rayner, Digital Audio Restoration (Springer, 1998)
1998
Earlier work this paper cites.
I. Kauppinen, J. Kauppinen, and P. Saarinen, “A Method for Long Extrapolation of Audio Signals,” J. Audio Eng. Soc. , vol. 49, no. 12, pp. 1167–1180 (2001 Dec.)
2001
Earlier work this paper cites.
I. Kauppinen and K. Roth, “Audio Signal Extrapolation—Theory and Applications,” in Proceedings of the International Conference on Digital Audio Effects (DAFx) , pp. 105–110 (Hamburg, Germany) (2002 Sep.)
2002
Earlier work this paper cites.
I. Kauppinen and J. Kauppinen, “Reconstruction Method for Missing or Damaged Long Portions in Audio Signal,” J. Audio Eng. Soc. , vol. 50, no. 7/8, pp. 594–602 (2002 Jul.)
2002
Earlier work this paper cites.
P. P. Ebner and A. Eltelt, “Audio Inpainting with Generative Adversarial Network,” arXiv preprint (2020 Mar.), 10.48550/arXiv.2003.07704
2003
Earlier work this paper cites.
P. A. A. Esquef, V. Välimäki, K. Roth, and I. Kauppinen, “Interpolation of Long Gaps in Audio Signals using the Warped Burg’s Method,” in Proceedings of the 6th International Conference on Digital Audio Effects (DAFx) , pp. 08–11 (London, UK) (2003 Sep.)
2003
Earlier work this paper cites.
P. A. A. Esquef and L. W. P. Biscainho, “An Efficient Model-Based Multirate Method for Reconstruction of Audio Signals Across Long Gaps,” IEEE Trans. Audio Speech Lang. Process. , vol. 14, no. 4, pp. 1391–1400 (2006 Jul.), 10.1109/TSA.2005.858018
2005
Earlier work this paper cites.
M. Lagrange, S. Marchand, and J.-B. Rault, “Long Interpolation of Audio Signals using Linear Prediction in Sinusoidal Modeling,” J. Audio Eng. Soc. , vol. 53, no. 10, pp. 891–905 (2005 Oct.)
2005
Earlier work this paper cites.
A. Hyvärinen and P. Dayan, “Estimation of Non-Normalized Statistical Models by Score Matching.” J. Mach. Learn. Res. , vol. 6, no. 4, p. 695–709 (2005 Dec.)
2005
Earlier work this paper cites.
R. Huber and B. Kollmeier, “PEMO-Q—A new Method for Objective Audio Quality Assessment using a Model of Auditory Perception,” IEEE Trans. Audio Speech Lang. Process. , vol. 14, no. 6, pp. 1902–1911 (2006 Nov.), 10.1109/TASL.2006.883259
2006
Earlier work this paper cites.
P. Smaragdis, B. Raj, and M. Shashanka, “Missing Data Imputation for Spectral Audio Signals,” in Proceedings of the IEEE International Workshop on Machine Learning for Signal Processing , pp. 1–6 (2009 Sep.), 10.1109/MLSP.2009.5306194
2009
Earlier work this paper cites.
C. Schörkhuber and A. Klapuri, “Constant-Q Transform Toolbox for Music Processing,” in Proceedings of the 7th Sound and Music Computing Conference , pp. 3–64 (Barcelona, Spain) (2010 Jul.)
2010
Earlier work this paper cites.
A. Adler, V. Emiya, M. G. Jafari, M. Elad, R. Gribonval, and M. D. Plumbley, “Audio Inpainting,” IEEE Trans. Audio Speech Lang. Process. , vol. 20, no. 3, pp. 922–932 (2012 Mar.), 10.1109/TASL.2011.2168211
2011
Earlier work this paper cites.
G. A. Velasco, N. Holighaus, M. Dörfler, and T. Grill, “Constructing an Invertible Constant-Q Transform with Non-Stationary Gabor Frames,” in Proceedings of the Interioantl Conference on Digital Audio Effects (DAFX) (Paris, France) (2011 Sep.)
2011
Earlier work this paper cites.
N. Holighaus, M. Dörfler, G. A. Velasco, and T. Grill, “A Framework for Invertible, Real-Time Constant-Q Transforms,” IEEE Trans. Audio Speech Lang. Process. , vol. 21, no. 4, pp. 775–785 (2012 Apr.), 10.1109/TASL.2012.2234114
2012
Earlier work this paper cites.
C. Schörkhuber, A. Klapuri, and A. Sontacchi, “Pitch Shifting of Audio Signals using the Constant-Q Transform,” in Proceedings of the International Conference on Digital Audio Effects (DAFx) (2012 Jul.)
2012
Earlier work this paper cites.
B.-K. Lee and J.-H. Chang, “Packet Loss Concealment based on Deep Neural Networks for Digital Speech Transmission,” IEEE/ACM Trans. Audio Speech Lang. Process. , vol. 24, no. 2, pp. 378–387 (2015 Dec.), 10.1109/TASLP.2015.2509780
2015
Earlier work this paper cites.
O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional Networks for Biomedical Image Segmentation,” in Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention , pp. 234–241 (2015 Nov.), 10.1007/978-3-319-24574-4_28
2015
Earlier work this paper cites.
ITU, “Method for the Subjective Assessment of Intermediate Quality Level of Audio Systems,” Rec. BS.1534-3, International Telecommunication Union, Geneva, Switzerland (2015 Oct.)
2015
Earlier work this paper cites.
J. Thickstun, Z. Harchaoui, and S. Kakade, “Learning Features of Music from Scratch,” in Proceedings International Conference on Learning Representations (ICLR) (2016 May)
2016
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, et al. , “Attention is All You Need,” in Proceedings of NeurIPS , vol. 30 (2017 Dec.)
2017
Cited alongside, same era.
F. Lieb and H.-G. Stark, “Audio Inpainting: Evaluation of Time-Frequency Representations and Structured Sparsity Approaches,” Signal Process. , vol. 153, pp. 291–299 (2018 Dec.), 10.1016/j.sigpro.2018.07.012
2018
Cited alongside, same era.
N. Perraudin, N. Holighaus, P. Majdak, and P. Balazs, “Inpainting of Long Audio Segments with Similarity Graphs,” IEEE/ACM Trans. Audio Speech Lang. Process. , vol. 26, no. 6, pp. 1083–1094 (2018), 10.1109/TASLP.2018.2809864
2018
Cited alongside, same era.
Z. Kong, W. Ping, J. Huang, K. Zhao, and B. Catanzaro, “DiffWave: A Versatile Diffusion Model for Audio Synthesis,” in Proceedings of the International Conference on Learning Representations (ICLR) (2021 May)
2021
Later among the works it cites.
A. Ragano, E. Benetos, and A. Hines, “Automatic Quality Assessment of Digitized and Restored Sound Archives,” J. Audio Eng. Soc. , vol. 70, no. 4, pp. 252–270 (2022 Apr.), 10.17743/jaes.2022.0002
2022
Later among the works it cites.
B. Kawar, M. Elad, S. Ermon, and J. Song, “Denoising Diffusion Restoration Models,” in Proceedings of NeurIPS , pp. 23593–23606 (2022 Dec.)
2022
Later among the works it cites.
O. Mokrỳ, P. Magron, T. Oberlin, and C. Févotte, “Algorithms for Audio Inpainting based on Probabilistic Nonnegative Matrix Factorization,” Signal Process. (2023 May), 10.1016/j.sigpro.2022.108905
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville, “Film: Visual Reasoning with a General Conditioning Layer,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 32 (2018 Mar.)
2018
Cited alongside, same era.
A. Marafioti, N. Perraudin, N. Holighaus, and P. Majdak, “A Context Encoder for Audio Inpainting,” IEEE/ACM Trans. Audio Speech Lang. Process. , vol. 27, no. 12, pp. 2362–2372 (2019 Dec.), https://doi.org/10.1109/TASLP.2019.2947232
2019
Cited alongside, same era.
O. Mokrỳ, P. Záviška, P. Rajmic, and V. Veselỳ, “Introducing SPAIN (Sparse Audio Inpainter),” in Proceedings of the 27th European Signal Processing Conference (EUSIPCO) , pp. 1–5 (2019 Sep.), 10.23919/EUSIPCO.2019.8902560
2019
Cited alongside, same era.
K. Kilgour, M. Zuluaga, D. Roblek, and M. Sharifi, “Fréchet Audio Distance: A Reference-Free Metric for Evaluating Music Enhancement Algorithms,” in Proceedings of INTERSPEECH , pp. 2350–2354 (2019 Sep.), http://dx.doi.org/10.21437/Interspeech.2019-2219
2019
Cited alongside, same era.
T. Bazin, G. Hadjeres, P. Esling, and M. Malt, “Spectrogram Inpainting for Interactive Generation of Instrument Sounds,” in Proceedings of the Joint Conference on AI Music Creativity (Stockholm, Sweden) (2020 Oct.)
2020
Cited alongside, same era.
J. Ho, A. Jain, and P. Abbeel, “Denoising Diffusion Probabilistic Models,” in Proceedings of NeurIPS , vol. 33, pp. 6840–6851 (2020 Dec.)
2020
Cited alongside, same era.
O. Mokrỳ and P. Rajmic, “Audio Inpainting: Revisited and Reweighted,” IEEE/ACM Trans. Audio Speech Lang. Process. , vol. 28, pp. 2906–2918 (2020 Oct.), 10.1109/TASLP.2020.3030486
2020
Cited alongside, same era.
G. Tauböck, S. Rajbamshi, and P. Balazs, “Dictionary Learning for Sparse Audio Inpainting,” IEEE J. Selected Topics Signal Process. , vol. 15, no. 1, pp. 104–119 (2021 Jan.), 10.1109/JSTSP.2020.3046422
2020
Cited alongside, same era.
L. Ou and Y. Chen, “Concealing Audio Packet Loss Using Frequency-Consistent Generative Adversarial Networks,” in Proceedings of the 5th International Conference on Pattern Recognition and Artificial Intelligence (PRAI) , pp. 826–831 (Paris, France) (2022 May)
2022
Later among the works it cites.
Z. Borsos, M. Sharifi, and M. Tagliasacchi, “SpeechPainter: Text-conditioned Speech Inpainting,” in Proceedings of INTERSPEECH (2022 Sep.)
2022
Later among the works it cites.
Y. Wang, J. Yu, and J. Zhang, “Zero-Shot Image Restoration Using Denoising Diffusion Null-Space Model,” in Proceedings of the International Conference on Learning Representations (ICLR) (2022 May)
2022
Later among the works it cites.
A. Lugmayr, M. Danelljan, A. Romero, et al. , “Repaint: Inpainting using Denoising Diffusion Probabilistic Models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pp. 11461–11471 (2022 Jun.)
2022
Later among the works it cites.
J. Song, A. Vahdat, M. Mardani, and J. Kautz, “Pseudoinverse-Guided Diffusion Models for Inverse Problems,” in Proceedings of the International Conference on Learning Representations (ICLR) (2022 May)
2022
Later among the works it cites.
J. Ho, T. Salimans, A. Gritsenko, W. Chan, M. Norouzi, and D. J. Fleet, “Video Diffusion Models,” in Proceedings of NeurIPS , pp. 8633–8646 (2022 Dec.)
2022
Later among the works it cites.
T. Karras, M. Aittala, T. Aila, and S. Laine, “Elucidating the Design Space of Diffusion-Based Generative Models,” in Proceedings of NeurIPS , pp. 26565–26577 (2022 Dec.)
2022
Later among the works it cites.
H. Kim, S. Kim, and S. Yoon, “Guided-TTS: A Diffusion Model for Text-to-Speech via Classifier Guidance,” in Proceedings of the 39th International Conference on Machine Learning (2022 Jul.)
2022
Later among the works it cites.
H. Chung, B. Sim, D. Ryu, and J. C. Ye, “Improving Diffusion Models for Inverse Problems using Manifold Constraints,” in Proceedings of NeurIPS , pp. 25683–25696 (2022 Dec.)
2022
Later among the works it cites.
E. Moliner, J. Lehtinen, and V. Välimäki, “Solving Audio Inverse Problems with a Diffusion Models,” in Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 1–5 (Rhodes Island, Greece) (2023 Jun.), https://doi.org/10.1109/ICASSP49357.2023.10095637
2023
Closest in time.
H. Chung, J. Kim, M. T. Mccann, M. L. Klasky, and J. C. Ye, “Diffusion Posterior Sampling for General Noisy Inverse Problems,” in Proceedings International Conference on Learning Representations (ICLR) (2023 May)
2023
Closest in time.
K. Liu, W. Gan, and C. Yuan, “MAID: A Conditional Diffusion Model for Long Music Audio Inpainting,” in Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 1–5 (Rhodes, Greece) (2023 Jun.), https://doi.org/10.1109/ICASSP49357.2023.10095769
2023
Closest in time.
K. W. Cheuk, R. Sawata, T. Uesaka, N. Murata, N. Takahashi, S. Takahashi, et al. , “DiffRoll: Diffusion-based Generative Music Transcription with Unsupervised Pretraining Capability,” in Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 1–5 (2023 Jun.), https://doi.org/10.1109/ICASSP49357.2023.10095935
2023
Closest in time.
2023
Closest in time.
J. Richter, S. Welker, J.-M. Lemercier, B. Lay, and T. Gerkmann, “Speech Enhancement and Dereverberation with Diffusion-Based Generative Models,” IEEE/ACM Trans. Audio Speech Lang. Process. , pp. 2351–2364 (2023 Jun.), https://doi.org/10.1109/TASLP.2023.3285241
2023
Closest in time.
H. Wu, K. Tan, B. Xu, A. Kumar, and D. Wong, “Rethinking Complex-Valued Deep Neural Networks for Monaural Speech Enhancement,” in Proceedings of INTERSPEECH (2023 Jan.)
2023
Closest in time.
ITU, “Method for Objective Measurements of Perceived Audio Quality,” Rec. BS.1387-2, International Telecommunication Union, Geneva, Switzerland (2023 May)
2023
Closest in time.
M. Schoeffler, S. Bartoschek, F.-R. Stöter, M. Roess, S. Westphal, B. Edler, et al. , “WebMUSHRA—A Comprehensive Framework for Web-Based Listening Tests,” J. Open Res. Softw., , vol. 6, no. 1, pp. 1–8 (2018 Feb.), 10.5334/jors.187
2023
Closest in time.