Fetching the paper…
Reading the bibliography…
A scattering transform defines a locally translation invariant representation which is stable to time-warping deformations.
L. Lucy, “An iterative technique for the rectification of observed distributions,” Astron. J. , vol. 79, p. 745, 1974
1974
Earlier work this paper cites.
S. Davis and P. Mermelstein, “Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences,” IEEE Trans. Acoust., Speech, Signal Process. , vol. 28, no. 4, pp. 357–366, 1980
1980
Earlier work this paper cites.
D. W. Griffin and J. S. Lim, “Signal estimation from modified short-time fourier transform,” IEEE Trans. Acoust., Speech, Signal Process. , vol. 32, no. 2, pp. 236–243, 1984
1984
Earlier work this paper cites.
W. Fisher, G. Doddington, and K. Goudie-Marshall, “The DARPA speech recognition research database: specifications and status,” in Proc. DARPA Workshop on Speech Recognition , 1986, pp. 93–99
1986
Earlier work this paper cites.
K.-F. Lee and H.-W. Hon, “Speaker-independent phone recognition using hidden markov models,” Acoustics, Speech and Signal Processing, IEEE Transactions on , vol. 37, no. 11, pp. 1641–1648, 1989
1989
Earlier work this paper cites.
M. Slaney and R. Lyon, Visual representations of speech signals . M. Cooke, S. Beet and M. Crawford (Eds.) John Wiley and Sons, 1993, ch. On the importance of time–a temporal representation of sound, pp. 95–116
1993
Earlier work this paper cites.
H. Hermansky, “The modulation spectrum in the automatic recognition of speech,” in Proc. IEEE ASRU , 1997, pp. 140–147
1997
Earlier work this paper cites.
T. Dau, B. Kollmeier, and A. Kohlrausch, “Modeling auditory processing of amplitude modulation. I. Detection and masking with narrow-band carriers,” J. Acoust. Soc. Am. , vol. 102, no. 5, pp. 2892–2905, 1997
1997
Earlier work this paper cites.
A. K. Halberstadt, “Heterogeneous acoustic measurements and multiple classifiers for speech recognition,” Ph.D. dissertation, Massachusetts Institute of Technology, 1998
1998
Earlier work this paper cites.
S. Mallat, A wavelet tour of signal processing . Academic Press, 1999
1999
Earlier work this paper cites.
P. Clarkson and P. J. Moreno, “On the use of support vector machines for phonetic classification,” in IEEE Trans. Acoust., Speech, Signal Process. , vol. 2. IEEE, 1999, pp. 585–588
1999
Earlier work this paper cites.
R. D. Patterson, “Auditory images: How complex sounds are represented in the auditory system,” Journal of the Acoustical Society of Japan (E) , vol. 21, no. 4, pp. 183–190, 2000
2000
Earlier work this paper cites.
M. S. Vinton and L. E. Atlas, “Scalable and progressive audio codec,” in Acoustics, Speech, and Signal Processing, 2001. Proceedings.(ICASSP’01). 2001 IEEE International Conference on , vol. 5. IEEE, 2001, pp. 3277–3280
2001
Earlier work this paper cites.
G. Tzanetakis and P. Cook, “Musical genre classification of audio signals,” IEEE Transactions on Speech and Audio Processing , vol. 10, no. 5, pp. 293–302, 2002
2002
Earlier work this paper cites.
J. K. Thompson and L. E. Atlas, “A non-uniform modulation transform for audio coding with increased time resolution,” in Acoustics, Speech, and Signal Processing, 2003. Proceedings.(ICASSP’03). 2003 IEEE International Conference on , vol. 5. IEEE, 2003, pp. V–397
2003
Earlier work this paper cites.
T. Chi, P. Ru, and S. Shamma, “Multiresolution spectrotemporal analysis of complex sounds,” J. Acoust. Soc. Am. , vol. 118, no. 2, pp. 887–906, 2005
2005
Earlier work this paper cites.
S. Schimmel and L. Atlas, “Coherent envelope detection for modulation filtering of speech,” in Proc. of ICASSP , vol. 1, 2005, pp. 221–224
2005
Earlier work this paper cites.
N. Mesgarani, M. Slaney, and S. A. Shamma, “Discrimination of speech from nonspeech based on multiscale spectro-temporal modulations,” IEEE Audio, Speech, Language Process. , vol. 14, no. 3, pp. 920–930, 2006
2006
Cited alongside, same era.
E. C. Smith and M. S. Lewicki, “Efficient auditory coding,” Nature , vol. 439, no. 7079, pp. 978–982, 2006
2006
Cited alongside, same era.
H.-A. Chang and J. R. Glass, “Hierarchical large-margin gaussian mixture models for phonetic classification,” in Proc. IEEE ASRU . IEEE, 2007, pp. 272–277
2007
Cited alongside, same era.
C. Lee, J. Shih, K. Yu, and H. Lin, “Automatic music genre classification based on modulation spectral analysis of spectral and cepstral features,” IEEE Transactions on Multimedia , vol. 11, no. 4, pp. 670–682, 2009
2009
Cited alongside, same era.
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-R. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath et al. , “Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,” Signal Processing Magazine, IEEE , vol. 29, no. 6, pp. 82–97, 2012
2012
Later among the works it cites.
E. J. Humphrey, T. Cho, and J. P. Bello, “Learning a robust tonnetz-space transform for automatic chord recognition,” in Proc. IEEE ICASSP , 2012, pp. 453–456
2012
Later among the works it cites.
E. Battenberg and D. Wessel, “Analyzing drum patterns using conditional deep belief networks,” in Proc. ISMIR , 2012
2012
Later among the works it cites.
I. Waldspurger and S. Mallat, “Recovering the phase of a complex wavelet transform,” CMAP, Ecole Polytechnique, Tech. Rep., 2012
2012
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2009
Cited alongside, same era.
Y. LeCun, K. Kavukvuoglu, and C. Farabet, “Convolutional networks and applications in vision,” in Proc. IEEE ISCAS , 2010
2010
Cited alongside, same era.
P. Hamel and D. Eck, “Learning features from music audio with deep belief networks,” in Proc. ISMIR , 2010
2010
Cited alongside, same era.
G. Sell and M. Slaney, “Solving demodulation as an optimization problem,” Audio, Speech, and Language Processing, IEEE Transactions on , vol. 18, no. 8, pp. 2051–2066, 2010
2010
Cited alongside, same era.
J. McDermott and E. Simoncelli, “Sound texture perception via statistics of the auditory periphery: Evidence from sound synthesis,” Neuron , vol. 71, no. 5, pp. 926–940, 2011
2011
Cited alongside, same era.
M. Ramona and G. Peeters, “Audio identification based on spectral modeling of bark-bands energy and synchronization through onset detection,” in Proc. IEEE ICASSP , 2011, pp. 477–480
2011
Cited alongside, same era.
D. Ellis, X. Zeng, and J. McDermott, “Classifying soundtracks with audio texture features,” in Proc. IEEE ICASSP , Prague, Czech Republic, May. 22-27 2011, pp. 5880–5883
2011
Cited alongside, same era.
R. Turner and M. Sahani, “Probabilistic amplitude and frequency demodulation,” in Advances in Neural Information Processing Systems , 2011, pp. 981–989
2011
Cited alongside, same era.
2012
Later among the works it cites.
I. Waldspurger, A. d’Aspremont, and S. Mallat, “Phase recovery, maxcut and complex semidefinite programming,” CMAP, Ecole Polytechnique, Tech. Rep., 2012
2012
Later among the works it cites.
B. L. Sturm, “An analysis of the GTZAN music genre dataset,” in Proceedings of the second international ACM workshop on Music information retrieval with user-centered and multimodal strategies . ACM, 2012, pp. 7–12
2012
Later among the works it cites.
V. Chudáček, J. Andén, S. Mallat, P. Abry, and M. Doret, “Scattering transform for intrapartum fetal heart rate characterization and acidosis detection,” in Proc. IEEE EMBC , 2013
2013
Closest in time.
L. Deng, O. Abdel-Hamid, and D. Yu, “A deep convolutional neural network using heterogeneous pooling for trading acoustic invariance with phonetic confusion,” in Proc. ICASSP , 2013
2013
Closest in time.
A. Graves, A.-R. Mohamed, and G. Hinton, “Speech recognition with deep recurrent neural networks,” Proc. ICASSP , 2013
2013
Closest in time.
J. Bruna and S. Mallat, “Invariant scattering convolution networks,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 35, no. 8, pp. 1872–1886, 2013
2013
Closest in time.
L. Sifre and S. Mallat, “Rotation, scaling and deformation invariant scattering for texture discrimination,” in Proc. CVPR , 2013
2013
Closest in time.
E. J. Candès, Y. C. Eldar, T. Strohmer, and V. Voroninski, “Phase retrieval via matrix completion,” SIAM Journal on Imaging Sciences , vol. 6, no. 1, pp. 199–225, 2013
2013
Closest in time.
C. Baugé, M. Lagrange, J. Andén, and S. Mallat, “Representing environmental sounds using the separable scattering transform,” in Proc. IEEE ICASSP , 2013
2013
Closest in time.
X. Chen and P. J. Ramadge, “Music genre classification using multiscale scattering and sparse representations,” in Proc. CISS , 2013
2013
Closest in time.
J. Andén, “Time and frequency scattering for audio classification,” Ph.D. dissertation, Ecole Polytechnique, 2014
2014
Closest in time.