Fetching the paper…
Reading the bibliography…
Self-supervised learning (SSL) is at the origin of unprecedented improvements in many different domains including computer vision and natural language processing.
1904
Earlier work this paper cites.
A. Caubrière, N. A. Tomashenko, et al., · 1906
Earlier work this paper cites.
Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter,
V. Sanh, L. Debut, J. Chaumond, T. Wolf, · 1910
Earlier work this paper cites.
R. De Mori, Spoken Dialogues with Computers, Academic Press, Inc., Orlando, FL, USA, 1997
1997
Earlier work this paper cites.
Long short-term memory,
S. Hochreiter, J. Schmidhuber, · 1997
Earlier work this paper cites.
2002
Earlier work this paper cites.
African accented french, slr57, 2003. Type: dataset, https://www.openslr.org/57/
2003
Earlier work this paper cites.
Statistical significance tests for machine translation evaluation,
P. Koehn, · 2004
Earlier work this paper cites.
The c-oral-rom corpus. a multilingual resource of spontaneous speech for romance languages,
E. Cresti, F. B. do Nascimento, A. M. Sandoval, J. Veronis, P. Martin, K. Choukri, · 2004
Earlier work this paper cites.
Présentation du corpus de référence du français parlé,
E. DELIC, S. Teston-Bonnard, J. Véronis, · 2004
Earlier work this paper cites.
A tutorial on text-independent speaker verification,
F. Bimbot, J.-F. Bonastre, C. Fredouille, G. Gravier, I. Magrin-Chagnolleau, S. Meignier, T. Merlin, J. Ortega-García, D. Petrovska-Delacrétaz, D. A. Reynolds, · 2004
Earlier work this paper cites.
On the use of finite state transducers for semantic interpretation,
C. Raymond, F. Béchet, R. De Mori, G. Damnati, · 2005
Earlier work this paper cites.
Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,
A. Graves, S. Fernández, F. Gomez, J. Schmidhuber, · 2006
Earlier work this paper cites.
Results of the French evalda-media evaluation campaign for literal understanding,
H. Bonneau-Maynard, C. Ayache, et al., · 2006
Earlier work this paper cites.
Corpus description of the ESTER evaluation campaign for the rich transcription of french broadcast news.,
S. Galliano, E. Geoffrois, G. Gravier, J.-F. Bonastre, D. Mostefa, K. Choukri, · 2006
Earlier work this paper cites.
The Nijmegen Corpus of Casual French,
F. Torreira, M. Adda-Decker, M. Ernestus, · 2009
Earlier work this paper cites.
Re-ranking models based-on small training data for spoken language understanding,
M. Dinarelli, A. Moschitti, G. Riccardi, · 2009
Earlier work this paper cites.
What’s in an ontology for spoken language understanding,
S. Quarteroni, G. Riccardi, M. Dinarelli, · 2009
Earlier work this paper cites.
The ESTER 2 evaluation campaign for the rich transcription of French radio broadcasts,
S. Galliano, G. Gravier, L. Chaubard, · 2009
Earlier work this paper cites.
The EPAC Corpus: Manual and Automatic Annotations of Conversational Speech in French Broadcast News,
Y. Estève, T. Bazillon, J.-Y. Antoine, F. Béchet, J. Farinas, · 2010
Earlier work this paper cites.
Lium spkdiarization: an open source toolkit for diarization,
S. Meignier, T. Merlin, · 2010
Earlier work this paper cites.
Comparing stochastic approaches to spoken language understanding in multiple languages,
S. Hahn, M. Dinarelli, et al., · 2010
Earlier work this paper cites.
fairseq S2T: fast speech-to-text modeling with fairseq,
C. Wang, Y. Tang, X. Ma, A. Wu, D. Okhonko, J. M. Pino, · 2010
Earlier work this paper cites.
Front-end factor analysis for speaker verification,
N. Dehak, P. J. Kenny, R. Dehak, P. Dumouchel, P. Ouellet, · 2010
Earlier work this paper cites.
S. Branca-Rosoff, S. Fleury, F. Lefeuvre, M. Pires, Discours sur la ville. Présentation du Corpus de Français parlé Parisien des années 2000 (CFPP2000), 2012. Http://cfpp2000.univ-paris3.fr/CFPP2000.pdf
2012
Earlier work this paper cites.
Un grand corpus oral "disponible" : le corpus d’Orléans 1968-2012,
I. Eshkol-Taravella, O. Baude, D. Maurel, L. Hriba, C. Dugua, I. Tellier, · 2012
Earlier work this paper cites.
Introducing the Geneva Multimodal Expression Corpus for Experimental Research on Emotion Perception,
T. Bänziger, M. Mortillaro, K. Scherer, · 2012
Earlier work this paper cites.
Robustesse et portabilités multilingue et multi-domaines des systèmes de compréhension de la parole : le projet PortMedia,
F. Lefèvre, D. Mostefa, L. Besacier, Y. Estève, M. Quignard, N. Camelin, B. Favre, B. Jabaian, L. Rojas-Barahona, · 2012
Earlier work this paper cites.
The ETAPE corpus for the evaluation of speech-based TV content processing in the French language,
G. Gravier, G. Adda, N. Paulsson, M. Carré, A. Giraudel, O. Galibert, · 2012
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing,
T. Kudo, J. Richardson, · 2012
Earlier work this paper cites.
ADADELTA: an adaptive learning rate method,
M. D. Zeiler, · 2012
Earlier work this paper cites.
The REPERE corpus: a multimodal corpus for person recognition.,
A. Giraudel, M. Carré, V. Mapelli, J. Kahn, O. Galibert, L. Quintard, · 2012
Earlier work this paper cites.
The role of appraisal in emotion,
A. Moors, K. R. Scherer, · 2013
Earlier work this paper cites.
The impact of emotion on perception, attention, memory, and decision-making,
T. Brosch, K. Scherer, D. Grandjean, D. Sander, · 2013
Earlier work this paper cites.
Introducing the recola multimodal corpus of remote collaborative and affective interactions,
F. Ringeval, A. Sonderegger, J. Sauer, D. Lalanne, · 2013
Earlier work this paper cites.
J. Carruthers, French oral narrative corpus, 2013. Commissioning Body / Publisher: Oxford Text Archive
2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate,
D. Bahdanau, K. Cho, Y. Bengio, · 2015
Earlier work this paper cites.
W. Chan, N. Jaitly, Q. V. Le, O. Vinyals, · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization,
D. P. Kingma, J. Ba, · 2015
Earlier work this paper cites.
D. Snyder, G. Chen, D. Povey, Musan: A music, speech, and noise corpus, 2015. arXiv:1510.08484
2015
Earlier work this paper cites.
Le projet orféo: un corpus d’étude pour le français contemporain,
C. Benzitoun, J.-M. Debaisieux, H.-J. Deulofeu, · 2016
Earlier work this paper cites.
V. André, Fleuron: Français langue Étrangère universitaire–ressources et outils numériques, 2016. URL: https://fleuron.atilf.fr/index.php?lg=fr
2016
Earlier work this paper cites.
Fabiole, a speech database for forensic speaker comparison,
M. Ajili, J.-F. Bonastre, J. Kahn, S. Rossato, G. Bernard, · 2016
Earlier work this paper cites.
Les parlers jeunes dans l’île-de-France multiculturelle,
G. Françoise, · 2017
Earlier work this paper cites.
Label-dependency coding in Simple Recurrent Networks for Spoken Language Understanding,
M. Dinarelli, V. Vukotic, C. Raymond, · 2017
Earlier work this paper cites.
Attention is all you need,
A. Vaswani, N. Shazeer, et al., · 2017
Earlier work this paper cites.
ICAR, Clapi, 2017. URL: https://hdl.handle.net/11403/clapi/v1 , ORTOLANG (Open Resources and TOols for LANGuage) –www.ortolang.fr
2017
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding,
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, S. Bowman, · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding,
A. v. d. Oord, Y. Li, O. Vinyals, · 2018
Earlier work this paper cites.
A Canadian French emotional speech dataset,
P. Gournay, O. Lahaie, R. Lefebvre, · 2018
Earlier work this paper cites.
Decoupled weight decay regularization,
I. Loshchilov, F. Hutter, · 2018
Cited alongside, same era.
Label-dependencies aware recurrent neural networks,
Y. Dupont, M. Dinarelli, I. Tellier, · 2018
Cited alongside, same era.
Towards end-to-end spoken language understanding,
D. Serdyuk, Y. Wang, et al., · 2018
Cited alongside, same era.
A call for clarity in reporting BLEU scores,
M. Post, · 2018
Cited alongside, same era.
CLESTHIA, Cfpp2000, 2018. URL: https://hdl.handle.net/11403/cfpp2000/v1 , ORTOLANG (Open Resources and TOols for LANGuage) –www.ortolang.fr
2018
Cited alongside, same era.
The jhu-mit system description for nist sre18 (2018)
J. Villalba, N. Chen, D. Snyder, D. Garcia-Romero, A. McCree, G. Sell, J. Borgstrom, F. Richardson, S. Shon, F. Grondin, et al., · 2018
Targeting the benchmark: On methodology in current natural language processing research,
D. Schlangen, · 2021
Later among the works it cites.
Phonetically motivated self-supervised speech representation learning.,
X. Yue, H. Li, · 2021
Later among the works it cites.
Speechbrain: A general-purpose speech toolkit,
M. Ravanelli, T. Parcollet, P. Plantinga, A. Rouhe, S. Cornell, L. Lugosch, C. Subakan, N. Dawalatabad, A. Heba, J. Zhong, et al., · 2021
Later among the works it cites.
VoxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,
C. Wang, M. Riviere, A. Lee, A. Wu, C. Talnikar, D. Haziza, M. Williamson, J. Pino, E. Dupoux, · 2021
Later among the works it cites.
Unsupervised Cross-Lingual Representation Learning for Speech Recognition,
A. Conneau, A. Baevski, R. Collobert, A. Mohamed, M. Auli, · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Superglue: A stickier benchmark for general-purpose language understanding systems,
A. Wang, Y. Pruksachatkun, N. Nangia, A. Singh, J. Michael, F. Hill, O. Levy, S. Bowman, · 2019
Cited alongside, same era.
fairseq: A fast, extensible toolkit for sequence modeling,
M. Ott, S. Edunov, A. Baevski, A. Fan, S. Gross, N. Ng, D. Grangier, M. Auli, · 2019
Cited alongside, same era.
Mpf, 2019. Https://hdl.handle.net/11403/mpf/v3, ORTOLANG (Open Resources and TOols for LANGuage) –www.ortolang.fr
2019
Cited alongside, same era.
An Unsupervised Autoregressive Model for Speech Representation Learning,
Y.-A. Chung, W.-N. Hsu, H. Tang, J. Glass, · 2019
Cited alongside, same era.
SLU FOR VOICE COMMAND IN SMART HOME: COMPARISON OF PIPELINE AND END-TO-END APPROACHES,
T. Desot, F. Portet, M. Vacher, · 2019
Cited alongside, same era.
R. Müller, S. Kornblith, G. Hinton, When Does Label Smoothing Help?, Curran Associates Inc., Red Hook, NY, USA, 2019
2019
Cited alongside, same era.
Where are we in semantic concept extraction for Spoken Language Understanding? ⋆ \star ,
S. Ghannay, A. Caubrière, et al., · 2021
Later among the works it cites.
End2End Acoustic to Semantic Transduction,
V. Pelloin, N. Camelin, et al., · 2021
Later among the works it cites.
Multilingual speech translation from efficient finetuning of pretrained models,
X. Li, C. Wang, Y. Tang, C. Tran, Y. Tang, J. Pino, A. Baevski, A. Conneau, M. Auli, · 2021
Later among the works it cites.
The multilingual tedx corpus for speech recognition and translation,
S. Elizabeth, W. Matthew, B. Jacob, R. Cattoni, M. Negri, M. Turchi, D. W. Oard, P. Matt, · 2021
Later among the works it cites.
Layer-wise analysis of a self-supervised speech representation model,
A. Pasad, J.-C. Chou, K. Livescu, · 2021
Later among the works it cites.
The IDLab VoxSRC-20 submission: Large margin fine-tuning and quality-aware score calibration in DNN based speaker verification,
J. Thienpondt, B. Desplanques, K. Demuynck, · 2021
Later among the works it cites.
The Energy and Carbon Footprint of Training End-to-End Speech Recognizers,
T. Parcollet, M. Ravanelli, · 2021
Later among the works it cites.
Self-supervised speech representation learning: A review,
A. Mohamed, H.-y. Lee, L. Borgholt, J. D. Havtorn, J. Edin, C. Igel, K. Kirchhoff, S.-W. Li, K. Livescu, L. Maaløe, et al., · 2022
Later among the works it cites.
Audio self-supervised learning: A survey,
S. Liu, A. Mallol-Ragolta, E. Parada-Cabaleiro, K. Qian, X. Jing, A. Kathan, B. Hu, B. W. Schuller, · 2022
Later among the works it cites.
Self-supervised pretraining improves self-supervised pretraining,
C. J. Reed, X. Yue, A. Nrusimha, S. Ebrahimi, V. Vijaykumar, R. Mao, B. Li, S. Zhang, D. Guillory, S. Metzger, et al., · 2022
Later among the works it cites.
S4rl: Surprisingly simple self-supervision for offline reinforcement learning in robotics,
S. Sinha, A. Mandlekar, A. Garg, · 2022
Later among the works it cites.
Self-supervised learning in remote sensing: A review,
Y. Wang, C. Albrecht, N. A. A. Braham, L. Mou, X. Zhu, · 2022
Later among the works it cites.
Self-supervised learning in medicine and healthcare,
R. Krishnan, P. Rajpurkar, E. J. Topol, · 2022
Later among the works it cites.
Self-supervised learning methods and applications in medical imaging analysis: A survey,
S. Shurrab, R. Duwairi, · 2022
Later among the works it cites.
Large-scale self-supervised speech representation learning for automatic speaker verification,
Z. Chen, S. Chen, Y. Wu, Y. Qian, C. Wang, S. Liu, Y. Qian, M. Zeng, · 2022
Later among the works it cites.
Toward Low-Cost End-to-End Spoken Language Understanding,
M. Dinarelli, M. Naguib, F. Portet, · 2022
Later among the works it cites.
On the use of semantically-aligned speech representations for spoken language understanding,
G. Laperriere, V. Pelloin, M. Rouvier, T. Stafylakis, Y. Esteve, · 2022
Later among the works it cites.
Investigating self-supervised learning for speech enhancement and separation,
Z. Huang, S. Watanabe, S.-w. Yang, P. García, S. Khudanpur, · 2022
Later among the works it cites.
Efficient personalized speech enhancement through self-supervised learning,
A. Sivaraman, M. Kim, · 2022
Later among the works it cites.
Superb-sg: Enhanced speech processing universal performance benchmark for semantic and generative capabilities,
H.-S. Tsai, H.-J. Chang, W.-C. Huang, Z. Huang, K. Lakhotia, S.-w. Yang, S. Dong, A. Liu, C.-I. Lai, J. Shi, et al., · 2022
Later among the works it cites.
Slue: New benchmark tasks for spoken language understanding evaluation on natural speech,
S. Shon, A. Pasad, F. Wu, P. Brusco, Y. Artzi, K. Livescu, K. J. Han, · 2022
Later among the works it cites.
A. Baevski, A. Babu, W.-N. Hsu, M. Auli, · 2022
Later among the works it cites.
Autoregressive predictive coding: A comprehensive study,
G.-P. Yang, S.-L. Yeh, Y.-A. Chung, J. Glass, H. Tang, · 2022
Later among the works it cites.
Self-supervised learning with random-projection quantizer for speech recognition,
C.-C. Chiu, J. Qin, Y. Zhang, J. Yu, Y. Wu, · 2022
Later among the works it cites.
S. Felice, S. Evain, F. Portet, Audiocité, 2022. URL: https://www.audiocite.net/
2022
Later among the works it cites.
Speech resources in the tamasheq language,
M. Zanon Boito, F. Bougares, F. Barbier, S. Gahbiche, L. Barrault, M. Rouvier, Y. Estève, · 2022
Later among the works it cites.
A study of gender impact in self-supervised models for speech-to-text systems,
M. Z. Boito, L. Besacier, N. Tomashenko, Y. Estève, · 2022
Later among the works it cites.
Deep versus Wide: An Analysis of Student Architectures for Task-Agnostic Knowledge Distillation of Self-Supervised Speech Models,
T. Ashihara, T. Moriya, K. Matsuura, T. Tanaka, · 2022
Later among the works it cites.
Match to win: Analysing sequences lengths for efficient self-supervised learning in speech and audio,
Y. Gao, J. Fernandez-Marques, T. Parcollet, P. P. de Gusmao, N. D. Lane, · 2022
Later among the works it cites.
Robust speech recognition via large-scale weak supervision,
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, I. Sutskever, · 2022
Later among the works it cites.
End-to-end spoken language understanding: Performance analyses of a voice command task in a low resource setting,
T. Desot, F. Portet, M. Vacher, · 2022
Later among the works it cites.
Toward Low-Cost End-to-End Spoken Language Understanding,
M. Dinarelli, M. Naguib, F. Portet, · 2022
Later among the works it cites.
ON-TRAC consortium systems for the IWSLT 2022 dialect and low-resource speech translation tasks,
M. Z. Boito, J. Ortega, H. Riguidel, A. Laurent, L. Barrault, F. Bougares, F. Chaabani, H. Nguyen, F. Barbier, S. Gahbiche, Y. Estève, · 2022
Later among the works it cites.
Multi-corpus affect recognition with emotion embeddings and self-supervised representations of speech,
S. Alisamir, F. Ringeval, F. Portet, · 2022
Later among the works it cites.
Multi-corpus affect recognition with emotion embeddings and self-supervised representations of speech,
S. Alisamir, F. Ringeval, F. Portet, · 2022
Later among the works it cites.
End-to-End Dependency Parsing of Spoken French,
A. Pupier, M. Coavoux, B. Lecouteux, J. Goulian, · 2022
Later among the works it cites.
Fine-tuning wav2vec2 for speaker recognition,
N. Vaessen, D. A. Van Leeuwen, · 2022
Later among the works it cites.
ML-SUPERB: Multilingual Speech Universal PERformance Benchmark,
J. Shi, D. Berrebbi, W. Chen, E.-P. Hu, W.-P. Huang, H.-L. Chung, X. Chang, S.-W. Li, A. Mohamed, H. yi Lee, S. Watanabe, · 2023
Closest in time.
Indicsuperb: A speech processing universal performance benchmark for indian languages,
T. Javed, K. Bhogale, A. Raman, P. Kumar, A. Kunchukuttan, M. M. Khapra, · 2023
Closest in time.
Speech self-supervised representation benchmarking: Are we doing it right?,
S. Zaiem, Y. Kemiche, T. Parcollet, S. Essid, M. Ravanelli, · 2023
Closest in time.
Reducing barriers to self-supervised learning: Hubert pre-training with academic compute,
W. Chen, X. Chang, Y. Peng, Z. Ni, S. Maiti, S. Watanabe, · 2023
Closest in time.
Comparative layer-wise analysis of self-supervised speech models,
A. Pasad, B. Shi, K. Livescu, · 2023
Closest in time.
Power hungry processing: Watts driving the cost of ai deployment?,
A. S. Luccioni, Y. Jernite, E. Strubell, · 2023
Closest in time.
CoVoST 2 and Massively Multilingual Speech Translation,
C. Wang, A. Wu, J. Gu, J. Pino, · 2027
Closest in time.