Fetching the paper…
Reading the bibliography…
In this paper, we develop a deep learning based semantic communication system for speech transmission, named DeepSC-ST.
C. E. Shannon and W. Weaver, The Mathematical Theory of Communication . Champaign, Il, USA: Univ. Illinois Press, 1949
1949
Earlier work this paper cites.
R. Carnap and Y. Bar-Hillel, “An outline of a theory of semantic information,” Res. Lab. Electron., Massachusetts Inst. Technol., Cambridge, MA, USA, RLE Tech. Rep. 247, Oct. 1952
1952
Earlier work this paper cites.
D. A. Huffman, “A method for the construction of minimum-redundancy codes,” Proc. the IRE , vol. 40, no. 9, pp. 1098–1101, Sep. 1952
1952
Earlier work this paper cites.
L. Rabiner, “A tutorial on hidden Markov models and selected applications in speech recognition,” Proc. the IEEE , vol. 77, no. 2, pp. 257–286, Nov. 1989
1989
Earlier work this paper cites.
E. Moulines and F. Charpentier, “Pitch-synchronous waveform processing techniques for text-to-speech synthesis using diphones,” Speech Commun. , vol. 9, no. 5-6, pp. 453–467, Dec. 1990
1990
Earlier work this paper cites.
A. Hunt and A. Black, “Unit selection in a concatenative speech synthesis system using a large speech database,” in Proc. IEEE Int. Conf. Acoustics, Speech, Signal Process. (ICASSP) , Atlanta, GA, USA, May 1996, pp. 373–376
1996
Earlier work this paper cites.
M. Schuster and K. Paliwal, “Bidirectional recurrent neural networks,” IEEE Trans. Signal Process. , vol. 45, no. 11, pp. 2673–2681, Nov. 1997
1997
Earlier work this paper cites.
N. Tishby, F. C. Pereira, and W. Bialek, “The information bottleneck method,” arXiv preprint physics/0004057 , Apr. 2000
2000
Earlier work this paper cites.
B. Bessette, R. Salami, R. Lefebvre, M. Jelinek, J. Rotola-Pukkila, J. Vainio, H. Mikkola, and K. Jarvinen, “The adaptive multirate wideband speech codec (AMR-WB),” IEEE Trans. Speech, Audio Process. , vol. 10, no. 8, pp. 620–636, Nov. 2002
2002
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,” in Proc. Int. Conf. Mach. Learning (ICML) , Pittsburgh, USA, Jun. 2006, pp. 369–376
2006
Earlier work this paper cites.
A.-r. Mohamed, G. E. Dahl, and G. Hinton, “Acoustic modeling using deep belief networks,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 20, no. 1, pp. 14–22, Jan. 2012
2012
Earlier work this paper cites.
N. Morgan, “Deep and wide: Multiple layers in automatic speech recognition,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 20, no. 1, pp. 7–13, Jan. 2012
2012
Earlier work this paper cites.
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, and B. Kingsbury, “Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,” IEEE Signal Process. Mag. , vol. 29, no. 6, pp. 82–97, Nov. 2012
2012
Earlier work this paper cites.
A. Graves, A.-r. Mohamed, and G. Hinton, “Speech recognition with deep recurrent neural networks,” in Proc. IEEE Int. Conf. Acoustics, Speech, Signal Process. (ICASSP) , Vancouver, BC, Canada, May 2013, pp. 6645–6649
2013
Earlier work this paper cites.
H. Ze, A. Senior, and M. Schuster, “Statistical parametric speech synthesis using deep neural networks,” in Proc. IEEE Int. Conf. Acoustics, Speech, Signal Process. (ICASSP) , Vancouver, BC, Canada, May 2013, pp. 7962–7966
2013
Earlier work this paper cites.
P. Basu, J. Bao, M. Dean, and J. Hendler, “Preserving quality of information by using semantic relationships,” Pervasive Mob. Comput. , vol. 11, pp. 188–202, Apr. 2014
2014
Earlier work this paper cites.
A. Graves and N. Jaitly, “Towards end-to-end speech recognition with recurrent neural networks,” in Proc. 31st Int. Conf. Mach. Learning (ICML) , Beijing, China, Jun. 2014, pp. 1764–1772
2014
Earlier work this paper cites.
H. Sak, A. Senior, and F. Beaufays, “Long short-term memory recurrent neural network architectures for large scale acoustic modeling,” in Proc. Interspeech , Singapore, Sep. 2014, pp. 338–342
2014
Earlier work this paper cites.
H. Zen and H. Sak, “Unidirectional long short-term memory recurrent neural network with recurrent output layer for low-latency speech synthesis,” in Proc. IEEE Int. Conf. Acoustics, Speech, Signal Process. (ICASSP) , South Brisbane, QLD, Australia, Apr. 2015, pp. 4470–4474
2015
Earlier work this paper cites.
Z. Wu, C. Valentini-Botinhao, O. Watts, and S. King, “Deep neural networks employing multi-task learning and stacked bottleneck features for speech synthesis,” in Proc. IEEE Int. Conf. Acoustics, Speech, Signal Process. (ICASSP) , South Brisbane, QLD, Australia, Apr. 2015, pp. 4460–4464
2015
Earlier work this paper cites.
A. Balatsoukas-Stimming, M. B. Parizi, and A. Burg, “LLR-based successive cancellation list decoding of polar codes,” IEEE Trans. Signal Process. , vol. 63, no. 19, pp. 5165–5179, Jun. 2015
2015
Earlier work this paper cites.
D. Amodei, S. Ananthanarayanan, R. Anubhai, and etc., “Deep speech 2: End-to-end speech recognition in English and Mandarin,” in Proc. 33rd Int. Conf. Mach. Learning (ICML) , New York, USA, Jun. 2016, pp. 173–182
2016
Cited alongside, same era.
Y. Zhang, G. Chen, D. Yu, K. Yao, S. Khudanpur, and J. Glass, “Highway long short-term memory rnns for distant speech recognition,” in Proc. IEEE Int. Conf. Acoustics, Speech, Signal Process. (ICASSP) , Shanghai, China, Mar. 2016, pp. 5755–5759
2016
Cited alongside, same era.
A. v. d. Oord et al. , “WaveNet: A generative model for raw audio,” in Proc. ISCA Workshop on Speech Synthesis Workshop (SSW 9) , 2016, pp. 125–125
2016
Cited alongside, same era.
E. Battenberg et al. , “Exploring neural transducers for end-to-end speech recognition,” in Proc. IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) , Okinawa, Japan, Dec. 2017, pp. 206–213
2017
Cited alongside, same era.
M. Bińkowski, J. Donahue, S. Dieleman, A. Clark, E. Elsen, N. Casagrande, L. C. Cobo, and K. Simonyan, “High fidelity speech synthesis with adversarial networks,” in Proc. Int. Conf. Learning Representations (ICLR) , Formerly Addis Ababa, Ethiopia, Apr. 2020
2020
Later among the works it cites.
Z. Weng, Z. Qin, and G. Y. Li, “Semantic communications for speech recognition,” in Proc. IEEE Global Commun. Conf. (GLOBECOM) , Madrid, Spain, Dec. 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
H. Xie, Z. Qin, G. Y. Li, and B.-H. Juang, “Deep learning enabled semantic communication systems,” IEEE Trans. Signal Process. , vol. 69, pp. 2663–2675, Apr. 2021
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Jose, M. Soroush, K. Kundan, S. J. Felipe, K. Kyle, C. Aaron, and B. Yoshua, “Char2wav: End-to-end speech synthesis,” in Proc. 5th Int. Conf. Learning Representations (ICLR) , Toulon, France, Apr. 2017
2017
Cited alongside, same era.
Y. Wang et al. , “Tacotron: Towards end-to-end speech synthesis,” in Proc. Interspeech , Stockholm, Sweden, Aug. 2017, pp. 4006–4010
2017
Cited alongside, same era.
A. Tamamori, T. Hayashi, K. Kobayashi, K. Takeda, and T. Toda, “Speaker-dependent WaveNet vocoder,” in Proc. Interspeech , Stockholm, Sweden, Aug. 2017, pp. 1118–1122
2017
Cited alongside, same era.
K. Ito and L. Johnson, “The LJ speech dataset,” https://keithito.com/LJ-Speech-Dataset/ , 2017
2017
Cited alongside, same era.
N. Farsad, M. Rao, and A. Goldsmith, “Deep learning for joint source-channel coding of text,” in Proc. IEEE Int. Conf. Acoustics, Speech, Signal Process. (ICASSP) , Calgary, Canada, Apr. 2018, pp. 2326–2330
2018
Cited alongside, same era.
W. Ping, K. Peng, A. Gibiansky, S. Ö. Arik, A. Kannan, S. Narang, J. Raiman, and J. Miller, “Deep voice 3: Scaling text-to-speech with convolutional sequence learning,” in Proc. 6th Int. Conf. Learning Representations (ICLR) , Vancouver, BC, Canada, Feb. 2018
2018
Cited alongside, same era.
J. Shen et al. , “Natural TTS synthesis by conditioning wavenet on MEL spectrogram predictions,” in Proc. IEEE Int. Conf. Acoustics, Speech, Signal Process. (ICASSP) , Calgary, AB, Canada, Apr. 2018, pp. 4779–4783
2018
Cited alongside, same era.
Z. Qin, H. Ye, G. Y. Li, and B.-H. F. Juang, “Deep learning in physical layer communications,” IEEE Wireless Commun. , vol. 26, no. 2, pp. 93–99, Apr. 2019
2019
Cited alongside, same era.
2021
Later among the works it cites.
Z. Weng and Z. Qin, “Semantic communication systems for speech transmission,” IEEE J. Sel. Areas Commun. , vol. 39, no. 8, pp. 2434–2444, Aug. 2021
2021
Later among the works it cites.
H. Tong, Z. Yang, S. Wang, Y. Hu, O. Semiari, W. Saad, and C. Yin, “Federated learning for audio semantic communication,” Frontiers Commun. and Netw. , vol. 2, Sep. 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
M. Jankowski, D. Gündüz, and K. Mikolajczyk, “Wireless image retrieval at the edge,” IEEE J. Sel. Areas Commun. , vol. 39, no. 1, pp. 89–100, Jan. 2021
2021
Later among the works it cites.
J. Shao, Y. Mao, and J. Zhang, “Learning task-oriented communication for edge inference: An information bottleneck approach,” IEEE J. Sel. Areas Commun. , vol. 40, no. 1, pp. 197–211, Nov. 2021
2021
Later among the works it cites.
W.-N. Hsu, Y.-H. H. Tsai, B. Bolte, R. Salakhutdinov, and A. Mohamed, “Hubert: How much can a bad teacher benefit asr pre-training?” in Proc. IEEE Int. Conf. Acoustics, Speech, Signal Process. (ICASSP) , Toronto, ON, Canada, Jun. 2021, pp. 6533–6537
2021
Later among the works it cites.
Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “Fastspeech 2: Fast and high-quality end-to-end text to speech,” in Proc. 9th Int. Conf. Learning Representations (ICLR) , Vienna, Austria, May 2021
2021
Later among the works it cites.
W. Tong and G. Y. Li, “Nine challenges in artificial intelligence and wireless communications for 6G,” IEEE Wireless Commun. , vol. 29, no. 4, pp. 140–145, May 2022
2022
Closest in time.
M. Sana and E. C. Strinati, “Learning semantics: An opportunity for effective 6G communications,” in Proc. IEEE Annual Consumer Commun. Netw. Conf. (CCNC) , Las Vegas, USA, Jan. 2022, pp. 631–636
2022
Closest in time.
2022
Closest in time.
D. Huang, F. Gao, X. Tao, Q. Du, and J. Lu, “Toward semantic communications: Deep learning-based image semantic coding,” IEEE J. Sel. Areas Commun. , vol. 41, no. 1, pp. 55–71, Nov. 2022
2022
Closest in time.
H. Xie, Z. Qin, X. Tao, and K. B. Letaief, “Task-oriented multi-user semantic communications,” IEEE J. Sel. Areas Commun. , vol. 40, no. 9, pp. 2584–2597, Jul. 2022
2022
Closest in time.
X. Kang, B. Song, J. Guo, Z. Qin, and F. R. Yu, “Task-oriented image transmission for scene classification in unmanned aerial systems,” IEEE Trans. Commun. , vol. 70, no. 8, pp. 5181–5192, Jun. 2022
2022
Closest in time.
2022
Closest in time.
P. Jiang, C.-K. Wen, S. Jin, and G. Y. Li, “Wireless semantic communications for video conferencing,” IEEE J. Sel. Areas Commun. , vol. 41, no. 1, pp. 230–244, Nov. 2023
2023
Closest in time.