Fetching the paper…
Reading the bibliography…
In this paper, we propose methods to build a powerful and efficient Image-to-Speech captioning (Im2Sp) model.
“Bleu: a method for automatic evaluation of machine translation”
Kishore Papineni et al · 2002
Earlier work this paper cites.
“Rouge: A package for automatic evaluation of summaries”
Chin-Yew Lin · 2004
Earlier work this paper cites.
“Im2text: Describing images using 1 million captioned photographs”
Vicente Ordonez, Girish Kulkarni and Tamara Berg · 2011
Earlier work this paper cites.
“Framing image description as a ranking task: Data, models and evaluation metrics”
Micah Hodosh, Peter Young and Julia Hockenmaier · 2013
Earlier work this paper cites.
“Microsoft coco: Common objects in context”
Tsung-Yi Lin et al · 2014
Earlier work this paper cites.
“Neural variational inference and learning in belief networks”
Andriy Mnih and Karol Gregor · 2014
Earlier work this paper cites.
“Meteor universal: Language specific translation evaluation for any target language”
Michael Denkowski and Alon Lavie · 2014
Earlier work this paper cites.
“Deep multimodal semantic embeddings for speech and images”
David Harwath and James Glass · 2015
Earlier work this paper cites.
“Deep visual-semantic alignments for generating image descriptions”
Andrej Karpathy and Li Fei-Fei · 2015
Earlier work this paper cites.
“Cider: Consensus-based image description evaluation”
Ramakrishna Vedantam, C Lawrence and Devi Parikh · 2015
Earlier work this paper cites.
“Spice: Semantic propositional image caption evaluation”
Peter Anderson et al · 2016
Earlier work this paper cites.
“Image2speech: Automatically generating audio descriptions of images”
Mark Hasegawa-Johnson et al · 2017
Earlier work this paper cites.
“Neural discrete representation learning”
Aaron Van and Oriol Vinyals · 2017
Earlier work this paper cites.
“Representations of language in a model of visually grounded speech signal”
Grzegorz Chrupała, Lieke Gelderloos and Afra Alishahi · 2017
Earlier work this paper cites.
“Attention is all you need”
Ashish Vaswani et al · 2017
Earlier work this paper cites.
“The LJ Speech Dataset”, https://keithito.com/LJ-Speech-Dataset/ , 2017
Keith Ito and Linda Johnson · 2017
Earlier work this paper cites.
“Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning”
Piyush Sharma et al · 2018
Earlier work this paper cites.
“MOSNet: Deep Learning-Based Objective Assessment for Voice Conversion”
Chen-Chou Lo et al · 2019
Cited alongside, same era.
“Discretalk: Text-to-speech as a machine translation problem”
Tomoki Hayashi and Shinji Watanabe · 2020
Cited alongside, same era.
“Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis”
Jungil Kong, Jaehyeon Kim and Jaekyoung Bae · 2020
Cited alongside, same era.
“An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale”
Alexey Dosovitskiy et al · 2020
Cited alongside, same era.
“wav2vec 2.0: A framework for self-supervised learning of speech representations”
Alexei Baevski et al · 2020
Cited alongside, same era.
“Synthesizing spoken descriptions of images”
“A survey on vision transformer”
Kai Han et al · 2022
Later among the works it cites.
“Textless Speech-to-Speech Translation on Real Data”
Ann Lee et al · 2022
Later among the works it cites.
“Findings of the IWSLT 2022 Evaluation Campaign.”
Anastasopoulos Antonios et al · 2022
Later among the works it cites.
“Lip-to-speech synthesis in the wild with multi-task learning”
Minsu Kim, Joanna Hong and Yong Ro · 2023
Closest in time.
“UnitY: Two-pass Direct Speech-to-speech Translation with Discrete Units”
Hirofumi Inaguma et al · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xinsheng Wang et al · 2021
Cited alongside, same era.
“On generative spoken language modeling from raw audio”
Kushal Lakhotia et al · 2021
Cited alongside, same era.
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units”
Wei-Ning Hsu et al · 2021
Cited alongside, same era.
“Learning transferable visual models from natural language supervision”
Alec Radford et al · 2021
Cited alongside, same era.
“Taming transformers for high-resolution image synthesis”
Patrick Esser, Robin Rombach and Bjorn Ommer · 2021
Cited alongside, same era.
“Vector-quantized Image Modeling with Improved VQGAN”
Jiahui Yu et al · 2021
Cited alongside, same era.
“Text-Free Image-to-Speech Synthesis Using Learned Segmental Units”
Wei-Ning Hsu et al · 2021
Cited alongside, same era.
Minsu Kim et al · 2023
Closest in time.
“Analysing discrete self supervised speech representation for spoken language modeling”
Amitay Sicherman and Yossi Adi · 2023
Closest in time.
“Generative spoken dialogue language modeling”
Tu Nguyen et al · 2023
Closest in time.
“Intelligible Lip-to-Speech Synthesis with Speech Units”
Jeongsoo Choi, Minsu Kim and Yong Ro · 2023
Closest in time.
“Exploration of Efficient End-to-End ASR using Discretized Input from Self-Supervised Learning”
Xuankai Chang et al · 2023
Closest in time.
“Lip Reading for Low-resource Languages by Learning and Combining General Speech Knowledge and Language-specific Knowledge”
Minsu Kim et al · 2023
Closest in time.
“Speechclip: Integrating speech with pre-trained vision and language model”
Yi-Jen Shih et al · 2023
Closest in time.
“Watch or Listen: Robust Audio-Visual Speech Recognition with Visual Corruption Modeling and Reliability Scoring”
Joanna Hong et al · 2023
Closest in time.
“SpeechLMScore: Evaluating speech generation using speech language model”
Soumi Maiti et al · 2023
Closest in time.
“SeiT: Storage-Efficient Vision Training with Tokens Using 1% of Pixel Storage”
Song Park et al · 2023
Closest in time.
“Show, attend and tell: Neural image caption generation with visual attention”
Kelvin Xu et al · 2057
Closest in time.