Fetching the paper…
Reading the bibliography…
Significant progress has been made recently on challenging tasks in automatic sign language understanding, such as sign language recognition, translation and production.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Continuous sign language recognition: Towards large vocabulary statistical recognition systems handling multiple signers
Oscar Koller, Jens Forster, and Hermann Ney · 2015
Earlier work this paper cites.
Attention-based bidirectional long short-term memory networks for relation classification
Peng Zhou, Wei Shi, Jun Tian, Zhenyu Qi, Bingchen Li, Hongwei Hao, and Bo Xu · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Neural sign language translation
Necati Cihan Camgoz, Simon Hadfield, Oscar Koller, Hermann Ney, and Richard Bowden · 2018
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson · 2018
Earlier work this paper cites.
How2: a large-scale dataset for multimodal language understanding
Ramon Sanabria, Ozan Caglayan, Shruti Palaskar, Desmond Elliott, Loïc Barrault, Lucia Specia, and Florian Metze · 2018
Earlier work this paper cites.
Openpose: Realtime multi-person 2d pose estimation using part affinity fields
Z. Cao, G. Hidalgo Martinez, T. Simon, S. Wei, and Y. A. Sheikh · 2019
Earlier work this paper cites.
Neural sign language translation based on human keypoint estimation
Sang-Ki Ko, Chang Jo Kim, Hyedong Jung, and Choongsang Cho · 2019
Earlier work this paper cites.
Mediapipe: A framework for building perception pipelines
Camillo Lugaresi, Jiuqiang Tang, Hadon Nash, Chris McClanahan, Esha Uboweja, Michael Hays, Fan Zhang, Chuo-Ling Chang, Ming Guang Yong, Juhyun Lee, et al · 2019
Earlier work this paper cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli · 2019
Earlier work this paper cites.
Monocular total capture: Posing face, body, and hands in the wild
Donglai Xiang, Hanbyul Joo, and Yaser Sheikh · 2019
Earlier work this paper cites.
On the continuity of rotation representations in neural networks
Yi Zhou, Connelly Barnes, Jingwan Lu, Jimei Yang, and Hao Li · 2019
Earlier work this paper cites.
Bsl-1k: Scaling up co-articulated sign language recognition using mouthing cues
Samuel Albanie, Gül Varol, Liliane Momeni, Triantafyllos Afouras, Joon Son Chung, Neil Fox, and Andrew Zisserman · 2020
Cited alongside, same era.
Multi-channel transformers for multi-articulatory sign language translation
Necati Cihan Camgoz, Oscar Koller, Simon Hadfield, and Richard Bowden · 2020
Cited alongside, same era.
Sign language transformers: Joint end-to-end sign language recognition and translation
Necati Cihan Camgöz, Oscar Koller, Simon Hadfield, and Richard Bowden · 2020
Cited alongside, same era.
Word-level deep sign language recognition from video: A new large-scale dataset and methods comparison
Dongxu Li, Cristian Rodriguez, Xin Yu, and Hongdong Li · 2020
Cited alongside, same era.
Stochastic fine-grained labeling of multi-state sign glosses for continuous sign language recognition
Zhe Niu and Brian Mak · 2020
Cited alongside, same era.
Self-mutual distillation learning for continuous sign language recognition
Aiming Hao, Yuecong Min, and Xilin Chen · 2021
Later among the works it cites.
Skeleton aware multi-modal sign language recognition
Songyao Jiang, Bin Sun, Lichen Wang, Yue Bai, Kunpeng Li, and Yun Fu · 2021
Later among the works it cites.
Visual alignment constraint for continuous sign language recognition
Yuecong Min, Aiming Hao, Xiujuan Chai, and Xilin Chen · 2021
Later among the works it cites.
Deep learning–based text classification: a comprehensive review
Shervin Minaee, Nal Kalchbrenner, Erik Cambria, Narjes Nikzad, Meysam Chenaghlu, and Jianfeng Gao · 2021
Later among the works it cites.
Body2hands: Learning to infer 3d hands from conversational gesture body dynamics
Evonne Ng, Shiry Ginosar, Trevor Darrell, and Hanbyul Joo · 2021
Later among the works it cites.
Tokenlearner: What can 8 learned tokens do for images and videos?
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Continuous sign language recognition through cross-modal alignment of video and text embeddings in a joint-latent space
Ilias Papastratis, Kosmas Dimitropoulos, Dimitrios Konstantinidis, and Petros Daras · 2020
Cited alongside, same era.
Everybody sign now: Translating spoken language to photo realistic sign language video
Ben Saunders, Necati Cihan Camgoz, and Richard Bowden · 2020
Cited alongside, same era.
Autsl: A large scale multi-modal turkish sign language dataset and baseline methods
Ozge Mercanoglu Sincan and Hacer Yalim Keles · 2020
Cited alongside, same era.
Neural sign language synthesis: Words are our glosses
Jan Zelinka and Jakub Kanis · 2020
Cited alongside, same era.
Progress in neural nlp: Modeling, learning, and reasoning
Ming Zhou, Nan Duan, Shujie Liu, and Heung-Yeung Shum · 2020
Cited alongside, same era.
Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text
Hassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui, and Boqing Gong · 2021
Cited alongside, same era.
BOBSL: BBC-Oxford British Sign Language Dataset
Samuel Albanie, Gül Varol, Liliane Momeni, Hannah Bull, Triantafyllos Afouras, Himel Chowdhury, Neil Fox, Bencie Woll, Rob Cooper, Andrew McParland, and Andrew Zisserman · 2021
Cited alongside, same era.
Michael S. Ryoo, A. J. Piergiovanni, Anurag Arnab, Mostafa Dehghani, and Anelia Angelova · 2021
Later among the works it cites.
Continuous 3d multi-channel sign language production via progressive transformers and mixture density networks
Ben Saunders, Necati Cihan Camgoz, and Richard Bowden · 2021
Later among the works it cites.
Mixed signals: Sign language production via a mixture of motion primitives
Ben Saunders, Necati Cihan Camgoz, and Richard Bowden · 2021
Later among the works it cites.
Skeletal graph self-attention: Embedding a skeleton inductive bias into sign language production
Ben Saunders, Necati Cihan Camgoz, and Richard Bowden · 2021
Later among the works it cites.
UNIK: A unified framework for real-world skeleton-based action recognition
Di Yang, Yaohui Wang, Antitza Dantcheva, Lorenzo Garattoni, Gianpiero Francesca, and François Brémond · 2021
Later among the works it cites.
Sign language video retrieval with free-form textual queries
Amanda Duarte, Samuel Albanie, Xavier Giró-i Nieto, and Gül Varol · 2022
Closest in time.
Perceiver IO: A general architecture for structured inputs & outputs
Andrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac, Carl Doersch, Catalin Ionescu, David Ding, Skanda Koppula, Daniel Zoran, Andrew Brock, Evan Shelhamer, Olivier J. Hénaff, Matthew M. Botvinick, Andrew Zisserman, Oriol Vinyals, and João Carreira · 2022
Closest in time.
Ben Saunders, Necati Cihan Camgoz, and Richard Bowden · 2022
Closest in time.
Open-domain sign language translation learned from online video
Bowen Shi, Diane Brentari, Greg Shakhnarovich, and Karen Livescu · 2022
Closest in time.
Multiview transformers for video recognition
Shen Yan, Xuehan Xiong, Anurag Arnab, Zhichao Lu, Mi Zhang, Chen Sun, and Cordelia Schmid · 2022
Closest in time.
Coca: Contrastive captioners are image-text foundation models
Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu · 2022
Closest in time.