手语识别作为跨越聋听沟通障碍的关键技术,其发展脉络清晰地映射了传感器技术与人工智能算法的演进历程。这一领域的研究已走过约四十年的历程,相关的研究可以分为基于数据手套和基于计算机视觉两个方向。
基于数据手套和计算机视觉的手语识别研究,核心区别在于数据源与交互方式。前者依赖物理传感器(如弯曲传感器、IMU),通过穿戴设备精准测量手部的关节角度和空间位置,优点是精度高、不受光照遮挡影响,但设备昂贵且会束缚用户;后者依赖光学镜头,通过算法解析图像或视频中的像素信息来推断手势,优点是设备普及、使用自然,但易受环境光线、复杂背景和手部自遮挡的干扰。简而言之,数据手套侧重捕捉手部内在的物理状态,而计算机视觉侧重于感知外部的视觉表象。
基于计算机视觉的手语识别技术主要通过摄像机获得手语图像或视频,而后使用机器学习技术或深度学习技术来实现手语识别。随着近年来智能手机的普及,很大程度上克服了数据手套难以获得的缺点,提高了手语识别系统能够应用于日常生活的可能性。
[1]国务院. 关于印发“十四五”残疾人保障和发展规划的通知: 国发〔2021〕10号[Z]. 北京: 国务院, 2021.
[2]计纬国. 王妍. 手语翻译系统在医院就医场景中的应用研究[J]. 中国听力语言康复科学杂志, 2022,20(2): 141-144.
[3]王向东. 李久江. 基于5G的远程手语翻译系统在无障碍公共服务中的应用[J]. 电信科学, 2020,36(12): 112-118.
[4]许巧仙. 从沟通无障碍看听障人士的就业支持[J]. 中国特殊教育, 2019(6): 33-37.
[5]Fels, S. Sidney, Hinton, Geoffrey E. Glove-Talk: a neural network interface between a data-glove and a speech synthesizer[J]. Neural Networics IEEE Transactions on 1993,4(1):2-8.
[6]Lee C, Xu Y. Online, interactive learning of gestures for human/robot interfaces[C]// IEEE International Conference on Robotics and Automation, 1996. Proceedings.IEEE,2002:2982-2987.
[7]Takahashi T.Kishino F.Hand gesture coding based on experiments usi-Ng a hand gesturinterface device[J].Acm Sigchi Bulletin, 1991, 23(2):67-74.
[8]R.H. Liang and M. Ouhyoung, A real-time continuous gesture recognition system for sign language. The Third International Conference on Automatic Face and Gesture Recognition, Nara, Japan, 1998:558-565.
[9]李勇,高文,姚鸿勋,基于颜色手套的中国手指语字母的动静态识别[J].计算机工程与应用,2002, 38(17):55-58.
[10]张亚新,原魁,杨学良,一种用于手语识别的新型数据手套[J].工程科学学报,2001, 23(4):000379-381.
[11]Bauer B, Hienz H, Kraiss K F. Video-based continuous sign language recognition using statistical methods. Proceedings of Pattern Recognition, 2000. Proceedings. 15th International Conference on, volume 2.IEEE, 2000.463-466.
[12]Kelly D, Mcdonald J, Markham C. A person independent system for recognition of hand postures used in sign language[!]. Pattern Recognition Letters, 2010, 31(1):1359-1368.
[13]杨文文,基于 HMM 与 Level Building 的连续中国手语识别系统研究[D].2016.
[14]JI Y, KIM S, and LEE K B. Sign language learning system with image sampling and convolutional neural network[C].The lst IEEE International Conference on Robotic Computing (IRC), Taichung, China, 2017: 371 -375.
[15]PU J, ZHOU W, LI H. Iterative alignment network for continuous sign language recognition[C]//IEEE Conference on Computer^sion and Pattern Recognition.2019:4165-4174.
[16]PU J, ZHOU W, LI H. Dilated convolutional network with iterastive optimization for continuous sign language recognition.[C]//Intemational Joint Conferences on Artificial Intelligence:volume 3. 2018:7.
[17]HUANG Jie, ZHOU Wengang, ZHANG Qilin, et al. Video-based sign anguage recognition without temporal segmentation[C]. The 32nd AAAI Conference on Artificial Intelligence, New Orleans, USA, 2018: 2257 -2264.
[18]GUO Dan, ZHOU Weiung, LI Mouyiang, et al.Hierarchiceal LSIM for sign language translation[C]. The 32nd AAAI Conference on Arificial Intelligence, the 30h innovative Applications of Artificial Intelligence (IAA:.18), and the 8th AAAI Symposium on Educational Advances in Artificial Intelligence, New Orleans, USA, 2018:6845 - 6852.
[19]CUI R,LIU H, ZHANG C. Recurrent convolutional neural networks for continuous sign language recognition by saged oplimization[C].//IEEE Conference on Computer Vision und Pattern Recognition. 2017:7361-7369.