Air-writing refers to virtually writing linguistic characters through hand gestures in three-dimensional space with six degrees of freedom. This paper proposes a generic video camera-aided convolutional neural network (CNN) based air-writing framework. Gestures are performed using a marker of fixed color in front of a generic video camera, followed by color-based segmentation to identify the marker and track the trajectory of the marker tip. A pre-trained CNN is then used to classify the gesture. The recognition accuracy is further improved using transfer learning with the newly acquired data. The performance of the system varies significantly on the illumination condition due to color-based segmentation. In a less fluctuating illumination condition, the system is able to recognize isolated unistroke numerals of multiple languages. The proposed framework has achieved 97.7%, 95.4% and 93.7% recognition rates in person independent evaluations on English, Bengali and Devanagari numerals, respectively.
翻译:空中手写是指通过具有六自由度的手势在三维空间中虚拟书写语言字符。本文提出了一种基于通用摄像头的卷积神经网络(CNN)空中手写框架。手势通过固定颜色的标记在通用摄像头前执行,随后采用基于颜色的分割方法识别标记并跟踪标记尖端的运动轨迹,进而利用预训练的CNN对手势进行分类。通过迁移学习结合新获取的数据,可进一步提升识别精度。由于基于颜色的分割方法,系统性能随光照条件变化显著。在光照波动较小的环境下,该系统能够识别多种语言的孤立单笔划数字。所提出的框架在英语、孟加拉语和天城文数字的独立于人评估中分别实现了97.7%、95.4%和93.7%的识别率。