GB/T 47517-2026Information technology - Technical specification for sign language digital humans (English PDF)
信息技术 手语数字人技术规范
Open the GB/T 47517-2026 preview as PDF
This is a limited preview
Buy now to download the full PDF (37 pages)
Issued by
SAMR; SAC
Level / Type
National · Recommended
Issue date
April 30, 2026
Implementation date
November 1, 2026
Scope
GB/T 47517-2026 is the English-translated version of 信息技术 手语数字人技术规范.
GB/T 47517-2026 is the Chinese national standard covering the animated signer that renders text or speech into Chinese Sign Language - the hand shapes, movements and facial expressions the avatar must produce, the linguistic accuracy required and the evaluation by deaf users rather than by hearing developers. First edition, in force since 1 November 2026. It was issued on 30 April 2026 and takes effect on 1 November 2026, as a first edition. The document is under the responsibility of the Standardization Administration of China. This page is published from the official record of the 2026 edition; the clause text of a standard this recent is not yet in circulation, and the figures, limits and tables it contains are those of the document itself, delivered in full with the English translation.
Document preview — GB/T 47517-2026
National Standard of the People's Republic of China
- ICS
- 35.240
- Classification
- L 70
Issued by: State Administration for Market Regulation; Standardization Administration of the PRC
Contents
- 6 Technical Requirements
- 6.1 Input Module
- 6.1.3 Voice Input
- 6.1.4 Image Input
- 6.2 Preprocessing Module
- 6.2.1 Speech Recognition
- 6.3 Sign Language-Driven Expression
- 6.4 Character Generation
- 7 Test Methods
- 7.2 Input Module
- 7.2.3 Voice Input
- 7.3 Preprocessing Module
- 7.3.1 Speech Recognition
- 7.4 Character Portrayal
- 7.5 Sign Language-Driven Expression
- 8 Unicode Transformation Format-8-bit (8-bit variable length) WAV. Waveform Audio
Foreword
GB/T 47517-2026 | Information technology - Technicalspecifications for sign language digital human
GB/T 47517-2026 English version. Information technology - Technical specifications for sign language digital human National Standards of the People's Republic of China ICS 35.240CCS L
70 Information Technology Sign Language Digital Human Technical Specifications Released on April 30, 2026 Implemented on 2026-11-
01 State Administration for Market Regulation The State Administration for Standardization issued a statement.
1.Scope This document specifies the technical framework and requirements for sign language digital humans, and describes the corresponding testing methods. This document applies to the development and testing of sign language digital humans.
4.Abbreviations The following abbreviations apply to this document. AAC. Advanced Audio Coding ASR. Automatic Speech Recognition BMP. Bitmap FLAC. Free Lossless Audio Codec JPEG. Joint Photographic Experts Group MP3.MPEG Audio Layer III OCR. Optical Character Recognition PNG. Portable Network Graphics UAC. USB Audio Class UTF-
6.1 Input Module
6.1.1 Supported Forms A sign language digital human should at least support text input, and preferably also support voice input via uploaded audio files, voice input for natural dialogue, and image input. The sign language digital human should support interaction in Simplified Chinese and Mandarin, and can support other languages and scripts on top of that.
Note. Unless otherwise specified, the technical requirements and test methods in this document are in Simplified Chinese and Mandarin.
6.1.2 Text Input The requirements are as follows:
a) It should at least support UTF-8 encoding format;
b) The maximum input length of supported text should be no less than.200 characters.
6.1.3 Voice Input
6.1.3.1 Voice Input for Uploaded Audio The requirements are as follows:
a) It should support at least the mainstream audio formats such as WAV, MP3, AAC, and FLAC;
b) It should support audio sampling rates in the range of at least 8 kHz to 48 kHz;
c) It should support at least mono audio;
d) The maximum supported audio input duration should be no less than 120 seconds.
6.1.3.2 Voice Input for Natural Dialogue The requirements are as follows:
a) Audio capture via UAC standard microphones (built-in or external) should be supported.
b) At least one of the following voice activation mechanisms should be supported. --Starts after the user clicks the "Start" button; --Silence detection activated; --Hot words awaken.
c) At least 120 seconds of continuous voice input should be supported.
6.1.4 Image Input
a) It should support at least the mainstream image formats such as BMP, JPEG, and PNG;
b) It should support image resolutions of at least 800 px × 600 px;
c) The maximum supported single file size should be no less than 5 MB.
6.2.1 Speech Recognition
6.2.1.1 Speech Recognition of Uploaded Audio If the sign language digitizer supports uploading audio input, the log should record the start time of the sign language digitizer's speech recognition and the completion time of the recognition. The time point of formation and the identification results. The speech recognition performance requirements for uploaded audio are as follows:
a) The real-time factor should be less than or equal to 1.0;
b) The accuracy rate should be no less than 85%.
6.2.1.2 Speech Recognition for Natural Dialogue If the sign language digital human supports natural dialogue input, then the time when the first audio frame arrives at the sign language digital human and the time when the audio frame ends should be recorded in the log. The time point when the frame arrives at the sign language digital human, the time point when speech recognition ends, and the recognition result. The speech recognition performance requirements for natural dialogue are as follows:
a) The real-time performance factor should not exceed 1.2;
b) The accuracy rate should be no less than 85%.
6.2.2 Optical Character Recognition If the sign language digital human supports image input with text, the time it takes for the image data to fully arrive at the sign language digital human should be recorded in the log. The time point of output for dot and optical character recognition results, and the recognition results. The performance requirements for optical character recognition are as follows:
a) The response efficiency ratio should not exceed 0.8;
b) The accuracy rate should be no less than 85%.
6.3 Sign Language-Driven Expression
6.3.1 Standardization of Sign Language Movements Each sign language gesture should be approximated to the level of accuracy required by Chapter 7 of GF 0020-2018.Fraser distance and normalization should be used. The dynamic time-normalized distance measurement, after optimization, is used to assess the standardization of sign language movements, with the following requirements.
a) The average Fraser distance for each sign language gesture, body posture, and facial expression should be less than 20;
b) The normalized dynamic time warp distance for each action should be less than 0.5.
6.3.2 Sign Language Comprehension The degree to which written translations into sign language can be understood by deaf people. Subjective evaluation was used, with the assessment team's average comprehension score and more than half of the test results being considered. The evaluators' comprehension score should be no less than 70 points.
6.3.3 Real-time performance When the input is less than or equal to.200 characters, the sign language digital human should start up with a delay of less than 3 seconds.
6.3.4 Speed Control It supports different sign language expression speeds, with at least three speed levels. fast, medium, and slow.
6.3.5 Video Stream Performance The sign language video stream should support at least 25 frames per second. In a 1280px × 720px scene, the network bandwidth requirement should not exceed
3 Mbit/s. In a 1920px × 1080px scenario, the network bandwidth requirement is no higher than
6.4 Character Generation
6.4.1 Face The facial requirements for sign language digital humans are as follows:
a) Facial contours should be clear and sharp, without jagged edges, blurriness, distortion, or missing parts;
b) It should have a standard facial structure, with harmonious proportions, and avoid facial distortion;
c) Basic facial expressions and mouth shape changes should be supported to complement sign language semantic expression;
d) There should be no missing, damaged, or obvious seams in the facial surface texture;
e) Facial expressions should be natural and smooth, avoiding pauses, misalignments, or stiffness.
6.4.2 Hands The hand gestures required for a sign language digital human are as follows:
a) The hand model should fully represent the structure of the five fingers, including the palm, the web between the thumb and index finger, and the five fingers. Each finger should have three joint segments to ensure... The structure is complete;
b) The hand model should support natural and accurate finger movements, including bending, extension, adduction, abduction, and rotation required by sign language. Common movements, clear trajectory, coordinated hand movements, accurate movements and positions, and no bone misalignment or twisting during the movement. Curves, clipping, or sudden changes in movement;
c) The surface of the hand should have sufficient visual detail, including nail shape, skin texture, fingertip creases, palm lines, etc.
d) The hands should always have clear edges and sharp contours, and should not have any defects such as blurriness, jaggedness, or broken surfaces;
e) The left and right hand models should have a high degree of consistency in terms of anatomical structure, scale, movement characteristics and texture representation.
6.4.3 Limbs The physical requirements for a sign language digital human are as follows:
a) It should have a realistic image that conforms to the true proportions of humans, and the overall modeling should be natural and harmonious;
7 Test Methods
7.1 Test Environment Setup The computer-side test environment settings are shown in Table 1, and the mobile-side test environment settings are shown in Table 2.
7.2 Input Module
7.2.1 Supported Formats Supported formal testing methods are as follows: Provide text input, audio file input, natural speech input, and image with text as input, respectively. Input the data and check its output; it meets the requirements of 6.1.1.
7.2.2 Text Input The text input test method is as follows:
a) Create a UTF-8 formatted text file and enter.200 Chinese characters (approximately 100 characters);
b) Input the file into the sign language digital human. If all text can be correctly recognized without garbled characters, missing words, or errors, and the digital sign language human can generate complete sign language videos, then it is in accordance with... It meets the requirements of 6.1.2.
7.2.3 Voice Input
7.2.3.1 Voice Input for Uploaded Audio The method for testing voice input when uploading audio is as follows:
a) Prepare audio files in four formats. WAV, MP3, AAC, and FLAC. For each format, prepare 8 kHz, 16 kHz, and... One file each of the four sampling rates (24 kHz, 48 kHz), for a total of 16 files. Each audio file is mono and 120 seconds in length. Left and right. The content is in clear Chinese audio.
b) Upload the 16 documents to the system sequentially for speech recognition. If the speech recognition result can be obtained correctly, it meets the requirements of 6.1.3.1.
7.2.3.2 Voice Input for Natural Dialogue The method for testing voice input in natural dialogue is as follows:
a) Prepare a microphone that conforms to the UAC standard and ensure that the system recognizes the microphone.
b) Observe whether the system starts by using the following methods in sequence. --Speak after clicking the "Start" button; --Speak directly; --Refer to the system manual and say wake-up words such as "Hello".
c) Observe the system after it starts up and speak continuously for 120 seconds. If the system starts under any of the conditions described in
b) and can correctly obtain the speech recognition result, it meets the requirements of 6.1.3.2.
7.2.4 Image Input The method for testing image input is as follows:
a) Prepare image files in three formats. BMP, JPEG, and PNG. For each format, prepare image files with resolutions of 800 pixels × 600 pixels. One image file each with dimensions of 1280px×720px and 1920px×1080px, for a total of 9 files. At least one image must be included. The size must be no less than 5 MB;
b) Upload the 9 documents to the system sequentially for optical character recognition. If the optical character recognition result can be obtained correctly, it meets the requirements of 6.1.4.
7.3.1 Speech Recognition
7.3.1.1 Speech Recognition for Uploaded Audio The method for testing the speech recognition of uploaded audio is as follows:
a) Prepare at least 20 audio tracks, with audio durations ranging from 5 seconds to 120 seconds.
b) Upload the audio file to the sign language digital human system and start recognition.
c) Collect the start time, completion time, and recognition result of speech recognition from the system log.
d) Calculate the real-time factor for each audio track using formula (1).
7.3.1.2 Speech Recognition for Natural Dialogue The speech recognition testing method for natural dialogue is as follows:
a) Select at least 5 evaluators, covering different age groups and genders.
b) Each evaluator speaks 5 sentences to the sign language digital human, with the duration of each sentence ranging from 5 seconds to 120 seconds.
c) Collect the start time, completion time, and recognition result of speech recognition from the system log.
d) Calculate the real-time coefficient using formula (3).
7.3.2 Optical Character Recognition The optical character recognition test method is as follows:
a) Prepare at least 20 images, covering the range of 16 to.200 characters.
b) Upload the image file to the sign language digital human system and start recognition.
c) Collect the start time, completion time, and recognition results of optical character recognition from the system log.
7.4 Character Portrayal
7.4.1 Face Facial assessment method. By reviewing the product manual, technical documents, and other materials of the sign language digital human, facial descriptions of the sign language digital human are obtained. The description and observation of the faces of the people in the generated video meet the requirements of 6.3.1.
7.4.2 Hands Hand assessment method. By reviewing the product manual, technical documents, and other materials of the sign language digital human, hand descriptions of the sign language digital human are obtained. The description and observation of the hands of the characters in the generated video meet the requirements of 6.3.2.
7.4.3 Limbs Basic requirements assessment method. By reviewing the product manual, technical documents, and other materials of the sign language digital human, the body language of the sign language digital human is obtained. Describe the figures visually and observe their limbs in the generated video, in accordance with the requirements of 6.3.3.
7.4.4 Clothing and Environment Clothing assessment method. Obtain the clothing description of the sign language digital human by reviewing its product manual, technical documents, and other materials. The description and observation of the clothing of the characters in the generated video meet the requirements of 6.3.4.
7.5 Sign Language-Driven Expression
7.5.1 Standardization of Sign Language Movements The method for testing the standardization of sign language movements is as follows:
a) Use the vocabulary from Chapter 7 of GF 0020-2018 as test items, requiring the sign language digital human to be evaluated to complete the corresponding sign language actions one by one. Verbal actions.
b) For a single action, extract the keypoint coordinates from each frame of animation, with the specific calibration positions as specified in Appendix A, to form a time sequence. Column data.
c) For a single action, use formula (5) to center the key points of each frame.
8 Unicode Transformation Format-8-bit (8-bit variable length) WAV. Waveform Audio
5.Technical Framework The technical framework for sign language digital humans is shown in Figure 1. Figure
1.Sign Language Digital Human Technology Framework It mainly includes the following contents.
a) Input module. Users input the text of the sentence to be translated, or upload an audio file containing natural language, or input the text via natural language. State what needs to be translated, or enter an image containing the text to be translated.
b) Preprocessing module. If the input is not text, the sign language digital human uses ASR and OCR technologies to convert the content to be translated into text. text.
c) Sign language driven expression. The sign language digital human translates text into sign language expressions through a series of processes.
d) Character generation. The sign language digital human generation presents the user with an image including face, hands, limbs, clothing, and environment.
......
This preview omits tables, figures, formulas and parts of the technical clauses. The complete document — 37 pages — is available in the English PDF.
How to Buy GB/T 47517-2026
- 1Add to cart. Click the "Buy GB/T 47517-2026" button on this page. You can add more standards before checkout.
- 2Checkout. Enter your email and billing details. Payment is processed securely by Stripe (cards, Apple Pay, Google Pay supported).
- 3Instant delivery (0–9 sec). Delivery is automatic: within seconds of payment you'll receive an email with a secure download link. The link stays valid for 72 hours.
- 4Invoice included. A tax invoice is attached to the confirmation email. Need a custom invoice? Contact us.
Related Standards
GB/T 47310-2026 — Determination of total silicon, aluminium, iron, potassium, sodium, calcium, magnesium, manganese, phosphorus, titanium and sulfur in soil - Monochromatic excitation energy dispersive X-ray fluorescence spectrometry
GB/T 47321-2026 — Specification for the warning data exchange of the national emergency early warning dissemination system
GB/T 47293-2026 — Determination of available mercury in soil
Secure payment via Stripe
Payments accepted
GB/T 47517-2026
$365.00