GB/T 44089-2024Information technology - General technical requirements of full duplex speech interaction system (English PDF)
信息技术 全双工语音交互系统通用技术要求
Open the GB/T 44089-2024 preview as PDF
This is a limited preview
Buy now to download the full PDF (15 pages)
Issued by
SAMR; SAC
Level / Type
National · Recommended
Issue date
May 28, 2024
Implementation date
May 28, 2024
Scope
GB/T 44089-2024 is the English-translated version of 信息技术 全双工语音交互系统通用技术要求.
GB/T 44089-2024 sets the general technical requirements for full duplex speech interaction systems, the voice assistants that listen while they speak rather than waiting for the user to finish. Full duplex is what makes a spoken interface feel like a conversation, because the user can interrupt, correct and continue without a wake word each time, and it is technically much harder than half duplex: the system must cancel its own output from the microphone signal, decide continuously whether incoming speech is addressed to it, and manage a dialogue whose turns overlap. The standard sets the system architecture and its functional components and then the requirements: the audio front end including acoustic echo cancellation, noise suppression and the wake and interruption detection, the speech recognition and endpointing under barge-in, the dialogue management and turn taking, the response latency, the multi-round and context handling, the rejection of speech not addressed to the system, and the requirements on performance testing, reliability, privacy and security. It took effect on 28 May 2024.
Document preview — GB/T 44089-2024
National Standard of the People's Republic of China
- ICS
- 35.240.01
- Classification
- L77
Issued by: State Administration for Market Regulation; Standardization Administration of the PRC
Contents
- 1 Scope1
- 2 Normative references1
- 3 Terms and Definitions1
- 4 Abbreviations2
- 5 System reference functional framework and interaction process2
- 5.1 System reference functional framework2
- 5.2 System Interaction Process4
- 6 Functional Requirements5
- 6.1 Core Requirements5
- 6.2 Acoustic Treatment Layer6
- 6.3 Speech Recognition Layer6
- 6.4 Dialogue Processing Layer6
- 6.5 Speech Synthesis Layer6
- 7 Performance Requirements6
- 7.1 Speech Recognition Layer6
- 7.2 Dialogue Processing Layer7
- 7.3 Speech Synthesis Layer7
- 10 Reference11
Foreword
This document is in accordance with the provisions of GB/T 1.1-2020 "Guidelines for standardization work Part
1.Structure and drafting rules for standardization documents" Drafting. This document was proposed and coordinated by the National Technical Committee for Information Technology Standardization (SAC/TC28). This document was drafted by: China Electronics Technology Standardization Institute, iFLYTEK Co., Ltd., Midea Group (Shanghai) Co., Ltd., Shenzhen Tencent Computer Systems Co., Ltd., Beijing Baidu Netcom Technology Co., Ltd., China Telecom Group Co., Ltd., Xiaomi Communications Technology Co., Ltd., China Mobile (Hangzhou) Information Technology Co., Ltd., Qingdao Haier Technology Co., Ltd., Fuzhou Data Technology Research Institute Co., Ltd., Shenzhen Intellifusion Technology Co., Ltd., Beijing Telecom Planning and Design Institute Co., Ltd., AISpeech Technology Co., Ltd., Hangzhou Fangde Intelligent Technology Co., Ltd., Antelope Industrial Internet Co., Ltd., Hefei Intelligent Voice Innovation Development Co., Ltd., Shenzhen Sihe Technology Co., Ltd. Company, Shanghai Intelligent Manufacturing Functional Platform Co., Ltd., Beijing Jitong Huasheng Technology Co., Ltd., and Mashang Consumer Finance Co., Ltd. The main drafters of this document are. Dong Jian, Xu Yang, Jia Yijun, Liu Ying, Song Wenlin, He Yongchun, Yu Lei, Sudan, Yuan Jie, E Lei, Cai Yasen, Mei Linhai, Zhao Pei, Liu Cong, Yang Zhen, Lei Zong, Gong Sheng, Fan Shuai, Hong Pengda, Huang Chao, Li Linlu, Fang Bin, Chen Ming, Hu Guoping, Yang Yifan, Liu Zhiqiang, Bi Shengnan, Ding Qiang, Gao Yu, and Li Xu. Information Technology Full-duplex Voice Interaction System General technical requirements
1 Scope
GB/T 44089-2024 sets the general technical requirements for full duplex speech interaction systems, the voice assistants that listen while they speak rather than waiting for the user to finish. Full duplex is what makes a spoken interface feel like a conversation, because the user can interrupt, correct and continue without a wake word each time, and it is technically much harder than half duplex: the system must cancel its own output from the microphone signal, decide continuously whether incoming speech is addressed to it, and manage a dialogue whose turns overlap. The standard sets the system architecture and its functional components and then the requirements: the audio front end including acoustic echo cancellation, noise suppression and the wake and interruption detection, the speech recognition and endpointing under barge-in, the dialogue management and turn taking, the response latency, the multi-round and context handling, the rejection of speech not addressed to the system, and the requirements on performance testing, reliability, privacy and security. It took effect on 28 May 2024.
This document specifies the reference functional framework, interaction process, functional requirements and performance requirements of a full-duplex voice interaction system. This document is applicable to the design, development, application, testing and maintenance of full-duplex voice interaction systems.
2 Normative references
This document has no normative references.
3 Terms and definitions
The following terms and definitions apply to this document.
3.1 Full-duplex A communication method capable of transferring data in both directions simultaneously.
3.2 functional unit A hardware or software entity that is capable of accomplishing a specific goal.
3.3 Speech recognition The process of converting human voice signals into text or instructions. [Source: GB/T 21023-2007, 3.1]
3.4 Enable the functional unit to understand the intention of human speech. [Source: GB/T 36464.1-2020, 3.11]
3.5 speech synthesisspeechsynthesis The process of synthesizing human language through mechanical and electronic methods. [Source: GB/T 21024-2007, 3.1]
3.6 Telephony The dialogue text content with certain logic used in the interaction process.
3.7 dialogue management Follow up the current conversation status and context input, update the conversation status, and generate the required actions based on the conversation processing logic. dialogue action.
......
This preview omits tables, figures, formulas and parts of the technical clauses. The complete document — 15 pages — is available in the English PDF.
Similar standards
GB 38031-2025|GB/T44089-2024|GB/T 1.1-2020|GB/T 21023-2007|GB/T 36464.1-2020|GB/T 21024-2007|GB/T 42450|GB/T 39788
How to Buy GB/T 44089-2024
- 1Add to cart. Click the "Buy GB/T 44089-2024" button on this page. You can add more standards before checkout.
- 2Checkout. Enter your email and billing details. Payment is processed securely by Stripe (cards, Apple Pay, Google Pay supported).
- 3Instant delivery (0–9 sec). Delivery is automatic: within seconds of payment you'll receive an email with a secure download link. The link stays valid for 72 hours.
- 4Invoice included. A tax invoice is attached to the confirmation email. Need a custom invoice? Contact us.
Related Standards
GB/T 47310-2026 — Determination of total silicon, aluminium, iron, potassium, sodium, calcium, magnesium, manganese, phosphorus, titanium and sulfur in soil - Monochromatic excitation energy dispersive X-ray fluorescence spectrometry
GB/T 47321-2026 — Specification for the warning data exchange of the national emergency early warning dissemination system
GB/T 47293-2026 — Determination of available mercury in soil
Secure payment via Stripe
Payments accepted
GB/T 44089-2024
$275.00