Valid

GB/T 44089-2024Information technology - General technical requirements of full duplex speech interaction system (English PDF)

信息技术 全双工语音交互系统通用技术要求

Open the GB/T 44089-2024 preview as PDF

Preview — first pages of GB/T 44089-2024 (full document: 15 pages)

This is a limited preview

Buy now to download the full PDF (15 pages)

Issued by

SAMR; SAC

Level / Type

National · Recommended

Issue date

May 28, 2024

Implementation date

May 28, 2024

Scope

GB/T 44089-2024 is the English-translated version of 信息技术 全双工语音交互系统通用技术要求.

GB/T 44089-2024 sets the general technical requirements for full duplex speech interaction systems, the voice assistants that listen while they speak rather than waiting for the user to finish. Full duplex is what makes a spoken interface feel like a conversation, because the user can interrupt, correct and continue without a wake word each time, and it is technically much harder than half duplex: the system must cancel its own output from the microphone signal, decide continuously whether incoming speech is addressed to it, and manage a dialogue whose turns overlap. The standard sets the system architecture and its functional components and then the requirements: the audio front end including acoustic echo cancellation, noise suppression and the wake and interruption detection, the speech recognition and endpointing under barge-in, the dialogue management and turn taking, the response latency, the multi-round and context handling, the rejection of speech not addressed to the system, and the requirements on performance testing, reliability, privacy and security. It took effect on 28 May 2024.

Document preview — GB/T 44089-2024

National Standard of the People's Republic of China

ICS
35.240.01
Classification
L77

Issued by: State Administration for Market Regulation; Standardization Administration of the PRC

Contents

  • 1 Scope1
  • 2 Normative references1
  • 3 Terms and Definitions1
  • 4 Abbreviations2
  • 5 System reference functional framework and interaction process2
  • 5.1 System reference functional framework2
  • 5.2 System Interaction Process4
  • 6 Functional Requirements5
  • 6.1 Core Requirements5
  • 6.2 Acoustic Treatment Layer6
  • 6.3 Speech Recognition Layer6
  • 6.4 Dialogue Processing Layer6
  • 6.5 Speech Synthesis Layer6
  • 7 Performance Requirements6
  • 7.1 Speech Recognition Layer6
  • 7.2 Dialogue Processing Layer7
  • 7.3 Speech Synthesis Layer7
  • 10 Reference11

Foreword

This document is in accordance with the provisions of GB/T 1.1-2020 "Guidelines for standardization work Part

1.Structure and drafting rules for standardization documents" Drafting. This document was proposed and coordinated by the National Technical Committee for Information Technology Standardization (SAC/TC28). This document was drafted by: China Electronics Technology Standardization Institute, iFLYTEK Co., Ltd., Midea Group (Shanghai) Co., Ltd., Shenzhen Tencent Computer Systems Co., Ltd., Beijing Baidu Netcom Technology Co., Ltd., China Telecom Group Co., Ltd., Xiaomi Communications Technology Co., Ltd., China Mobile (Hangzhou) Information Technology Co., Ltd., Qingdao Haier Technology Co., Ltd., Fuzhou Data Technology Research Institute Co., Ltd., Shenzhen Intellifusion Technology Co., Ltd., Beijing Telecom Planning and Design Institute Co., Ltd., AISpeech Technology Co., Ltd., Hangzhou Fangde Intelligent Technology Co., Ltd., Antelope Industrial Internet Co., Ltd., Hefei Intelligent Voice Innovation Development Co., Ltd., Shenzhen Sihe Technology Co., Ltd. Company, Shanghai Intelligent Manufacturing Functional Platform Co., Ltd., Beijing Jitong Huasheng Technology Co., Ltd., and Mashang Consumer Finance Co., Ltd. The main drafters of this document are. Dong Jian, Xu Yang, Jia Yijun, Liu Ying, Song Wenlin, He Yongchun, Yu Lei, Sudan, Yuan Jie, E Lei, Cai Yasen, Mei Linhai, Zhao Pei, Liu Cong, Yang Zhen, Lei Zong, Gong Sheng, Fan Shuai, Hong Pengda, Huang Chao, Li Linlu, Fang Bin, Chen Ming, Hu Guoping, Yang Yifan, Liu Zhiqiang, Bi Shengnan, Ding Qiang, Gao Yu, and Li Xu. Information Technology Full-duplex Voice Interaction System General technical requirements

1 Scope

GB/T 44089-2024 sets the general technical requirements for full duplex speech interaction systems, the voice assistants that listen while they speak rather than waiting for the user to finish. Full duplex is what makes a spoken interface feel like a conversation, because the user can interrupt, correct and continue without a wake word each time, and it is technically much harder than half duplex: the system must cancel its own output from the microphone signal, decide continuously whether incoming speech is addressed to it, and manage a dialogue whose turns overlap. The standard sets the system architecture and its functional components and then the requirements: the audio front end including acoustic echo cancellation, noise suppression and the wake and interruption detection, the speech recognition and endpointing under barge-in, the dialogue management and turn taking, the response latency, the multi-round and context handling, the rejection of speech not addressed to the system, and the requirements on performance testing, reliability, privacy and security. It took effect on 28 May 2024.

This document specifies the reference functional framework, interaction process, functional requirements and performance requirements of a full-duplex voice interaction system. This document is applicable to the design, development, application, testing and maintenance of full-duplex voice interaction systems.

2 Normative references

This document has no normative references.

3 Terms and definitions

The following terms and definitions apply to this document.

3.1 Full-duplex A communication method capable of transferring data in both directions simultaneously.

3.2 functional unit A hardware or software entity that is capable of accomplishing a specific goal.

3.3 Speech recognition The process of converting human voice signals into text or instructions. [Source: GB/T 21023-2007, 3.1]

3.4 Enable the functional unit to understand the intention of human speech. [Source: GB/T 36464.1-2020, 3.11]

3.5 speech synthesisspeechsynthesis The process of synthesizing human language through mechanical and electronic methods. [Source: GB/T 21024-2007, 3.1]

3.6 Telephony The dialogue text content with certain logic used in the interaction process.

3.7 dialogue management Follow up the current conversation status and context input, update the conversation status, and generate the required actions based on the conversation processing logic. dialogue action.

......
This preview omits tables, figures, formulas and parts of the technical clauses. The complete document — 15 pages — is available in the English PDF.

Similar standards

How to Buy GB/T 44089-2024

  1. 1Add to cart. Click the "Buy GB/T 44089-2024" button on this page. You can add more standards before checkout.
  2. 2Checkout. Enter your email and billing details. Payment is processed securely by Stripe (cards, Apple Pay, Google Pay supported).
  3. 3Instant delivery (0–9 sec). Delivery is automatic: within seconds of payment you'll receive an email with a secure download link. The link stays valid for 72 hours.
  4. 4Invoice included. A tax invoice is attached to the confirmation email. Need a custom invoice? Contact us.

Related Standards

English PDF
15 pages
Instant delivery (0–9 sec)
Invoice included
View Cart

Secure payment via Stripe

Payments accepted

VisaMastercardAmerican ExpressApple PayGoogle PayStripe

GB/T 44089-2024

$275.00

$235.00for partners