GB/T 47478.1-2026Public security - Voiceprint recognition applications - Part 1: Technical requirements for capture, recognition and database construction (English PDF)
公共安全 声纹识别应用 第1部分:采集识别建库技术要求
Open the GB/T 47478.1-2026 preview as PDF
This is a limited preview
Buy now to download the full PDF (24 pages)
Issued by
SAMR; SAC
Level / Type
National · Recommended
Issue date
April 30, 2026
Implementation date
November 1, 2026
Scope
GB/T 47478.1-2026 is the English-translated version of 公共安全 声纹识别应用 第1部分:采集识别建库技术要求.
GB/T 47478.1-2026 is the Chinese national standard covering voiceprint recognition in policing - how a voice sample is captured and what quality it must have, how the comparison is made and scored, and how the reference database is built and maintained. Part 1 of the series, first edition, in force since 1 November 2026, under the Ministry of Public Security. It was issued on 30 April 2026 and takes effect on 1 November 2026, as a first edition. This page is published from the official record of the 2026 edition; the clause text of a standard this recent is not yet in circulation, and the figures, limits and tables it contains are those of the document itself, delivered in full with the English translation.
Document preview — GB/T 47478.1-2026
National Standard of the People's Republic of China
- ICS
- 13.310
- Classification
- A 91
Issued by: State Administration for Market Regulation; Standardization Administration of the PRC
Contents
- 1 Scope
- 4 Source of voiceprint audio files
- 4.2 Requirements for Voiceprint Speech Sample Collection
- 5 Voiceprint data quality inspection and quality grading
- 5.1 General Requirements for Quality Inspection
- 5.3 Voiceprint Data Quality Classification
- 5.3.2 Level
- 5.3.3 Level
- 5.3.4 Level
- 6 Requirements for building a voiceprint audio file library
- 6.2 General Requirements for Database Construction
- 6.3 Voiceprint Algorithm Test Library
- 6.4 Voiceprint Identity Database
- 6.5 Voiceprint-Specific Application Library
Foreword
GB/T 47478.1-2026 | Public security - Voiceprint recognition applications - Part 1: Technical requirements for capture recognition and database construction
GB/T 47478.1-2026 English version. Public security - Voiceprint recognition applications - Part
1 Scope
GB/T 47478.1-2026 is the Chinese national standard covering voiceprint recognition in policing - how a voice sample is captured and what quality it must have, how the comparison is made and scored, and how the reference database is built and maintained. Part 1 of the series, first edition, in force since 1 November 2026, under the Ministry of Public Security. It was issued on 30 April 2026 and takes effect on 1 November 2026, as a first edition. This page is published from the official record of the 2026 edition; the clause text of a standard this recent is not yet in circulation, and the figures, limits and tables it contains are those of the document itself, delivered in full with the English translation.
91 National Standards of the People's Republic of China Public Security Voiceprint Recognition Application Part
1.Data Collection Identify database construction technical requirements Published on 2026-04-
30 Implemented on 2026-11-
01 State Administration for Market Regulation The State Administration for Standardization issued a statement.
1.Scope This document specifies the sources of voiceprint voice files, voiceprint data quality detection and quality classification, and requirements for voiceprint voice file database construction. This document applies to the acquisition of voiceprint files, voiceprint data quality inspection, construction of voiceprint file databases, and voiceprint recognition. application.
4 Source of voiceprint audio files
4.1 Source Classification Depending on how the voiceprint voice file is obtained, the sources of voiceprint voice files are generally divided into two categories.
a) Direct acquisition. After collecting voiceprint speech samples on-site or remotely using voice acquisition equipment, the voiceprint speech files are saved;
b) Indirect acquisition. After reading existing electronic data files and performing non-tampering operations such as truncation and merging, the voiceprint is re-saved to obtain the voiceprint. document.
4.2 Requirements for Voiceprint Speech Sample Collection
4.2.1 Factors to be considered during data collection When collecting voiceprint samples, the following factors should be considered and specific requirements should be put forward according to the needs of actual applications.
a) Data collection environment. environmental reverberation conditions, ambient noise levels, presence of other human voices, etc.
b) Acquisition equipment. factors such as frequency response, sensitivity, waveform distortion, directivity, and whether a microphone array is used;
c) Acquisition parameters. sampling rate, quantization accuracy, number of channels, etc.
d) Data collection constraints. Whether there is supervision at the data collection site, whether the data subject cooperates, and the data subject's speaking style (e.g., reading aloud, speaking casually, etc.). Factors such as dialogue;
e) Content to be collected. Factors such as whether prompts or specified voice content are used, and the required effective voice duration to be collected;
f) Other specific factors required by the application. gender, age, language or dialect, accent, mood, speaking speed, volume, etc. of the person being collected.
4.2.2 Requirements for Voiceprint Identity Database Collection The voiceprint files for the voiceprint identity database should be obtained directly. Guidance should be provided to the person being collected before voiceprint sample collection. Understand the precautions for voiceprint and speech sample collection. The following should be considered regarding the collection site environment, personnel deployment, collection method, and collection equipment. Require.
a) Verify the identity information of the person being collected before data collection;
b) The data collection site was quiet with no significant noise;
c) The voice acquisition device has no obvious echo, electrical noise, or howling sound;
d) Use a sampling frequency of 8000Hz or 16000Hz, with a quantization precision of 16 bits;
e) Point the effective acquisition direction of the voice acquisition device towards the person being acquired, and maintain an appropriate distance between the voice acquisition device and the person being acquired. When stationary, disable the automatic noise reduction function of the acquisition device and software;
f) Do not collect speech when the subject is in a state of obvious physical discomfort;
5.1 General Requirements for Quality Inspection
5.1.1 Based on the acquisition parameters (including sampling rate, quantization accuracy, number of channels, etc.) during the acquisition of voiceprint speech samples, the voiceprint speech can be directly analyzed. Speech quality is tested using audio samples.
5.1.2 For voiceprint-based speech files, the audio data should be read according to the corresponding audio file format, and then the audio data should be processed for voice quality control. Quantity detection.
5.1.3 During the voice quality detection process, the voiceprint data being detected should not be altered.
5.2 Voiceprint Speech Quality Testing Content Voiceprint speech quality testing should include at least the following.
a) Upper/lower cutoff frequencies. Calculate the upper/lower cutoff frequencies of the speech data based on the sampling rate and sampled values. Note
1.The upper/lower cutoff frequencies are the high/low frequency values located 3dB below the average energy of the passband.
b) Amplitude and waveform distortion. Determine whether there are amplitude and waveform distortion problems based on the changes in the sampled values.
c) Spraying phenomenon.
Note 2."Pop-up" refers to the noise phenomenon caused by the airflow being too strong and directly hitting the voice acquisition device during the voice acquisition process.
d) Speech mixing phenomenon. Note
3.Speech mixing refers to the phenomenon of multiple voices in a single audio clip. For example, multiple people speaking at different times, or multiple people speaking simultaneously at the same time.
e) Sudden noise phenomenon. Note
4.Sudden noise refers to noise that is not controlled by the speaker, and that occurs suddenly from the acquisition equipment or acquisition environment, has a short duration, and a high intensity.
f) Signal-to-noise ratio, tested according to formula
5.3 Voiceprint Data Quality Classification
5.3.1 Quality Classification Based on the occurrence of issues such as amplitude truncation, waveform distortion, speech interference, sudden noise, and obvious popping sounds, the voiceprint data quality can be classified from low to high. The high-level classification is divided into three levels. Level 1, Level 2, and Level 3, as shown in Table 1, which are applicable to different applications.
5.3.2 Level
1 The upper cutoff frequency of the voiceprint data is greater than or equal to 3000Hz and the lower cutoff frequency is less than or equal to 500Hz, but does not meet the requirements of... Quality requirements for Level 2.
Note. This level of voiceprint data is suitable for voiceprint recognition applications that do not require high recognition accuracy, such as using voiceprint recognition to achieve personalized services for voice assistants.
5.3.3 Level
2 The upper cutoff frequency of voiceprint data is greater than or equal to 3400Hz and the lower cutoff frequency is less than or equal to 300Hz, indicating the signal strength of the speech segment. Audio with a noise ratio greater than or equal to 25dB, exhibiting issues such as amplitude clipping, waveform distortion, speech interference, sudden noise, and noticeable popping, but still possessing quality problems. The total duration of each segment is less than or equal to 10% of the total audio duration; or the signal-to-noise ratio of the speech segment is less than or equal to 25 dB and greater than or equal to 15 dB. Furthermore, it does not suffer from issues such as amplitude clipping, waveform distortion, speech mixing, sudden noise, or noticeable popping.
Note. This level of voiceprint data is suitable for voiceprint recognition applications that require a certain level of accuracy, such as identifying the user checking their bank account balance. Authentication.
5.3.4 Level
3 The upper cutoff frequency of voiceprint data is greater than or equal to 7000Hz and the lower cutoff frequency is less than or equal to.200Hz, indicating the signal strength of the speech segment. The noise ratio is greater than or equal to 25dB, and there are no issues such as clipping, waveform distortion, speech mixing, sudden noise, or obvious popping.
Note. This level of voiceprint data is suitable for voiceprint recognition applications that require high accuracy, such as high-security vaults authenticating operators.
6 Requirements for building a voiceprint audio file library
6.1 Database Classification Based on their different uses, voiceprint databases are divided into the following three categories.
a) Voiceprint Algorithm Test Library. The voiceprint speech files contained in the database are mainly obtained indirectly; the voiceprint speech files contained in the database... The document typically specifies whether it is for voiceprint recognition registration or voiceprint recognition testing;
b) Voiceprint Identity Database. The voiceprint audio files contained in the database should meet the requirements of 4.2.2; the database should only contain those used for voiceprint registration. Voiceprint audio files;
c) Voiceprint-Specific Application Library. The database contains voiceprint audio files obtained directly and indirectly, depending on the specific application being studied. The application specifies how voiceprint voice files should be used.
6.2 General Requirements for Database Construction
6.2.1 Database Composition Voiceprint audio files are compiled into a voiceprint audio file database based on application scenario requirements. A single speaker can have multiple voiceprints. Voice files can correspond to the same application scenario, or different voiceprint voice files can correspond to different application scenarios. Newly acquired voiceprint voice files... The samples are saved as voiceprint speech files according to the application scenario requirements and then stored in the voiceprint speech file database.
6.2.2 General Requirements The general requirements for building a voiceprint audio file library are as follows:
a) Voice quality testing should be performed on the voiceprint files to ensure they meet the corresponding quality requirements for database entry;
b) The voiceprint file used for voiceprint registration should contain only the voice of one speaker and no other non-human audio.
c) It should be confirmed that the content of the voiceprint file meets the requirements of the application scenario;
d) The documentation for the voiceprint and speech file database should be maintained and updated regularly, including at least the purpose of database creation, time of creation, and speech data. Quality requirements, types of audio content, and number of audio files, etc.;
e) It is advisable to collect or detect relevant information for each voiceprint speech file, and correctly fill in each information item, as shown in Table 2, where voiceprint speech Document identifiers should be unique.
6.2.3 Data Security Requirements The construction of voiceprint audio files should comply with the requirements of GB/T 41807-2022, and adopt technologies such as data encryption, logical isolation, or physical isolation. Measures are taken to ensure the confidentiality and integrity of voiceprint files during storage and transmission; the filenames of voiceprint files should not contain any characters that could identify the speaker. Information about the speaker.
6.3 Voiceprint Algorithm Test Library
6.3.1 Basic Requirements for Database Construction The database should be built according to the following basic requirements.
a) The number of voiceprint audio files and the number of speakers should be designed.
b) Voiceprint files containing both target and non-target speakers; each target speaker should include voiceprint registration information. Audio files and voice files used for voiceprint recognition.
c) The attribute distribution of voiceprint audio files should be considered.
---Language/Dialect;
---Source of voiceprint audio file;
---Voiceprint data quality level;
d) Other factors affecting voiceprint recognition evaluation should be considered according to the requirements of algorithm research, such as noise level, number of speakers, and speech patterns. A person's physical condition, etc.
6.3.2 General Voiceprint Algorithm Test Library In algorithm research or application scenarios where there are no specific requirements regarding the number and distribution of speakers, or the number and distribution of voiceprint files. The following requirements should be met when constructing a voiceprint algorithm test library.
a) Contains voiceprint samples of the target speaker and optional voiceprint files of non-target speakers, with the target speaker and non-target speaker being separate entities. The total number of speakers is greater than or equal to 10,000, and the target number of speakers is greater than or equal to 2,000.
b) A single target speaker has voiceprint speech files for voiceprint registration and voiceprint recognition, and the voiceprint speech files used for voiceprint recognition... There are 5 or more items.
c) The distribution of the target speakers should take the following factors into consideration.
---Speaker gender. Male and female each accounted for (50±5)%;
---Speaker age. Under 18 years old accounted for (17±3)%; 18-60 years old accounted for (70±5)%; Over 60 years old accounted for (13±3)%;
6.4 Voiceprint Identity Database
6.4.1 Identity Information Association and Storage Voiceprint files in the voiceprint identity database are associated with the speaker's identity information; the two should be logically or physically isolated. The techniques are stored separately.
6.4.2 Specific requirements for database construction Based on meeting the general requirements of 6.2.2, the voiceprint voice files in the voiceprint identity database must meet the following conditions.
a) Data is extracted and saved from the collected voiceprint speech samples, and the collection of voiceprint speech samples should meet the requirements of 4.2.2. Requirements.
b) Use standard uncompressed WAV format, with a sampling rate of 8000Hz or 16000Hz and a quantization precision of 16 bits. The vocal tract.
c) The voice data quality reaches level 2 or level 3.
d) Fill in the speaker identifier in the voiceprint audio file information field.
e) The audio content meets the following requirements.
---When the content is long text, the coverage and balance of phonemes should be considered;
---When the content is a fixed short text, it should have 6 or more syllables, and at least 3 syllables should have different pronunciations;
---When the content is a pure string of numbers, it should contain 40 or more numbers, with each number appearing 4 or more times.
Note. If the same speaker has multiple voiceprint audio files, the audio content of these files can be merged to meet the requirements.
6.5 Voiceprint-Specific Application Library
6.5.1 Personal Information Association Voiceprint-specific voice files or voice segments within files in the voiceprint-specific application library should be associated with speaker information, which can be used for identity verification. The information can also be other information that identifies the speaker, such as a virtual identity ID.
6.5.2 Factors affecting database construction Based on meeting the general requirements of 6.2.2, the construction of voiceprint-specific application libraries should be tailored to the specific application domain or scenario requirements. Consider the following influencing factors.
a) The acquisition conditions of directly obtained voiceprint files should meet the application requirements;
b) The data processing method for indirectly obtained voiceprint audio files should meet application requirements;
c) The speaker identifier in the voiceprint file information field should be filled in;
d) Voiceprint data quality level;
e) Valid voice duration;
f) Language/Dialect;
g) Audio content;
h) Encoding method;
......
This preview omits tables, figures, formulas and parts of the technical clauses. The complete document — 24 pages — is available in the English PDF.
How to Buy GB/T 47478.1-2026
- 1Add to cart. Click the "Buy GB/T 47478.1-2026" button on this page. You can add more standards before checkout.
- 2Checkout. Enter your email and billing details. Payment is processed securely by Stripe (cards, Apple Pay, Google Pay supported).
- 3Instant delivery (0–9 sec). Delivery is automatic: within seconds of payment you'll receive an email with a secure download link. The link stays valid for 72 hours.
- 4Invoice included. A tax invoice is attached to the confirmation email. Need a custom invoice? Contact us.
Related Standards
GB/T 47310-2026 — Determination of total silicon, aluminium, iron, potassium, sodium, calcium, magnesium, manganese, phosphorus, titanium and sulfur in soil - Monochromatic excitation energy dispersive X-ray fluorescence spectrometry
GB/T 47321-2026 — Specification for the warning data exchange of the national emergency early warning dissemination system
GB/T 47293-2026 — Determination of available mercury in soil
Secure payment via Stripe
Payments accepted
GB/T 47478.1-2026
$260.00