Valid

GB/T 40035-2021Basic requirements for bilingual parallel corpus processing services (English PDF)

双语平行语料加工服务基本要求

Open the GB/T 40035-2021 preview as PDF

Preview — first pages of GB/T 40035-2021 (full document: 19 pages)

This is a limited preview

Buy now to download the full PDF (19 pages)

Issued by

SAMR; SAC

Level / Type

National · Recommended

Issue date

April 30, 2021

Implementation date

November 1, 2021

Scope

GB/T 40035-2021 is the English-translated version of 双语平行语料加工服务基本要求.

China's national standard for the basic requirements of bilingual parallel corpus processing services. It specifies the terms and definitions, the general provisions and the basic requirements. A parallel corpus is a body of text in two languages with the segments aligned to each other, and it is the raw material of machine translation and of professional translation memory. Building one is a service industry in its own right: text is collected, cleaned, segmented into sentences, aligned across the languages, checked, and annotated. The quality of what results determines the quality of every model trained on it, and the failure modes are specific - misaligned segments teach a model to translate one sentence as another, and a corpus scraped without attention to its provenance carries the machine-translated text of earlier systems, so a model trained on it learns their errors. Standardising the process is how a buyer of such a service can specify what they are getting.

Document preview — GB/T 40035-2021

National Standard of the People's Republic of China

ICS
03.080.99;35.240.30
Classification
A 10

Issued by: State Administration for Market Regulation; Standardization Administration of the PRC

Contents

  • 1 Scope1
  • 2 Normative references1
  • 3 Terms and definitions1
  • 4 General2
  • 5 Basic requirements3
  • 5.1 Service provider3
  • 5.2 Corpus Processing Staff3
  • 5.3 Service environment3
  • 5.4 Processing content3
  • 5.5 Processing results3
  • 5.5.1 Completeness3
  • 5.5.2 Accuracy3
  • 5.5.3 Availability4
  • 5.5.4 Normative4
  • 5.6 Corpus Processing Tools4
  • 5.6.1 Reliability4
  • 5.6.2 Ease of use4
  • 5.6.2.1 Localization interface4
  • 5.6.2.2 Operation function4
  • 5.6.2.3 Help system5
  • 5.6.2.4 Efficiency5
  • 5.6.3 Compatibility5
  • 6 Processing flow5
  • 6.1 Pretreatment5
  • 6.1.1 Corpus preparation5
  • 6.1.2 Cleaning5
  • 6.1.3 Deduplication5
  • 6.1.4 Desensitization5
  • 6.2 Corpus Alignment6
  • 6.3 Corpus review6
  • 7 Service content6
  • 7.1 Demand communication6
  • 7.2 Customer Agreement6
  • 7.3 Project Management6
  • 7.4 Processing link6
  • 7.5 Delivery content7
  • 7.6 Warranty period7
  • 7.7 Service evaluation and improvement7
  • 8 Data Security7
  • 8.1 Data backup7
  • 8.2 Document Management and Log7
  • 14 Reference15

Foreword

This document is in accordance with the provisions of GB/T 1.1-2020 "Guidelines for Standardization Work Part

1.Structure and Drafting Rules of Standardization Documents" Drafting. Please note that some of the contents of this document may involve patents. The issuing agency of this document is not responsible for identifying patents. This document was proposed and managed by the National Technical Committee for Language and Terminology Standardization (SAC/TC62). Drafting organizations of this document. China National Institute of Standardization, China Translation Association, Shanghai Yizhe Information Technology Co., Ltd., Shanghai Youyi Information Division Technology Co., Ltd., GTCOM Technology Co., Ltd., Beijing Yueer Information Technology Co., Ltd., Suzhou Lianyue Technology Co., Ltd., Sichuan Yanqiao Information Technology Co., Ltd., Beijing Baidu Netcom Technology Co., Ltd., Shenyang Yayi Network Technology Co., Ltd., Shanghai Zhishanhe Network Division Technology Co., Ltd., Beijing Language and Culture University, Beijing University of Posts and Telecommunications. The main drafters of this document. Liu Zhiyang, Zhang Jing, Ye Jian, Chai Ying, Huang Baorong, Luo Huifang, Meng Yongye, Zhu Li, Zhang Xuetao, Wang Haitao, Zhu Xianchao, Han Lintao, Zheng Chunping, He Zhongjun, Yu Limei, Zhang Chunliang, Gan Keqin, Zhang Baolin. Basic requirements for bilingual parallel corpus processing services

1 Scope

China's national standard for the basic requirements of bilingual parallel corpus processing services. It specifies the terms and definitions, the general provisions and the basic requirements. A parallel corpus is a body of text in two languages with the segments aligned to each other, and it is the raw material of machine translation and of professional translation memory. Building one is a service industry in its own right: text is collected, cleaned, segmented into sentences, aligned across the languages, checked, and annotated. The quality of what results determines the quality of every model trained on it, and the failure modes are specific - misaligned segments teach a model to translate one sentence as another, and a corpus scraped without attention to its provenance carries the machine-translated text of earlier systems, so a model trained on it learns their errors. Standardising the process is how a buyer of such a service can specify what they are getting.

This document specifies the basic requirements, processing procedures, service content and data security of bilingual parallel corpus processing services. This document is applicable to the digital bilingual corpus processing service that takes the original text and the translation as the object and the text as the expression form, and other digital The corpus processing of the text can also be used as a reference, and it is also suitable for the evaluation of corpus alignment tools.

2 Normative references

There are no normative references in this document.

3 Terms and definitions

The following terms and definitions apply to this document.

3.1 Text Data formed by the arrangement of characters, symbols, words, phrases, paragraphs, sentences, tables or other characters, used to express meaning, and its interpretation is basically It depends on the reader's knowledge of a certain natural language or artificial language. [Source: GB/T 4894-2009, 4.1.1.2.4]

3.2 Corpus Language materials or information.

3.3 Bilingual parallel corpus bilingualparalelcorpus A corpus composed of two languages and aligned in parallel in chapters, paragraphs, sentences or other levels (3.2).

3.4 Original sourcelanguagetext Source language text (3.1). [Source: GB/T 19363.1-2008, 3.4, with modification]

3.5 Translation targetlanguagetext Target language text (3.1). [Source: GB/T 19363.1-2008, 3.5, with modification]

3.6 Client Accept individuals or organizations that provide products or services according to their requirements. [Source: GB/T 19000-2016, 3.2.4, with modification]

3.7 Metadata Descriptive data about the content, quality, condition, and other characteristics of the data.

......
This preview omits tables, figures, formulas and parts of the technical clauses. The complete document — 19 pages — is available in the English PDF.

How to Buy GB/T 40035-2021

  1. 1Add to cart. Click the "Buy GB/T 40035-2021" button on this page. You can add more standards before checkout.
  2. 2Checkout. Enter your email and billing details. Payment is processed securely by Stripe (cards, Apple Pay, Google Pay supported).
  3. 3Instant delivery (0–9 sec). Delivery is automatic: within seconds of payment you'll receive an email with a secure download link. The link stays valid for 72 hours.
  4. 4Invoice included. A tax invoice is attached to the confirmation email. Need a custom invoice? Contact us.

Related Standards

English PDF
19 pages
Instant delivery (0–9 sec)
Invoice included
View Cart

Secure payment via Stripe

Payments accepted

VisaMastercardAmerican ExpressApple PayGoogle PayStripe

GB/T 40035-2021

$300.00

$255.00for partners