Skip to content
DCC · DATA COMPLIANCE CHINA China data law, for overseas counsel.
§ LAW · GB/T 45652

Cybersecurity Technology — Security Specification for Generative Artificial Intelligence Pre-training and Fine-tuning Data (GB/T 45652-2025).

网络安全技术 生成式人工智能预训练和优化训练数据安全规范 (GB/T 45652-2025)

DCC summary, not a translation. GB/T 45652-2025 is a copyrighted national standard. The structured summary below is DCC’s own paraphrase grounded in the published text; specific clauses should be checked against the standard.

Published by: State Administration for Market Regulation and Standardization Administration of China; proposed and administered by the National Information Security Standardization Technical Committee (SAC/TC260).
Published April 25, 2025. Implemented November 1, 2025. Recommended national standard.

Scope

GB/T 45652-2025 specifies the security requirements for generative-AI pre-training and fine-tuning data and for the activities that process it, and describes the corresponding evaluation methods. It applies to generative-AI service providers carrying out pre-training and fine-tuning data processing and security self-assessment, and to third-party bodies assessing the security of such data. It normatively references GB/T 35273 (personal information security specification) and GB/T 41479-2022 (network data processing security requirements). Pre-training is training on large-scale data to give a model general knowledge; fine-tuning (优化训练) is training on domain data, on top of pre-training, to give the model domain-service capability; the two data categories are defined accordingly.

Key contents

General security requirements (clause 4)

Providers shall: (a) adopt a security-management policy for pre-training and fine-tuning data covering the protection organization, classification and grading rules, processing-activity security and incident response; (b) keep redundant backups in storage; (c) encrypt data in transit; (d) isolate training data by batch and label batches so that training content is traceable; (e) comply with Chapter 5 of GB/T 41479-2022; (f) process any personal information in accordance with GB/T 35273 and preferably anonymize or de-identify it; (g) protect training data with reasonable safeguards and tools; (h) preferably run training systems at MLPS Level 3 or above; (i) adopt and enforce a deletion policy with defined targets, approval and logging, deleting on data-subject request within the prescribed time; (j) preferably make deleted data unrecoverable (repeated overwrite, multiple formatting, physical destruction); (k) preferably establish a data-security team and supervisory function with defined roles; (l) assess training-data security periodically, respond to incidents, and train and examine staff in key positions; (m) protect industry data according to sector rules and standards; (n) preferably detect and repair or filter poisoned data — both bulk mislabeled or irrelevant data inserted to degrade the model and targeted data inserted to trigger specific wrong outputs; and (o) assess the authenticity of training data.

Pre-training data (clause 5)

Collection (5.1). Assess and record data at collection; unlawful or harmful content (the 29 risks in GB/T 45654 Annex A.1–A.4) must not exceed 5%. Do not self-collect data that others have expressly barred. Comply with open-source licences or hold authorization. Record sources for externally collected data: the URL for websites; the dataset name and source organization, with a legally effective contract, cooperation agreement, licence or authorization, for organizations or individuals; and the service name and user identifier, with an authorization record, for user-sourced data. Each data type (code, images, audio, video, text in a given language) must have multiple sources, each accounting for at least 1%. Obtain consent for personal information and separate consent for sensitive personal information unless another legal basis applies. Review the data, undertakings and supporting materials supplied by trading or cooperation partners. Comply with cross-border data-security rules where collection crosses borders.

Pre-processing (5.2). Sample-verify each source and do not use any source in which more than 5% of content is unlawful or harmful. Secure the pre-processing environment with platforms and tools matched to the data grade, preferably tested periodically for vulnerabilities; use secure transfer and storage. Preferably de-identify personal information and store re-identification keys separately under tighter access control. Add provenance metadata to every sample — existing source information, the URL of the sample or its page, the dataset and organization name, or the service name and user identifier as applicable. Filter all training data of each type before training by keywords, classifiers and manual sampling, and record the results. Manage intellectual-property risk with a policy and named person, do not train on data found to infringe (with particular attention to copyright works), run a complaints channel and update the policy, and warn users in the service agreement. Build modality-specific filters — sensitive-word and ambiguity libraries and text classifiers for privacy and sensitive information; sensitive-image classifiers; audio recognition and filtering; and video methods combining both — and language-specific tools, assessing translation and conversion scenarios for semantic consistency and preferably tailoring filters to each language’s grammar and context.

Use (5.3). Reduce the likelihood of the model being induced to generate risky content, including by fully filtering identified risky samples. When foreign-sourced training data is used, pair it in the same training batch with a reasonable proportion of domestic-sourced data. Where data is displayed, assess the necessity and safety of display and do not display unlawful content.

Fine-tuning data (clause 6)

Collection and pre-processing must meet clauses 5.1 and 5.2, and use must meet 5.3, with additions. When collecting data generated by a generative-AI service, record the provider, version, acquisition time and data identifier. Check fine-tuning data for consistency with the tuning objective (compliance, domain fit), for erroneous knowledge and inappropriate expression, and, across iterations, record and publish the time and content of each optimization round. Vertical-domain data must comply with sector rules, pass a domain-specific check for erroneous, misleading, fabricated or tampered content, and preferably come from industry-certified or authoritatively certified sources. Filtering must follow the established mechanism, leave no known risky content, cover unfiltered personal information with subject authorization, and delete intermediate and temporary files promptly; a value-alignment check should identify and dispose of data contrary to human values or ethics across prompts, annotations and knowledge-distillation data. When AI-generated content is used as training data, establish a hallucination-risk assessment to catch misleading knowledge; select data by objective — domain-appropriate data for fine-tuning, values-aligned data for alignment — and preferably evaluate fine-tuning data quality.

Evaluation methods (clause 7)

Each requirement is paired with an evaluation method, an expected result and a pass rule. General requirements are checked through policy documents, system logs, design documents, organization charts, assessment and training records and interviews; results 1–7, 9, 12, 13 and 15 are mandatory, while MLPS Level 3, irreversible deletion, the security team and poisoning detection are optional items. Collection is checked by sampling at least 10% of source-assessment records and verifying the 5% rule, collection records, licences, source records, source diversity (at least two sources per type, each at 1% or more), consent records, partner-review records and cross-border compliance. Pre-processing is checked by sampling at least 1,000 samples per source type for provenance metadata, at least 1,000 for filtering records and unlawful content, at least 1,000 for IP risk records, by testing the IP complaints channel, and by reviewing modality- and language-specific filters; secure platforms, HTTPS and encrypted storage, and de-identification (pseudonymization, k-anonymity) are among the expected results. Use is checked by manually sampling at least 4,000 items (at least 96% free of risky content) and technically sampling at least 10% (at least 98%), plus domestic-pairing records and display assessments. Fine-tuning collection, pre-processing and use are checked by 10%-or-1,000-record sampling for objective consistency, error checks, iteration records, sector compliance, domain checks and certification, filtering records, value-alignment records, hallucination-risk records and data-selection rationale.

How it fits the regime

The standard unpacks clause 4 of GB/T 45654 into a data-engineering specification that a filing or assessment body can audit line by line, and it borrows its general controls from the network-data and personal-information standards (GB/T 41479 and GB/T 35273). Its significance for overseas model developers lies in four places: the provenance-metadata rule, which effectively requires a per-sample record of where every training example came from; the source-diversity floor, which makes single-source or majority-foreign corpora non-compliant; the personal-information and cross-border rules, which import the PIPL consent regime and the Cross-border Data Flows Provisions into dataset assembly; and the fine-tuning provisions on synthetic data, which require recording which model generated the data and assessing hallucination risk. Annotation of the same data is governed by GB/T 45674.

§ RELATED LAWS

See also.

§ COMMENTARY

Briefs on this law.

No briefs filed yet under this law.

§ SUBSCRIBE

The Monday brief.

One short email every Monday. New briefs on Chinese data-compliance rules from the previous week, with the source law cited.

Opt-in only. Unsubscribe anytime by replying "unsubscribe" to any issue.

SUPPORT DCC

Keep the publication free to read. Suggested support is $19.99, or choose your own amount.

Support →