DCC summary, not a translation. GB/T 45654-2025 is a copyrighted national standard. The structured summary below is DCC’s own paraphrase grounded in the published text; specific clauses should be checked against the standard.
Published by: State Administration for Market Regulation and Standardization Administration of China; proposed and administered by the National Information Security Standardization Technical Committee (SAC/TC260).
Published April 25, 2025. Implemented November 1, 2025. Recommended national standard.
Scope
GB/T 45654-2025 specifies requirements for generative-AI services in three areas — training-data security, model security and security measures — and applies to service providers carrying out generative-AI service activities, while serving as a reference for competent authorities and third-party assessment bodies. Its introduction states that it is aimed principally at services with public-opinion attributes or social-mobilization capability and is meant to support filing management and testing and assessment. It is designed to be read with two companions: GB/T 45674-2025 on data annotation and GB/T 45652-2025 on pre-training and fine-tuning data. “Service provider” means an organization or individual providing generative-AI services through an interactive interface or a programmable interface; “training data” means all data used directly as model-training input, including pre-training and fine-tuning data; and annotation is split into functional annotation (for task capability) and security annotation (for safer outputs).
Key contents
Training-data security (clause 4)
Source selection. Before collecting from a data source (a domain, a data supplier, or an open dataset), sample it for security; if more than 5% of the content is unlawful or harmful — defined by reference to the 29 risks in Annex A.1–A.4 — do not collect from it. After collection, sample each source again and, if the 5% threshold is exceeded, do not use it for training. Sampling may be manual, keyword-based or classifier-based.
Source mix and traceability. Use multiple sources for each language and each modality, and pair any foreign-sourced training data with a reasonable proportion of domestic data. For open-source data, comply with the licence or hold authorization; for self-collected data, keep collection records and do not collect what others have expressly barred (robots.txt, technical restrictions, refused personal-information consent); for commercial data, hold enforceable contracts and supplier undertakings on source, quality and security, and review them; for user inputs used in training, hold user authorization records.
Content management. Filter all training data of every modality before training using keywords, classifiers and manual sampling. Maintain an intellectual-property management policy with a named person in charge, do not infringe others’ IP, run an IP complaints channel and update the policy in response to national policy and complaints, and warn users in the service agreement of the IP risks of using generated content. Obtain consent before using personal information in training, and separate consent for sensitive personal information, unless another lawful basis applies.
Data annotation. Train annotators on the law, task rules, tools, quality and security verification and data-security management; examine them and grant, periodically renew, and where necessary suspend qualifications; separate the roles of annotation executor and reviewer so that no person performs both on the same task. Annotation rules must specify objective, format, method and quality indicators, be written separately for functional and security annotation, and — for security annotation — cover the main risks, preferably all 31 in Annex A. Every batch of functional annotation is manually sampled (re-annotate if inaccurate, discard the batch if it contains unlawful content); every item of security annotation is reviewed by at least one reviewer; security annotation data should be stored in isolation.
Model security (clause 5)
Training. Make the security of generated content one of the principal criteria for evaluating outputs — for example by maintaining a security-risk test bank used to optimize and re-test the model after each update, and by building a security-annotation dataset for safety fine-tuning. Audit development frameworks and code regularly for open-source and vulnerability issues. Test regularly for model backdoors and remediate by fine-tuning or unlearning.
Output. Maintain a generated-content pass rate (samples free of the 31 Annex A risks) of at least 90%. Take technical measures to improve accuracy (consistency with scientific consensus and mainstream understanding) and reliability (useful, well-structured output). Refuse clearly extreme prompts and prompts that clearly induce unlawful content, but answer all other prompts. Label images, video and other generated content in accordance with national rules and standards.
Monitoring, updates and environment. Continuously monitor inputs for injection, data-theft and adversarial attacks; keep routine monitoring and emergency-management measures and fix problems through targeted instruction fine-tuning or reinforcement learning; adopt a security-management policy for model updates and re-run a self-assessment after every important update or upgrade; and isolate the training environment from the inference environment, physically or logically.
Security measures (clause 6)
Use cases. Fully justify the necessity, suitability and safety of applying generative AI in each field within the service scope. Where the service is used for critical information infrastructure or important settings such as social governance, public security, automatic control, medical information services, psychological counselling or financial information services, apply safeguards proportionate to the risk. Services open to minors must let guardians set anti-addiction limits, must not sell paid services beyond a minor’s civil capacity, and should actively show beneficial content; services not for minors must keep minors out by technical or management means.
Transparency. Interactive services must publish, in a prominent place such as the homepage, the intended users, settings and uses (and preferably the base model used), and, in the homepage or service agreement, the service’s limitations, a summary of the models and algorithms used, and the personal information collected and its use; API services must put the same in their documentation.
User inputs, complaints and service. Where user inputs are used for training, give users a convenient way to opt out — no more than four clicks from the main interface — and prominently disclose the training status and the opt-out. Provide complaint and reporting channels (telephone, email, in-app window, SMS) with published handling rules and time limits. Detect user inputs with keywords and classifiers, publish and apply a rule that repeated or cumulative daily entry of unlawful content triggers suspension, and staff content monitors in numbers matching the service’s scale, whose duties include tracking national policy and analyzing third-party complaints. Keep backups and recovery strategies for data, models, frameworks and tools.
On-device models. Require official activation on first use and push security-policy updates when the device is online; ship an on-device security module that audits generated content with a keyword library, keeps security logs that can be uploaded or exported, and refreshes its keyword library and configuration when online; and maintain an update mechanism that patches vulnerabilities and repeatedly warns users who have not updated before a major model update.
Annexes
Annex A lists the main security risks of training data and generated content in five groups: A.1 content contrary to the core socialist values (eight items, from inciting subversion and endangering national security to terrorism, ethnic hatred, violence and pornography, and false and harmful information); A.2 discriminatory content (nine items — ethnicity, belief, nationality, region, gender, age, occupation, health and other); A.3 commercial violations (IP infringement, breach of business ethics, disclosure of trade secrets, monopoly and unfair competition through algorithms, data or platform advantage); A.4 infringement of others’ lawful rights (physical and mental health, portrait, reputation, honor, privacy, personal-information rights); and A.5 inability to meet the safety needs of specific service types such as CII, automatic control, medical, psychological and financial services (inaccurate or unreliable content). The 29 risks in A.1–A.4 define “unlawful or harmful information” for the 5% source threshold; all 31 define the pass-rate test.
Annex B describes the assessment method: a keyword library of at least 10,000 entries covering the 17 risks in A.1–A.2 (at least 200 keywords per A.1 risk, 100 per A.2 risk, updated weekly); a generated-content test bank of at least 2,000 questions covering all modalities and languages and all 31 risks (at least 50 questions per A.1–A.2 risk, 20 for the others, updated monthly); should-refuse and should-answer test banks of at least 500 questions each (the should-answer bank covering China’s system, beliefs, image, culture, customs, ethnicity, geography, history and heroes, and gender, age, occupation and health, at least 20 questions per topic); and classifiers covering all 31 risks. It then sets test methods, expected results and pass criteria for every requirement — for example manual sampling of at least 4,000 training-data items per modality with a 96% pass rate and technical sampling of 10% with 98%; at least 1,000 test questions for the 90% output pass rate; and 300 questions each for the 95% refusal / 5% over-refusal thresholds.
How it fits the regime
GB/T 45654 is the operational text behind Article 17 of the Interim Measures for Generative AI Services, which requires providers of services with public-opinion attributes or social-mobilization capability to undergo security assessment and algorithm filing under the Algorithmic Recommendation Provisions. The CAC’s assessment bodies have used its predecessor, the TC260-003 practice guide, as the checklist since 2024; elevation to a GB/T in 2025 stabilizes the thresholds — the 5% source cut-off, the 90% output pass rate, the 95%/5% refusal ratios, the four-click opt-out — that providers must evidence in their self-assessment reports. The two companion standards particularize the data chapter: GB/T 45652 for pre-training and fine-tuning data and GB/T 45674 for annotation, while labeling of outputs is governed by the mandatory GB 45438 and the CAC Labeling Measures. For overseas developers, three provisions bite hardest: the pairing rule for foreign-sourced training data, the personal-information consent requirements for training data, and the transparency and opt-out duties that apply equally to API services.