---
title_en: "AI Safety Governance Framework 2.0"
title_zh: "人工智能安全治理框架 2.0"
abbreviation: "AI Safety Governance Framework 2.0"
hierarchy: "handbook"
issuing_body: "National Information Security Standardization Technical Committee (TC260); National Computer Network Emergency Response Technical Team/Coordination Center of China (CNCERT), under the guidance of the Cyberspace Administration of China"
adopted_date: 2025-09-15
effective_date: 2025-09-15
status: "effective"
source_url: "https://www.cac.gov.cn/2025-09/15/c_1759653448369123.htm"
related_laws: ["genai-services-interim-measures", "ai-content-labeling-measures", "gb-45438-ai-content-labeling-method", "algorithmic-recommendation-provisions", "deep-synthesis-provisions", "gbt-45654-genai-basic-security-requirements", "gbt-45652-genai-training-data-security", "gbt-45674-genai-data-annotation-security", "tc260-agent-deployment-security-guide", "ai-agent-standardization-innovation-opinion", "sci-tech-ethics-review-measures", "ai-ethics-review-service-measures", "ai-anthropomorphic-interaction-measures", "pipl", "dsl"]
domains: ["ai-governance", "data-security"]
url: https://datacompliancechina.com/laws/ai-safety-governance-framework-2/
summary: "Released on September 15, 2025 at the National Cybersecurity Awareness Week by TC260 and CNCERT under CAC guidance, the AI Safety Governance Framework 2.0 replaces the September 2024 version 1.0 as China's policy-level map of AI risk and response, and the document from which the 2025–2026 wave of AI standards, agent guidance and sector guidelines draws its vocabulary. It states five principles — inclusive prudence, risk-oriented agile governance, integration of technology and management, open cooperation, and trustworthy application with prevention of loss of control — and reorganizes risk into three tiers: inherent technical risks (model and algorithm: explainability, bias, robustness, hallucination, adversarial attack, defect propagation through open-source base models; data: unlawful collection, poisoned or infringing content, poor annotation, leakage through model parameters), application risks (network and systems: components and compute, expanded attack surface from local deployment and agents, supply-chain cut-offs, AI-enabled attacks; content: unlawful output, unlabeled deepfakes, ecosystem pollution; physical: CII failures, criminal misuse, CBRN knowledge; cognitive: filter bubbles and cognitive warfare), and derivative risks (labor, resources, bias and the intelligence gap, education, research ethics, anthropomorphic dependency, social order, and loss of control). Each risk is paired with technical countermeasures and fourteen governance measures, including AI safety legislation, ethics rules, open-source and supply-chain security, application classification with five risk grades, content traceability, sector deployment guides, a testing system, threat-information sharing, data and personal-information rules, and loss-of-control consensus. Chapter 6 gives operational guidance for developers, deployers, operators and users — including six-month log retention, human review in key scenarios and circuit-breakers for autonomous systems. Annex 1 sets grading factors (scenario, level of intelligence, scale) and five risk grades; Annex 2 states eight trustworthy-AI principles led by ultimate human control and respect for national sovereignty."
---

> **Source: Data Compliance China** — https://datacompliancechina.com/laws/ai-safety-governance-framework-2/ · English rendering and annotations by DCC; the Chinese original governs. Cite as: Data Compliance China, "AI Safety Governance Framework 2.0", https://datacompliancechina.com/laws/ai-safety-governance-framework-2/
> *DCC summary, not a translation.* The Framework is published by TC260
> and CNCERT with an official English version; DCC's page is a structured
> summary keyed to the Chinese text rather than a reproduction of either
> version. Specific passages should be checked against the published
> document.

**Published by:** National Information Security Standardization Technical Committee (全国网络安全标准化技术委员会, TC260) and the National Computer Network Emergency Response Technical Team/Coordination Center of China (国家计算机网络应急技术处理协调中心, CNCERT), organized under the guidance of the Cyberspace Administration of China.  
**Released September 15, 2025** at the main forum of the 2025 National Cybersecurity Awareness Week; supersedes version 1.0 of September 2024. Acknowledged contributors include CESI, the Chinese Academy of Cyberspace Studies, the CAC Data and Technology Support Center, the Zhongguancun and Shanghai AI laboratories, CAICT, Peking, Tsinghua and Zhejiang universities, Huawei, Alibaba and MiniMax.

## Scope and structure

The Framework is a policy document rather than a standard: it implements the Global AI Governance Initiative and the "people-centered, AI for good" orientation, and is offered as a basis for consensus among governments, industry, institutions and the public at home and abroad. Its preface explains why 1.0 needed replacing within a year — high-performance reasoning models, open-source lightweight models lowering deployment barriers, the shift from chatbots to agents embedded in business processes, and embodied AI and brain–computer interfaces bringing human–machine integration into view. It has six chapters (principles; framework composition; risk classification; technical countermeasures; comprehensive governance measures; and safety guidelines for R&D and application), a mapping table, and three annexes (risk-grading principles, basic principles for trustworthy AI, and terminology).

## Key contents

### Principles (chapter 1)

The governance outlook is "common, comprehensive, cooperative and sustainable" security with equal weight on development and safety, innovation as the first priority and risk prevention as the point of departure. Five principles follow: **inclusive prudence, ensuring safety** — tolerance for error through pilots in controllable environments, with a hard floor for national security, public interest and citizens' rights; **risk orientation, agile governance** — tracking risks across technology, application and social impact, exploring grading by scenario, level of intelligence and scale, and responding promptly where government regulation is genuinely needed; **technology with management, coordinated response** — allocating responsibility among model and algorithm developers, service providers and system users, and combining regulation, industry self-discipline and public supervision, with explicit attention to the open-source model ecosystem; **open cooperation, joint governance** — international platforms and a globally consensual governance system; and **trustworthy application, preventing loss of control** — multi-layer trustworthy-AI principles covering technical protection, value alignment and collaborative governance, so that AI "always remains under human control."

### Risk classification (chapter 3)

**Inherent technical risks (3.1).** *Model and algorithm:* insufficient explainability; bias and discrimination introduced in design or from training data, including discriminatory output on ethnicity, belief, nationality, region and gender; weak robustness; unreliable output — "hallucination"; external adversarial attacks that steal or tamper with parameters, structure or function or exhaust resources; and defect propagation from base models to downstream fine-tuned models and applications, accelerated by open-sourcing, which also facilitates the training of "malicious models." *Data:* unlawful collection and use of data and personal information in training and interaction; improper training content — false, biased or infringing — and poisoning that distorts value alignment and probability distributions; non-standard annotation; and leakage of data and personal information latent in model parameters through weak protection, un-"forgotten" sensitive information, induced interaction or attack.

**Application risks (3.2).** *Network and system:* defects, vulnerabilities and backdoors in frameworks, platforms and compute, malicious consumption of compute, and cross-boundary propagation across heterogeneous compute; an expanded attack surface from local deployment and from agents that call terminal files, permissions, interfaces and tools; supply-chain risk, framed as unilateral technology monopolies and export controls threatening chip, software and tool supply; and AI-enabled attacks, including synthetic media that defeats face and voice authentication. *Information content:* unlawful and harmful output driven by weak model safety, weak application guardrails or malicious prompting; confusion and misleading of users through unlabeled and deepfake content; and pollution of the online content ecosystem through low-quality output and model self-citation. *Physical:* new challenges to economic and social operation where hallucination, wrong decisions, misuse or attack affect AI in energy, telecommunications, finance, transport and other CII sectors; use by criminal activity (terrorism, violence, gambling, drugs); and loss of control over nuclear, biological, chemical and missile knowledge and capability through broad training corpora and retrieval augmentation. *Cognitive:* intensified "information cocoons" through precise profiling of users' needs, intentions and ideological currents; and cognitive warfare — propaganda, interference in other states' internal affairs, and social bots seizing discourse and agenda-setting power.

**Derivative risks (3.3).** *Social and environmental:* disruption of employment structures as capital, technology and data displace labor; and pressure on power, land and water from disorderly compute construction, fragmented lightweight deployment and duplicative model development. *Ethical:* systemic discrimination and a widening intelligence gap from classifying and treating groups differently; dependence that erodes learning, research and creative capacity; lowered barriers to high-ethical-risk research in biology and genetics; emotional dependence on anthropomorphic interaction; challenges to established views of employment, child-bearing and education; and the possibility of a sudden, unexpected leap in capability leading to autonomous resource acquisition, self-replication, "self-awareness" and a contest with humans for control.

### Technical countermeasures (chapter 4)

For inherent risks: explainability and transparency about internal structure, reasoning, interfaces and outputs; better architectures, larger and more diverse data and human oversight against bias; secure development norms and adversarial training against prompt injection; assessment of defect propagation from base and open-source models; compliance with data-collection and personal-information rules across the lifecycle of training and interaction data; true, accurate, objective, diverse and lawfully sourced data filtered of false, biased, stale and erroneous content and of sensitive CBRN material; standardized annotation; data-security management for sensitive personal information and important data, with synthetic data encouraged in place of personal features; and IP protection in data selection and output. For application risks: disclosure of principles, capabilities, scenarios and risks; permission management, disabled non-essential services and access control on platforms aggregating multiple models; secure deployment and maintenance with vulnerability tracking, scanning and patching; supply-chain attention to chips, software, tools, compute and data; redundancy and disaster recovery; guardrails filtering inputs and outputs against injection, unlawful content and leakage of sensitive personal information and important data; content labeling for identifiability and traceability; capability boundaries that trim abusable functions; verification, fault-tolerance and correction for decisions; "circuit-breakers" and "one-key control" wherever highly autonomous execution is introduced; extreme-condition testing of perception systems for driving and drones; end-use traceability against CBRN scenarios; detection of unexpected, untrue or inaccurate outputs; strict prevention of misuse of systems that profile users' identity, preferences and ideological tendencies; and detection technology against cognitive-warfare content. For derivative risks: green AI standards and low-power computing; data filtering, value alignment and output verification against discrimination; efficient emergency controls for systems in government, CII and public-safety and health settings; and transparent, explainable models.

### Comprehensive governance measures (chapter 5)

Fourteen measures: (1) AI safety legislation covering infrastructure protection, classified and graded regulation, testing, end-use management and key-scenario application, with room for local experimentation; (2) AI science-and-technology ethics rules and orderly ethics review in life and health, dignity, employment, environment and sustainability, with an ethics-service system and SME support; (3) full-lifecycle security capability — reliability, trustworthiness, transparency, fault tolerance, privacy and value alignment, tested adversarially; (4) open-source ecosystem and supply-chain security, including open-source rules that impose disclosure duties on model providers toward downloading users and specify "prohibited" uses; (5) application classification and risk grading, with registration and filing of AI systems used in CII (Annex 1); (6) global promotion of the content-labeling traceability paradigm across production, dissemination and distribution; (7) baseline security guides for large-model deployment in key industries — model selection, deployment, operation and decommissioning — followed by sector guides for energy, telecommunications, finance, transport, education and industry; (8) a three-layer testing system (model and algorithm, general application, specific scenario) and crowdsourced vulnerability testing; (9) an AI vulnerability database and threat-information sharing among developers, providers and technical institutions, with international cooperation; (10) data-security and personal-information rules for each stage of training, annotation, use and output, de-identification of personal data, and protection of important and core data in government and finance; (11) consensus on loss-of-control risk through end-use management in CBRN scenarios, trustworthy-AI principles (Annex 2) and periodic developer testing for loss-of-control potential; (12) talent cultivation from basic to higher education and in frontier fields; (13) society-wide awareness, industry self-discipline above regulatory requirements, and a public reporting channel; and (14) international cooperation through the UN, APEC, G20, SCO, BRICS, Belt and Road and Global South partners and the Global AI Governance Action Plan.

### Safety guidelines for R&D and application (chapter 6)

**Model and algorithm development (6.1).** Design in reliability, fairness, transparency, explainability, privacy and value alignment; assess bias and build alignment algorithms; secure training environments; manage data sources with cleaning, annotation and review; filter errors and unlawful content with classifiers and sampling; use cross-annotation and result audits; protect data, personal information and IP with de-identification; follow open-source licences and audit frameworks and code; run periodic graded security testing with diverse datasets; keep version control with rollback for commercial releases; test in sandboxes and, for commercial developers, produce detailed test reports; tell providers and downstream developers the model's tolerance limits, scope, cautions and contraindications; disclose audits and anomaly handling periodically; and contribute governance tools to open-source communities.

**Application construction and deployment (6.2).** Assess necessity and long-term impact, grade risk by scenario importance, level of intelligence and scale, and audit accordingly; obtain model files, frameworks and libraries from official channels in stable versions with integrity checks and security testing; scan hardware, software and third-party tools for unpatched exploitable vulnerabilities and trace supply-chain backdoors; harden configuration, disable unnecessary ports and services and fix default passwords; authenticate and authorize human and API interfaces with least privilege, rate limits, disabled high-risk operations for ordinary users and suspension or blocking of malicious users; limit data access, plan backup and recovery and audit data flows; and deploy guardrails against unlawful content and prompt injection.

**Operation and management (6.3).** Establish management and supervision with clear responsibility, human review in key scenarios and decisions that are transparent, controllable and under human authorization; enforce least privilege and account security with encryption for sensitive data; monitor operation with warning thresholds, incident plans, the ability to fall back to manual or traditional systems, and regular drills; add explicit or implicit labels and deploy deepfake detection for government disclosure and judicial evidence; adopt content-interaction norms, security operations, complaint mechanisms and technical protection against generation and spread of false and harmful information; keep operation logs of system and user behavior for at least six months and audit them; monitor risk in real time; publish capabilities, limitations, intended users and scenarios; explain goal attainment and deviation and give explanations for significant decisions; inform users of scope, cautions and contraindications in plain contract language and support human supervision and control rights; identify the party responsible for data ownership and algorithm defects; assess and manage data-leakage and unlawful-collection risk across the data lifecycle; test resilience under faults and attacks for minimum viable function; train staff; reserve contractual rights to correct or terminate on misuse; and design for the usability and safety of minors, the elderly and special groups.

**Access and use (6.4).** Choose reputable applications, read contracts and privacy policies and calibrate expectations, avoid entering sensitive information unnecessarily, understand data-handling practices, guard against the application becoming an attack target, and prevent addiction and overuse by children and adolescents.

### Annexes

**Annex 1 — risk grading.** Three grading factors: application scenario (purpose, sector, environment, users, social and economic impact); level of intelligence (from advisory systems requiring human decisions to fully autonomous operation); and scale (from internal or regional tools to systems with large user bases or deep embedding in key sectors such as assisted driving, city operations, industrial scheduling and financial risk models). Five grades: low, ordinary, relatively major, major and especially major — the last defined as catastrophic, systemic, subversive or irreversible. Grading is to proceed through a national standard on AI application classification and grading, sector rules instantiating the factors and weights, and sector-organized grading exercises focused on major and relatively major risks.

**Annex 2 — trustworthy AI.** Eight principles: ultimate human control (thresholds, kill switches, intervention windows); respect for national sovereignty (compliance with host-country law and no interference in other states' affairs); value alignment with peace, development, fairness, justice, democracy and freedom; transparency about goals, logic, models, data sources and decision basis; objective verifiability through testing and certification; safety protection through risk modeling, testing and lifecycle audit; forward-looking prevention and response; and global collaborative governance through the UN.

**Annex 3 — terminology.** Fifteen definitions including explainability, synthetic data, annotation, pre-training and fine-tuning, alignment, reinforcement learning, inference, explicit and implicit labels, data poisoning, adversarial attack, agent, and guardrails.

## How it fits the regime

The Framework is not binding, but it is the document that explains the shape of what is. Its 2024 predecessor was cited by the CAC when it drafted the [Labeling Measures](/laws/ai-content-labeling-measures/) and [GB 45438](/laws/gb-45438-ai-content-labeling-method/); version 2.0's data chapter is the policy statement behind the three generative-AI standards of 2025 ([GB/T 45654](/laws/gbt-45654-genai-basic-security-requirements/), [45652](/laws/gbt-45652-genai-training-data-security/), [45674](/laws/gbt-45674-genai-data-annotation-security/)); its agent-risk paragraphs and "circuit-breaker" countermeasure anticipate the [TC260 agent deployment guide](/laws/tc260-agent-deployment-security-guide/) of July 2026; measure 5.2 points to the [Science and Technology Ethics Review Measures](/laws/sci-tech-ethics-review-measures/) and the [AI ethics review measures](/laws/ai-ethics-review-service-measures/); and measures 5.5 and 5.7 announce the classification-and-grading standard and the sector deployment guides that are now in preparation. For overseas counsel three signals matter: the Framework's insistence that base-model and open-source providers carry disclosure and "prohibited use" duties toward downstream users; its treatment of export controls on chips as a supply-chain *risk* to be countered by domestic ecosystems; and Annex 2's "respect for national sovereignty" principle, which frames compliance with Chinese law as a trustworthiness criterion for any AI service operating in China.
