Skip to content
DCC · DATA COMPLIANCE CHINA China data law, for overseas counsel.
§ LAW · SZDEX GENAI TRADING GUIDELINES

Guidelines for Compliance Assessment of Transactions Involving Generative Artificial Intelligence Services.

生成式人工智能服务交易合规评估指引

Issued by: Shenzhen Data Exchange (深圳数据交易所). Guiding organizations: Cyberspace Administration of Shenzhen (深圳市互联网信息办公室) · Shenzhen Municipal Administration of Government Services and Data (深圳市政务服务和数据管理局). Editor-in-chief: Wang Qinglan, General Manager, Compliance Department, Shenzhen Data Exchange. Drafting organizations include: Shenzhen Data Exchange · Futian District Administration of Government Services and Data · Shenzhen Beipeng Institute of Frontier Technology Law · the Third Research Institute of the Ministry of Public Security · Shenzhen Huaao Data Technology · Shenzhen Institute of Standards and Technology · UBTECH · ZTE · Shenzhen TCL New Technology · China Unicom Smart City Research Institute. Current as of: 9 February 2026.

Editor’s Note — DCC.

The Interim Measures for Generative AI Services regulate a provider serving the public. They do not tell you what happens when the model itself, or the corpus behind it, changes hands as a commercial asset. That is the gap this document addresses, and DCC has not seen another Chinese instrument that addresses it operationally.

Its organizing idea is worth stating up front, because it changes how due diligence has to be scoped: generative AI is not one tradeable thing but three, and liability runs upward through the stack. Buy a service and you inherit the compliance state of the model; buy a model and you inherit the compliance state of its training data. A buyer who diligences only the layer being purchased has, on this framework, diligenced nothing.

Like its cross-border companion, this is exchange-level guidance rather than law — written by the operator of Shenzhen’s data market under the supervision of the Shenzhen cyberspace and data authorities, and filed in 2024 as a Shenzhen local-standard project. Its authority is practical: it is what a Chinese data exchange will ask for.

Summaries below are DCC’s distillation. The Guidelines permit non-commercial reproduction with attribution; DCC does not reproduce the full text.

The three tradeable objects, and the upward-tracing rule

The Guidelines split generative AI transactions into three types and set the assessment scope for each:

What is being tradedWhat must be assessed
Training dataThe training data
A large modelThe model and its training data
A GenAI serviceThe service and the model behind it and that model’s training data

This is the document’s central contribution and the reason it is useful outside Shenzhen. It converts a vague intuition — that model provenance matters — into a defined diligence scope. For an overseas buyer licensing a Chinese model, or a Chinese buyer licensing a foreign one, it says exactly how far back the enquiry has to reach.

The rest of the framework mirrors the cross-border trading guide: three sequential stages — subject (the trading parties), object (the thing traded), circulation (the transfer) — each reviewed for legality, security, integrity and rights protection, across preparation, execution and reporting phases. Annex A is a report template; Annex B is the evidence list per cell.

Training data

Provenance. Sources must be lawful, non-infringing and traceable, with a lawful basis for any personal information — or with the PI anonymized so the model processes none. Every source is verified before and after collection. Then the number that makes the section operational:

A source whose training-data content contains more than 5% illegal or harmful information should not be collected or used.

That is a source-rejection rule, not a filtering rule, and it is checked per source rather than across the corpus. Authorization evidence is required by source type: an open-source licence for open data; collection records for self-collected data, excluding anything the holder has marked uncollectable, non-crawlable web data, and PI where the individual refused; and enforceable contracts for purchased data — with a hard instruction not to use commercial data where the supplier cannot produce provenance, quality and security undertakings. Illegal training data found in the corpus must be taken out of use immediately.

Content. Before training: filter illegal and harmful information by keyword, classifier and human spot-check; verify accuracy and correct or delete what fails; identify IP infringement risk, maintain an IP strategy updated against third-party complaints, and delete infringing content; and identify personal information, confirming consent or another lawful condition. The corpus must also be relevant and sufficiently representative for the intended service, scenario and user behaviour — a data-quality test, not only a legality test.

Annotation. Following GB/T 45674-2025, functional and safety annotation get separate rulebooks, each covering at least four dimensions: annotation target, data format, annotation method, and quality metric. Annotator roles are defined; annotators are trained and examined before starting — on law, on the rules themselves, on tooling, on risk judgement and on data security — and re-assessed continuously. Confidentiality agreements are required where annotation touches personal information, important data, public data or trade secrets. Annotated output is sampled manually, automatically or both, with failures re-annotated or deleted. Where an automated annotation tool is itself a GenAI service, it must independently satisfy the GenAI rules.

The model

Design and development. The mechanism and rationale of the service must be articulated, the application scenario shown to be lawful, and the resource envelope specified. Ethics review is required, with escalated review for algorithms carrying public-opinion or social-mobilization capability and for highly autonomous decision systems in safety- or health-critical settings — the review track set out in the AI Ethics Review and Service Measures. Base models must come from a lawful, vetted source. Access to training data involving important data, personal information, trade secrets or public data requires authentication and authorization, with logs retained at least six months.

Third-party models — the rule foreign vendors should read first. Where a third-party model is used, its source must be lawful and its use within the terms of the agreement, model card or open-source licence. Then two bright lines:

  • accessed by API, the third-party model must be one that has completed regulatory filing (备案);
  • deployed locally — whether third-party open-source or otherwise — the model must be one developed by a domestic entity.

The practical effect is that a locally deployed foreign open-weights model does not clear this framework, however permissive its licence. Any Chinese buyer procuring a model for on-premises deployment will run into this, and it is better surfaced in term-sheet discussions than at assessment.

Compute. Infrastructure should in principle sit inside China. Where offshore deployment is genuinely necessary for the service, the data-export and export-control requirements of the circulation section attach. Hardware and network must be green and trustworthy, and the compute agreement must allocate billing, security, PI protection, IP and confidentiality responsibilities. DCC has covered how compute providers are classified in Compute centres and MaaS: five regulatory identities.

Validation and monitoring. Generation and refusal test banks are built to probe the Annex C risk taxonomy; pre-launch testing covers mechanism, content risk, effectiveness and quality. In operation: an update and upgrade security policy, periodic security audit of development frameworks and open-source code with audit working papers retained, and periodic sampling of generated content per GB/T 45654-2025 — with the second operational number:

Sampled generated content must reach a pass rate of at least 90%, content free of illegal and harmful information counting as a pass.

Fine-tuning must not degrade the safety mechanisms already in place.

The service

Labelling. Explicit labels go on text (start, end or an appropriate interior position, or prominently in the interface), audio (voice or rhythm cue), images, video (opening frame, around the player, end and interior), and virtual scenes (opening frame and during the session) — and must survive download, copy and export. Implicit labels go into the content and file metadata, with log information kept at least six months. Format and content follow GB 45438-2025 and the AI Content Labelling Measures. Malicious deletion, alteration, forgery or concealment of either label type is listed as a circulation-stage violation.

Identity. Users are authenticated by mobile number, ID document number, unified social credit code, or the national network identity authentication public service.

Minors. Where the service is not for minors, the user agreement and interface must say so. Where it is: age-banded spending limits with guardian confirmation for out-of-range purchases; age-banded limits on session timing, frequency and duration; guardian access to top-up and usage records with the ability to set limits including on content domains; and stricter content safety.

User data. Minimum-necessary collection; no unlawful retention of inputs and usage records that identify the user, and where retention is lawful, only for the shortest period necessary; no unlawful onward provision. And the one that shows up in every enterprise procurement negotiation:

Where user personal information, inputs or usage records are used for model training or optimization, the user must be expressly informed, must consent, and must be given a channel to opt out.

Security, and the enterprise-facing thresholds

The security dimension runs across four layers — data, model, system, use. Data security covers encrypted storage with access restricted to authorized personnel, periodic facility assessment, backups, encrypted transport with integrity verification, multi-factor access control with periodic entitlement review, classification and grading per GB/T 43697-2024, audit logging of all data operations, and irreversible destruction by hard or soft means.

Model security requires robustness testing against adversarial examples and data poisoning, vulnerability monitoring with periodic audit, scanning and adversarial testing, and a content-objection channel: acknowledge the objector, judge whether the content breaches law or public morality, and if so stop generation, stop transmission, remove, and remediate by retraining — then report back to the objector and, where warranted, to the regulator.

System security covers vulnerability management with a response plan, secure coding with pre-launch security review, third-party and open-source component review under GB/T 43848-2024, redundancy for fast recovery, and IDS/IPS with drills. The Guidelines ask for measurable indicators of control effectiveness — unauthorized access attempts, inference, bypass, extraction, exfiltration, provenance verification.

At the circulation stage the threshold that most often decides an architecture:

Systems processing important data, or the personal information of more than 1,000,000 people, should in principle meet Level 3 or above of the cybersecurity multi-level protection scheme.

Two further circulation-stage points bear on cross-border work: a GenAI service with public-opinion or social-mobilization attributes must have completed security assessment and filing before the trade; and export of model source code, object code or executable code is subject to export-control law — a licensing question distinct from data export, and easy to miss when the asset being sold is a model rather than a dataset.

Annexes

Annex C — risk taxonomy for training data and generated content. Five families: content contrary to core socialist values; discriminatory content across nine listed attributes; commercial violations including IP infringement, trade-secret disclosure and algorithm- or data-enabled monopoly and unfair competition; infringement of others’ lawful rights and interests, from portrait and reputation to privacy, personal information and voice; and failure to meet the safety needs of high-stakes service types — automated control, medical information, financial information, critical information infrastructure — where content is inaccurate against scientific consensus or simply unreliable.

Annex D — scenario-based violation catalogue. Mobile and smart devices (multi-user devices without permission management leaking trade secrets or household members’ personal information; continuous behavioural monitoring without separate consent). Automotive (unassessed use in cabin assistants, navigation and automated driving that misjudges driver intent, signals, pedestrians or vehicle positions before generating content or triggering commands). Financial (marketing without consent, failure of suitability obligations, unfiltered and misleading marketing content, failure to retain audio-video or system records when explaining risk, credit-scoring products deployed without internal testing and without explainable rules and traceable sources). Medical (patient records used as training data without anonymization and without separate consent; no ethics review; no professional review of clinical training-data accuracy; generated content leaking identifiable patient identity or reproducing unlicensed copyrighted medical material; no labelling to warn that output is not direct medical advice).

And a section DCC expects to be cited well beyond Shenzhen — AI agents: unlawfully using the pages, apps or user data an agent browses during service as training data; executing instructions that auto-fill CAPTCHAs, mass-grab tickets, brute-force passwords or blast email and SMS; reading other data on a user’s device without consent; and maliciously executing operations beyond the user’s instructed intent. Set against DCC’s coverage of the agent rules’ risk taxonomy, the convergence is notable.

Annex E — filings and licences. A single table mapping seven procedures to the subject, the deadline and the legal basis: ICP filing for non-commercial internet information services and a value-added telecom licence for commercial ones; mobile app filing before operating through an app; public-security filing within 30 days of network connection; algorithm filing within 10 working days of launch for recommendation services with public-opinion or social-mobilization capability; GenAI service filing within 10 working days on the same trigger; and security assessment before launch for internet information services with those attributes.

One item in the subject-side legality test is easy for foreign investors to miss: where the trading party involves foreign investment, its qualifications must also satisfy foreign-investment law — and internet news information services, internet culture operation (music excepted) and internet public information publishing are all on the foreign investment negative list. A foreign-invested entity may therefore be barred from holding the very permit its GenAI service requires.

How to use it

  1. Identify which of the three objects you are buying — data, model, or service — and scope diligence upward from there. This is the step most Western AI-procurement checklists get wrong against Chinese counterparties.
  2. Ask the model-provenance question early. API access requires a filed model; local deployment requires a domestically developed one. Both are architecture decisions, not paperwork.
  3. Pull Annex E and confirm which filings and licences the target actually holds, and whether foreign investment blocks any of them.
  4. Read Annex D for your sector. The violation catalogue is more concrete than the operative clauses and maps directly to diligence questions.

On the liability side, DCC has covered how Chinese courts are treating provider fault in Generative AI provider fault and statutory duties.

Companion document

The Shenzhen Data Exchange published a parallel guide to the same framework for cross-border trades: the Guidelines for Compliance Assessment of Cross-border Data Transactions. Deals that export training data or license a model across the border engage both. Both build on the China–Singapore Joint Data Compliance Guide.

Source

Original document: 《生成式人工智能服务交易合规评估指引》 (Guidelines for Compliance Assessment of Transactions Involving Generative Artificial Intelligence Services), Shenzhen Data Exchange, 55 pages, content current as of 9 February 2026. Guided by the Cyberspace Administration of Shenzhen and the Shenzhen Municipal Administration of Government Services and Data.

The Guidelines were filed as a Shenzhen local-standard project in the 2024 Shenzhen Local Standards Project Plan, and have been described by the Shenzhen Data Exchange as the first of their kind nationally, covering more than 500 risk-identification points across the chain from collection and annotation through to trading.

Normative references include GB/T 45654-2025 (basic security requirements for generative AI services), GB/T 45674-2025 (generative AI data-annotation security), GB 45438-2025 (labelling method for AI-generated synthetic content), GB/T 45288-2025 (large models, Part 1: general requirements), GB/T 37932-2025 (data transaction service security requirements) and the Shenzhen local standard DB4403/T 564-2024 (specification for compliance assessment of data transactions).

The document is a public-interest publication. It states that it is for informational reference only, is not to be used for commercial purposes, and must be attributed with a note of its public-interest character when quoted or circulated. DCC reproduces no full text; the above is distillation and commentary.

§ RELATED LAWS

See also.

§ COMMENTARY

Briefs on this law.

No briefs filed yet under this law.

§ SUBSCRIBE

The Monday brief.

One short email every Monday. New briefs on Chinese data-compliance rules from the previous week, with the source law cited.

Opt-in only. Unsubscribe anytime by replying "unsubscribe" to any issue.

SUPPORT DCC

Keep the publication free to read. Suggested support is $19.99, or choose your own amount.

Support →