Skip to content
DCC · DATA COMPLIANCE CHINA China data law, for overseas counsel.
§ LAW · GB/T 47949

Asset Management — Classification and Codes for Data Assets (GB/T 47949-2026).

资产管理 数据资产分类与代码 (GB/T 47949-2026)

FILED UNDER · Data Economy

DCC summary, not a translation. GB/T 47949-2026 is a copyrighted national standard. The structured summary below is DCC’s own paraphrase of the standard’s framework, together with the classification code table, prepared for overseas compliance teams who need to understand how Chinese organizations are expected to classify data held as an asset.

Standard number: GB/T 47949-2026 (recommended national standard). ICS / CCS: 35.040 / A 24. Issued: 2 July 2026. Effective: 1 September 2026. Issued by: State Administration for Market Regulation and the Standardization Administration of China. Proposed and administered by: National Technical Committee on Asset Management Standardization (SAC/TC 583). Drafting units: Asset Management Department of the Ministry of Finance; China National Institute of Standardization.


Scope

GB/T 47949-2026 establishes the classification principles for data assets, specifies the coding method and code structure, and sets out the classification-and-code table.

It applies to data assets that are managed as intangible assets, for the purposes of asset configuration (配置), registration (登记), inventory-taking (清查) and reporting (报告).

The standard carries an express limiting note: the classification and codes focus on the basic attributes of data assets in order to meet the needs of asset informatization management. They do not alter the definition or classification of data-asset accounting subjects under the existing accounting-standards regime. In other words, this is an asset-register taxonomy, not an accounting-recognition rule — the recognition question continues to be governed by MOF’s Interim Provisions on the Accounting Treatment of Enterprise Data Resources (财会〔2023〕11号) and the related data-asset management notices.

Two standards are normatively referenced: GB/T 14885 (basic classification and codes for fixed assets and other assets) and GB/T 33172 (asset management — overview, principles and terminology).

The definition of “data asset”

The standard adopts a single defined term (Clause 3.1):

Data asset (数据资产) — a data resource that is formed by an organization’s past transactions or events, is lawfully owned or controlled by the organization, and is expected to generate service potential or bring an inflow of economic benefits.

This is the MOF-lineage definition — the same three-limb test (past transaction or event / lawful ownership or control / expected service potential or economic benefit) that runs through China’s data-asset management policy. It is worth noting what the test does not require: it does not require that the data be tradable, that data property rights have been registered, or that the asset already sit on the balance sheet. The registration and confirmation mechanics are handled by GB/T 47950; the property-rights layer is handled by the Data Property Rights Registration Work Guide (Trial).

Classification principle and code structure

Classification principle (Clause 4). Classification follows the principles of GB/T 14885 and proceeds by the basic attributes of the object — that is, by what the data technically is, not by what sector it comes from or what it is used for.

Coding method (Clause 5). The standard follows GB/T 14885’s line classification method (线分类法) and hierarchical coding, dividing data-asset classification codes into five levels: division (门类), major class (大类), medium class (中类), minor class (小类) and sub-class (细类).

A code consists of one uppercase Latin letter plus ten Arabic digits. The leading letter is A, from “Asset”. The fixed prefix positions are:

PositionLevelValueMeaning
LetterPrefix identifierAAsset
Digits 1–2Division (门类)08Intangible assets
Digits 3–4Major class (大类)06Information-and-data intangible assets
Digits 5–6Medium class (中类)02Data
Digits 7–8Minor class (小类)variesData-asset minor class
Digits 9–10Sub-class (细类)variesData-asset sub-class

So every data asset in China’s asset-classification system now sits under the block A0806020000.

One drafting convention matters in practice: 99 in the sub-class position is the catch-all (收容项), used for data assets not yet enumerated in the table. Each of the three minor classes therefore terminates in an “other …” sub-class.

The classification code table (Clause 6, Table 1)

The table divides data assets into three minor classes by data structure, each with its own sub-classes and prescribed units of measure. English category names below are DCC’s rendering; the units are as prescribed.

A0806020100 — Structured data. Data assembled from data elements where every record shares the same structure and can be effectively described by a relational model.

CodeCategoryUnit of measure
A0806020101Database table — the basic storage unit of a relational database, used to store same-type data in structured formMB, GB, record count
A0806020102Spreadsheet — tabular data with a single header row, homogeneous columns and a contiguous data region, with no embedded objectsMB, sheet
A0806020103Delimiter-separated data — in-row delimited data (CSV, TSV, fixed-width text, etc.)MB, record
A0806020199Other structured dataMB

A0806020200 — Semi-structured data. Data models that do not conform to a relational database or other data-table form, but that contain related markup used to separate semantic elements and to impose a hierarchy on records and fields.

CodeCategoryUnit of measure
A0806020201Markup-language data — general markup data (XML, HTML, etc.) and lightweight markup data (JSON, YAML, Markdown, etc.)MB, copy
A0806020202Log-file data — fixed fields per line, but with extensible field length and variable contentMB, item
A0806020299Other semi-structured dataMB

A0806020300 — Unstructured data. Data with no predefined model, or not organized in a predefined manner.

CodeCategoryUnit of measure
A0806020301Text data — plain text (TXT etc.) and rich text (RTF etc.)MB, copy
A0806020302Image data — raster images, vector graphics, etc.MB, image
A0806020303Audio data — speech, music, ambient sound effects, etc.MB, track
A0806020304Video data — file-based video, streaming video, etc.MB, title
A0806020305Compound-document data — office documents, e-books, etc.MB, item
A0806020306Spatial-relationship data — spatial vector data files, spatial raster data files, three-dimensional terrain data files, building-information-model data files, etc.MB, item
A0806020307AI-training multimodal data — multimodal pre-training and fine-tuning data used to train artificial-intelligence models, including image-text pairs, video-text pairs, audio-text pairs and other multimodal aligned samples, as well as synthetic labelled data and multi-source heterogeneous dataMB, token
A0806020399Other unstructured dataMB

The A0806020307 entry is the one to flag. It is, so far as DCC is aware, the first time a Chinese national asset-classification standard has given AI training corpora a dedicated asset code — and it is measured in tokens as well as megabytes. For any organization capitalizing training data, this is the code line its asset register will use, and the token count is the quantity the register expects.

Extending the table (Clause 6 and Annex A)

Data-asset class divisions follow the classification dimensions of GB/T 14885 and may be extended and refined according to management needs and practical application; the principles and methods for extension and mapping are given in Annex B of GB/T 14885-2022.

Annex A (informative) adds two points of practical guidance:

  • The main classification code is divided according to the basic attributes of the data asset (structured / semi-structured / unstructured).
  • Organizations may add other auxiliary codes as needed, forming a multidimensional classification-coding system for data assets. This is the hook for combining the structural code with, for example, industry codes under GB/T 4754.

The annex gives four worked examples of code assignment: crop-species monitoring data stored as database tables (A0806020101); XML-format publication data produced by structurally processing books in the publishing industry (A0806020201); animation video source material modelled and rendered by an animation studio (A0806020304); and approval-result data — certificates, licences and official replies — stored and transmitted by administrative organs as compound documents (A0806020305).

How it fits the regime

GB/T 47949 is infrastructure, not obligation. It creates no compliance duty of its own; what it does is give the data-asset policy stack a stable identifier layer it had been missing.

The policy stack above it is the Ministry of Finance’s: the Guiding Opinions on Strengthening Data Asset Management (财资〔2023〕141号), the Interim Provisions on the Accounting Treatment of Enterprise Data Resources (财会〔2023〕11号), and the Notice on Strengthening Data Asset Management of Administrative Institutions (财资〔2024〕1号) — all of which direct organizations to inventory, register and report data assets without saying what codes to use. The Regulations on the Administration of State-Owned Assets of Administrative Institutions (State Council Order No. 738) supplies the statutory frame for the public-sector half of that population.

The standard’s immediate companion is GB/T 47950-2026, the data-asset registration guidance, which normatively references this standard: an organization registering a data asset takes the asset classification and the unit of measure for its data-asset card from this table. The two were issued the same day by the same technical committee and take effect together on 1 September 2026.

Read alongside the Data Property Rights Registration Work Guide (Trial), the picture is a two-track registration architecture that overseas counsel should not conflate: the NDA track registers rights in data (hold / use / operate) and issues a property-rights certificate; the MOF/SAC track registers data as an asset on the organization’s own books. The classification code here belongs to the second track. DCC’s briefs on identifying and defining data assets and on what a data-asset ABS actually securitises work through why the distinction between accounting recognition (入表) and rights confirmation (确权) keeps causing trouble in practice — this standard sits squarely on the 入表 side of that line.

§ RELATED LAWS

See also.

§ COMMENTARY

Briefs on this law.

1 brief references this law.

  • § 01 · DATA-ASSETS

    Two Registrations, One Word: China's New Data-Asset Standards and the Line Between 登记 and 登记

    On 2 July 2026 China issued two national standards for data as an asset — GB/T 47949-2026 (classification and codes) and GB/T 47950-2026 (registration guidance) — both effective 1 September 2026. They give data assets a fixed place in the asset-classification code system (block A0806020000, including a first-ever asset code for AI-training multimodal data measured in tokens) and a step-by-step model for putting data on an organization's own books. The trap for overseas counsel is the word 登记 (registration): these MOF/SAC standards register data as an asset internally, while the National Data Administration's Data Property Rights Registration Work Guide (Trial), finalized 1 July 2026, registers rights in data externally through a certificated institution. Same word, two regimes, two artifacts, two purposes. This DCC brief separates them, reads the two standards for what they require, and explains why the 入表 (balance-sheet entry) vs 确权 (rights confirmation) distinction keeps tripping up data-asset deals.

    data-assets · data-property-rights · data-registration
§ SUBSCRIBE

The Monday brief.

One short email every Monday. New briefs on Chinese data-compliance rules from the previous week, with the source law cited.

Opt-in only. Unsubscribe anytime by replying "unsubscribe" to any issue.

SUPPORT DCC

Keep the publication free to read. Suggested support is $19.99, or choose your own amount.

Support →