Editor’s Note — DCC.
This is Part 3 of 洪延青 (Hong Yanqing)‘s four-part commentary on Beijing’s Several Measures on Accelerating Agent-Led Development (京发改〔2026〕1185号). Part 1 argued the Measures answer supply and leave adoption open; Part 2 made security governance the precondition of every other measure. Part 3 turns both arguments into an instrument: a five-level maturity scale for how far an agent has actually reshaped an enterprise process. Part 4 closes the series with the evidence chain for measuring the value those processes create.
For overseas counsel the scale is more than industrial-policy analysis. Levels 3 and up are where the compliance questions concentrate — agent identity, scoped permissions, human approval checkpoints, audit trails, takeover and rollback — and Hong’s insistence that maturity is measured by responsibility reallocation, not automation rate, is the same logic that runs through the national AI Agent Implementation Opinions’ decision-authority boundaries. Clients localizing agent products for China can use the scale directly: it predicts which deployments will attract benchmark-project scrutiny, what acceptance evidence will be demanded, and where a human decision-maker is likely to remain mandatory regardless of technical capability.
Article 3 of the Several Measures of Beijing Municipality on Accelerating Agent-Led Development encourages innovators to “restructure underlying architecture and interaction paradigms” around agents, calls on software enterprises to “restructure product architecture and service models,” supports sector-native solutions, and aims to “reshape core business processes, product forms, and service models.” The document proposes an AI operating system, AI-for-science assistants and “AI scientists,” and on-site co-creation by forward-deployed engineers (FDEs). Article 5 takes up the organizational changes agents bring; Article 6 proposes Agent-as-a-Service, Results-as-a-Service, and value-based billing.
All of it points one direction: an agent’s value should not stop at generating content or raising personal efficiency. It should enter enterprise processes and progressively change how enterprises complete tasks, allocate resources, and bear responsibility.
But “the enterprise uses an agent” and “the enterprise’s processes have been reshaped by agents” are not the same thing. Giving employees a general-purpose assistant, adding a smart feature to existing software, letting an agent complete a task segment independently, orchestrating a complete cross-departmental business, and redesigning the organization around a new human–agent division of labor are entirely different depths of application.
Without that distinction, enterprises equate connecting a model, launching an assistant, building a platform, or deploying several agents with “process restructuring” — and policy projects get accepted on agent counts, model call volume, and system complexity, with no way to tell what actually changed in how the enterprise runs.
Beijing’s push for agent-native applications and benchmark scenarios therefore needs a maturity framework for process reshaping. Its purpose is not to sort enterprises into advanced and backward, but to answer three questions: how much of a process does the agent currently carry; what has the enterprise changed in data, permissions, roles, and responsibility to support that; and what technical support, security requirements, and policy resources should come next.
Hong’s scale has five levels: tool assistance (工具辅助), step embedding (环节嵌入), bounded closed loop (有界闭环), end-to-end orchestration (端到端编排), and native restructuring (原生重构).
1. Maturity is measured by reallocated responsibility, not deployed agents
An enterprise’s agent maturity cannot be read off model capability, agent count, automation rate, or architectural complexity.
Ten agents are not necessarily more mature than one. A multi-agent system that needs constant human correction, redundant calls, and error-passing between components may deliver less than a cleanly structured single-agent workflow. Conversely, a system that reserves final judgment to a qualified professional is not thereby less mature: for high-risk settings — medicine, government affairs, finance, industrial control — a human-machine division with clear boundaries, complete evidence, and stable operation is often more mature than nominal “full autonomy.”
Judging how far a process has been reshaped takes at least five observations:
- Has the task unit changed? Does the agent help one employee with one operation, complete a process segment continuously, or coordinate systems and departments around a final business outcome?
- Have data and system permissions changed? Does it process only what the user pastes in, or can it read enterprise databases, call business software, modify system state, control terminal devices?
- Has the human role changed? Do people still execute step by step with the agent as aid — or set goals, write rules, approve key items, and handle exceptions?
- Has the evaluation unit changed? Is the project measured by user counts, call volume, and output quantity — or task completion rates, business cycle times, customer outcomes, and cost per valid completed task?
- Have organization and accountability adjusted? Is the agent stacked onto existing positions, or has the enterprise re-drawn process owners, data permissions, approval nodes, and departmental duties?
The essence of process maturity is thus not how autonomous the agent looks technically, but how much process responsibility the agent can carry under controlled conditions — and whether the enterprise has reconfigured people, systems, permissions, metrics, and accountability around that shift.
By that standard, agents enter enterprises through a progression.
2. Level 1 — Tool assistance: the person in the process, the agent outside it
At the first level the agent is mainly a personal productivity tool: retrieval, summarization, drafting, translation, coding, diagramming, first-cut proposals.
Strictly, part of this product class is still closer to a generative AI assistant than to an agent that executes tasks autonomously. It typically does not enter the enterprise’s formal business systems and does not change business state on its own. The employee feeds material in, takes the output, and carries it back into the original process by copying, editing, keying in, and submitting.
An employee drafts a contract with an agent — but the contract still runs the existing business, legal, and sign-off route. A researcher uses an agent for literature search — but experiment planning, data recording, and project management are unchanged. A support rep generates suggested replies — but confirms and sends each one manually. A programmer generates code — but it ships through the existing develop-test-release process.
The signature of this stage: the way individuals work has changed; the enterprise’s processes have not.
Article 1’s push on model reasoning, long context, and real task completion, and Article 3’s AI-for-science assistants and intelligent upgrades to basic and transactional software, will produce many applications that start here. Article 8’s low-cost standardized computing for SMEs and one-person companies — computing vouchers, Token vouchers, agent service vouchers — lowers the entry price.
Tool assistance has real value: it builds usage habits, validates model capability, surfaces candidate scenarios, and raises knowledge-work efficiency. Its limits are equally clear: outputs still travel by hand; the agent never holds full business context; different employees re-enter the same information; nothing accumulates into enterprise-level process capability.
Evaluation at this stage should target personal and single-task effect: processing time, output quality, less repetitive labor, reasonable cost. Large account counts and heavy Token consumption do not establish process restructuring. Policy support should stay correspondingly light — training, vouchers, general-purpose tools, basic compliance guidance — without demanding that every level-1 enterprise build an agent platform, a complex data middle platform, or a multi-agent system.
3. Level 2 — Step embedding: the agent enters positions and single workflow nodes
At the second level the agent stops being an external personal tool and enters the software, positions, or specific business steps the enterprise already runs.
It can usually read defined business data, process inputs by fairly clear rules, and produce structured outputs that flow directly to the next step. It may propose suggestions, assemble items for approval, or auto-complete low-risk operations within limits. But hand-offs between workflow nodes still run through people; the agent does not yet own a complete task.
The finance system extracts invoice data, checks reimbursement rules, and drafts review conclusions. The CRM digests communication records, flags sales opportunities, and generates follow-up tasks. The service desk classifies tickets and queues draft replies. Industrial software proposes parameter adjustments against design requirements. Security software aggregates alerts, correlates them, and suggests dispositions. A research system organizes experimental data and drafts the next protocol.
The signature: the agent is inside the process, but carries one position or one node.
Article 3’s support for intelligent technical upgrades of basic software, industrial software, transactional software, and new security software — and its call for software enterprises to restructure product architecture and service models around agents — lands mainly here. Article 4’s embedding of agent capability in phones, glasses, earphones, robots, and vehicles may also first appear as node-level enhancement of single functions.
Unlike level 1, level 2 forces the real enterprise-adaptation problem. Generic model capability does not automatically understand an enterprise’s business fields, approval rules, role divisions, and exception types. The Measures’ forward-deployed engineers — on-site co-creation, continuous iteration feeding back into agent capability — exist to close exactly this gap between general technology and specific business.
But FDEs cannot stop at model tuning and interface plumbing. Stable embedding requires jointly fixing the node’s input conditions, output formats, data scope, error boundaries, human-review rules, and responsible parties. Otherwise the agent is connected to the software but still produces semi-finished output no one can use directly.
Evaluation should shift from personal use to the process node: is average processing time down, first-pass rate up, duplicate entry reduced, human review shrinking, and are errors caught and corrected? An agent that only helps employees produce material faster — without reducing downstream review, entry, and coordination costs — has not maturely transformed the node.
4. Level 3 — Bounded closed loop: the agent completes a process segment on its own
At the third level the agent stops processing isolated nodes and continuously completes multiple steps around a clearly bounded task.
It typically holds an identifiable identity; obtains scoped data and tool permissions; manages task state; and, within defined limits, modifies business systems or executes operations. People no longer perform each action, intervening instead at key approval points, exceptions, and high-risk steps.
An after-sales agent receives a claim, verifies the order, checks return-exchange conditions, assigns responsibility, generates a resolution, and completes the refund within its amount limit. A procurement agent collects supplier materials, runs formal review and risk screening, sends deficiency notices, and forwards qualifying items to the procurement owner for decision. An equipment-maintenance agent reads telemetry, spots anomalies, pulls repair history, drafts a service plan, and schedules the work order. A research agent adjusts parameters from results, drives the next experimental round on the instruments, and calls for human help when budget or anomaly limits trip.
The signature: people stop operating the process step-by-step, and instead set boundaries, approve key actions, and handle exceptions.
Article 1’s complex reasoning and planning, tool calling, memory, and multi-agent collaboration, and Article 2’s context engineering, task persistence, system scalability, long-horizon execution stability, and secure-controllable technology stack, are largely aimed at making bounded closed loops possible. Without stable task-state management the agent cannot resume after interruption; without tool calling it can only advise; without memory and context management it loses the task’s conditions; without permission, budget, retry, and termination mechanisms the “closed loop” degenerates into abnormal loops, duplicate operations, and boundary violations.
Level 3 is where the security governance discussed in Part 2 of this series becomes a production precondition. The enterprise must fix whom the agent represents, what tasks it accepts, which data it may access, which tools it may call, what state it may modify, and when it must stop or seek renewed approval.
“Bounded” does not mean every step is human-confirmed. It means automatic execution is confined to defined tasks, counterparties, time limits, amounts, and operation types. Low-risk, repetitive, reversible actions run continuously; high-risk actions — payment, contracting, external data transmission, permission changes, equipment control — carry special approval requirements set according to risk.
Evaluation at level 3 targets the complete task: completion rate, first-pass success rate, average completion time, human-takeover rate, abnormal retries, cost per valid completed task, and recovery after error.
Article 6’s shift from Token-consumption billing to value-based billing only becomes realistic from this level. Once an agent completes a relatively whole task, what the enterprise buys is no longer model calls and software features but a verifiable task result.
5. Level 4 — End-to-end orchestration: the agent reorganizes systems and departments around outcomes
At the fourth level the agent carries not a local process segment but a complete business outcome, coordinating multiple systems, departments, and roles.
It may connect CRM, contracts, procurement, inventory, production, finance, and logistics; track task state continuously; and adjust plans as reality shifts. The human role moves further from execution and item-by-item review toward goal setting, rule making, exception handling, and final responsibility.
An order-fulfillment agent does not merely generate the order: it verifies customer requirements, checks inventory, schedules production, coordinates procurement, tracks logistics, handles delay, and reports back to the customer. An R&D agent coordinates requirements analysis, solution generation, simulation, materials selection, costing, and production validation. An enterprise-services agent organizes a complete response to a customer problem across sales, support, finance, and after-sales. A manufacturer’s agent re-plans production and supply-chain arrangements against order, equipment, material, and staffing state.
The signature: the agent no longer belongs to a position — it reorganizes positions and systems around the business outcome.
Article 3’s enterprise- and sector-level AI operating systems — breaking ecosystem walls, raising cross-device, cross-ecosystem, cross-scenario capability — and its agent-driven restructuring of electronics R&D and production point mainly here. Article 2’s cross-model and cross-chip middleware stack, multi-agent collaboration, interconnection protocols, and development frameworks matter most at this level too: no single model connects every data source, tool, and system, so orchestration, permission, and runtime management have to organize many components into one task.
But multi-agent architecture is not a requirement of level 4 — and must not become a proxy for project sophistication. A task that one agent plus a structured workflow completes stably should not be split up to manufacture “swarm intelligence.” Maturity turns on whether the complete business improved, not on how many agent roles the system contains.
Past level 4 the binding constraint is organizational, not technical. (For the analogous exercise on the government side, Hong points to his own piece on rebuilding government capability and service delivery around an agent-native architecture, using Guangzhou’s Nansha district as the example.) Conventional enterprises set targets by department: sales watches signings, production watches output, procurement watches price, logistics watches delivery, finance watches cost. Orchestrate end-to-end and those local metrics start colliding; local efficiency gains stop implying whole gains. Faster order generation that piles up production and delivery; procurement optimization that raises quality and delivery risk; shorter service calls that raise repeat complaints — none of that is maturity.
Level 4 therefore requires naming an end-to-end process owner, re-coordinating departmental KPIs and data permissions, and lifting the evaluation unit to the complete outcome: order cycle time, customer wait, first-delivery success, inventory turns, cross-departmental hand-offs, recovery time. What the Measures call “reshaping core business processes” is precisely this — not raising each position’s intelligence, but removing dead hand-offs, duplicate approvals, and departmental walls so the enterprise runs one continuous process around the outcome.
6. Level 5 — Native restructuring: the enterprise reorganizes around the new division of labor
At the fifth level the enterprise stops taking existing positions, departments, and software as the starting point. It first fixes what outcome it must deliver to customers or society, then decides afresh which tasks people perform, which agents execute, and which require joint human-machine judgment.
Existing processes formed under conditions of limited human processing capacity and costly information transfer: hierarchy, departments, approvals, reports, and layers of intermediate roles all exist to preserve control and coordination. Once agents can continuously process information, call systems, manage tasks, and coordinate resources, part of that inheritance is no longer optimal.
Native restructuring is therefore not adding automation nodes to existing processes — it is redesigning the process itself.
Traditional software ships functions by department — finance, HR, sales, procurement. Agent-native software may organize capability directly around “deliver one project,” “keep one small business running,” or “meet one customer’s complete need.” A traditional research pipeline moves serially through people — literature, design, operation, analysis; an autonomous laboratory organizes those steps continuously around the research goal, with scientists concentrated on problem formulation, scientific judgment, and exception handling. Traditional professional services bill hours and headcount; Results-as-a-Service re-assembles people, models, tools, and responsibility around one verifiable task.
The signature: not the agent adapting to the enterprise, but the enterprise reorganizing around the new human–agent division of labor.
Article 3’s “native AI software and super software driven by demand intelligence” and its self-discovering, self-validating research paradigm correspond mainly to this level. So does Article 5: an AI-augmented one-person company is not one person with more tools — it is the tasks of many former roles recombined into one accountable natural person plus a set of agents. Article 6’s Agent-as-a-Service and Results-as-a-Service change the trading unit accordingly: the customer no longer buys licenses, model calls, or work hours but a defined, verifiable, attributable business result — with the vendor absorbing more orchestration, quality control, and risk management. Article 8’s data flywheel only becomes a process-improvement mechanism at these higher levels — task results, takeover records, anomalies, and customer feedback continuously reshaping workflows, rules, and product design — and even then through version control, regression testing, and responsibility review, never “autonomous evolution” of a production system without governance. Article 9’s open protocols, frameworks, operating systems, and internationalization matter most here too: at level 5 a process may cross enterprise boundaries, and open interfaces and portability decide whether the enterprise avoids lock-in to a single model, tool, or platform.
Native restructuring does not mean pursuing the unmanned enterprise. Human value moves from repetitive entry, information transport, and procedural execution toward goal setting, professional judgment, relationship coordination, rule design, exception handling, and bearing responsibility. Matters touching major rights and interests, public power, and professional qualification still end in a named, accountable person’s final decision.
Nor can level 5 be evaluated by headcount reduction. What counts is whether the enterprise forms new customer value, product forms, revenue sources, and cost structures; whether it can operate sustainably; and whether business continuity survives changes of model and vendor.
7. Five levels — but not a race to maximum automation
The framework’s most likely misreading is as a ladder every enterprise and every process must climb.
In fact, different tasks have different legitimate endpoints. Document processing, retrieval, and internal knowledge lookup can move quickly from tool assistance to bounded closed loop. Customer service, supply chain, equipment maintenance, and standardized administration can reach end-to-end orchestration when conditions allow. Medical disposition, administrative decisions, financial signing, and educational evaluation — matters of professional judgment, public power, or major rights — may properly retain a human final decision-maker indefinitely, however capable the technology.
That does not make those processes immature. A system that already carries information gathering, evidence assembly, rule checking, solution generation, and process recording — and hands a qualified person a complete, verifiable basis for decision — can rank very high on process maturity even though a human decides. Conversely, a system that completes tasks automatically but with unclear goals, over-broad permissions, unmanageable exceptions, and no one accountable for results is not mature, however little human involvement it needs.
Maturity is therefore not automation rate. The real standard: under the relevant business and risk conditions, has the enterprise formed the most effective, stable, controllable division of labor between people and agents?
Maturity should also be assessed per process, not stamped on a whole enterprise. The same company may run document processing at level 3, customer service at level 4, and major investment decisions at level 2 or 3 — simultaneously, and correctly.
Nor does maturity equal technical complexity. The Measures support world models, swarm intelligence, multi-agent collaboration, and AI operating systems; these may underpin complex processes, but they should not become required forms. Policy should stay technology-route-neutral, judging maturity by task effect and process change — never inviting the inference that more agents and heavier architecture mean a more advanced project.
8. Allocating scenarios, funding, and governance by maturity level
Read against the five levels, the Measures’ provisions each serve different stages of process reshaping.
Articles 1–2 supply the model, harness-layer, and middleware capability that carries enterprises from tool assistance and step embedding toward bounded closed loops and end-to-end orchestration — capability that converts into maturity only when it enters a concrete process and carries its responsibility.
Article 3 is the core process-reshaping provision: general assistants and software upgrades serve levels 1–2; FDE co-creation carries enterprises from step embedding to bounded closed loop; sector-native solutions and the AI operating system aim at levels 4–5.
Article 4 extends reshaping to terminals and the physical world. Once agents sense environments and control vehicles, robots, and production equipment, maturity evaluation must add physical-state change and physical safety; terminal projects need grading of permissions, abnormal termination, and human takeover — not just model performance and sales.
Articles 5–6 reflect the organizational and business-model changes of high maturity. OPCs, Agent-as-a-Service, and Results-as-a-Service are concept packaging unless built on mature task definition, process evidence, and responsibility mechanisms. Tokens and call volume can measure resource input; from level 3 upward, policy evaluation should shift to task results and business value.
Article 7’s security governance runs through all five levels, tiered: level 1 mainly governs input data and output use; level 2 adds system access and business data; level 3 requires identity, permissions, runtime control, and human takeover; level 4 adds cross-system, cross-department responsibility and end-to-end recovery; level 5 adds organizational responsibility, external collaboration, continuous updating, and systemic impact.
Article 8’s computing, data, and talent support should follow maturity too: level-1 enterprises rarely need dedicated infrastructure; levels 3–4 raise real demands on latency, stable scheduling, data isolation, task persistence, and continuous operation. “Token factories” should be judged on effective load from real tasks and cost per valid task, not nominal capacity.
Article 9’s open source and interoperability lower the cost of switching models, tools, and vendors — most important at levels 4–5, where processes cut deepest. Publicly funded platforms and sector systems should not become ecosystems that admit only their own agents.
Article 10’s key projects and fiscal support are the direct lever for making the framework operational. Beijing can require scenario-demand lists to state not “build an agent” or “connect a large model” but the specific process to be transformed, its current baseline, the target maturity level, the task scope the agent will carry, and the expected business result. Level-1 projects get popularization, training, and cheap tools. Level-2 projects get software upgrades, data preparation, standard interfaces, and real-business validation, judged by node improvement. Level-3 projects get FDE co-creation, production pilots, tool access, permission management, and security testing, accepted on whole-task success, cost, and stability. Level-4 projects get help breaking cross-system and cross-department walls, must name end-to-end process owners, and are judged on complete business outcomes rather than local departmental metrics. Level-5 projects pilot native software, OPCs, outcome billing, liability insurance, sustained data use, and cross-organization collaboration — few in number, but each proving sustainable operation and non-subsidy revenue.
Benchmark scenarios would then be classified not only horizontally by industry — healthcare, education, government, manufacturing, culture — but vertically by depth of process transformation, letting policymakers tell a personal tool application from a software upgrade, a local closed loop, an end-to-end reconstruction, or a native organizational change, and match funding, evaluation, and security requirements accordingly.
9. From “more enterprises using agents” to “more processes substantively changed”
Agents transform enterprises not in one jump from “no agent” to “agent deployed,” but along a deepening path: first people work in existing processes with agents as external tools; then agents enter positions and software nodes; then they complete bounded process segments; then they coordinate departments and systems around business outcomes; finally the enterprise redesigns process, organization, and business model around the new human–agent division of labor.
The five levels are not product generations, and not a universal endpoint. They are a measuring rod: how much has the process actually changed, what responsibility does the agent carry, and have people, data, permissions, and evaluation adjusted with it.
The Measures already contain the forward-looking policy vocabulary — native applications, process restructuring, the AI operating system, FDEs, OPCs, Results-as-a-Service, the data flywheel. The next step is converting those concepts into identifiable, evaluable process maturity, so that each project receives support and governance matched to its actual depth.
The goal of Beijing’s agent policy should not be more enterprises owning agents, connecting models, or burning Tokens. It should be more enterprises moving, within appropriate risk boundaries, from personal tool use to stable process capability — and finally to verifiable, replicable, sustainable business results.
Whether agents have truly reshaped an enterprise is measured not by how many agents it deployed, but by how much its processes for completing tasks, its division of labor, its responsibility structure, and its way of creating value have actually changed.
This brief translates and condenses 洪延青 (Hong Yanqing), “智能体如何逐级 重塑企业流程——评《北京市关于加快智能体引领发展的若干措施》之三,” published on the 网安寻路人 WeChat Official Account on 6 August 2026 (original). Details of the Measures (document number, issuing bodies, dates, and article structure) are taken from the official text published at beijing.gov.cn on 23 July 2026. Parts 1, 2, and 4 of the series are translated in separate DCC briefs.
— Not legal advice.