Skip to content
DCC · DATA COMPLIANCE CHINA China data law, for overseas counsel.
§ 093 · AI-AGENTS

Why Would an Enterprise Dare Hand Tasks to an Agent? Hong Yanqing on Security Governance in Beijing's Agent Measures (Part 2 of 4)

Article 7 of 京发改〔2026〕1185号 carries the security agenda: graded-and-categorized agent regulation, a trusted sandbox, security ranges, 'using models to govern models.' Hong's argument is structural — security governance cannot live in one article out of ten, because every other article (self-evolving models, task persistence, long-term memory, tool calling, Token pricing, data flywheels, open source) presupposes it. The unit of agent security is not the model but the delegation.

Part 2 of Hong Yanqing's commentary on the Several Measures of Beijing Municipality on Accelerating Agent-Led Development (北京市关于加快智能体引领发展的若干措施, 京发改〔2026〕1185号). The Measures assign security governance to Article 7 — graded-and-categorized regulation, regularized crackdowns on malicious misuse, AI industry legislation, security-service platforms, ranges, and a trusted sandbox. Hong argues security cannot be one measure among ten: an enterprise that adopts an agent is not buying content-generation software but delegating tasks, data, system permissions, and the power to act externally to a technical system, and that delegation only continues if the agent's action boundary can be limited, its running state observed, its abnormal behavior halted, its errors remedied, its key steps traced, and its final responsibility assigned. He walks the other nine articles showing how each presupposes this 'trusted delegation' — self-evolution needs version governance and rollback; task persistence needs budget caps, retry limits, and human takeover; long-term memory is data processing and data residency, not a free 'data flywheel'; tool calling turns identity and permissions into the core problem; multi-agent skill markets stretch the responsibility chain — and proposes four foundational institutions: action-and-consequence-based agent classification, a trusted-delegation baseline for production agents, security capability as public infrastructure, and security evidence as a condition of fiscal support, procurement, and benchmark-scenario acceptance. Security governance, he concludes, is itself a form of productive capacity: it is what makes enterprises willing to open data, systems, and permissions at all.

Editor’s Note — DCC.

This is Part 2 of 洪延青 (Hong Yanqing)‘s four-part commentary on Beijing’s Several Measures on Accelerating Agent-Led Development (京发改〔2026〕1185号, issued 21 July 2026). Part 1 mapped the gap between agent supply and enterprise adoption; Part 3 builds a five-level maturity scale for agent-reshaped processes; Part 4 supplies the value-measurement evidence chain that turns both into acceptance criteria.

Part 2 is the installment most directly relevant to data-compliance practice. Hong’s “trusted delegation” frame — in whose name the agent acts, what data and tools it may touch, which actions auto-execute, which require renewed approval, who can halt it, who answers for the result — is the same architecture the national AI Agent Implementation Opinions sketched in May (decision-authority boundaries, agent identity, tiered governance) and that TC260’s agent-deployment security guidelines turned into a lifecycle checklist in July. Read together, the three documents show a consistent direction of travel: Chinese agent governance is converging on delegation, not generation, as the unit of regulation — and Hong’s memory-as-data-processing and execution-data points preview compliance questions PIPL and DSL counsel will be asked about agent deployments well before any dedicated statute arrives.

The Several Measures of Beijing Municipality on Accelerating Agent-Led Development places “comprehensively raising security-governance capability” in Article 7: exploring graded-and-categorized regulation of agents (智能体分级分类监管), regularized crackdowns on malicious misuse, accelerating AI industry legislation, and building security-service platforms, security ranges (靶场), a trusted sandbox (可信沙箱), and security-testing capability. The direction is clear, and it shows how seriously Beijing takes agent security.

But by the internal logic of the industry, security governance cannot be understood as one measure among ten, or carried mainly by Article 7. The online learning, self-evolution, and long-term memory of Article 1; the tool calling, task persistence, and skill marketplace of Article 2; the business-process restructuring of Article 3; the smart terminals and physical devices of Article 4; the AI-augmented entrepreneurship of Article 5; the agent services and outcome-based pricing of Article 6; the data flywheel and public computing power of Article 8; the open-source community and overseas expansion of Article 9 — all of it rests on security governance as a precondition.

The reason: an enterprise that adopts an agent is not merely buying software that generates content. It is delegating tasks, data, system permissions — and, within some scope, the power to act externally — to a technical system. Whether the agent keeps receiving that delegation depends on whether it can demonstrate that its action boundary can be limited, its running state observed, its abnormal behavior halted, its erroneous results remedied, its key steps traced, and its final responsibility assigned.

In that sense, security governance is not a restraint bolted on after the industry matures. It is the common foundation on which model capability, application scenarios, business models, data utilization, and an open ecosystem can stand at all.

1. Agents move the security question from “what is generated” to “acting on whose behalf”

Agents are built on generative AI, but their industrial significance goes beyond generation.

Ordinary generative AI supplies text, images, code, or plans; the user sees the result and can still decide whether to use it. An agent enters the execution of the task itself. It decomposes goals, makes plans, reads data, calls tools, connects systems, and takes successive actions based on running state. Generating content is one step; what matters is whether the agent then sends the file, modifies the account, submits the filing, schedules production, initiates payment, or operates the device.

The object of security governance changes accordingly — in four ways.

From model output to real-world action. Whether an agent is safe is not only whether it generates unlawful, harmful, or wrong content, but whether it has authority to perform a given operation and whether that operation remains within the user’s or enterprise’s authorization.

From single interactions to continuous operation. A generation tool returns one result per request. An agent may run for hours or longer, completing a task across many steps. The longer it runs and the more intermediate states accumulate, the likelier its actions drift from the original goal.

From the model to the whole system. An agent’s actual behavior is determined not only by the base model but by context, long-term memory, workflow, external tools, identity permissions, plug-in components, business rules, and the terminal environment. Proving the model passed testing does not prove the agent system is safe once inside a real process.

From after-the-fact correction to runtime control. A user can simply decline to use a wrong paragraph. A wrong payment, a wrong data transfer, a wrong equipment operation may be unrecoverable once discovered. Agent governance must move the control point to before the action and into the run.

The core of agent security is therefore not a blanket “human in the loop,” and not the user re-confirming every intermediate step. It is a trusted delegation mechanism (可信委托机制): clarity on whom the agent represents, what task it carries, which data and tools it may obtain, what impact it can produce, which actions execute automatically, which situations require renewed approval, when it must pause or terminate, and who answers for the final result.

Only when that delegation can be clearly defined and continuously enforced does an agent turn from a generation tool into a production tool an enterprise can rely on.

2. Articles 1–2: technical capability becomes production capability only when controllable

Article 1 of the Measures proposes raising models’ online learning, continual learning, and self-evolution; optimizing ultra-long-horizon task, complex-reasoning, and planning algorithms; and developing world models, swarm intelligence, long context, tool calling, multi-agent collaboration, and memory. Article 2 proposes harness-layer engineering (驾驭层工程) — strengthening context engineering, task persistence, multi-agent collaboration, system scalability, and long-horizon execution stability — plus an agent development platform, skill marketplace, software store, and a secure, controllable technology stack.

These are the right technical priorities. They also show that security has moved inside agent capability, not outside it.

Online learning and self-evolution let an agent adjust behavior on new data and feedback — which also means a tested, approved system can change after deployment. Without data review, version management, regression testing, and rollback, an enterprise cannot confirm that today’s system is still the one it evaluated. In high-risk settings — healthcare, government affairs, finance, industrial control — continuous updating cannot mean production behavior changing without governance.

Ultra-long-horizon tasks and task persistence let agents handle interruptions, waits, and multi-stage work — and bring goal drift, abnormal loops, duplicate execution, and runaway resource use. An agent may retry endlessly after a failed tool call, keep executing a plan whose premises have changed, or gradually breach its original cost, time, and operational boundaries. Task persistence must ship with state observation, budget limits, retry caps, termination conditions, human takeover, and recovery.

Long-term memory improves continuity and personalization — but memory is itself data processing and data residency. What may be written to memory, kept how long, whether it involves third parties who appear in the user’s interactions, whether the user can inspect and correct it, whether it is deleted when authorization is withdrawn — none of this can be left to the model’s own judgment. Without memory governance, the “data flywheel” becomes unbounded accumulation and opaque use.

Tool calling gives agents real execution capability — and makes identity and permissions the central problem. Calling a search tool and calling payment, email, a database, or production equipment are different risks. Control cannot rest on a static whitelist; it must weigh the current user, the specific task, the counterparty, the data scope, the amount, and the operation type. Low-risk, reversible actions run continuously within prior authorization; high-risk, irreversible, or plan-exceeding actions trigger renewed approval.

Multi-agent collaboration and the skill marketplace stretch the responsibility chain. One task may pass through a planning agent, a retrieval agent, an execution agent, and several third-party skills. The system must know what data each component obtained, what tools it called, and how it shaped the final result — otherwise every participant answers for a component and no one answers for the task.

The “harness layer” of Article 2 therefore cannot only solve how models are called and orchestrated. It must also carry identity management, permission control, runtime monitoring, abnormal termination, evidence recording, and version governance. Cross-model, cross-chip, cross-framework interoperability likewise cannot stop at functional compatibility: authorization boundaries, security rules, and audit information must survive model switches, component substitutions, and cross-platform calls.

For agents, capability that cannot be effectively controlled is not yet capability fit for production. Security governance is thus an internal component of whether Articles 1–2’s research agenda converts into real productive capacity.

3. Articles 3–4: scenarios and terminals expand demand only through trustworthy action

Article 3 proposes native AI software, AI-for-science assistants, “AI scientists” and autonomous laboratories, an AI operating system, and benchmark scenarios (标杆场景) across science, healthcare, education, government affairs, manufacturing, and culture. Article 4 pushes agent capability deep into smart phones, glasses, earphones, wearables, robots, and intelligent vehicles.

These measures move agents from software environments into core enterprise processes, public services, and the physical world. The more real the scenario and the weightier the task, the less security can be deferred until after the system is built.

Healthcare, education, government affairs, manufacturing, and culture are all “agent application scenarios,” but the action types and consequences differ. An agent that retrieves research literature or assembles meeting materials cannot be evaluated and regulated the same way as one that participates in medical treatment, educational evaluation, administrative approval, or production-equipment control.

Classification should look at what the agent does in the process, not the industry label or product name. Does the system generate information, propose suggestions, prepare items for approval, or execute directly? Can it change accounts, business systems, or physical state? Do its consequences take effect externally; can they be withdrawn and recovered? Those factors reflect real risk better than “does it use a large model” or “is it multi-agent.”

Article 3’s forward-deployed-engineer co-creation model suits complex enterprise processes, but co-creation cannot mean technical adaptation alone. The vendor and enterprise must jointly identify the process’s data boundaries, permission nodes, approval requirements, exception cases, and responsible parties — not connect the agent to as many systems as possible first and leave the security team to add restrictions at the end.

Benchmark-scenario acceptance likewise cannot stop at whether the system launched, how many models it connects, or whether the demo runs. A real benchmark should prove: the system runs continuously in real business; task success rates and human-review costs are acceptable; permission boundaries hold; abnormal operations can be stopped; updates can be re-verified; errors can be recovered and responsibility established. Projects without those properties are technology showcases, not replicable production applications.

Article 4’s smart terminals sharpen the point. Phones, glasses, earphones, vehicles, and robots sense the environment continuously and may reach cameras, microphones, location, contacts, screen content, and device control. What a terminal captures does not necessarily belong only to its user — it may include third parties who never chose to interact. (DCC has covered the same tension in the on-device agent fight over reverse interoperability.)

Joint hardware–software design of agents and terminals therefore cannot optimize only model performance, chip power draw, and response latency. It must design permission prompts, physical state indicators, switches for sensitive capabilities, local-versus-cloud data boundaries, abnormal termination, and secure update in parallel. The Measures propose adding qualifying new products to the new-purchase subsidy program for digital and smart products; “qualifying” should not mean only technical sophistication and sales volume, but data protection, permission management, sustained update, and incident-handling capability.

From the demand side: whether enterprises and consumers adopt agents depends on the convenience and value created; whether they keep opening data, systems, and terminal permissions depends on whether risk is controlled. Security governance does not shrink scenarios — it is the precondition for agents to advance from peripheral assistance (retrieval, summarization, drafting) into core processes and execution.

4. Articles 5–6: new organizations and business models need security to establish trading credit

Article 5 supports new entrepreneurship models represented by the one-person company (OPC), public service platforms around computing power, incubation, science-and-technology finance, IP, and policy consulting, and agent-supported lightweight entrepreneurship. Article 6 proposes developing the Token (词元) economy — Token-as-a-Service, Agent-as-a-Service, Results-as-a-Service — and moving from billing by Token consumption to value-based billing.

Both articles reach into how agents restructure enterprise organization and market transactions. The newer the organizational form and business model, the more it needs security governance as a stable credit foundation.

AI-augmented solo entrepreneurship can cut operating costs dramatically: one founder, with agents, covers market research, content, customer service, bookkeeping, contract review, software development, and daily administration. But the execution, review, and supervision functions that separate people perform in a conventional firm now concentrate in one natural person and a set of technical systems.

If multiple agents share one set of accounts and long-term credentials; if payment, contracting, filings, and customer communication all run continuously through systems; if the founder is momentarily unavailable — errors can propagate without human counterweight. The public service platforms and OPC full-cycle service stations of Article 5 should therefore offer more than registration, computing power, and coaching: they should provide permission management, data governance, compliance review, liability insurance, security testing, and business-continuity services scaled to lightweight operators.

And where law requires professional judgment, statutory signature, and responsibility from qualified natural persons, no small company size or strong agent capability substitutes. Agents may assist lawyers, accountants, and doctors; they may not become a route around statutory qualification and responsibility requirements.

Article 6’s Results-as-a-Service and value-based billing equally presuppose security and verifiability. Billing by Token, the vendor mainly proves that model resources were consumed. Billing by result, it must further show the task was actually completed, the result met quality requirements, the actions stayed within authorization, and no unagreed risks or losses were produced.

Moving from Token-consumption billing to value billing is therefore not a change of price unit. It requires new mechanisms for task definition, quality evaluation, process evidence, and responsibility allocation. A system that completes a task cheaply by over-opening data permissions, skipping required approvals, or shifting high-risk errors onto the enterprise has not created value.

The Measures propose exploring Token service-quality assessment and billing norms, with an evaluation system centered on intelligence quality, call volume, and conversion efficiency. Future quality metrics should not stop at response speed, call counts, and task completion rates; they should include human-takeover rates, abnormal retries, error losses, permission violations, external data transmission, and system recovery capability. Tokens can serve as an input and cost-analysis metric — not as a direct measure of industrial value or policy performance.

Service vouchers, Token vouchers, and fiscal support should follow the same logic. Public money should not cash out automatically because an enterprise generated call volume; it should tie to real business tasks, verifiable results, stable operation, and security performance. That prevents subsidy-driven junk calls, and pushes vendors from selling computation toward delivering reliable results.

Security governance is thus not a constraint on the new business models — it is the market foundation on which outcome-based pricing, professional services, insurance, and risk pricing become possible at all. Without verifiable process and attributable results, agent services cannot form long-term trades.

5. Articles 8–9: data, computing power, and open source scale only on a common security baseline

Article 8 proposes a multi-tier computing-power supply system, low-cost standardized public computing services, support for “Token factories” (词元工厂), and pooling sectoral knowledge and agent execution data into a data flywheel. Article 9 proposes open-source communities and platforms, open interconnection protocols, development frameworks, terminal operating systems and reference hardware, and support for agent products and services expanding overseas.

These inputs lower development and usage costs. But the more open the inputs and the more numerous the participants, the more a common set of security rules is needed to lower the market’s identification and transaction costs.

First, agent execution data is not a freely poolable, endlessly reusable production input. Execution traces may contain personal instructions, enterprise business information, operation credentials, customer materials, approval records, error histories, and third-party information. Derived objects — long-term memory, retrieval indexes, embedding vectors, inference caches — can retain the sensitivity of the originals.

Article 8’s data flywheel should therefore rest on classification, authorization, and traceable use. A project must be clear which data serves only the current task, which may improve the service, and which may enter subsequent training; whether memory, vectors, and caches are processed in step when source data is deleted; whether new versions formed by online learning are validated and can be rolled back. Without those rules, the flywheel grows data volume in the short run while eroding the willingness of enterprises and users to keep providing data.

Second, public computing power is a shared runtime environment, not just a resource. Pooling, virtualization, and elastic supply cut SME costs — and concentrate different users, models, data, and agent tasks in one place, raising the stakes on data isolation, permission configuration, model-asset protection, and cross-tenant access. Public computing services need identity authentication, resource isolation, access audit, data cleanup, incident response, and service continuity built in step. “Token factories” should be judged on schedulable real capability, cost per valid task, operational reliability, and isolation — not nominal capacity and Token output. Computing scale becomes industrial capability only when it converts into reliable, usable task results.

Third, open source does not mean lower security requirements. An agent system may simultaneously depend on open-source models, frameworks, plug-ins, skills, external interfaces, and terminal components. The more open components, the more the enterprise needs to know each one’s developer, version, license, permission scope, third-party dependencies, and vulnerability-response mechanism.

Article 9 proposes an open-source contribution evaluation system and project incubation. Support for open-source projects should weigh maintenance continuity, version management, dependency transparency, vulnerability remediation, and emergency-withdrawal capability — not only code volume, downloads, and contributor counts. Components in the skill marketplace and software store should declare the data and permissions they require, be re-reviewed on major permission changes, and be promptly disabled — with users notified — on serious risk.

Agent expansion overseas (出海) also runs on security capability. The Measures propose country-by-country research on expansion paths plus compliance training, IP services, and risk warning. Agent export is not only model-and-software export: it can involve cross-border access to enterprise data, calls to overseas tools, remote task execution, and continuous updates. Security, data, permission, and supply-chain evidence will directly shape whether overseas customers adopt Chinese agent products — and whether those products reach high-value regulated sectors abroad.

Security governance and open data, shared computing, and open source are therefore not opposed. Uniform identity, permission, version, logging, and vulnerability-response rules reduce every enterprise’s cost of reviewing every component, make market entry easier for small vendors, and let an open ecosystem expand under predictable conditions.

6. Article 7 should be the horizontal governance hub — not the sole carrier of security

Article 7’s contents — graded-and-categorized regulation, industry self-discipline, misuse crackdowns, industry legislation, security infrastructure, the trusted sandbox, “using models to govern models” (以模治模) — are foundational. But they concentrate on regulatory method, public platforms, and external testing, and cannot alone carry agent security from development through operation.

Effective agent security governance cannot be a single pre-launch test, nor rely mainly on an external security platform spotting problems. It must enter each of the activities Articles 1–9 support:

  • Technical-research projects under Articles 1–2 should build identity, permissions, memory governance, runtime observation, abnormal termination, and rollback into agent engineering itself.
  • Benchmark scenarios and smart terminals under Articles 3–4 should set security requirements by action type and consequence, and put real operation, human takeover, and incident recovery into acceptance criteria.
  • Entrepreneurship models under Article 5 need lightweight but present permission, tax, data, insurance, and business-continuity services.
  • Token and agent services under Article 6 should be evaluated — for quality and for fiscal support — on secure, reliable task results, not raw call volume.
  • Data flywheels, public computing, and Token factories under Article 8 need data provenance, use control, tenant isolation, and operational audit built in step.
  • Skill marketplaces, software stores, and open-source communities under Article 9 need component identity, permission declarations, version management, vulnerability response, and exit rules.
  • Article 10’s instruments — fiscal funds, government investment funds, and key-project support — are the most direct lever: security should not be a boilerplate undertaking at application, but an evidentiary condition of project selection, staged disbursement, and final acceptance.

On this horizontal logic, Hong proposes four foundational institutions for Beijing’s next step:

  1. A classification mechanism centered on action permissions and consequences. Every agent project states where it sits in the workflow — generating, recommending, requesting approval, or executing — what data and tools it can access, whether it changes system or physical state, and whether consequences are reversible. Higher risk brings stricter authorization, evaluation, recording, and human oversight; low-risk applications get fast development and deployment.
  2. A trusted-delegation baseline for production-grade agents. Each project specifies whom the agent represents, what task it performs, which data and tools it uses, which actions auto-execute, which require renewed approval, who can pause and take over, and who bears business and technical-operations responsibility respectively.
  3. Security capability as public infrastructure. The ranges, trusted sandbox, testing, and certification platforms of Article 7 should grow beyond one-off testing into unified test methods, reference architectures, incident-information sharing, and public security services for SMEs — exactly the common capability no single market player can neutrally build.
  4. Security evidence wired into fiscal support, government procurement, and benchmark acceptance. Judge projects not by whether the platform stands, the model connects, and call volume grows — but by sustained stable operation, effective permission limits, stoppable and recoverable anomalies, traceable key actions, re-verification after updates, and migration and exit capability after incidents.

That would make Article 7 not one security measure alongside nine development measures, but the governance hub running through all of them.

7. Security governance is itself productive capacity

In traditional industrial policy, security reads as a compliance cost that development must bear. In the agent industry, security governance is productive.

It lowers adoption costs. Unified classification, testing, and security standards spare each enterprise a bespoke review and negotiation with every vendor.

It expands what enterprises will authorize. Only when data access can be bounded, tool calls controlled, actions traced, and anomalies halted will enterprises hand agents more of their core tasks.

It supports the new business models. Outcome-based pricing, Agent-as-a-Service, and lightweight organizations all require clear task definition, responsibility allocation, and operational evidence.

It promotes open competition. When third-party agents can demonstrate security capability against uniform standards, platforms, software, and terminals can no longer justify fully closed ecosystems by pointing at abstract security risk.

Security governance, in short, is not an external constraint on agent development. It is the institutional infrastructure by which agents convert from technical capability into tradable products, from pilot projects into production tools, from point applications into an open ecosystem.

The Measures already lay out Beijing’s agent industry across technology, scenarios, terminals, entrepreneurship, pricing, computing power, data, and open source. The real next step is not to keep expanding Article 7’s security workstream, but to embed security requirements into every form of technical support, scenario building, public platform, and fiscal instrument the other nine articles create.

In the end, the true threshold for agent development is not whether the model can complete the task — it is whether enterprises and society are willing to hand the task over. How much data an enterprise will surrender, how many systems it will open, how much authority to act it will grant: that determines how far agents become real productive capacity.

What security governance builds is precisely the condition of trust under which that delegation can occur and continue. Security is not a defensive line to be added after the industry develops — it is the common foundation on which the model capabilities, application scenarios, business models, data flywheels, and open ecosystem of the Beijing Measures can actually land.


This brief translates and condenses 洪延青 (Hong Yanqing), “企业为何敢把任务 交给智能体:评《北京市关于加快智能体引领发展的若干措施》之二,” published on the 网安寻路人 WeChat Official Account on 4 August 2026 (original). Details of the Measures (document number, issuing bodies, dates, and article structure) are taken from the official text published at beijing.gov.cn on 23 July 2026. Parts 1, 3, and 4 of the series are translated in separate DCC briefs.

— Not legal advice.

— Not legal advice.


§ SUBSCRIBE

The Monday brief.

One short email every Monday. New briefs on Chinese data-compliance rules from the previous week, with the source law cited.

Opt-in only. Unsubscribe anytime by replying "unsubscribe" to any issue.

SUPPORT DCC

Keep the publication free to read. Suggested support is $19.99, or choose your own amount.

Support →