Editor’s Note — DCC.
This is the fourth and final installment of 洪延青 (Hong Yanqing)‘s commentary on Beijing’s Several Measures on Accelerating Agent-Led Development (京发改〔2026〕1185号, issued 21 July 2026). Part 1 mapped the gap between agent supply and enterprise adoption; Part 2 made security governance the precondition of every other measure; Part 3 built a five-level maturity scale for agent-reshaped processes. Part 4 supplies the measuring rod the other three presuppose: what counts as evidence that an agent created value at all.
The practical payload for overseas counsel sits in two places. First, the quality-and-risk correction: Hong argues that a task “completed” by over-opening data access, skipping approvals, or leaking risk onto the enterprise is not a valid task — which folds Article 7’s security agenda, and by extension PIPL and DSL compliance, directly into the price of agent services. Second, the procurement design: staged evidence gates, composite pricing (base fee + resource fee + outcome performance fee), and recorded refusals and human-takeover shares are the acceptance mechanics Chinese government and SOE buyers are likely to write into agent contracts. Vendors selling into China should expect value-billed deals to carry exactly this evidence burden.
Article 6 of the Several Measures of Beijing Municipality on Accelerating Agent-Led Development encourages a “Token economy”: cultivating Token-as-a-Service, Agent-as-a-Service, and Results-as-a-Service, exploring Token service-quality assessment and billing norms, and moving from billing by Token consumption toward value-based billing. Article 8 adds support for “Token factories” (词元工厂) and pilots of Token vouchers and agent service vouchers. Article 3 requires benchmark scenarios built through on-site co-creation by forward-deployed engineers; Article 5 develops one-person-company (OPC) entrepreneurship; Article 10 commits support to technical research, common platforms, and demonstration applications.
Taken together, these provisions mark a notable extension of Beijing’s policy horizon: from model training to scaled inference, from software sales to agent services, from technology supply to task delivery and outcome billing. The extension is genuinely forward-looking — policy is beginning to face the problems that arise after agents enter real production, rather than attending only to the model.
But extending the policy horizon does not mean a measure of value has formed by itself. On the contrary, it pushes a more basic question to the front: what should actually measure the real value agents create?
Tokens are the easiest thing to count. Model calls, agent counts, platform registrations, and theoretical computing capacity all produce intuitive numbers. The problem is that ease of measurement is not accuracy of evaluation. A system that consumed a mass of Tokens has not thereby completed a mass of valid tasks. A batch of completed tasks does not mean enterprise processes improved. A locally faster process does not necessarily become sustained revenue, new products, or public value.
A boundary has to be drawn first. Tokens meter input. Calls describe the running process. Task completion is direct output. Process improvement is the business result. Sustained operation and better public services are value on a longer cycle. The layers connect — but none substitutes for another. What Beijing’s agent economy needs to build is precisely this evidence chain from resource input to realized value.
1. Tokens can meter input — they cannot serve as the measure of value
The Token is the basic technical unit by which a model meters information processing. Agents work longer contexts, sustain longer task chains, and call external tools constantly, so Token consumption is a real variable in the cost structure of intelligent services.
Article 6’s efficiency agenda — heterogeneous collaboration, storage-compute collaboration, intelligent scheduling, inference-cache reuse, task routing — points the right way. As inference costs fall, the same computing power supports more tasks, and the threshold for SMEs and solo founders drops. In that sense Token efficiency is genuinely part of the industry’s competitiveness.
But that speaks to input efficiency, not value.
The same task tokenizes differently in different models. Parameter scale, inference mechanics, and infrastructure differ, so the real computing cost behind each Token differs too. Adding up Tokens across models and tasks does not yield a stable, comparable indicator of industrial value.
Nor is a task’s real cost anywhere near Token-only. External tool calls — search, databases, maps, payment, professional software — plus vector retrieval and long-term-memory storage, network and terminal execution, human review and exception handling, and the extra losses from failure, retry, and rework all enter a task’s full cost. Counting Tokens leaves a good share of it outside the statistics.
Most important, Token consumption bears no fixed ratio to task value. For the same contract review, an inefficient system may re-read the materials repeatedly, spin up several agents to debate one another, and revise its conclusion multiple times; a more mature system, using structured retrieval, clear rules, and sensible caching, reaches the same or a better result on fewer Tokens. The former consumed more intelligent resources; it cannot on that account be credited with creating more value.
The distinction matters for incentives. If Token call volume becomes the measure of industrial scale or project performance, then the more bloated the context, the more the system loops, and the more it retries after failure, the more prosperous the measured “Token economy” — not the policy’s intent, but potentially its executed result.
The more accurate relation is this: Tokens measure the consumption of intelligent means of production (智能生产资料), not the value of intelligent products.
Tokens work for cost accounting, infrastructure planning, and system-efficiency comparison. They do not say what tasks an agent completed, still less convert into enterprise revenue, process improvement, or social benefit. Beijing’s “Token economy” is better read as a new service economy built on scaled inference and intelligent task execution — not as Token volume itself constituting a new form of value.
What developing a Token economy really has to raise is the rate at which each unit of inference resource converts into valid tasks and business results.
2. Between Tokens and real value stand five layers of evidence
The Measures’ own structure implies a chain from technical input to industrial value: Articles 1–2 support models and the harness layer, Articles 3–4 push scenarios and terminals, Articles 5–6 address organizations and business models, Articles 8–9 supply inputs and ecosystem, Article 10 configures fiscal and project tools. What remains is to say what each layer of that chain can actually prove.
At the front sits resource input: computing power, Tokens, model calls, storage, network, external tools, data processing, development and deployment, human review. It answers: how much was spent to build and run the agent.
Input must first become system capability: response speed, stability in continuous operation, tool-call success rates, recovery after interruption, cross-model and cross-chip portability, and permission control, abnormal termination, and rollback. This layer answers whether the system works stably and controllably.
A running system has not yet completed anything. The third layer is valid task results: task completion rate, first-pass success rate, share meeting quality requirements, human-takeover rate, failures and retries, erroneous operations and permission breaches. Only here does evaluation begin to answer whether the agent actually completed what was delegated.
Completed tasks still have to show up as process results: is the full business cycle shorter, are human hand-offs fewer, are error and complaint rates down, is customer waiting time down, is inventory turning faster, are R&D and time-to-market cycles compressed? These changes — not individual task successes — show whether the enterprise’s process as a whole got better.
Last comes enterprise and social value: new products and revenue, customers who keep paying and renew, improved operating costs and cash flow, better public-service quality, and a replicable, sustainable operating model. This layer answers whether process change converts into longer-cycle economic and social value.
The five layers are causally linked, but the links require evidence — the jump cannot be made by naming an indicator. A given Token volume proves an inference load, not a quantity of valid tasks. Rising task counts do not automatically prove the enterprise’s overall process improved. Even a shorter processing time does not guarantee higher profit, because demand, pricing, reorganization, and other factors move at the same time.
Attribution deserves particular care. Agent projects usually run alongside process standardization, data governance, and management change. When business results move, one must separate how much came from model capability, how much from process restructuring, how much from changed management. Benchmark scenarios should establish pre-launch business baselines wherever possible, record exactly which step the agent entered, and judge through sustained operation whether the change holds. Important projects can add phased pilots and comparisons across similar business units to build a reasonable reference.
At the task layer, Hong proposes a unit metric — cost per valid completed task (单位有效任务成本):
Cost per valid completed task = (model and tool costs + human review costs + failure and retry costs + expected risk losses) ÷ number of tasks that met quality requirements and closed the loop.
The word that matters is valid. A task nominally completed but needing heavy human rework is not valid; neither is one whose cost was lowered by widening data access, skipping required approvals, or raising error risk. Otherwise “efficiency” is just cost and risk moved outside the statistics.
Hence the basic principle of agent-value evaluation: record layer by layer, prove layer by layer — never substitute input for output, and never substitute a local effect for final value.
3. Outcome billing still needs quality and risk correction
Article 6’s three service models differ in more than billing labels. They deliver different objects, and vendor responsibility shifts with the object.
Token-as-a-Service delivers inference resources; the vendor answers mainly for availability, response speed, and baseline service quality. Agent-as-a-Service delivers a system able to complete defined tasks; the vendor answers additionally for tool integration, task orchestration, and operational stability. Results-as-a-Service extends further, moving the trading unit from technical resources to verifiable task results — and with it, more direct quality and delivery responsibility.
Moving from Token-consumption billing to value billing is the right direction. It pushes vendors from selling call volume to delivering task results, and forces continuous optimization of inference cost, tool calling, and human review. But the problem does not end there, because “the result” is not self-defining.
Pay a customer-service agent per inquiry handled, and it learns to end conversations quickly — while repeat calls and complaints rise. Grade a procurement agent on price reduction alone, and quality, delivery, and supply-chain resilience drop out. Judge a government-affairs agent on speed alone, and procedure, fairness, and deliberate judgment suffer. Set the outcome metric too narrowly and the agent optimizes the metric, not the goal.
A subtler bias: vendors may take the easy tasks and route complex customers and high-risk matters to humans. The agent’s success rate looks excellent; the human team’s residual workload becomes harder and heavier. Evaluation must record not only what the agent completed, but what it refused, transferred, and failed — and how much residual work people actually absorbed.
Results therefore need correcting for quality, risk, and long-run effect. Customer service: not just volume, but first-contact resolution, repeat calls, complaints. Production scheduling: not just output, but quality, energy use, equipment wear. Sales agents: not just lead counts, but real closings, retention, refunds. Government services: not just speed, but procedural legality, error correction, and remedies for the parties.
Article 7’s security governance also belongs directly in the value calculation — the argument of Part 2 of this series. A system that lowers surface costs while adding personal-information leakage, wrong payments, or unauthorized external data transmission has not created higher value. Only value corrected for quality and risk can be booked as agent output.
Which means Results-as-a-Service is not a new price label but a re-definition of task, quality, risk, evidence, and responsibility. For projects whose task boundaries are still unstable, or whose results cannot be cleanly attributed to the vendor, pure outcome payment should wait. The sounder arrangement blends a base service fee, a resource usage fee, and an outcome performance fee — raising the outcome share as task boundaries and result standards stabilize.
4. Different policy objects need metrics from different layers
The Measures cover models, the harness layer, scenarios, terminals, entrepreneurship, Tokens, computing power, open source, and overseas expansion. These objects sit at different points on the value chain; one shared indicator set will judge some too early and others too shallowly.
Articles 1–2 (base models, harness layer, middleware) should be evaluated on system capability and resource efficiency: long-horizon task success, tool-call stability, interruption recovery, cross-model portability, cost per valid completed task. Parameter counts, maximum context length, and theoretical Token throughput can still be recorded — they describe partial technical capability and substitute for no real-task test.
Articles 3–4 (benchmark scenarios, FDE co-creation, smart terminals) move the weight to tasks and processes: state the pre-existing business baseline, prove sustained operation in real production, and record what happened to completion rates, human review, error rates, and the full business cycle. Terminals must distinguish product sales, feature activation, and sustained effective use — “ships with agent features” is not yet created value.
Articles 5–6 (OPC, Agent-as-a-Service, Results-as-a-Service) require operating-level evidence: real customers, sustained revenue, renewals, a stable compliance record, and business continuity when the operator is briefly unable to act. Those say more about whether a new organization or service actually stands than registration counts and call volume ever will.
Article 8 (Token factories, public computing, vouchers) should be judged on effective load, not nominal scale. Equipment built, theoretical capacity formed, Token output up — that proves supply increased. What has industrial meaning is how much capacity schedules stably, how much carries real production tasks, and how much non-subsidized customers keep paying for.
Article 9 (open source, overseas expansion) should move from distribution counts to sustained adoption. Open-source projects: not just code volume, downloads, and registered developers, but maintenance continuity, adoption in real products, and lowered integration and migration costs. Agents expanding overseas: sustained foreign usage and renewals, stable service revenue, and long-run conformity with local data, security, and sectoral requirements.
The process-maturity scale of Part 3 maps onto the same chain: tool assistance is judged on personal task efficiency; step embedding on the single process node; the bounded closed loop on complete task results; end-to-end orchestration on whole-process effect; native restructuring on sustained change in organization, product, and business model.
Evidence requirements should move with maturity. Judging a technical prototype by enterprise revenue is premature; judging a production system by model call volume is plainly not enough. The measuring rod must travel down the chain as the project matures.
5. Fiscal money should enter with value evidence — and exit when evidence stops
Article 10 coordinates fiscal funds, government investment funds, and market funds behind key projects, supporting technical research, common platforms, and demonstration applications. The way to make that money efficient is not to demand final commercial proof from every project on day one, but to set evidence gates matched to project stage.
At the technical prototype stage, the question is whether the basic capability exists at all. Early exploration is uncertain; failure must be allowed, and mature commercial results must not be demanded.
At production pilot, the rod moves: prove stable task completion in real business; record takeovers, errors, retries, and security incidents; compare against the pre-existing process baseline. Good results in a demo environment do not justify scaled rollout.
At demonstration and replication, prove the project works beyond one special team or environment — in a second department, factory, institution, or adjacent scenario — at an acceptable replication cost.
At commercialization, prove customers keep paying, the project survives as subsidies taper, and the vendor can maintain, migrate, and handle incidents over time.
The stages need different evidence, but one requirement is constant: each new stage demands new evidence. Early technical indicators cannot keep answering later production and operating questions. Projects that never reach real production, permanently depend on the resident team, have no stable user, or stop running the moment subsidies stop should exit promptly.
Exit is not punishment for failed innovation. It is the normal mechanism by which a high-risk innovation policy keeps allocating resources efficiently. Allowing failure also means allowing funding to stop for projects that can no longer produce new evidence.
Token vouchers and agent service vouchers should carry the same staging: subsidize inference and service costs at first adoption to lower the threshold; tie ongoing subsidy to the enterprise’s own investment, real tasks, quality results, and stable operation once production begins; taper fiscal support out at commercial maturity rather than substituting for real market demand indefinitely.
Government and SOE first-purchase, first-use (首购首用) programs should likewise migrate from procuring abstract “agent platforms” and Token quotas toward procuring defined task capability. Procurement documents should state the business problem, the existing process baseline, the quality boundaries, and the acceptance evidence. Where pure outcome payment is premature, use a base service fee plus a task-quality performance fee — but never accept “system launched” or “model connected” as the standard of completion.
6. What ultimately needs proving is that agents became productive capacity
Beijing’s Token economy reflects a real shift: the AI industry moving from concentrated training to scaled inference, from model capability to agent services. The policy judgment is well grounded.
But what Beijing really needs to build is not a Token statistics system. It is an evaluation system that runs the whole chain — resource input, system capability, valid tasks, process results, final value.
Tokens can say how much intelligent resource was used; calls can say how many times the system ran. Only quality- and risk-corrected valid tasks show direct output. Only tasks that go on to improve enterprise processes, form sustained revenue, or raise public-service quality prove that agents became productive capacity.
Different policy measures should carry indicators from different layers; projects at different maturity should carry different layers of value evidence. Fiscal support entering with the evidence and exiting when the evidence stays absent is what converts the evaluation system into an executable policy mechanism.
Beijing’s agent economy will of course watch scale. But meaningful scale is not how many Tokens were produced, models connected, or platforms built in isolation. It is how much inference resource converted into stable services, how many calls into valid tasks, how many tasks into process improvement — and how much process improvement finally became sustained enterprise and public value.
Beijing needs to watch Tokens; it cannot stop at Tokens. Tokens meter input, calls describe process. Only tasks reliably completed under controlled conditions — and the process improvement and sustained value they bring — prove that agents truly became productive capacity.
This brief translates and condenses 洪延青 (Hong Yanqing), “如何衡量智能体 创造的真实价值:评《北京市关于加快智能体引领发展的若干措施》之四,” published on the 网安寻路人 WeChat Official Account on 7 August 2026 (original). Details of the Measures (document number, issuing bodies, dates, and article structure) are taken from the official text published at beijing.gov.cn on 23 July 2026. Parts 1, 2, and 3 of the series are translated in separate DCC briefs.
— Not legal advice.