---
title: "Tokens Meter Input, Not Value: Hong Yanqing on Beijing's Agent Measures (Part 4 of 4)"
author: "DCC Editorial"
published: 2026-08-07T04:30:00.000Z
url: https://datacompliancechina.com/posts/beijing-agent-measures-token-to-value/
description: "Part 4, closing Hong Yanqing's commentary on the Several Measures of Beijing Municipality on Accelerating Agent-Led Development (北京市关于加快智能体引领发展的若干措施, 京发改〔2026〕1185号). The Measures' Article 6 proposes a Token (词元) economy — Token-as-a-Service, Agent-as-a-Service, Results-as-a-Service, and a shift from billing by Token consumption to value-based billing; Article 8 funds 'Token factories' and Token vouchers. Hong draws the line the policy still needs: Tokens measure the consumption of intelligent means of production, not the value of intelligent products. Tokenization differs across models; a task's full cost includes tool calls, memory storage, human review, and failed retries; and Token volume has no fixed ratio to task value — so treating Token throughput as industrial performance rewards long contexts, loops, and retries. His alternative is a five-layer evidence chain (resource input → system capability → valid task results → process results → enterprise and social value), a cost-per-valid-completed-task formula that counts review, retries, and expected risk losses, and an attribution discipline of pre-launch baselines and phased pilots. Outcome billing must be corrected for quality and risk — narrow metrics make customer-service agents rush calls and procurement agents chase price cuts, and vendors can cream-skim easy tasks while humans absorb the hard residue — so projects with unstable task boundaries should blend base, resource, and performance fees rather than jump to pure Results-as-a-Service. Different policy objects need different-layer metrics, mapped onto Part 3's five maturity levels, and fiscal support should pass staged evidence gates: prototypes may fail, pilots must beat baselines in real business, demonstrations must replicate at acceptable cost, and commercial-stage projects must survive subsidy taper — with prompt exit for projects that stop producing new evidence."
tags: ["ai-agents", "beijing", "token-economy", "value-measurement", "industrial-policy", "hong-yanqing", "commentary", "智能体"]
laws_cited: ["beijing-agent-led-development-measures"]
domains: ["ai-governance", "data-economy"]
account: "wangan-xunluren"
original_title: "如何衡量智能体创造的真实价值：评《北京市关于加快智能体引领发展的若干措施》之四"
original_author: "洪延青 (Hong Yanqing)"
original_publication: "网安寻路人 WeChat Official Account"
original_url: "https://mp.weixin.qq.com/s/vDhEhk2UTrQ1SOb_xNEjrw"
source_language: "zh"
---

> **Source: Data Compliance China** — https://datacompliancechina.com/posts/beijing-agent-measures-token-to-value/ · China data law, translated and annotated for overseas counsel. Cite as: Data Compliance China, "Tokens Meter Input, Not Value: Hong Yanqing on Beijing's Agent Measures (Part 4 of 4)", https://datacompliancechina.com/posts/beijing-agent-measures-token-to-value/
> *Editor's Note — DCC.*
>
> This is the fourth and final installment of 洪延青 (Hong Yanqing)'s
> commentary on Beijing's [Several Measures on Accelerating Agent-Led
> Development](/laws/beijing-agent-led-development-measures/)
> (京发改〔2026〕1185号, issued 21 July 2026).
> [Part 1](/posts/beijing-agent-measures-adoption-gap/) mapped the gap
> between agent supply and enterprise adoption;
> [Part 2](/posts/beijing-agent-measures-security-as-foundation/) made
> security governance the precondition of every other measure;
> [Part 3](/posts/beijing-agent-measures-process-maturity/) built a
> five-level maturity scale for agent-reshaped processes. Part 4 supplies
> the measuring rod the other three presuppose: what counts as evidence
> that an agent created value at all.
>
> The practical payload for overseas counsel sits in two places. First, the
> quality-and-risk correction: Hong argues that a task "completed" by
> over-opening data access, skipping approvals, or leaking risk onto the
> enterprise is not a valid task — which folds Article 7's security agenda,
> and by extension PIPL and DSL compliance, directly into the price of
> agent services. Second, the procurement design: staged evidence gates,
> composite pricing (base fee + resource fee + outcome performance fee),
> and recorded refusals and human-takeover shares are the acceptance
> mechanics Chinese government and SOE buyers are likely to write into
> agent contracts. Vendors selling into China should expect value-billed
> deals to carry exactly this evidence burden.

Article 6 of the *Several Measures of Beijing Municipality on Accelerating
Agent-Led Development* encourages a "Token economy": cultivating
Token-as-a-Service, Agent-as-a-Service, and Results-as-a-Service, exploring
Token service-quality assessment and billing norms, and moving from billing
by Token consumption toward value-based billing. Article 8 adds support for
"Token factories" (词元工厂) and pilots of Token vouchers and agent service
vouchers. Article 3 requires benchmark scenarios built through on-site
co-creation by forward-deployed engineers; Article 5 develops
one-person-company (OPC) entrepreneurship; Article 10 commits support to
technical research, common platforms, and demonstration applications.

Taken together, these provisions mark a notable extension of Beijing's
policy horizon: from model training to scaled inference, from software
sales to agent services, from technology supply to task delivery and
outcome billing. The extension is genuinely forward-looking — policy is
beginning to face the problems that arise after agents enter real
production, rather than attending only to the model.

But extending the policy horizon does not mean a measure of value has
formed by itself. On the contrary, it pushes a more basic question to the
front: **what should actually measure the real value agents create?**

Tokens are the easiest thing to count. Model calls, agent counts, platform
registrations, and theoretical computing capacity all produce intuitive
numbers. The problem is that ease of measurement is not accuracy of
evaluation. A system that consumed a mass of Tokens has not thereby
completed a mass of valid tasks. A batch of completed tasks does not mean
enterprise processes improved. A locally faster process does not
necessarily become sustained revenue, new products, or public value.

A boundary has to be drawn first. Tokens meter input. Calls describe the
running process. Task completion is direct output. Process improvement is
the business result. Sustained operation and better public services are
value on a longer cycle. The layers connect — but none substitutes for
another. What Beijing's agent economy needs to build is precisely this
evidence chain from resource input to realized value.

## 1. Tokens can meter input — they cannot serve as the measure of value

The Token is the basic technical unit by which a model meters information
processing. Agents work longer contexts, sustain longer task chains, and
call external tools constantly, so Token consumption is a real variable in
the cost structure of intelligent services.

Article 6's efficiency agenda — heterogeneous collaboration,
storage-compute collaboration, intelligent scheduling, inference-cache
reuse, task routing — points the right way. As inference costs fall, the
same computing power supports more tasks, and the threshold for SMEs and
solo founders drops. In that sense Token efficiency is genuinely part of
the industry's competitiveness.

But that speaks to input efficiency, not value.

The same task tokenizes differently in different models. Parameter scale,
inference mechanics, and infrastructure differ, so the real computing cost
behind each Token differs too. Adding up Tokens across models and tasks
does not yield a stable, comparable indicator of industrial value.

Nor is a task's real cost anywhere near Token-only. External tool calls —
search, databases, maps, payment, professional software — plus vector
retrieval and long-term-memory storage, network and terminal execution,
human review and exception handling, and the extra losses from failure,
retry, and rework all enter a task's full cost. Counting Tokens leaves a
good share of it outside the statistics.

Most important, Token consumption bears no fixed ratio to task value. For
the same contract review, an inefficient system may re-read the materials
repeatedly, spin up several agents to debate one another, and revise its
conclusion multiple times; a more mature system, using structured
retrieval, clear rules, and sensible caching, reaches the same or a better
result on fewer Tokens. The former consumed more intelligent resources; it
cannot on that account be credited with creating more value.

The distinction matters for incentives. If Token call volume becomes the
measure of industrial scale or project performance, then the more bloated
the context, the more the system loops, and the more it retries after
failure, the more prosperous the measured "Token economy" — not the
policy's intent, but potentially its executed result.

The more accurate relation is this: **Tokens measure the consumption of
intelligent means of production (智能生产资料), not the value of
intelligent products.**

Tokens work for cost accounting, infrastructure planning, and
system-efficiency comparison. They do not say what tasks an agent
completed, still less convert into enterprise revenue, process
improvement, or social benefit. Beijing's "Token economy" is better read
as a new service economy built on scaled inference and intelligent task
execution — not as Token volume itself constituting a new form of value.

What developing a Token economy really has to raise is the rate at which
each unit of inference resource converts into valid tasks and business
results.

## 2. Between Tokens and real value stand five layers of evidence

The Measures' own structure implies a chain from technical input to
industrial value: Articles 1–2 support models and the harness layer,
Articles 3–4 push scenarios and terminals, Articles 5–6 address
organizations and business models, Articles 8–9 supply inputs and
ecosystem, Article 10 configures fiscal and project tools. What remains is
to say what each layer of that chain can actually prove.

At the front sits **resource input**: computing power, Tokens, model
calls, storage, network, external tools, data processing, development and
deployment, human review. It answers: how much was spent to build and run
the agent.

Input must first become **system capability**: response speed, stability
in continuous operation, tool-call success rates, recovery after
interruption, cross-model and cross-chip portability, and permission
control, abnormal termination, and rollback. This layer answers whether
the system works stably and controllably.

A running system has not yet completed anything. The third layer is
**valid task results**: task completion rate, first-pass success rate,
share meeting quality requirements, human-takeover rate, failures and
retries, erroneous operations and permission breaches. Only here does
evaluation begin to answer whether the agent actually completed what was
delegated.

Completed tasks still have to show up as **process results**: is the full
business cycle shorter, are human hand-offs fewer, are error and complaint
rates down, is customer waiting time down, is inventory turning faster,
are R&D and time-to-market cycles compressed? These changes — not
individual task successes — show whether the enterprise's process as a
whole got better.

Last comes **enterprise and social value**: new products and revenue,
customers who keep paying and renew, improved operating costs and cash
flow, better public-service quality, and a replicable, sustainable
operating model. This layer answers whether process change converts into
longer-cycle economic and social value.

The five layers are causally linked, but the links require evidence — the
jump cannot be made by naming an indicator. A given Token volume proves an
inference load, not a quantity of valid tasks. Rising task counts do not
automatically prove the enterprise's overall process improved. Even a
shorter processing time does not guarantee higher profit, because demand,
pricing, reorganization, and other factors move at the same time.

Attribution deserves particular care. Agent projects usually run alongside
process standardization, data governance, and management change. When
business results move, one must separate how much came from model
capability, how much from process restructuring, how much from changed
management. Benchmark scenarios should establish pre-launch business
baselines wherever possible, record exactly which step the agent entered,
and judge through sustained operation whether the change holds. Important
projects can add phased pilots and comparisons across similar business
units to build a reasonable reference.

At the task layer, Hong proposes a unit metric — **cost per valid
completed task** (单位有效任务成本):

> **Cost per valid completed task = (model and tool costs + human review
> costs + failure and retry costs + expected risk losses) ÷ number of
> tasks that met quality requirements and closed the loop.**

The word that matters is *valid*. A task nominally completed but needing
heavy human rework is not valid; neither is one whose cost was lowered by
widening data access, skipping required approvals, or raising error risk.
Otherwise "efficiency" is just cost and risk moved outside the statistics.

Hence the basic principle of agent-value evaluation: **record layer by
layer, prove layer by layer — never substitute input for output, and never
substitute a local effect for final value.**

## 3. Outcome billing still needs quality and risk correction

Article 6's three service models differ in more than billing labels. They
deliver different objects, and vendor responsibility shifts with the
object.

Token-as-a-Service delivers inference resources; the vendor answers mainly
for availability, response speed, and baseline service quality.
Agent-as-a-Service delivers a system able to complete defined tasks; the
vendor answers additionally for tool integration, task orchestration, and
operational stability. Results-as-a-Service extends further, moving the
trading unit from technical resources to verifiable task results — and
with it, more direct quality and delivery responsibility.

Moving from Token-consumption billing to value billing is the right
direction. It pushes vendors from selling call volume to delivering task
results, and forces continuous optimization of inference cost, tool
calling, and human review. But the problem does not end there, because
"the result" is not self-defining.

Pay a customer-service agent per inquiry handled, and it learns to end
conversations quickly — while repeat calls and complaints rise. Grade a
procurement agent on price reduction alone, and quality, delivery, and
supply-chain resilience drop out. Judge a government-affairs agent on
speed alone, and procedure, fairness, and deliberate judgment suffer. Set
the outcome metric too narrowly and the agent optimizes the metric, not
the goal.

A subtler bias: vendors may take the easy tasks and route complex
customers and high-risk matters to humans. The agent's success rate looks
excellent; the human team's residual workload becomes harder and heavier.
Evaluation must record not only what the agent completed, but what it
refused, transferred, and failed — and how much residual work people
actually absorbed.

Results therefore need correcting for quality, risk, and long-run effect.
Customer service: not just volume, but first-contact resolution, repeat
calls, complaints. Production scheduling: not just output, but quality,
energy use, equipment wear. Sales agents: not just lead counts, but real
closings, retention, refunds. Government services: not just speed, but
procedural legality, error correction, and remedies for the parties.

Article 7's security governance also belongs directly in the value
calculation — the argument of
[Part 2 of this series](/posts/beijing-agent-measures-security-as-foundation/).
A system that lowers surface costs while adding personal-information
leakage, wrong payments, or unauthorized external data transmission has
not created higher value. Only value corrected for quality and risk can be
booked as agent output.

Which means Results-as-a-Service is not a new price label but a
re-definition of task, quality, risk, evidence, and responsibility. For
projects whose task boundaries are still unstable, or whose results cannot
be cleanly attributed to the vendor, pure outcome payment should wait. The
sounder arrangement blends a base service fee, a resource usage fee, and
an outcome performance fee — raising the outcome share as task boundaries
and result standards stabilize.

## 4. Different policy objects need metrics from different layers

The Measures cover models, the harness layer, scenarios, terminals,
entrepreneurship, Tokens, computing power, open source, and overseas
expansion. These objects sit at different points on the value chain; one
shared indicator set will judge some too early and others too shallowly.

**Articles 1–2** (base models, harness layer, middleware) should be
evaluated on system capability and resource efficiency: long-horizon task
success, tool-call stability, interruption recovery, cross-model
portability, cost per valid completed task. Parameter counts, maximum
context length, and theoretical Token throughput can still be recorded —
they describe partial technical capability and substitute for no real-task
test.

**Articles 3–4** (benchmark scenarios, FDE co-creation, smart terminals)
move the weight to tasks and processes: state the pre-existing business
baseline, prove sustained operation in real production, and record what
happened to completion rates, human review, error rates, and the full
business cycle. Terminals must distinguish product sales, feature
activation, and sustained effective use — "ships with agent features" is
not yet created value.

**Articles 5–6** (OPC, Agent-as-a-Service, Results-as-a-Service) require
operating-level evidence: real customers, sustained revenue, renewals, a
stable compliance record, and business continuity when the operator is
briefly unable to act. Those say more about whether a new organization or
service actually stands than registration counts and call volume ever
will.

**Article 8** (Token factories, public computing, vouchers) should be
judged on effective load, not nominal scale. Equipment built, theoretical
capacity formed, Token output up — that proves supply increased. What has
industrial meaning is how much capacity schedules stably, how much
carries real production tasks, and how much non-subsidized customers keep
paying for.

**Article 9** (open source, overseas expansion) should move from
distribution counts to sustained adoption. Open-source projects: not just
code volume, downloads, and registered developers, but maintenance
continuity, adoption in real products, and lowered integration and
migration costs. Agents expanding overseas: sustained foreign usage and
renewals, stable service revenue, and long-run conformity with local
data, security, and sectoral requirements.

The process-maturity scale of
[Part 3](/posts/beijing-agent-measures-process-maturity/) maps onto the
same chain: tool assistance is judged on personal task efficiency; step
embedding on the single process node; the bounded closed loop on complete
task results; end-to-end orchestration on whole-process effect; native
restructuring on sustained change in organization, product, and business
model.

Evidence requirements should move with maturity. Judging a technical
prototype by enterprise revenue is premature; judging a production system
by model call volume is plainly not enough. The measuring rod must travel
down the chain as the project matures.

## 5. Fiscal money should enter with value evidence — and exit when evidence stops

Article 10 coordinates fiscal funds, government investment funds, and
market funds behind key projects, supporting technical research, common
platforms, and demonstration applications. The way to make that money
efficient is not to demand final commercial proof from every project on
day one, but to set evidence gates matched to project stage.

At the **technical prototype** stage, the question is whether the basic
capability exists at all. Early exploration is uncertain; failure must be
allowed, and mature commercial results must not be demanded.

At **production pilot**, the rod moves: prove stable task completion in
real business; record takeovers, errors, retries, and security incidents;
compare against the pre-existing process baseline. Good results in a demo
environment do not justify scaled rollout.

At **demonstration and replication**, prove the project works beyond one
special team or environment — in a second department, factory,
institution, or adjacent scenario — at an acceptable replication cost.

At **commercialization**, prove customers keep paying, the project
survives as subsidies taper, and the vendor can maintain, migrate, and
handle incidents over time.

The stages need different evidence, but one requirement is constant: each
new stage demands new evidence. Early technical indicators cannot keep
answering later production and operating questions. Projects that never
reach real production, permanently depend on the resident team, have no
stable user, or stop running the moment subsidies stop should exit
promptly.

Exit is not punishment for failed innovation. It is the normal mechanism
by which a high-risk innovation policy keeps allocating resources
efficiently. Allowing failure also means allowing funding to stop for
projects that can no longer produce new evidence.

Token vouchers and agent service vouchers should carry the same staging:
subsidize inference and service costs at first adoption to lower the
threshold; tie ongoing subsidy to the enterprise's own investment, real
tasks, quality results, and stable operation once production begins; taper
fiscal support out at commercial maturity rather than substituting for
real market demand indefinitely.

Government and SOE first-purchase, first-use (首购首用) programs should
likewise migrate from procuring abstract "agent platforms" and Token
quotas toward procuring defined task capability. Procurement documents
should state the business problem, the existing process baseline, the
quality boundaries, and the acceptance evidence. Where pure outcome
payment is premature, use a base service fee plus a task-quality
performance fee — but never accept "system launched" or "model connected"
as the standard of completion.

## 6. What ultimately needs proving is that agents became productive capacity

Beijing's Token economy reflects a real shift: the AI industry moving from
concentrated training to scaled inference, from model capability to agent
services. The policy judgment is well grounded.

But what Beijing really needs to build is not a Token statistics system.
It is an evaluation system that runs the whole chain — resource input,
system capability, valid tasks, process results, final value.

Tokens can say how much intelligent resource was used; calls can say how
many times the system ran. Only quality- and risk-corrected valid tasks
show direct output. Only tasks that go on to improve enterprise processes,
form sustained revenue, or raise public-service quality prove that agents
became productive capacity.

Different policy measures should carry indicators from different layers;
projects at different maturity should carry different layers of value
evidence. Fiscal support entering with the evidence and exiting when the
evidence stays absent is what converts the evaluation system into an
executable policy mechanism.

Beijing's agent economy will of course watch scale. But meaningful scale
is not how many Tokens were produced, models connected, or platforms
built in isolation. It is how much inference resource converted into
stable services, how many calls into valid tasks, how many tasks into
process improvement — and how much process improvement finally became
sustained enterprise and public value.

Beijing needs to watch Tokens; it cannot stop at Tokens. Tokens meter
input, calls describe process. Only tasks reliably completed under
controlled conditions — and the process improvement and sustained value
they bring — prove that agents truly became productive capacity.

---

*This brief translates and condenses 洪延青 (Hong Yanqing), "如何衡量智能体
创造的真实价值：评《北京市关于加快智能体引领发展的若干措施》之四," published
on the 网安寻路人 WeChat Official Account on 7 August 2026
([original](https://mp.weixin.qq.com/s/vDhEhk2UTrQ1SOb_xNEjrw)). Details of
the Measures (document number, issuing bodies, dates, and article
structure) are taken from the official text published at beijing.gov.cn on
23 July 2026. Parts 1, 2, and 3 of the series are translated in separate
DCC briefs.*

— Not legal advice.
