TREE NEWS update: Beijing’s economy and information technology bureau and its development and reform commission issued an action plan for the city’s token economy covering 2026-2028. The plan sets tiered evaluation standards for ‘token factories’ covering model compatibility, token throughput, first-token latency, cache hit rate and power usage effectiveness, and backs inference-specific chips, model lightweighting and edge deployment to lift token production capacity.
Beijing issues action plan to grade and boost ‘token factories’
Beijing's move is notable less for the token-economy framing than for treating inference capacity as planned infrastructure, with tiered metrics that let officials rank and steer operators. That shifts the policy lever from model development to throughput, latency, cache efficiency and power — the cost drivers that determine whether inference scales economically. Who feels it most: domestic compute and model-serving providers, plus edge and chip vendors positioned for inference-specific demand. The open question is whether these standards stay evaluative or become a de facto gate for support and deployment.
Generated by AI for reference only.
Share on WeChat
Open WeChat → Scan → then tap "…" to send to a chat or Moments.
Tap "…" in the top-right corner to send to a chat or share to Moments.