TREE NEWS update: Anthropic released Claude Haiku 5.5, which it calls its cheapest, fastest and most capable small model, targeting high-volume uses such as summarization, context compression, database queries, classification, sub-agents, real-time customer service and browser operation. Haiku 5.5 is the first Haiku model with adjustable reasoning effort. Average running costs are about 75% below Haiku 4.5, with input and output priced at $0.10 and $0.50 per million tokens for prompts under 100,000 tokens.
Anthropic launches Claude Haiku 5.5, cutting average running cost about 75%
Cheap, fast small models are becoming the workhorse layer for agentic and high-volume pipelines, and a 75% cost cut on the successor to Haiku 4.5 pushes that layer closer to commodity pricing. The more telling detail is adjustable reasoning effort: it lets builders trade latency and spend against quality per task, which matters most for sub-agents and browser operation where many calls, not one, drive cost. Whether this compresses margins across the small-model segment, and whether rivals match the pricing, is the open question worth watching.
Generated by AI for reference only.
Share on WeChat
Open WeChat → Scan → then tap "…" to send to a chat or Moments.
Tap "…" in the top-right corner to send to a chat or share to Moments.