PuppyIP Resource Center
AI Tool Updates 7 min Published 2026-09-29 Updated 2026-10-08

Claude Sonnet 5.5 cache reads now cost less: pricing and migration changes

Sonnet 5.5 cache reads were cut from $0.20 to $0.10 per million tokens, affecting the cache-hit portion. The official top table and Prompt caching text now show the new rate. Check your channel’s bill, migration parameters, and fallback route; whole requests do not automatically become half-price.

Claude Sonnet 5.5 Cache-read price cut Sonnet 4.5 retirement API pricing 400 errors Model migration

Service eligibility and regional restrictions

PuppyIP serves only compliant overseas businesses and their authorized personnel. Proxy services are not available in mainland China. The service may only be used for lawful business activities outside mainland China. Use of this service within mainland China is prohibited.

Hosting a proxy IP or server overseas does not change these restrictions. The service must not be provided to end users in mainland China through relaying, forwarding, sharing or resale. Before use, read the Terms of Service.

Key Takeaways

  • October 7’s announcement halves cache-read prices; the current table and text agree. Channel billing needs separate verification.
  • New input, cache writes, and output are separate charges. The greater your cache-read share, the greater the effect on overall cost.
  • Sonnet 4.5 retires November 30 on applicable platforms including Claude API. Sonnet 5 remains Active; do not share deadlines between them.
  • Check platform IDs, thinking, tool selection, and fallback paths. Do not resend unchanged 400 requests; validate tool outcomes as well.

What benefits from the new $0.10 cache-read price?

In the October 7, 2026 Haiku 5.5 announcement, Anthropic cut Sonnet 5.5 cache reads from $0.20 to $0.10 per million tokens effective that day. Caching reuses previously processed prompt content; only actual cache reads receive this rate.

At this October 8 check, both the pricing model table and Prompt caching section list $0.10, with the latter specifying 0.05 times input price. The September 28 launch page retains its historical $0.20 price. Follow current pricing and channel conditions for current bills.

The claim of roughly 20% lower costs for most agent tasks is Anthropic’s workload estimate, not a per-request guarantee. Calculate from your cache reads; it does not establish the same rate for Sonnet 5, all models, or every cloud channel.

How much does a request save? A hypothetical bill

Other Claude API Standard rates remain $2 input and $10 output per million tokens, with five-minute/one-hour cache writes at $2.50/$4. Cache read is only one item; first writes, uncached input, and output still count separately.

Suppose calls total 1 million cache-read tokens, 100,000 new input, and 50,000 output. Under Standard/global routing, excluding cache writes, Batch, tools, taxes, and contract discounts, the old bill is $0.20+$0.20+$0.50=$0.90.

Changing only cache reads gives $0.10+$0.20+$0.50=$0.80: $0.10 or about 11.1% saved. This is a calculation, not a measurement. With few hits, this price cut barely affects total cost.

Pricing says caching multipliers can combine with Batch and residency modifiers; Bedrock and Google Cloud have separate rates. This example is not their bill. Recalculate regional endpoints, tools, and special contracts under applicable conditions.

Before migration, identify the actual platform

Sonnet 5.5 launched September 28, 2026. Claude API uses claude-sonnet-5-5; Bedrock anthropic.claude-sonnet-5-5; Google Cloud lists claude-sonnet-5-5. Public availability still requires your organization, region, and entry permissions.

Claude Platform on AWS differs from Amazon Bedrock; AWS billing does not make endpoints interchangeable. Foundry’s API model field uses the actual deployment name, which may differ from the public model. Inspect deployment details/application mapping before edits.

Inventory each business’s platform, endpoint, actual model, SDK, configuration location, and owner, including shared gateways, batches, scheduled tasks, and fallbacks. Do not replace every Sonnet-containing string or store keys in the inventory.

Who is affected by Sonnet 4.5’s November 30 deadline?

Anthropic marked claude-sonnet-4-5-20250929 Deprecated on September 30 with retirement November 30, 2026 and recommended claude-sonnet-5-5. Deprecated means still callable but due for migration; after retirement requests fail. No exact shutdown time/timezone was specified.

The schedule covers Claude API, Claude Platform on AWS, and Microsoft Foundry; check Bedrock/Google Cloud separately. Sonnet 5 remains Active and does not inherit 4.5’s deadline.

After changing model, which 400-causing parameters need changes?

For Messages API, Sonnet 5.5 rejects thinking.type=disabled. To disable upfront thinking use between_tools with low, medium, or high effort. xhigh/max requires adaptive; omitting thinking also enables adaptive.

tool_choice any/tool produces 400: use auto and verify tools still execute as intended. Bedrock Sonnet 5.5 does not support strict tool-input constraints, so validate in your application. HTTP success does not mean a tool action completed.

From 4.5 remove thinking.type=enabled/budget_tokens, assistant prefill, and nondefault temperature/top_p/top_k. Set output_config.effort explicitly. Handle old beta headers/tool versions using the migration checklist for your starting model.

Parse content blocks by type, return thinking blocks unchanged in tool loops, and keep conversations append-only. For 400 save request ID/error field, fix configuration, and avoid identical retries. Keep acceptance examples for multi-turn, streaming, and tool loops.

How do you establish that migration helps your own tasks?

Claude Console Usage→Export produces CSV usage by API key/model. Cross-check calls with code configuration. Before a low-frequency task has run a cycle, no old-model calls today does not prove migration complete.

Compare identical tasks/success criteria on new input, cache writes/reads, output, retries, and tool fees. September 28’s up to 30% per-task savings claim and October 7’s cache price cut are separate; do not add them as your savings.

Validate response model, format, tools, and costs on a small isolated representative set before expanding. Assign owners and cover batch/scheduled cycles. Stop expansion and diagnose quality decline, wrong actions, or budget overruns.

Can a fallback still point to the old model?

Validate a usable replacement and parameters early, removing soon-retired fallback routes. Without an acceptable replacement, pause affected tasks instead of failure loops. Check fallbacks separately when starting from Sonnet 5 or 4.5 rather than copying old configuration.

Claude API’s server-side safety-refusal fallback is an explicitly configured beta for particular safety classifications. Rate limits, overload, and service errors still return; it is not generic failover or a substitute for validating application fallbacks.

Claude app subscriptions, Copilot billing, and API bills have different definitions. Cache cuts do not imply more subscription allowance. Separate network timeouts from parameter/platform-permission errors; changing exits does not grant model eligibility.

Sources

Frequently Asked Questions

If billing still shows $0.20, can I demand an automatic refund?

No personal bill was verified, and the pricing announcement gives no automatic refund process. Confirm model, call time, channel, and cache usage with platform support; seeing $0.20 alone does not establish overcharging.

Can Sonnet 5 use the $0.10 rate too?

Not on the basis of this announcement. It names 5.5 and migration docs distinguish their cache-read rates. Check Sonnet 5’s own pricing/bill.

Can teams combine records from several starting models?

Share representative tasks, but group configuration/results by starting model and platform with owners. This exposes gateways, schedules, and fallbacks still using old parameters.

Is TypeError before an HTTP response a service incident?

Not necessarily. Lifecycle docs say Python SDK v1.0 onward removed temperature/top_p/top_k request parameters; passing them gives TypeError. Check SDK/error location and configuration instead of treating local errors as network incidents.