Service eligibility and regional restrictions
PuppyIP serves only compliant overseas businesses and their authorized personnel. Proxy services are not available in mainland China. The service may only be used for lawful business activities outside mainland China. Use of this service within mainland China is prohibited.
Hosting a proxy IP or server overseas does not change these restrictions. The service must not be provided to end users in mainland China through relaying, forwarding, sharing or resale. Before use, read the Terms of Service.
Key Takeaways
- Anthropic announced availability October 7. Claude Platform uses claude-haiku-5-5; verify cloud region and account access separately.
- Prompts up to 100K tokens cost $0.10 input and $0.50 output per million. Above 100K, rates are $0.50 and $2.50.
- Haiku gains adjustable effort, but reasoning effort and tokenization affect actual costs. A 75% reduction is not guaranteed for every task.
- API credits roll out to eligible Max and Team subscriptions this week and require linking an organization. SDK, claude -p, and third-party applications can still use existing subscription allowances; this is not a universal chat-limit reset.
Available now: Which workloads fit?
Anthropic released Claude Haiku 5.5 on October 7, 2026 for frequent, cost- and speed-sensitive tasks. Developers can use claude-haiku-5-5 on Claude Platform.
AWS, Google Cloud, and Microsoft Azure availability was also announced. That platform statement does not mean every region, account, or third-party tool already exposes it. Cloud users must verify catalogs and permissions.
Evaluate summaries, classification, context compaction, and clearly scoped subtasks first. Compare Sonnet or Opus for complex long-running coding. A cheaper small model is not a reason to migrate every task.
The $0.10 rate requires prompts no longer than 100K
Official pricing for prompts up to 100,000 tokens is $0.10 per million input and $0.50 output tokens. Above 100K, rates become $0.50 and $2.50 respectively. This is a pricing threshold, not a context limit established by this guide.
Cache reads are also tiered at $0.01 or $0.05 per million tokens. Five-minute writes are $0.125 or $0.625; one-hour writes $0.20 or $1. The first rate in each pair is for prompts up to 100K, the second above it. Cache duration affects costs too.
Hypothetical calculation: 20,000 uncached input and 2,000 output tokens cost $0.002 plus $0.001, totaling $0.003 for these two items alone. This is not $0.10 per request and excludes other tools or services.
Adjustable effort still needs a fresh cost estimate
Haiku 5.5 first introduces adjustable effort to the Haiku family: how much reasoning work the model invests. It adds choice, but higher effort does not automatically retain the lowest cost. Compare results, latency, and actual tokens on your workload.
The vendor's approximately 75% average runtime-cost reduction accounts for request-length distribution and tokenization changes. Migration documentation warns identical text may use more tokens with the new tokenizer. A unit-price reduction is not each task's bill reduction.
Sonnet 5.5 cache reads also fell from $0.20 to $0.10 per million tokens that day. Savings depend on cache-read share; halving that item does not halve all Sonnet call costs.
A new model ID does not make old parameters compatible
Try claude-haiku-5-5 in a separate test request and recheck input length, output limits, and response contents. The ID has no date suffix. Changing its name alone does not preserve request format or usage.
Migration guidance replaces older enabled plus budget_tokens thinking configuration with adaptive. Adaptive thinking may put a thinking block first; extract text by content-block type instead of assuming the first block is the answer.
Check assistant prefill, which returns 400 on the new model, and remove older temperature, top_p, and top_k parameters as instructed. Computer and browser tools have separate toolset requirements. Parameter errors are not network failures.
Max and Team credits: Eligibility before delivery
The announcement proposes monthly credits rolling out this week: $100 for Max 5x and $200 for Max 20x. Team pools $20 per Standard seat and $100 per Premium seat, capped at $500 monthly in total, not $500 per account.
The help center requires an active, good-standing subscription at least seven days old. Free, Pro, and Enterprise are excluded; discounted Team plans qualify. Eligible users link a Claude Console organization through API credits in web Billing, including iOS/Android subscribers. A missing entry does not prove credits arrived.
The Max subscriber or Team Owner/Primary Owner performs the link, also needing Console organization Owner, Admin, or Billing permissions. Being a Team administrator alone does not authorize linking any API organization.
One plan links to one organization, and one organization to one plan, without arbitrary self-service relinking. Credits refresh by billing cycle and do not roll over. Linked API keys share the balance, consuming promotional credits before purchased credits.
Credits cover Claude API, Managed Agents, and Agent SDK on Claude Platform, not Bedrock, Vertex, or Foundry. They cannot pay for interactive Claude Code or extra Claude, Claude Code, or Cowork usage. A prepaid organization with no remaining balance or auto-reload stops API requests after credits run out; it does not charge the Claude subscription.
The new benefit does not remove subscription-based paths. The October 7 help update says Claude Agent SDK, claude -p, and third-party applications can still use subscription allowances under existing limits. Distinguish integration and billing paths instead of claiming everything must now be API-billed.
Turn low prices into actual savings
Choose one frequent, clearly scoped task and keep the current model as a comparison. Use the same authorized samples, recording correct results, retries, total tokens, and costs. One fast response does not establish cheaper workflow completion.
Apply the tier matching actual prompt length and calculate cache items from actual hits. Include rework and larger-model escalation when the small model cannot finish, rather than displaying only the lowest input rate.
This guide made no model calls, claimed no credits, and measured no speed. Release dates, prices, and conditions come from opened official sources. Verify cloud regions, account menus, and actual delivery in your own platform.
Sources
Frequently Asked Questions
Is Haiku 5.5 free?
No. The API charges by tokens. Eligible Max/Team credits are a separate periodic entitlement with access restrictions, not unlimited free model usage.
Do API credits increase weekly Claude chat limits?
No. The help center says existing limits remain. Credits apply to the linked Claude Platform organization's API and must be distinguished from chat, interactive Claude Code, and extra usage.
Must Agent SDK and claude -p always pay separately now?
No. The October 7 help center says they and third-party apps can still use subscription allowances. Claude Platform API calls use the organization's balance. Monthly credits neither raise subscription limits nor automatically switch every integration.
Is it faster than Opus Fast Mode?
The vendor footnote compares standard speeds and says Haiku 5.5 remains slower than Opus Fast Mode. Task, effort, tools, and platform affect actual latency; this guide has no speed tests.