PuppyIP Resource Center
AI Tool Updates 7 min Published 2026-09-27

Are Claude refusals billed? Three pre-output categories and fallback cost boundaries

On September 24, 2026, Anthropic resumed billing for three categories of pre-output refusal. Developers using Claude Fable or Opus safety classifiers should check stop_reason, stop_details.category, and usage before deciding whether an empty response is billed and whether fallback adds a second request's cost.

Claude API Anthropic Refusal billing fallback API costs

Service eligibility and regional restrictions

PuppyIP serves only compliant overseas businesses and their authorized personnel. Proxy services are not available in mainland China. The service may only be used for lawful business activities outside mainland China. Use of this service within mainland China is prohibited.

Hosting a proxy IP or server overseas does not change these restrictions. The service must not be provided to end users in mainland China through relaying, forwarding, sharing or resale. Before use, read the Terms of Service.

Key Takeaways

  • Direct answer: only pre-output refusals in bio, frontier_llm, and reasoning_extraction resume billing at the rates of the model actually run. Other or null categories remain unbilled but count toward rate limits.
  • For refusals partway through streaming, input tokens and tokens already output were already billed at normal rates. Empty final content or HTTP 200 does not establish the bill.
  • If fallback is enabled, a billable first refusal and the subsequent fallback request are billed separately. fallback credit compensates only for prompt-cache misses on the subsequent request; it does not make both model calls free.
  • Anthropic documentation applies this billing rule to Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud, and Microsoft Foundry. The server-side fallbacks parameter remains available only in Claude API beta.
  • Cost-sensitive teams should reconcile bills separately by model, category, token usage, and each retry. This article has no account-level tests; false-classification rates, actual cost increases, and the precise effective time remain unknown. PRODUCT_FIT=NONE: network exits cannot change refusal categories or model billing.

The cost question: an empty response is not necessarily free

Anthropic's September 24, 2026 Claude Platform update says pre-output refusals with stop_details.category of bio, frontier_llm, or reasoning_extraction resume billing at the rate of the model run. This concerns safety-classifier refusals, not every API request that generates no text. The announcement gives a date without an exact transition time or time zone.

Relevant documentation for Claude Fable 5.1, Fable 5, Opus 5.5, and Opus 5 represents classifier refusals as normal HTTP 200 responses with stop_reason set to refusal, potentially empty content, and tokens still listed in usage. Retain the response model name, stop_reason, category, and usage before assessing charges. Check current pricing for model rates; for Opus 5.5 pricing and migration boundaries, see Claude Opus 5.5 pricing and migration.

Impact matrix: calculate refusal stage, category, and fallback separately

Pre-output refusals: bio, frontier_llm, and reasoning_extraction are billed for the model actually run. cyber, general_harms, other categories, and null categories are not billed but still count toward rate limits. Even when output_tokens is 0, do not label the first group free. Conversely, input tokens shown in usage do not prove the second group was charged.

Mid-stream refusals: processed input and already-streamed output are billed at normal rates, and the partial output should be treated as incomplete and discarded. Fallback: if the first refusal occurs mid-stream or belongs to one of the three billable categories, the first call and the fallback model that subsequently actually runs incur separate charges. fallback credit covers prompt-cache-miss cost on the retry; it does not cancel all token charges for both requests.

Affected platforms and choosing a fallback approach

Anthropic lists the refusal-billing rule as applying to Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud, and Microsoft Foundry. Platform coverage does not imply identical feature entry points. The server-side fallbacks parameter is currently available only in Claude API beta and requires the server-side-fallback-2026-07-01 header. Bedrock, Google Cloud, and Microsoft Foundry users need to evaluate SDK middleware or client-side retries.

For Anthropic's recommended category-based fallback within one Claude API call, evaluate server-side default mode. For a fixed, reviewable backup model or consistent cross-platform behavior, evaluate a client-side approach. In either case, record the model that actually answers and usage.iterations. Safety-classifier fallback does not automatically resolve rate limits, overload, or server errors. Do not treat every HTTP failure as a refusal to retry.

Reconcile costs in existing call chains without assuming account outcomes

First distinguish stop_reason=refusal from ordinary 4xx/5xx in existing logs. Record category, response model, input/output tokens, whether output occurred mid-stream, whether fallback_message actually appeared, and the matching billing window only within lawful retention limits. Group by category and model, compare samples on the same basis before and after September 24, and reconcile them with provider bills. If category fields or cross-platform billing mappings are missing, or bills are unsettled, leave the conclusion unknown.

If cost or refusal rate exceeds the team's threshold, pause expansion of automatic retries, retain original request and response IDs, and investigate classifier refusals separately from application exceptions. Do not assume changing IPs or accounts can bypass billing or safety rules. For ordinary 400 parameter errors, see Claude Fable 5.1 migration errors, but do not include those errors in refusal-cost statistics.

What the official announcement does not establish

Official documentation provides no per-customer actual bills, classifier false-positive rate, affected-request frequency, or cost increase for your application. The announcement uses a calendar date, which cannot be extrapolated into midnight in any particular time zone. Category lists may also change with Anthropic's later measurements. Base budgets and alerts on the current official classification table and real account bills.

Sources

Frequently Asked Questions

Are all Claude pre-output refusals billed?

No. Under the September 24, 2026 announcement, bio, frontier_llm, and reasoning_extraction are billed at the rate of the model run. Other or null pre-output refusal categories are unbilled but count toward rate limits.

With HTTP 200 and empty content, how do I identify a billable refusal?

Check whether stop_reason is refusal, then inspect stop_details.category, the response model, and usage. HTTP status or empty content alone is insufficient. Reconcile the final amount with account billing.

If a refusal occurs halfway through streaming, are earlier tokens billed?

Yes. Processed input tokens and already-streamed output tokens are billed at normal rates. Treat the partial output as incomplete.

Does enabling fallback offset the first refusal's charge?

No. If the first refusal is in a billable category or occurs mid-stream, it is billed separately from the subsequent fallback request that actually runs. fallback credit compensates only for prompt-cache misses on the subsequent request.

Can Bedrock or Google Cloud directly use Claude API's server-side fallbacks parameter?

Do not assume so. Refusal-billing rules apply to these platforms, but the server-side fallbacks parameter is currently a Claude API beta feature. Other platforms should check their SDK middleware or client-side retry support.

Can changing network exits avoid Claude refusal billing?

No. Refusal classification and billing are determined by the model request and platform rules. A network exit does not change categories, rates, or account bills.