PuppyIP Resource Center
AI Tool Updates 7 min Published 2026-10-02

Cloudflare Clef decision models: hosted calls, open weights, and fine-tuning boundaries

Cloudflare released Clef and Clef-flash on October 1, 2026: models that choose among predefined answers and return associated probabilities. Hosting and open weights have been announced; the self-service platform for collection, training, and redeployment is still in progress. This article checked the official announcement and Clef model documentation on October 2 Beijing time, without account calls, fine-tuning, or independent performance tests.

Cloudflare Clef Workers AI Decision models Reinforcement learning

Service eligibility and regional restrictions

PuppyIP serves only compliant overseas businesses and their authorized personnel. Proxy services are not available in mainland China. The service may only be used for lawful business activities outside mainland China. Use of this service within mainland China is prohibited.

Hosting a proxy IP or server overseas does not change these restrictions. The service must not be provided to end users in mainland China through relaying, forwarding, sharing or resale. Before use, read the Terms of Service.

Key Takeaways

  • Clef and Clef-flash are decision models for finite options. Probability output is not a guarantee of application accuracy.
  • Cloudflare announced Workers AI hosting and open weights on Hugging Face under Apache 2.0.
  • Hands-on collaborative fine-tuning is currently available through discussions with Cloudflare's FDE team. Self-service collection, training, and deployment are future plans.
  • Define the task, labels, and cost of mistakes first, then compare quality, latency, costs, and fallback behavior with independent samples.

What changed: a decision interface with finite answers

Cloudflare's October 1 official announcement introduces Clef and Clef-flash and outlines reinforcement-learning services and a platform direction for decision tasks. Instead of generating free-form text, decision models take context, predefined questions, and candidate answers, returning choices and probabilities that applications can process. Classification, routing, and next-step selection are possible use cases, but these examples do not establish accuracy sufficient for automatic execution in any business.

This article records October 1 as the publication date shown on the official page; verification occurred on October 2 Beijing time. Announcement timing, model availability, actual account-call results, and completion of the future platform are recorded separately, not combined into a claim that every feature has launched.

What is available now: hosted models and open weights

The announcement says Clef and Clef-flash are hosted on Workers AI, with Apache 2.0 open weights available through Cloudflare's Hugging Face page. Hosted calls and downloaded weights are different paths. For hosting, check account permissions, interfaces, and billing; for downloaded weights, also assess deployment requirements, operating costs, and license obligations. See the official model and licensing information.

The current Clef model reference lists full model ID @cf/cloudflare/clef, 27B size, a 65,536-token context window, and vision input. These fields apply to Clef and cannot be copied directly to Clef-flash or treated as an unconditional maximum for every request. This article has tested neither hosted calls nor self-deployment.

Hands-on fine-tuning is open for discussion; self-service training remains planned

The announcement describes the currently available fine-tuning path as a service in collaboration with the Forward Deployed Engineering (FDE) team. Development teams must contact Cloudflare separately to confirm fit, data and task scope, and service conditions. An inquiry channel does not mean every account has fine-tuning eligibility. See the current service scope.

The self-service platform for collecting feedback, training, and redeployment, along with related trainers and bring-your-own-model capabilities, remains under development in the announcement. Without a clear opening date or evidence of general availability, treat it as a future plan. Available hosting or weights do not mean the full self-service reinforcement-learning platform is open.

Prepare integration: define answers before checking model input

First express the goal as a clearly bounded question, such as routing a request to one of several defined queues. Specify each label's meaning, exclusions, and cases requiring human handling. Second, gather authorized representative samples, expected labels, and the cost of misclassification. Separate training or tuning examples from independent acceptance samples, and do not begin by enabling actions that write to production systems.

Third, check state, questions, and permitted answer types against the official input schema. For images and other inputs, check format, size, and count limits. Do not reuse an ordinary chat model's free-text request format directly, or assume vision support permits arbitrary remote image URLs.

Fourth, record decisions only in an isolated workflow and check missing fields, out-of-bounds inputs, timeouts, repeated requests, and human fallback after failures. This article offers preimplementation checks and has not made account requests or verified any SDK's compatibility. For general API parameter checks, see the API Key, Base URL, and Model verification guide.

Evaluation: record probability, benchmarks, and application quality separately

The official interface provides candidate-answer probabilities. A probability does not prove calibration on your data and is not application accuracy. Use independent acceptance samples to measure misclassifications, misses, and human handoffs by class. Check behavior under low confidence, distribution shifts, and invalid input before deciding on thresholds or human review.

Speed and quality comparisons in the announcement are vendor tests that this article has not independently reproduced. Keep inputs, labels, acceptance criteria, and calling conditions identical during comparisons, and retain latency distributions, request usage, retries, and full costs. Calculate hosting, self-deployment, and hands-on fine-tuning costs separately. If quality is insufficient, interface behavior unstable, or fallback cannot complete, retain the original process and stop expanding automatic execution.

Unconfirmed boundaries and follow-up checks

This verification did not establish specific account-call results, full commercial terms for fine-tuning, a formal opening date for the self-service platform, or independent results for any enterprise task. Confirm these separately through official model documentation and Cloudflare service channels when needed. Keep unknowns explicit where evidence is absent; a platform vision in an announcement is not an implemented feature.

For missing models, permission denials, or quota errors, check model ID, account authorization, and service status first, handling connection-layer issues separately. Changing IP cannot provide eligibility for a model or fine-tuning service. When official guidance changes the schema, model version, or self-service availability, materially update this same guide and rerun the original acceptance samples.

Sources

Frequently Asked Questions

Is Clef an ordinary chat model?

It targets decision tasks with predefined questions and finite candidate answers. Applications should organize inputs and results according to its schema rather than treating free-text chat interfaces as the same format.

Where are Clef and Clef-flash available now?

The official announcement offers Workers AI hosting and open weights on Hugging Face under Apache 2.0. Verify actual account calls, deployment conditions, and costs for the chosen path.

Can @cf/cloudflare/clef specifications be applied directly to Clef-flash?

No. The 27B size and 65,536-token context listed here come from Clef's model reference. Check the other model's own current documentation.

Are self-service reinforcement-learning training and redeployment fully available?

There is no public basis for that claim. The announced self-service platform is still being built. Current inquiries concern hands-on fine-tuning with the FDE team, with eligibility and terms confirmed separately.

Can a high model probability directly trigger a business action?

First verify misclassifications, misses, calibration, and failure fallback using independent samples, then decide on human review or thresholds according to task risk. High probability is not an accuracy guarantee.

Did this article verify calls, training, or vendor speed benchmarks?

No. It checks the publicly documented scope and offers preimplementation checks. It has not performed account calls, training, paid services, or independent performance tests.