PuppyIP Resource Center
AI Tool Updates 7 min Published 2026-10-02

Cloudflare AI Search GA: checking November billing, OCR, and image retrieval

Cloudflare announced AI Search general availability on October 1, 2026, with billing starting November 1. Before connecting scanned PDFs or an image knowledge base, check the actual model, OCR settings, and account-level free pools, then validate indexing and retrieval with a small sample.

Cloudflare AI Search OCR RAG Billing

Service eligibility and regional restrictions

PuppyIP serves only compliant overseas businesses and their authorized personnel. Proxy services are not available in mainland China. The service may only be used for lawful business activities outside mainland China. Use of this service within mainland China is prohibited.

Hosting a proxy IP or server overseas does not change these restrictions. The service must not be provided to end users in mainland China through relaying, forwarding, sharing or resale. Before use, read the Terms of Service.

Key Takeaways

  • General availability and billing start are separate dates: the official announcement is October 1, and billing is scheduled from November 1. These pages provide no exact time or time zone.
  • Each account receives 5 million ingestion tokens, 10 GB-month of storage, 1,000 semantic searches, and 1,000 full-text searches per month. The two search allowances are separate.
  • OCR is off by default. With OCR enabled, the PDF limit is 10 MiB; without it, the limit remains 4 MiB. Changing the OCR setting triggers a full reindex.
  • Qwen3-VL can be selected for native image embeddings, but it is not the default model. Check answer-generation, query-rewriting, and external-model service charges separately.

What changed and who should check first

Cloudflare's GA announcement combines native image embeddings, PDF OCR, larger-file handling, and billing arrangements in one update. This article verified public information on October 2 Beijing time and did not run a real account or index for readers. Existing knowledge bases should estimate reindexing costs first; new projects can test retrieval quality with a small number of texts, scans, and images.

The official getting-started documentation offers dashboard, Wrangler, Workers, and Python/REST paths. The product is available to Workers Free and Paid accounts, but instance, file-count, and other limits vary by plan. Further audio and video support is a future direction in the announcement, not a currently validated capability.

Separate free pools from overage rates

According to the official Limits and pricing, each account has monthly pools of 5 million ingestion tokens, 10 GB-month of storage, 1,000 semantic searches, and 1,000 full-text searches. Semantic and full-text search are separate pools, not one interchangeable allowance of 2,000 searches. Images and OCR share the same ingestion-token pool.

Base ingestion overage costs $0.75 per million tokens, with an additional $0.50 per million tokens for image processing. Storage overage costs $2 per GB-month. Semantic search, whether vector or hybrid, costs $0.75 per thousand requests, and full-text search costs $0.10 per thousand. Ingestion is measured on final chunks, including overlapping tokens. OCR-extracted text incurs base ingestion plus the image-processing surcharge.

In a hypothetical text-only month with 6 million ingestion tokens, 12 GB-month of storage, 3,000 semantic searches, and 3,000 full-text searches, subtracting each free pool gives AI Search overages of $0.75 + $4 + $1.50 + $0.20 = $6.45 across these four items. This is neither an actual bill nor total cost: it excludes answer generation, query rewriting, external model services, taxes, and other account usage. Prices reflect public documentation on October 2; check again before actual settlement.

Not every PDF gets 10 MiB, and OCR is not automatically enabled

The data-source documentation distinguishes PDF limits of 10 MiB with OCR and 4 MiB without OCR. Plain text and code allow 10 MiB; other rich-text formats allow 4 MiB. Oversized files do not index successfully. Check error logs rather than assuming uploaded information is searchable.

OCR is available to all accounts but disabled by default. Set indexing_options.use_ocr to true when creating or updating an instance. Changing this setting triggers a full reindex. For large existing libraries, record file count, chunk volume, and budget before scheduling the change. Test scan recognition directly; OCR does not guarantee complete accuracy for tables, charts, or low-resolution images.

Confirm the embedding model before retrieving images

The supported-model table lists @cf/qwen/qwen3-vl-embedding-2b as the native image model, supporting images, 1024-dimensional vectors, and 32768-token input. The default @cf/qwen/qwen3-embedding-0.6b does not support native image embeddings. Enabling the product does not mean a multimodal model is selected.

With a text embedding model, images can follow a text-description path; native multimodal vectors retain a different representation. Compare retrieved sources, misses, and false matches using identical questions rather than judging quality by model name alone. Eligibility and fees for external providers' models in the table are separate and do not inherit the billing boundaries of built-in Workers AI models.

Complete a prelaunch acceptance run with a small sample

First, create a test instance using an official getting-started path, selecting built-in upload, your own R2, or a website data source. Include only material you are authorized to process. Prepare a small scanned PDF, an ordinary text document, and a few images with clear content, recording file sizes and expected answers.

Second, enable OCR as needed and select an embedding model that supports images. Check indexing status and error logs, identifying oversized, failed, and unindexed files separately. Third, retrieve using both keywords and semantic questions, verify original excerpts and sources in the answers, and include counterexamples that should receive an “unknown” response.

Fourth, record ingestion tokens, storage, and counts for both search types, and check answer-generation and query-rewriting model usage separately. Fifth, decide whether to expand the dataset or rebuild an old index. If quality or cost misses expectations, retain test results and adjust configuration instead of repeatedly ingesting the entire library.

Included costs and charges that need separate checks

Official billing documentation says Workers AI embeddings and reranking are included in AI Search charges, with AI Search storage, vector indexing, and website crawling also within the corresponding scope. Answer generation, query rewriting, and external models are still billed by their respective services. Existing R2 data-source buckets are not automatically emptied under the new storage arrangement, and legacy objects may continue to incur R2 costs. Review their purpose and retention needs before deleting anything; do not remove the only copy.

For poor retrieval, disabled OCR, oversized files, or exhausted free pools, first check the relevant configuration and metering evidence. Perform separate network troubleshooting only for clear DNS, TLS, or connection-timeout errors. Network changes cannot increase free allowances or remove product limits.

Sources

Frequently Asked Questions

Does AI Search charge immediately upon general availability?

The October 1, 2026 official announcement says billing starts November 1. The dates differ, and the pages provide no exact transition time or time zone. Recheck the current billing page before use.

Can the 1,000 semantic and 1,000 full-text searches borrow from each other?

Current public documentation lists two separate free pools and does not say they are interchangeable. Hybrid search belongs to the semantic-search category.

Can all PDFs be uploaded up to 10 MiB?

No. PDFs with OCR enabled allow 10 MiB; PDFs without OCR remain limited to 4 MiB. OCR is off by default, and changing the setting triggers a full reindex.

Can the default embedding model process images directly?

The default @cf/qwen/qwen3-embedding-0.6b does not support native image embeddings. Select the image-capable @cf/qwen/qwen3-vl-embedding-2b according to the official model table and test retrieval quality.

Does AI Search's free allowance include all model inference costs?

No. Built-in Workers AI embeddings and reranking have a defined included scope. Check answer generation, query rewriting, and external-model charges separately.