Service eligibility and regional restrictions
PuppyIP serves only compliant overseas businesses and their authorized personnel. Proxy services are not available in mainland China. The service may only be used for lawful business activities outside mainland China. Use of this service within mainland China is prohibited.
Hosting a proxy IP or server overseas does not change these restrictions. The service must not be provided to end users in mainland China through relaying, forwarding, sharing or resale. Before use, read the Terms of Service.
Key Takeaways
- Perplexity released 0.6B/9B weights, with MIT license on official model cards.
- Their shared space allows 0.6B queries against 9B documents, not arbitrary old embedding or single-vector index compatibility.
- Text, images and visual documents are supported; PDFs need page-image rendering. No OCR does not mean no preprocessing.
- More vectors preserve detail but increase storage/scoring cost. API rollout is planned; current price and endpoint were not verified.
Choose document and query model sizes separately
Perplexity announced pplx-embed-v2-late October 7, 2026 with official research. Blog heading and citation both say October 7; model-file creation is not release time.
Embedding models encode comparable representations to retrieve information, rather than generate final chat answers. The 0.6B and 9B variants work independently or together.
Build an offline document index with 9B and encode real-time questions with 0.6B against the same space. Indexing can concentrate computation while queries need latency, letting teams choose each side's costs separately.
PDF retrieval still needs page rendering
OCR extracts text from images before traditional retrieval; recognition errors and lost layout affect tables/charts and relationships. The new model can retrieve rendered page images with text questions.
Bypassing text extraction is not dropping arbitrary PDFs into chat. Rendering, input processing, document encoding, indexing and retrieval still remain, with subsequent reading/answering needed after a hit.
Hypothetical unexecuted example: index chart/table-heavy manuals as pages and ask for a specification, returning original pages for verification. No guarantee is made for every low-resolution scan.
Multi-vector indexing differs from old single vectors
Single-vector retrieval compresses a passage to one vector. This model retains 128-dimensional vectors for multiple tokens and scores MaxSim: each query part finds the closest document part before combining similarity.
Late interaction keeps details separately matchable, but storage and scoring differ. Vectors grow with document length and increase index and candidate-scoring costs. Fewer model parameters do not guarantee lower total retrieval memory.
Shared space means these two variants, not arbitrary models. An old single-vector ANN index does not become compatible through a name change. Check multi-vector support, scoring and rebuilding before migration.
Compare three supported arrangements
For fully local work, consider 0.6B documents and queries. With more query compute, evaluate both at 9B. To reduce real-time costs, 9B documents plus 0.6B queries is the third supported option.
Official evaluation says mixed sizes improve over both at 0.6B, while both at 9B retain quality advantage. That is their datasets/workflow, not your recall, latency or cost guarantee.
Use a small authorized set of text, scans and images with expected pages. Compare hits, false retrievals, index size and query time before projecting the full corpus, rather than parameter counts or rankings alone.
Cards provide no universal minimum VRAM/hardware requirement. Lightweight-query design does not prove smooth operation on every phone. No weights were downloaded or devices tested here.
Released weights do not establish hosted API availability
Official Hugging Face cards are public, nonprivate and ungated with MIT license. Check model and dependency licensing before use. Public weights do not provide free compute or hosting.
Compatibility currently requires sentence-transformers 6.0.0+ and transformers 5.4.0+. Cards show text/image encoding but do not support mixing text and images as one input; use corresponding paths separately.
The blog plans gradual API Platform support for late-interaction, dense and contextual embeddings. Exact callable endpoints, prices and qualifications for these sizes were not verified; Search API pricing does not transfer.
Choose local weights or waiting for hosting after reading both cards. For local integration, have retrieval developers verify index/runtime; for API, use actual availability and billing docs.
Sources
Frequently Asked Questions
Does it generate PDF answers itself?
It primarily represents content for retrieval. Answering usually requires another model and evidence-reading stage; a page hit is not a correct fully cited answer.
Does local-model choice guarantee no upload?
The whole deployment determines where encoding, indexing, retrieval and answering occur. Local/cloud combinations are discussed; choosing 0.6B does not prove all processing local.
What about a question containing both image and text?
The current cards do not support one mixed text/image input. Define separate encoding/retrieval rather than borrowing another model's interleaved capabilities.