Perplexity's open embedders search scanned PDFs and slide decks as page images with no OCR or chunking, but each page becomes hundreds of stored vectors, so your index can grow tenfold or more. The two models, pplx-embed-v2-late-9b and pplx-embed-v2-late-0.6b, landed on Hugging Face on October 7, 2026 under the MIT license. There is no hosted API for them yet, so a team that wants them today has to self-host the model and run a vector store that supports multi-vector search. That decision rests on two numbers: what the bigger index costs you, and whether the vendor's retrieval scores hold on your own documents.
If your RAG pipeline today is OCR, then chunking, then one embedding per chunk, this release is worth a pilot. Nothing published so far justifies a straight swap.
What Perplexity Actually Released
Perplexity released two late-interaction retrievers that share one embedding space, so you can build the index with the large model and query it with the small one. A late-interaction model (the ColBERT family) keeps a separate vector for every token or image patch in a document, then scores a query by matching each query token to its best document token and summing those maxima, an operation called MaxSim. A single-vector model squeezes the whole chunk into one vector before any query arrives.
Per the 9B model card, both models are built on Qwen3.5 with bidirectional attention, emit "one 128-dimensional vector per token," and handle text, images and visual documents through a vision encoder. The 9B model runs 7.4B active parameters; the 0.6B card lists 340M active. Both were distilled from an internal 18B ColBERT teacher.
The shared space matters most for cost. You pay the 9B model's compute once, at indexing time, and run the cheap 0.6B model on every query. Perplexity's community announcement, posted October 7, describes it as "a shared embedding space across model sizes" and says both models are "publicly available on Hugging Face."
What is missing: the 9B card says no inference provider currently deploys the model, and Perplexity's embeddings API docs list only the four v1 single-vector models, priced from $0.004 to $0.05 per million tokens. A hosted v2-late price does not exist yet, so the self-host math is the only math available.
How Good Are the Benchmark Numbers?
The scores published on the model card are solid, and they put the 9B model level with the best models of late 2025 rather than clearly ahead of them. On ViDoRe v3, the 9B card reports 65.2% nDCG@10 on page images and 64.7% on markdown text; the 0.6B model scores 62.3% and 61.2%. nDCG@10 is a ranking metric: it rewards a retriever for putting the relevant pages near the top of the first ten results.
For context, the ViDoRe v3 launch post from November 2025 describes a benchmark of 26,000 pages and 3,099 human-verified queries across finance, HR, energy, pharma and other enterprise domains, and reports nemo-colembed-3b at 0.656 on the English average. That comparison is loose. The launch post averaged seven English datasets, and the model card does not say which subset or languages produced its 65.2%, so the most the two figures support is that the models are in the same range.
The headline figure in trade coverage is 92.4% on MADQA, a question-answering benchmark over PDF pages. AI Weekly reports it along with 64.0% on BrowseComp+ when paired with GPT-OSS-120B. Those numbers do not appear on the model card, and Perplexity's own blog post blocked our fetch, so treat them as vendor claims relayed by a third party. AI Weekly flags the open question in the same piece: whether anyone reproduces the MADQA score on private corpora.
The small model gives up about three points on ViDoRe v3. For many document-search workloads that is a reasonable price for running roughly a twentieth of the active parameters on every query, but you only know that after you measure it on your own query set.
What Does Skipping OCR Save You?
Skipping OCR saves you a per-page fee and most of the indexing latency, and both are modest next to the storage you take on. Amazon's Textract pricing page lists Detect Document Text at $1.50 per 1,000 pages for the first million pages in US West (Oregon), and the Tables feature, which includes layout, at $15 per 1,000 pages. A one-million-page archive therefore costs $1,500 to $15,000 to OCR once, before you count re-runs when the parser changes.
Speed is the bigger operational win. The ColPali paper, which introduced this style of page-image retrieval, timed a conventional pipeline (layout detection, OCR, captioning, then embedding) at 7.22 seconds per page on an NVIDIA L4, against 0.39 seconds for ColPali to encode the page image directly. Those figures are for ColPali, not for Perplexity's models, and the 9B model is far larger, so expect slower encoding than 0.39 seconds. The structural point carries over: one model pass replaces a chain of four tools, each of which can mangle a table or drop a footnote.
Chunking also goes away for page retrieval. You index the page as an image, so you stop tuning chunk sizes and overlap, and charts and tables stay intact. We compared the OCR vendors in Textract vs Azure vs Gemini; late interaction removes the OCR step from retrieval, but you may still need extraction if a downstream system wants structured fields.
How Big Does a Multi-Vector Index Get?
In the two published measurements below, a multi-vector index ran about 20 to 64 times larger than a dense one before compression, and that is the number to put in front of your platform team. A 128-dimension float32 vector takes 512 bytes, and late interaction stores one per token.
- For text passages, Hugging Face's multi-vector encoder guide, published August 18, 2026, encoded 4,874 Natural Questions passages into 608,414 token vectors: 311.5 MB as raw float32, against 15.0 MB for a gte-modernbert-base dense index. That is about 20x. The same vectors took 92 MB as a compressed fast-plaid index, and pooling tokens by a factor of 2 cut the raw index to 156.4 MB.
- For page images, the ColPali paper reports 257.5 KB per page at float16, from 1,030 patch vectors. At that rate a million pages is roughly a quarter of a terabyte, against about 4 GB for one 1,024-dimension float32 vector per page, a gap of roughly 64x.
Perplexity's model card does not state how many vectors a page image produces, so neither figure is a v2-late estimate. Encode a thousand representative pages and measure the bytes yourself before anyone signs off on a migration.
Compression narrows the gap. The ColPali authors report that pooling with a factor of 3 removed 66.7% of vectors while keeping 97.8% of retrieval performance, and that binary quantization or centroid pooling "can reduce storage costs by two orders of magnitude." The Hugging Face guide recommends starting with a pool factor of 2 and measuring with an evaluator before going further.
Your vector store has to support it. The guide lists Qdrant, Weaviate, Vespa, LanceDB, Milvus and VectorChord as storing multi-vectors natively, and notes that OpenSearch and Elasticsearch can only rescore with MaxSim, not retrieve with it. Qdrant's vector docs also cap a single point at sub-vector count times dimension under 1,048,576, which at 128 dimensions means 8,191 vectors per point, ample for a page but a limit on stuffing a whole document into one point. Our vector database pricing piece explains why a storage multiple turns directly into a bill on most managed tiers.
The Index Now Leaks the Document
A multi-vector index of page images holds enough information to reconstruct much of the page, so it needs the same access controls as the source documents. A paper posted to arXiv the same day as Perplexity's release, "Inverting Multi-Vector Visual Document Indices", reports that pages inverted from raw indices recovered 47% of words and 45% of sensitive tokens on ViDoRe v3. The authors tested other multi-vector visual document retrievers, not pplx-embed-v2-late, but the attack targets the architecture.
The same paper found that token pooling or shuffling cut word recall to about 8%, though an attacker who restores a shuffled index's order raised source-page re-identification from 3.8% back to 93.5%. Pooling held up better, and the authors call inverting pooled indices an open problem. So the compression you apply to save storage also buys some privacy. If your security team classified embeddings as low-risk derived data, that classification does not hold for this index.
What to Do Before You Swap
Run a bounded pilot against your existing pipeline, and decide on storage cost and measured recall rather than on the release-day benchmarks.
This Week:
- Pull 1,000 pages that your current OCR pipeline handles badly: scanned contracts, invoices with tables, slide decks with charts. Encode them with the 9B model from Hugging Face (the card requires sentence-transformers 6.0.0 and transformers 5.4.0 or later, and text and image batches must be encoded separately) and record bytes per page and pages per GPU hour.
- Ask your security lead to classify the resulting index at the same sensitivity level as the source PDFs, citing the inversion paper.
This Month:
- Score the pilot on 200 hand-written queries, the method we laid out in RAG Eval Dataset: 200 Hand-Written Questions, comparing recall@10 for 9B-indexed and 0.6B-queried against your current dense index. If the small model's loss on your queries is near the three points ViDoRe v3 shows, query with the small one.
- Test pool factors of 2 and 3 on the same query set and price the surviving index on your current vector store. If the store can't retrieve with MaxSim natively, cost the migration as part of the decision.
Before Renewal:
- If you pay per page for OCR mainly to feed retrieval, put the pilot's numbers next to that contract. Keep OCR where a downstream system needs extracted fields; drop it where it only feeds search.
The Bottom Line
The top English scores in the ViDoRe v3 launch post came from ColBERT-style models such as nemo-colembed and colnomic, and storage is the usual reason teams pass on them. Perplexity's release lowers the query-side cost with a small model that reads a large model's index, and the MIT license removes the legal friction. The storage side is unchanged, and the scores you can verify today are level with existing open models, not ahead of them. Retriever choices hold up when they are measured on your own corpus, the way JPMorganChase ranked 62 retrievers for $800.
Measure bytes per page on your own documents before you plan the migration.
Continue Reading
- OpenAI vs Cohere vs Qwen3: Re-Embedding 10M Chunks Costs $520
- Textract vs Azure vs Gemini: Split OCR From Extraction
- Self-Host the Vector DB for Residency. Not for the Bill.
- GraphRAG vs Vector RAG: The 1,000x Index Bill Is the Small One
- RAG Build vs Buy: Buy the Index. Build the Eval Set.
- What RAG Actually Costs: $1,308 a Month at 10M Tokens/Day
