Back
Cohere's Embed 5 lets teams index with Pro and query with Fast
SiTech AI Team3 წთ. საკითხავი

Cohere's Embed 5 lets teams index with Pro and query with Fast

Cohere released Embed 5, letting teams index data with the Pro model and query the same vectors with the cheaper Fast model. In its tests, Fast queries scored 98.4 against a Pro-to-Pro baseline of 100.

Cohere released Embed 5 on Wednesday. The model lets teams index data with Embed 5 Pro and query the same vectors with the cheaper Embed 5 Fast, with no second index. Cohere recommends Pro for indexing and Fast for queries, especially in RAG and agent workloads where the same data is searched repeatedly and latency compounds.

There is a tradeoff in retrieval quality, though Cohere's testing suggests it is small. Across 40 datasets covering text, images and parsed documents, Fast queries against a Pro index scored 98.4 against a Pro-to-Pro baseline of 100; using Fast for indexing too dropped the score to 96.6.

One embedding space, two models

Pro and Fast share an embedding space, so teams can switch between them without re-embedding the corpus. Both produce compatible vectors at the same dimensions, and can be mixed with Matryoshka truncation or int8 quantization. Pro costs $0.12 per million tokens and Fast $0.08, at about 2.4 times the document throughput, so Pro handles documents entering the index while Fast absorbs the heavier query traffic.

Shrinking vectors with Matryoshka

Both support six vector dimensions from 256 to 2,048 in float32, int8 and binary formats. A 2,048-dimensional float32 vector takes 8 KB, or roughly 819 GB for 100 million chunks; a 1,024-dimensional int8 vector cuts that to about 102 GB, and a 256-dimensional binary vector to roughly 3.2 GB. Cohere recommends 1,024-dimensional int8 for most deployments, while binary compression suits an initial retrieval stage before high-precision reranking.

Retrieval beyond plain text

Embed 5 supports text, images and fused text-image inputs across more than 100 languages with a 128K-token context window. Cohere reports 82.3 for Pro on its five-dataset fused text-image evaluation, against 81.2 for Fast and 61.3 for Google's Gemini Embedding 2; on ViDoRe V3, Pro averaged 85.8 and Fast 84.5, against 77 for the earlier Embed 4. Multilingual results are less one-sided: Pro leads Cohere's five-language European average with 77, though it trails Gemini Embedding 2 on nine tests.

Reading the benchmark fine print

Embed 5 is Cohere's first model family evaluated with RCP-nDCG@10, which uses query-specific relevance criteria instead of fixed labels. Cohere says it can catch relevant results the original labels missed, but the metric measures reranking over a fixed candidate set, not first-stage retrieval from the full corpus, so scores are not directly comparable across evaluations.

Embed 5 also treats indexing and serving as separate decisions: one corpus can be indexed for retrieval quality while the query path is optimized for throughput and latency. Cohere's 98.4 score suggests little quality loss, but production systems still need to benchmark the pair on their own corpus. Embed 5 Pro and Fast ship through Cohere's API and Model Vault, Microsoft Foundry and Amazon SageMaker, with on-premises deployment via vLLM.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.