Back
Qwen 3.8 27B Joins Cerebras Public Endpoints at ~1,850 Tokens per Second
SiTech AI Team3 წთ. საკითხავი

Qwen 3.8 27B Joins Cerebras Public Endpoints at ~1,850 Tokens per Second

Cerebras has added Alibaba's 27B dense multimodal model Qwen 3.8 27B to its model catalog: a 64K/128K context window, $0.99 per million input tokens and roughly 1,850 tokens per second on public endpoints.

Cerebras now lists Qwen 3.8 27B among the models served on its public inference endpoints. Alibaba's 27-billion-parameter model sits alongside OpenAI's gpt-oss-120b in the company's Model Catalog, where it is quoted at roughly 1,850 tokens per second.

What the catalog lists

The model ID is qwen-3.8-27b. Cerebras describes it as Alibaba's 27B dense multimodal model for agentic coding, tool use, research and long-running workflows: it accepts text and image inputs and supports configurable reasoning. The context window is 64K tokens (65,536) on the Free Trial tier and 128K tokens (131,072) on paid tiers, with a maximum output of 32K (32,768) and 40K (40,960) tokens respectively. Pricing is listed at $0.99 per million input tokens and $1.49 per million output tokens. For comparison, the same catalog lists gpt-oss-120b at 120 billion parameters, a 65k/131k context window and about 3,000 tokens per second.

Rate limits and capabilities

On the Free Trial tier the limits are 5 requests per minute, 30K input tokens per minute, 90K total tokens per minute, 1M tokens per day and up to 2 images per request. The Developer tier allows 300 requests per minute, 150K input tokens and 450K total tokens per minute, up to 10 images per request and no listed daily token cap. Supported features include image inputs, reasoning, streaming, sampling controls, structured outputs, tool calling, parallel tool calling and prompt caching. Two endpoints are available: Chat Completions and Completions.

Limits worth knowing

Reasoning is enabled by default at the "high" setting and can be turned off by setting reasoning_effort to none. Image inputs work only through Chat Completions and must be base64-encoded PNG or JPEG data URIs; external image URLs, image detail controls, image generation, video and audio are not supported. The Completions endpoint returns text only and does not support images, reasoning controls, structured chat messages or tools.

Unpruned models and compression

The documentation states that all models on the public endpoints are the original, unpruned versions. Weight quantization is selective and used only for storage, in the 16-bit, 8-bit and 4-bit range; sensitive layers are kept at full precision with dequantization performed on the fly, while activations, attention and the KV cache remain fully unquantized. REAP pruned models, the result of the company's research into router-weighted expert activation pruning, are published on Hugging Face but are not served through the API. Cerebras says it will not change an existing model's architecture without notice, and that any future pruning would be offered as separate, clearly named endpoints.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.