
Qwen 3.8 27B Joins Cerebras Public Endpoints at ~1,850 Tokens per Second
Cerebras has added Alibaba's 27B dense multimodal model Qwen 3.8 27B to its model catalog: a 64K/128K context window, $0.99 per million input tokens and roughly 1,850 tokens per second on public endpoints.
Cerebras now lists Qwen 3.8 27B among the models served on its public inference endpoints. Alibaba's 27-billion-parameter model sits alongside OpenAI's gpt-oss-120b in the company's Model Catalog, where it is quoted at roughly 1,850 tokens per second.
What the catalog lists
The model ID is qwen-3.8-27b. Cerebras describes it as Alibaba's 27B dense multimodal model for agentic coding, tool use, research and long-running workflows: it accepts text and image inputs and supports configurable reasoning. The context window is 64K tokens (65,536) on the Free Trial tier and 128K tokens (131,072) on paid tiers, with a maximum output of 32K (32,768) and 40K (40,960) tokens respectively. Pricing is listed at $0.99 per million input tokens and $1.49 per million output tokens. For comparison, the same catalog lists gpt-oss-120b at 120 billion parameters, a 65k/131k context window and about 3,000 tokens per second.
Rate limits and capabilities
On the Free Trial tier the limits are 5 requests per minute, 30K input tokens per minute, 90K total tokens per minute, 1M tokens per day and up to 2 images per request. The Developer tier allows 300 requests per minute, 150K input tokens and 450K total tokens per minute, up to 10 images per request and no listed daily token cap. Supported features include image inputs, reasoning, streaming, sampling controls, structured outputs, tool calling, parallel tool calling and prompt caching. Two endpoints are available: Chat Completions and Completions.
Limits worth knowing
Reasoning is enabled by default at the "high" setting and can be turned off by setting reasoning_effort to none. Image inputs work only through Chat Completions and must be base64-encoded PNG or JPEG data URIs; external image URLs, image detail controls, image generation, video and audio are not supported. The Completions endpoint returns text only and does not support images, reasoning controls, structured chat messages or tools.
Unpruned models and compression
The documentation states that all models on the public endpoints are the original, unpruned versions. Weight quantization is selective and used only for storage, in the 16-bit, 8-bit and 4-bit range; sensitive layers are kept at full precision with dequantization performed on the fly, while activations, attention and the KV cache remain fully unquantized. REAP pruned models, the result of the company's research into router-weighted expert activation pruning, are published on Hugging Face but are not served through the API. Cerebras says it will not change an existing model's architecture without notice, and that any future pruning would be offered as separate, clearly named endpoints.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.