Back
Ember-1 vs. Kimi K3: nearly identical results at 3.4 times the speed
SiTech AI Team3 წთ. საკითხავი

Ember-1 vs. Kimi K3: nearly identical results at 3.4 times the speed

Fireworks Research launched Ember-1 on September 23 as a research preview built on Moonshot's open-weight Kimi K3. Across three progressively harder test sets the two models scored nearly identically, but Ember-1 finished 3.4 times faster.

Fireworks Research launched Ember-1 on September 23 as a research preview. The model is built on Moonshot's open-weight Kimi K3, and Fireworks claims it matches Kimi K3's quality using roughly 40% fewer tokens. The company says Ember-1 "learned to cut unnecessary reasoning while keeping the thinking that matters." Reasoning tokens bill as output, so a model that thinks less should cost less.

There is a catch. On OpenRouter, Ember-1 costs $3 per million input tokens and $15 per million output tokens. Fireworks charges the same for Kimi K3, but other providers sell Kimi for as little as $1 and $9.

How the tests were run

The New Stack ran both models through OpenRouter with identical prompts and their default reasoning settings, routing both to Fireworks so the speed comparison stayed fair. Each test ran five times per model, with reasoning tokens logged separately.

The three test sets grew progressively harder: logic puzzles with 4, 5 and 7 engineers; a deploy schedule covering 12 services with dependencies, one deploy per team at a time and two blackout windows, where the fastest possible rollout is 17 hours; and five probability questions about a retry system whose server flips between healthy and degraded, answered as exact fractions.

Where the two models landed

Both models solved all three logic puzzles on all five runs. Ember-1 averaged 13,630 reasoning tokens, 3 minutes 46 seconds and $0.27 per run; Kimi K3 averaged 16,679 reasoning tokens, 12 minutes 26 seconds and $0.34, which is 18% more reasoning tokens. Kimi K3's slowest run took nearly 20 minutes.

On deploy scheduling, both models found the 17-hour plan every time and every schedule passed the checker. Ember-1 averaged 6,543 reasoning tokens, 1 minute 29 seconds and $0.10; Kimi K3 averaged 7,792 tokens, 4 minutes 46 seconds and $0.13, or 16% fewer reasoning tokens for Ember-1.

The probability set produced the only miss. Kimi K3 answered all five questions correctly on every run, while Ember-1 got them right on four of five. On the fifth run it made a small arithmetic slip on the first question, writing 0.94619 instead of 0.94629, and the error carried into two other answers.

Faster, but not more accurate

Across all runs Kimi K3 was perfect 15 of 15 times, Ember-1 14 of 15. At Fireworks pricing ($3/$15) the full test set cost $2.48 with Ember-1 and $3.26 with Kimi K3, making Ember-1 about 24% cheaper. That edge is provider-dependent: at Kimi's lowest listed price ($1/$9) the same runs would have cost $1.96, less than Ember-1.

One difference was not in the marketing claims. Ember-1 finished each set of tests 3.4 times faster than Kimi K3 and used 23% fewer reasoning tokens overall. The two models are close in accuracy: Ember-1 is the faster choice for near-identical results, while a user who wants to spend less can route Kimi K3 to a cheaper provider on OpenRouter.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.