Fireworks AI said its evaluation of open-source Kimi K3 and closed-source Fable 5 covered about 1,030 agent tasks across SWE, terminal operations, algorithms, multilingual coding, and legal scenarios. According to ChainCatcher, both models posted similar SWE benchmark accuracy, with K3 at 92.4% and Fable 5 at 92.6%.
The report said each model showed strengths in different areas: K3 led in symbolic math and developer tools, while Fable performed better in web and data visualization tasks. In long-horizon terminal tasks, K3 independently solved 11 tasks that Fable could not complete.
Fireworks AI also said a task-level routing strategy could dynamically allocate work between the two models and reach 93% accuracy, exceeding either single model alone. The report added that, on the Fireworks platform, K3 had a substantial cost advantage and could be up to 50 times cheaper than Fable for long-running agent tasks.
The report recommended using an open-source model as the default option and continuously learning the best task-model match through a router to balance quality and cost.