Fine-Tuning Frontier Open-Weight Models
Most fine-tuning quietly overwrites a model's reasoning in exchange for surface-level task compliance — you gain a format and lose a mind. Our approach protects the model's native thinking by construction, and spends the training budget on data that actually improves the model: demonstrations are screened to remove examples that game the task rather than solve it. The same recipes run on our own GPU clusters and on hosted training APIs alike, and the strongest results go back to the community on Hugging Face.
The Smaug line is the result. The motivation came from our own product: we build self-improving agents that automate complex work, and running them at scale kept hitting the same wall — long agentic loops with large contexts, repeated tool calls, and prompt caches that break down across time gaps are expensive and fragile. So we developed a fine-tuning methodology that combines human-curated, real-world agentic traces with synthetic data grounded in hard examples, and applied it to three open-weight bases at different points on the capability–efficiency curve: Smaug Flash on DeepSeek V4 Flash, Smaug Mini on Qwen3.8 27B, and Smaug Agentic on Kimi K3. The results are consistent across all three: meaningful lifts on the benchmarks that matter for production agents — agentic coding, real-world tool use, automation, and long-context reasoning and instruction following.