What we're
imagining.
Three artifacts compose the Kōzu lab — fine-tuned models, the apparatus that trains them, and the datasets that shape them.
Models
Intelligence, distilled.
Deimos R1
Sharper reasoning. Fewer tokens.
Our Satellite-class reasoning specialist — carrying the Deimos line's terse, token-efficient chain-of-thought forward into a full release.
Hugging FaceEuropa B1
More intelligence per parameter.
Our medium-sized model, exploring how intelligence scales per parameter while further refining the experimental reasoning techniques from Deimos.
Hugging FaceGanymede A1
More correct answers. Less reasoning overhead.
Our flagship model — scaling the techniques we've refined into a model tested and validated for real-world scenarios where reasoning efficiency is key.
Hugging FaceTools
Instruments of the craft.
Hadron
An LLM distillation framework built around NousResearch's AutoReason tournament refinement. A single teacher answers, critiques, adversarially revises, synthesizes, and blind-Borda ranks itself until "do nothing" wins twice. Distillation labels measurably beat the teacher's own single-shot output. Full reasoning traces per role — drop them straight into process-supervision fine-tuning.
GitHubTokamak
Extracts reasoning traces from LLM conversations and compresses them into a super token-efficient stream of internal chain-of-thought and concise outputs. Built for generating tight, high-signal training data from long, branching dialogues — without losing the reasoning that got you there.
GitHubStellarator
A control plane for fine-tuning and reinforcement-learning workloads on Tinker. Sandbox runs feed a structured pre-flight gate before promotion to scale, with cost projections, budgets, and live alert streams threaded through every step. A Rust supervisor handles per-job polling and websocket fan-out; an integrated research subsystem cites HF papers, arXiv, and code examples per run.
GitHubDatasets
Signal, isolated.
Quark
Our first dataset — built for concise chain-of-thought reasoning (CCoT) and token efficiency. Packs additional reasoning steps inside the same output footprint, so models think further per token instead of spending tokens to think.
Hugging FacePhoton
A compliance-hardening supervised fine-tuning (SFT) dataset family for large language models. Rather than bolting on guardrails at inference time, Photon bakes decision-tree-structured compliance rationales directly into model weights during fine-tuning — making compliant behavior a first-class model capability, not an afterthought. Each domain variant is called an isotope. Isotopes share a common schema and training philosophy but target distinct regulatory and security surfaces.