Escape frontier costs
Frontier models are brilliant — and metered. Run them around the clock and the bill climbs with every token. Keikaku lets you move the bulk of the grind onto your own hardware at a fixed cost, and keep frontier models on tap only when you want them.
What it is
A hedge against per-token pricing. Instead of paying a frontier provider for every task an agent runs, you run open coding models on a GPU you already own (or rent once). The marginal cost of another task drops to roughly zero, and your spend stops scaling with usage.
How it works in Keikaku
- Your GPU, open models. Agents run open coding models via Ollama — no per-token charge, and your code never leaves your machine.
- Right-size the model. The built-in benchmark measures real models on your hardware and recommends one that fits your VRAM, so you get the best quality your GPU can actually run.
- Mix in frontier, on your terms. Keep a frontier option for the hard problems — connect your own Claude (Code/Desktop) over MCP, or point an agent at a frontier model with your own API key. You pay your provider directly; we never mark it up.
- Route by role. Per-project agent roles let you send routine implement/ test work to local models and reserve a frontier agent for planning or tricky reviews.
Where it shines
High-volume, long-running work where frontier costs would dominate — big refactors, test-writing campaigns, migrations, or simply keeping a loop running 24/7. Predictable hardware cost instead of an unbounded metered bill.