26 Downloads Updated 3 days ago
ollama run igovet/kimi-k2.6-opencode
Updated 3 days ago
3 days ago
f5b9947d146a · 418B ·
Modelfile & provider preset tuned for stable work inside OpenCode over Ollama Cloud.
When using Ollama Cloud models from OpenCode, several issues surface out of the box:
The settings below were arrived at empirically and raise the stability of Ollama Cloud + OpenCode to roughly 95%. The remaining edge cases look like Ollama Cloud throughput / model overload bugs, not configuration problems.
⚠️ These settings are experimental. They are not endorsed by Ollama or OpenCode — they are what happened to work best in our environment.
| Field | Value |
|---|---|
| Base model | kimi-k2.6:cloud |
| Context window | 262144 |
| Output limit (in OpenCode) | 262144 (equals context) |
| Temperature | 1.0 |
| Top-p | 0.95 |
| Repeat penalty | 1.15 (last 2048 tokens — slightly stronger than the other presets) |
| Variants | (no variants block — model has no reasoningEffort surface) |
FROM kimi-k2.6:cloud
PARAMETER num_ctx 262144
PARAMETER num_predict 16384
PARAMETER temperature 1.0
PARAMETER top_p 0.95
PARAMETER repeat_penalty 1.15
PARAMETER repeat_last_n 2048
{
"$schema": "https://opencode.ai/config.json",
"provider": {
"ollama": {
"npm": "@ai-sdk/openai-compatible",
"name": "Ollama",
"options": {
"baseURL": "http://localhost:11434/v1",
"timeout": 1200000,
"headerTimeout": 1200000
},
"models": {
"igovet/kimi-k2.6-opencode": {
"_launch": false,
"name": "Kimi K2.6 OpenCode",
"limit": {
"context": 262144,
"output": 262144
}
}
}
}
}
}
output is set equal to contextOllama does not honor a separate output (max output tokens) reliably through the OpenAI-compatible surface that @ai-sdk/openai-compatible speaks to. If you set output near 16384 you will see the stream cut off mid-response with no error.
The workaround used here is to make output formally equal to context. The model still decides when to stop on its own; we just stop clipping it on the client side.
timeout and headerTimeout are bumped to 20 minutes (1200000 ms). Cloud models occasionally queue for several minutes during peak load, and the default AI SDK timeouts will fire long before that.
Kimi K2.6 tends to echo recent tool outputs verbatim when given large tool traces. repeat_penalty 1.15 (vs. 1.1 for most other presets) tames this without hurting normal generation quality.
This preset does not declare variants, so the OpenCode reasoning-effort selector will be unavailable for this model. Use the default model selection.
If you want explicit effort control anyway, you can add the standard block — Ollama will ignore unknown values rather than error out:
"variants": {
"high": { "reasoningEffort": "high" },
"medium": { "reasoningEffort": "medium" },
"low": { "reasoningEffort": "low" },
"none": { "reasoningEffort": "none" }
}
Recommended use inside OpenCode (when variants are added):
| Variant | Use it for |
|---|---|
high |
Multi-step orchestrator tasks, planning chains, complex bug hunts across files. |
medium |
Default sub-agent workload (backend-developer, refactorer, full-stack-developer). |
low |
code-reviewer, qa-engineer, security-auditor on small scopes. |
none |
Inline completions / quick classification. |