211 Downloads Updated 3 weeks ago
ollama run brnpistone/Qwen-3.8-27B-AgentCoder-Q5-K-M
ollama launch claude --model brnpistone/Qwen-3.8-27B-AgentCoder-Q5-K-M
ollama launch opencode --model brnpistone/Qwen-3.8-27B-AgentCoder-Q5-K-M
ollama launch hermes --model brnpistone/Qwen-3.8-27B-AgentCoder-Q5-K-M
ollama launch openclaw --model brnpistone/Qwen-3.8-27B-AgentCoder-Q5-K-M
Qwen-3.8-27B-AgentCoder-Q5-K-M is a fine-tuned version of the Qwen/Qwen3.8-27B model, optimized for: - ๐งฎ Complex reasoning tasks - ๐งฐ Tool calling - ๐ป Code generation
The model was developed through sequential fine-tuning, followed by a Direct Preference Optimization (DPO) post-training stage to improve alignment, coherence, and reasoning accuracy.
Qwen-3.8-27B-AgentCoder-Q5-K-M can be used directly for: - โ Tool calling in complex reasoning tasks - โ Code generation for Python, JS, and other languages - โ Multi-domain reasoning (math, logic, Q&A)
After sequential fine-tuning and preference alignment, the model underwent a GRPO phase in which an LLM judge โ not a static preference dataset โ supplied the reward signal. For each prompt the policy samples a group of completions, the judge scores each one against a rubric, and the advantage is computed relative to the group mean, so no value network is required.
Reward model (RLAIF)
- Judge: moonshotai/Kimi-K3 via the Amazon Bedrock Converse API
- Rubric-scored 0โ10, normalised to 0โ1; rewards evidence-gathering, correct tool choice, root-cause reasoning and verification, and penalises speculation, wrong-tool calls and symptom patching
- Unscorable completions (content filter, malformed verdict) are excluded from the group baseline rather than scored zero
GRPO parameters
- Generations per prompt (group size): 8
- Beta (KL coefficient): 0.04
- Loss type: dapo
- Reward scaling: group (unit variance within each group)
- Sampling temperature: 1.0
- Max completion length: 640 tokens, truncated completions masked out of the loss
Training parameters
- Learning rate: 1e-5, cosine schedule with a 10% floor
- Warmup steps: 46
- Per-device batch size: 2 ร 16 GPUs ร gradient accumulation 4 = 128 completions (16 prompts) per step
- Epochs: 1 โ 230 optimiser steps
- Max grad norm: 0.5
- LoRA: rank 16, alpha 32, dropout 0.1, applied to both the full-attention (q/k/v/o_proj) and linear-attention (in_proj_qkv, in_proj_z, out_proj) projections
- Precision: bf16, FSDP full-shard
- Hardware: 2 ร ml.p5.48xlarge (16 ร H100), 21h47m wall clock
Objective - Improve clarity, correctness, and helpfulness - Reduce hallucinations and verbosity
Outcome - Mean judge reward rose from 0.615 (first 25 steps) to 0.689 (final 25), +12.0% - Mean completion length and truncation rate both fell, i.e. the gain came from concision, not padding - Policy KL from the reference model rose from ~2.6e-4 to ~8.0e-3, confirming the policy actually moved
Hardware
- GPU: NVIDIA H100 (80 GB VRAM)
- System RAM: 2 TiB
- Memory per vCPU: 10.67 GiB
Software
- Python: 3.12
- Transformers: 5.3.0
- Libraries: bitsandbytes, safetensors, torch, trl, scikit-learn, tokenizers, psutil, py7zr