351 Downloads Updated 2 weeks ago
ollama run brnpistone/Qwen3.5-4B-AgentCoder-q6-k
Updated 2 weeks ago
2 weeks ago
c73a04c05f45 ยท 3.5GB ยท
Qwen3.5-4B-AgentCoder-Q6-K is a fine-tuned version of the Qwen/Qwen3.5-4B model, optimized for: - ๐งฎ Complex reasoning tasks - ๐งฐ Tool calling - ๐ป Code generation
The model was developed through sequential fine-tuning, followed by a Direct Preference Optimization (DPO) post-training stage to improve alignment, coherence, and reasoning accuracy.
Qwen3.5-4B-AgentCoder-Q6-K can be used directly for: - โ Tool calling in complex reasoning tasks - โ Code generation for Python, JS, and other languages - โ Multi-domain reasoning (math, logic, Q&A)
ollama run brnpistone/Qwen3-4B-AgentCoder-q6-k
After sequential fine-tuning, the model underwent a DPO phase to enhance response alignment, reasoning robustness, and factual consistency.
3e-6141sigmoid27Objective - Encourage the model to prefer chosen completions - Improve clarity, correctness, and helpfulness - Reduce hallucinations and verbosity
Hardware
- GPU: NVIDIA H100 (80 GB VRAM)
- System RAM: 2 TiB
- Memory per vCPU: 10.67 GiB
Software
- Python: 3.12
- Transformers: 5.3.0
- Libraries: bitsandbytes, safetensors, torch, trl, scikit-learn, tokenizers, psutil, py7zr
BibTeX
@article{qwen3.5-4b-thinking-2507-toolcode,
title={Qwen3.5-4B-AgentCoder-Q6-K: A Fine-Tuned Model for Enhanced Tool Calling, Code Generation, and Reasoning},
author={Bruno Pistone},
year={2025},
journal={Hugging Face Model Hub}
}
APA
Bruno Pistone. (2026). Qwen3.5-4B-AgentCoder-Q6-K: A Fine-Tuned Model for Enhanced Tool Calling, Code Generation, and Reasoning.
๐ง Qwen3.5-4B-AgentCoder โ created by Bruno Pistone
Enhanced reasoning, tool calling, and code generation โ refined with DPO alignment