51 5 months ago

Deterministic 4B RAG extractor. 128k context, zero-hallucination, verbatim-only, and native PII shield. Built for Seed 42 reproducibility. Secure, ultra-fast, production-ready.

4b
ollama run lokeshjothiram/rag-distiller-v1:4b

Details

5 months ago

fa0afa7e3db7 · 2.5GB

phi3
·
3.84B
·
Q4_K_M
Microsoft. Copyright (c) Microsoft Corporation. MIT License Permission is hereby granted, free of ch
{{ if .System }}<|system|> {{ .System }}<|end|> {{ end }}{{ if .Prompt }}<|user|> CONTEXT: --- {{ .P
### ROLE: You are 'RAG-DISTILLER-CORE-V1' — a high-speed, zero-hallucination data extraction engin
{ "num_ctx": 8192, "num_predict": 1024, "repeat_penalty": 1, "seed": 42, "stop":

Readme

RAG-DISTILLER Logo

RAG-DISTILLER-CORE-V1 (Optimized)

High-Speed Factual Extraction & Zero-Hallucination Firewall

RAG-DISTILLER-CORE-V1 is a security-hardened utility model designed to act as the “Factual Firewall” in RAG pipelines. This version is performance-optimized for low-latency extraction, making it ideal for real-time applications like STOCKU and Voice AI.

⚡ Performance Specifications

Metric Specification Status
Architectural Base Phi-4-mini (3.8B) ✅ Verified
Context Window 8,192 Tokens (Fast-Path) 🚀 Optimized
Latency ~2.5s (Standard Hardware) ⚡ Sub-5s Goal
Reproducibility Seed 42 🔒 Deterministic
Security PII Shield & Injection Defense 🛡️ Hardened

🛠️ Key Enhancements (V1 Update)

  • Context Optimization: Reduced from 128k to 8k to eliminate VRAM swapping and reduce pre-fill latency by ~80%.
  • Deterministic Logic: Temperature set to 0.0 with simplified sampling for character-perfect verbatim extraction.
  • Refusal Precision: When information is missing, the model returns exactly: NO_RELEVANT_DATA.

🛡️ Security & Integrity

  • Privacy Shield: Automatic redaction of Passwords, API Keys, and IP Addresses as [REDACTED].
  • Verbatim Integrity: Preserves technical units and markdown citations (e.g., [Source A]) word-for-word.
  • Prompt Injection Defense: Neutralizes “jailbreak” instructions hidden within retrieved context chunks.

🚀 Quick Start

ollama pull lokeshjothiram/rag-distiller-v1:4b

💻 Implementation Pattern

Ideal for Agentic Workflows, Local Search, and Voice AI where speed and factual grounding are non-negotiable.

Example CLI Usage

ollama run lokeshjothiram/rag-distiller-v1:4b "Find the server status. CONTEXT: The backend is healthy. [1] API Latency is 20ms. [2]"