46 3 weeks ago

Fastest For Quick Answers

cloud
ollama run treyleo16/haiku-4-5

Details

3 weeks ago

75e7bec42099 · 15kB ·

`<claude_behavior>` `<product_information>` Here is some information about Claude and Anthropic's pr

Readme

Claude Haiku 4.5

Claude Haiku 4.5 is Anthropic’s ultra-fast, high-efficiency model designed for near-instant responsiveness, high-throughput scaling, and cost-optimized execution. Delivering performance comparable to flagship models from previous generations at a fraction of the latency and operational cost, Haiku 4.5 brings extended reasoning, vision, and agentic tool execution to light-tier deployments.


Model Overview

Haiku 4.5 is optimized for real-time interactions, large-scale parallel processing, and multi-agent systems where execution speed and low per-token cost are critical. It bridges the gap between lightweight models and frontier reasoning capabilities.

  • Developer: Anthropic
  • Model Tier: Haiku (Lightweight / Ultra-Fast Efficiency)
  • Architecture: Autoregressive Transformer with Native Vision and Extended Thinking Support
  • Primary Target Use Cases: Real-time conversational AI, high-volume classification/summarization, agentic sub-task orchestration, computer use automation, and rapid code completion.

Core Capabilities & Performance Profile

Ultra-Low Latency & High Throughput

Haiku 4.5 is built for real-time responsiveness. It streams responses rapidly with sub-second time-to-first-token (TTFT), making it suitable for live customer support, streaming agent turn-taking, and sub-agent task loops.

Extended Thinking

For the first time in the Haiku model family, Haiku 4.5 incorporates Extended Thinking. When configured, the model can pause to generate hidden or summarized chain-of-thought tokens prior to outputting its final response, enabling higher accuracy on coding, logic puzzles, and multi-step deduction tasks.

Agentic Tool & Computer Use

  • Sub-Agent Orchestration: Ideal for acting as worker nodes in hierarchical multi-agent architectures (e.g., Sonnet/Opus handling global planning while Haiku instances run execution loops in parallel).
  • Computer Use Capabilities: Supports screenshot analysis, click navigation, keyboard inputs, and terminal interface interactions.
  • Native Tool Integration: Executes parallel tool calling, JSON API formatting, and bash command execution with low error rates.

Context Awareness

Haiku 4.5 tracks its own context consumption within its 200,000-token window, allowing agents to dynamically budget prompt space, prune context windows, or trigger self-summarization loops.


Technical Specifications

Parameter Specification
Model Type Multimodal Large Language Model (MLLM)
Context Window 200,000 tokens
Max Output Tokens Up to 64,000 tokens
Input Modalities Text, Code, Images, Structured Documents (PDFs)
Output Modalities Text, Code, Structured JSON, Reasoning Tokens
Features Extended Thinking, Computer Use, Context Awareness, Parallel Tool Calling

Alignment & Deployment Efficiency

  • Cost Efficiency: Engineered for cost-sensitive, high-volume workloads, dramatically reducing inference costs compared to flagship models.
  • Safety & Alignment: Trained with Anthropic’s Constitutional AI protocols, achieving low misaligned behavior rates and strict prompt-injection guardrails.
  • Safety Metrics: Evaluated to show significantly lower rates of misaligned behavior relative to earlier generations, balancing rapid execution with strict compliance. *