Anthropic’s Budget Model Excels at Ignoring Hidden Commands
Anthropic’s Claude Haiku 5.5 model improves vulnerability identification and exploit crafting while balancing security and performance.
Overview of Claude Haiku 5.5
Anthropic’s Claude Haiku 5.5 model, optimized for rapid, repetitive tasks and latency-critical applications, demonstrates improved capability in identifying vulnerabilities and crafting exploits compared to its predecessor. The system incorporates enhanced cybersecurity measures compared to Haiku 4.5 but maintains less stringent protections than its advanced counterparts, which retain superior offensive capabilities.
Testing and Performance
Performance in Vulnerability Identification
Testing conducted by the company evaluated the model’s performance under controlled conditions without active safeguards. In a series of experiments targeting known vulnerabilities in Chrome’s V8 engine, Haiku 5.5 achieved arbitrary code execution in 4 out of 410 trials.
Multi-Stage Cyber Operations
A separate assessment of multi-stage cyber operations revealed a pre-release version completing 3.3% of challenges, significantly lower than the 46.1% success rate of Sonnet 5.5 and 67.6% for Opus 5.5. These results do not reflect typical user scenarios under standard security configurations.
Security Protocols and Response Rates
Response to Harmful Queries
Eligible security professionals may request adjusted restrictions through Anthropic’s Cyber Verification Program. Evaluations of the model’s response to harmful queries, sensitive topics, and manipulated conversations covered areas such as weapons, extremist activities, surveillance, child safety, mental health, and election integrity.
Refusal Rates for Harmful Requests
With a near-final production system prompt, Haiku 5.5 achieved a 99.71% rate of harmless responses to malicious requests, improving from 3.05% refusal of non-harmful queries in Haiku 4.5 to 0.82%. In extended interactions, the model showed enhanced performance in influence operations and surveillance-related tasks but experienced reduced effectiveness in weapon-related scenarios when tested via API without system prompts.
Resistance to Attacks
Prompt Injection Attacks
Anthropic states that Haiku 5.5 represents the most resilient iteration of the Haiku series against prompt injection attacks, which conceal malicious instructions within external content to manipulate AI behavior. While its resistance matched frontier models in adaptive attack scenarios, it lagged behind Sonnet 5.5 and Opus 5.5 on a Gray Swan benchmark, with vulnerabilities primarily observed in graphical computing tasks.
Pricing and Cost Efficiency
Pricing Details
Pricing for Haiku 5.5 is reduced compared to Haiku 4.5, with the most significant cost savings for shorter requests. Input and output token pricing is 90% lower for prompts up to 100,000 tokens, which accounted for the majority of Haiku 4.5 usage. An adjustable configuration option allows users to prioritize cost efficiency or computational power.
Cost Savings and Features
A developer at Asana reported a 30% reduction in latency and 2.5x faster inference speeds when using Haiku 5.5 for AI agent tasks such as bug triage and project setup. The model can function as a subagent for Opus 5.5 and Sonnet 5.5. Anthropic has also reduced caching costs for Sonnet 5.5 by 50%, lowering overall AI agent operational expenses by approximately 20%.
Additional Features and Availability
Cloud Platforms and SDKs
The model is accessible across major cloud platforms including AWS, Google Cloud, and Microsoft Azure, with the identifier claude-haiku-5-5. The company plans to introduce monthly API credits for Max and Team Anthropic subscribers and is expanding beta support for computer and browser interactions in its Python and TypeScript SDKs.
Anthropic noted persistent challenges in handling conversations involving self-harm and eating disorders, advising developers to implement supplementary protections.
Anthropic states that Haiku 5.5 represents the most resilient iteration of the Haiku series against prompt injection attacks, which conceal malicious instructions within external content to manipulate AI behavior.
Anthropic has also reduced caching costs for Sonnet 5.5 by 50%, lowering overall AI agent operational expenses by approximately 20%.
