Elena' s AI Blog

Haiku 5.5 Is 90% Cheaper. The Real Cost Is in the Task

09 Oct 2026 (updated: 09 Oct 2026) / 26 minutes to read

Elena Daehnhardt

Generated with ChatGPT / OpenAI. Prompt: A futuristic AI control room showing a robot analysing a software vulnerability, a frontier model training system being slowed down, agent-oriented code infrastructure, and European AI governance. Deep blue, teal, violet and coral tech colours.
I am still working on this post, which is mostly complete. Thanks for your visit!


TL;DR:
  • Cheap small models are now priced for subagent fleets, but the capability gap is wide: Haiku 5.5 scores 39.2% on Terminal-Bench 4.0 against Sonnet 5.5's 70.6%, so routing between models is the design question, not replacement.
  • "Open-weight" is turning into a staged claim: Mistral Large 4 is a hosted preview with weights promised for the end of October and no named licence, while Google's EmbeddingGemma 2 shipped weights under Apache 2.0 on day one.
  • The agent failures reaching incident reports are mundane (sandbox edits and millions of API requests at Wikimedia), while access to cyber-capable models is hardening into tiers and New York City has a bill in committee that would require 24-hour incident reporting from AI contractors.

Introduction

Anthropic made its smallest model ten times cheaper on short prompts, Mistral announced an open-weight model you cannot yet download, and Wikimedia reported that some very keen AI agents had been poking at its tools. It was a week about the unglamorous parts of deploying models: price lists, licences, rate limits and who gets asked about it afterwards.

I have grouped the stories by theme. Last week’s issue is here if you need the background on Sonnet 5.5, Argon and the Senate bill.

In this issue:

  1. Claude Haiku 5.5: A Tenth of the List Price, Not a Tenth of the Model
  2. Mistral Large 4: Open-Weight on Paper, a Preview in Practice
  3. EmbeddingGemma 2: One Vector Space for Text, Images, Video and Audio
  4. OpenAI’s Decisions API Reaches Public Beta
  5. Wikimedia Says OpenAI Agents Probed Its Tools
  6. Anthropic Turns Cyber Access into Three Tiers
  7. New York City Asks the Labs to Promise Safety Under Oath

Frontier Models

1. Claude Haiku 5.5: A Tenth of the List Price, Not a Tenth of the Model

Anthropic favicon Claude Haiku 5.5 — Anthropic, 7 October 2026

Let's Data Science favicon Anthropic Releases Claude Haiku 5.5 for High-Volume Tasks — Let's Data Science, 7 October 2026

Anthropic released Claude Haiku 5.5 on 7 October. Its list price is $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens, rising to $0.50/$2.50 above that. Haiku 4.5 was $1/$5, so the headline rate is a 90% cut for those prompts, whereas Anthropic’s own claim, as relayed by AI Weekly, is “around 75% less to run on average”. Anthropic’s footnote explains part of the gap: the discount falls to 50% above 100,000 tokens, Haiku 5.5 has an updated tokenizer that uses slightly more tokens per task, and the average accounts for both. The full arithmetic is not published, so your own savings will depend on prompt length, output length, caching and how many tokens the model needs to finish a task. The model ID is claude-haiku-5-5, and it is available on AWS, Google Cloud and Microsoft Azure. It is also the first Haiku-class model with an adjustable effort setting, which gives you one more dial for balancing capability against cost.

Anthropic’s own table shows what the lower price buys. Haiku 5.5 reaches 72.4% on the offline subset of OSWorld 2.1 (Haiku 4.5: 15.7%) and 39.2% on Terminal-Bench 4.0, against 70.6% for Sonnet 5.5, which is the same figure I quoted for Sonnet last week. These are vendor-reported numbers. Two smaller changes arrived alongside: Sonnet 5.5 cache reads were halved from $0.20 to $0.10 per million tokens (Anthropic estimates about 20% cheaper on most agentic work), and Max and Team plans gain a monthly API credit of $100 (Max 5x), $200 (Max 20x) or up to $500 pooled (Team).

Why this matters

Haiku 5.5 is not a cheaper Sonnet. It is a different tool, and the Terminal-Bench gap says so plainly. You would not ask a very capable junior colleague to run the quarterly architecture review, but you would happily let them sort the post, and that is the job description Anthropic itself gives the model: summarisation, classification, compaction and subagent work.

The practical consequence is routing. If your agent fans out into twenty subagents that each read a file and report back, a 90% list-price cut changes the economics of the whole design, but if one of them has to debug a difficult build failure, the saving can vanish when a lower success rate means repeated attempts or escalation to a larger model. The benchmark gap does not prove that Haiku cannot fix a build; it says you should measure.

The number to track is cost per successful task: total cost of all attempts ÷ number of successful tasks, with retries, escalations and validation included. A cheap model that needs five attempts can cost more than a capable one that finishes first time, rather like sending the junior colleague back to redo the post until it is sorted. Measure it on your own workload before moving anything, and include the effort setting as one of the variables, because the 90% and 75% figures prove only that the answer depends on the token mix.


Open Models

2. Mistral Large 4: Open-Weight on Paper, a Preview in Practice

Mistral favicon Mistral Large 4 — Mistral AI, 6 October 2026

Mistral favicon Mistral Large model documentation — Mistral AI

Let's Data Science favicon Mistral Opens Public Preview of Large 4, With Weight Release Planned Later This Month — Let's Data Science, 6 October 2026

Local AI Zone favicon October 2026 AI Model Updates — Local AI Zone

On 6 October Mistral opened a public preview of Mistral Large 4 through Mistral Studio. It is a natively multimodal mixture-of-experts model of roughly one trillion parameters in total, with 52 billion active parameters, a 1.6 billion-parameter vision encoder and a one-million-token context window according to Mistral’s model documentation (the exact total is 1.05 trillion). A secondary report gives 49 billion active parameters, but I would go with the vendor’s own figure.

Mistral calls the model open-weight, but what shipped is a hosted preview: the weights are planned for the end of October and no licence has been named. The preview API is priced at $0.68 per million input tokens, $0.07 for cached input and $2.09 for output, which Mistral shows as discounted from $1.36, $0.14 and $4.18, so I would not budget against the preview rates. Vendor-reported benchmarks include 61.7% on DeepSWE v1.1. In a human code evaluation by Surge AI, which Mistral reports, the model scored 3.74 out of 5, second of five models behind Claude Opus 5 at 4.22. Independent tracking describes the preview as guardrailed and lists no architecture details beyond the headline numbers.

Why this matters

An open-weight model is only open once you can download it and read the licence. Until then it is a menu, and I would not plan around a dish that is “coming at the end of the month”. The licence matters as much as the weights: “open-weight” has covered everything from Apache 2.0 to terms that exclude commercial use, and the answer decides whether you can fine-tune and self-host.

The size is its own hurdle. By my arithmetic, 1.05 trillion parameters need roughly 2.1 TB of weights at 16-bit precision, 1.05 TB at 8-bit and around 525 GB at 4-bit, before quantisation metadata, caches and runtime overhead. The 52 billion active parameters reduce the compute per token, but you still have to store the inactive experts, so active parameters are not a memory requirement. For most readers the realistic uses are the hosted API now and a quantised variant, if one appears. A trillion parameters is the sort of number that arrives with its own shipping container.


3. EmbeddingGemma 2: One Vector Space for Text, Images, Video and Audio

Google favicon EmbeddingGemma 2 — Google, 6 October 2026

Google released EmbeddingGemma 2 on 6 October under Apache 2.0, with weights on Hugging Face and Kaggle. At 740 million parameters in total, it maps text, code, images, video and audio into one shared embedding space; text-only use needs just the 270M text core, with optional 170M vision and 300M audio encoders. Google reports that MTEB Code rises from 68.76 to 78.68 over the first EmbeddingGemma, and the context window grows to 8K tokens, four times the original.

For on-device use, Google says quantised text-only weights need about 191 MB of active RAM on a Pixel 11 Pro and the full multimodal model about 567 MB. Matryoshka truncation lets you cut the 768-dimension vectors to 512, 256 or 128, which leaves two thirds, a third or a sixth of the vector storage at the same numeric precision. The first EmbeddingGemma passed 20 million downloads, which is some evidence people will actually use this one.

Why this matters

A single embedding space for audio, video and text is useful, because a query such as “find the clip where someone mentions the invoice” no longer needs a separate index for each medium. It will not give you a transcript or an exact quote, though, so for precise spoken words you may still want speech recognition. The tooling is already there: Google lists transformers, sentence-transformers, MLX, vLLM, llama.cpp, SGLang, Ollama and LMStudio for serving, and Qdrant for storing the vectors, which suits local retrieval projects. My caution is migration cost: vectors from different embedding models are not comparable, so adopting this means re-embedding your whole corpus, a bit like re-cataloguing a library after changing the classification scheme. Run the numbers on your own data first, since a gain on MTEB Code does not guarantee a gain on your support tickets.

The Matryoshka option is the part I would try on day one. Truncating to 256 or 128 dimensions and checking whether retrieval quality survives takes an afternoon and tells you how much of your index is padding. Bear in mind that this saves vector storage only, not IDs, metadata or index structures.


Developer Tooling

4. OpenAI’s Decisions API Reaches Public Beta

OpenAI favicon Decisions API guide — OpenAI developer documentation

AI Weekly favicon OpenAI's Decisions API charges only for input, at $0.10 per million tokens — AI Weekly, 7 October 2026

Last week I described the Decisions API as a limited preview; it opened as a public beta on 6 October. It runs only on gpt-6-luna through POST /v1/decisions and returns typed answers instead of prose: a predicate (probability from 0 to 1 that a condition is true), a choice (probabilities across options you supply) or a score against ordered levels. Input costs $0.10 per million tokens and output, cache reads and cache writes are not charged, although the guide adds that regional processing premiums and long-context input multipliers apply, so check the pricing page for your configuration. It accepts text and images (images only as inline base64 data URLs), and offers Zero Data Retention, HIPAA eligibility for eligible customers, and data residency in the US and EEA/Switzerland.

OpenAI says general availability is expected “in the coming weeks”. OpenAI’s guide claims typed answers roughly 10x faster than the Responses API, a vendor claim with no published method or accuracy benchmarks, and AI Weekly notes that /v1/decisions is missing from the Luna model page’s endpoint support table, which lists only Chat Completions, Responses and Batch.

Why this matters

A classifier that hands back probabilities slots into a pipeline far more easily than a paragraph that has to be parsed. At $0.10 per million input tokens with free output, routing tickets or flagging damaged products becomes cheap enough that cost stops being the argument against using a model.

What I cannot tell you is whether the probabilities are calibrated, meaning whether “0.8” is right about eight times in ten. Nobody has published that, so I would test it on a labelled sample of your own before you set a threshold on it. Suppose you gather 100 support tickets for which the API returned about 0.8 for “urgent”. If it were calibrated, roughly 80 would be urgent; if only 55 were, the model is overconfident in that range and a 0.8 threshold would flood your queue (those numbers are invented, for illustration). Decisions is also not a replacement for Structured Outputs: it answers constrained questions, whereas Structured Outputs can produce an object that follows your own JSON schema. Also expect a beta to move: an endpoint missing from the documentation table is a sign the paperwork has not caught up with the launch.


Security

5. Wikimedia Says OpenAI Agents Probed Its Tools

Wikimedia Foundation favicon OpenAI “rogue” agent activities found on Wikimedia projects — Wikimedia Foundation, 5 October 2026

BleepingComputer favicon Rogue OpenAI agents behind potentially malicious Wikipedia edits — BleepingComputer, 6 October 2026

The Hacker News favicon Wikimedia says OpenAI agents tried to hack its tools — The Hacker News, 6 October 2026

On 5 October the Wikimedia Foundation, in a statement by its chief product and technology officer Selena Deckelmann, said it had found edits it believes came from AI agents operated by OpenAI. Almost all were test edits in sandbox areas and none appeared on pages that readers see. A few touched the configuration of a citation tool in what the Foundation called potentially malicious edits intended to misuse it as a proxy for fetching data from remote services, and attempts to compromise its public Etherpad note-taking tool, including attempts to use it as a proxy, failed. None of the community approvals that Wikipedia requires for bots were sought.

The traffic was larger than the edits. The agents made millions of automated requests to Wikimedia’s public APIs, crawled millions of Wikidata and Commons pages and sent hundreds of thousands of queries to the Wikidata Query Service, and the Foundation says the traffic may have contributed to a partial outage of that service in May. Wikimedia found no evidence that its systems were used for coordination between agents, or of compromised systems or data. OpenAI told The Verge, as relayed by The Hacker News, that it is working with the Foundation to review the activity.

Why this matters

This is a different sort of agent incident from the ones that make headlines. This was not a confirmed breach, but it was not harmless experimentation either: unauthorised edits, attempted use of Wikimedia’s tools as a proxy and heavy automated traffic cost volunteers and engineers time on infrastructure maintained for the public. An agent can cause real operational problems without ever compromising anything. The outage link is “may have contributed”, so I would not treat it as proven cause.

For anyone building agents the lesson is dull and actionable: identify your agent in the user-agent string, respect rate limits and robots rules, and cap requests per task in code, not in the prompt. At a minimum, a production agent should have:

  • a global request budget per task, covering retries and subagents;
  • per-domain concurrency limits, with exponential backoff on HTTP 429 and 503 responses;
  • domain allowlists, and approval for state-changing actions;
  • logs that link every request to the run responsible;
  • a circuit breaker that stops the agent when traffic or error rates pass a limit.

Assume that any agent you run will eventually appear in somebody else’s incident report, and make sure your logs can tell them whose it was.


6. Anthropic Turns Cyber Access into Three Tiers

Anthropic favicon Cyber Verification Program — Anthropic, 6 October 2026

CSO Online favicon Anthropic widens access to AI cyber capabilities for vetted security teams — CSO Online, 7 October 2026

The Hacker News favicon Anthropic expands Claude access for vetted security teams — The Hacker News, 7 October 2026

Anthropic expanded its Cyber Verification Program, now folded together with Project Glasswing, into three access tiers: Defense Access (incident response, malware reverse engineering, vulnerability validation), Red Team Access (authorised penetration testing, open to organisations rather than individual researchers) and Specialized Access, reserved for a limited set of verified organisations testing critical systems, with applicants reviewed alongside the US government. Covered models are Claude Opus 5.5, Sonnet 5.5 and Mythos 5.1. Members get reduced safeguards, and their data is retained so Anthropic can monitor for misuse. Anthropic’s announcement documents one exception that matters to teams handling confidential code: until Enterprise Frontier Safeguards arrives later this autumn, organisations with zero data retention can use the programme with Claude Fable 5.1 or Mythos 5.1.

Anthropic reports that Glasswing partners found at least 129,000 verified software vulnerabilities between April and July 2026, and a further 5,500 came from its own open-source scanning between April and October. It says more than 33,000 have so far been rated critical or high severity, and I would not assign that count to either group alone. In Anthropic’s CyScenarioBench test of Opus 5.5, all 50 trials were blocked at the first prompt without CVP access; under Defense Access 46 were blocked and 4 succeeded, and under Red Team Access none were blocked and 34 were completed. That is a vendor-run test. For balance, The Hacker News notes VulnCheck’s finding that only 2 of 300 Anthropic- or Glasswing-discovered vulnerabilities (0.67%) have been exploited in the wild.

Why this matters

Last week Argon went to defenders first; this week Anthropic published the paperwork that makes the same idea a product. I expect tiered, verified access to become standard for any model with serious cyber capability, which means security teams will need to apply for access, and the application process will become part of your tooling budget.

The 129,000 figure is a count of verified finds, not fixes. Discovery is now fast and cheap, while triage and patching still sit with maintainers who are, in my experience, rarely enjoying their week. VulnCheck’s observation that only two of 300 sampled vulnerabilities had been exploited in the wild is useful context, but it is not a measure of severity or of what attackers will do next; a flaw can be serious long before anyone is seen exploiting it. For defenders the hard part is prioritising, validating, assigning owners and getting fixes deployed.


Governance

7. New York City Asks the Labs to Promise Safety Under Oath

NYC Council favicon New York City Council Convenes Hearing on Artificial Intelligence Risks — NYC Council, 5 October 2026

Fox News favicon OpenAI, Anthropic, Meta, Google stop short of AI safety guarantee — Fox News live coverage, 6 October 2026

Unite.AI favicon NYC Council hearing puts Anthropic, OpenAI, Google and Meta under oath — Unite.AI, October 2026

NY1 favicon Major tech companies to testify at City Council hearing on AI safety — NY1, 5 October 2026

On Monday 5 October, representatives of OpenAI, Anthropic, Meta and Google were questioned under oath by the New York City Council at a rare Committee of the Whole hearing. Speaker Julie Menin asked each to quantify the worst-case catastrophic risk of AI, and to assure the public that its agents would follow safeguards. As Fox News reports, none would guarantee it: OpenAI’s Morgan Dwyer said no technology is without risk, Anthropic’s Logan Graham called controlling these systems an unresolved scientific problem, and Google’s Alice Friend said perfection is not achievable for any product. Menin replied that the Council wanted accountability and transparency, not perfection. Elon Musk’s SpaceXAI had been subpoenaed after not responding to requests to participate.

The agenda matters more than the testimony. Two of the bills are now on the record. Int 1100-2026 would bar marketing, selling or deploying an AI model in the city without third-party validation and a human-operated shutdown capability, with civil penalties of up to $25,000. Int 1109-2026 would require city contractors and agencies to report qualifying AI safety incidents to the city’s Cyber Command within 24 hours, and Cyber Command to disclose them publicly within 24 hours, with exclusions for trade secrets, cybersecurity, public safety and national security. A further proposal would create a private right of action where an AI provider failed to implement reasonable safeguards and someone was harmed. Both numbered bills were introduced on 8 October and are in committee; neither is law, and their obligations would fall on the providers, agencies and contractors they cover, not on every developer.

Why this matters

Refusing to promise zero risk is the correct answer, since no software team can sign that oath, but “we cannot guarantee it” will not survive many more hearings as a complete reply. The question is moving to what you can show: logs, shutdown paths and incident reports.

The bill I would read first is Int 1109-2026, because it is the one an ordinary developer team could meet indirectly, through a city contract. It is not law, and its direct obligations would apply to the entities it covers. But it raises a plain question for everyone building agents: could you reconstruct, within a day, what your system did, which tools it called and who authorised it? That is a useful engineering test whether or not New York passes the bill, and the Wikimedia story suggests that many teams could not pass it.


Closing Thoughts

Haiku 5.5 makes small models cheap enough to run in fleets, Mistral Large 4 and EmbeddingGemma 2 show two very different meanings of “open”, and the Decisions API turns a model call into a typed answer you can put in a pipeline. Wikimedia’s report shows what an untended agent looks like in practice, Anthropic’s tiers show access being rationed to the verified, and New York’s hearing shows the questions that follow. Put together, the week says that models are becoming parts of ordinary infrastructure, and the work now is metering, licensing and accounting for them. Let me know what you think.


References

  1. Claude Haiku 5.5 — Anthropic
  2. Anthropic Releases Claude Haiku 5.5 for High-Volume Tasks — Let’s Data Science
  3. Anthropic releases Claude Haiku 5.5, cutting prices by 75% — AI Weekly
  4. Mistral Large 4 — Mistral AI
  5. Mistral Large model documentation — Mistral AI
  6. Mistral Opens Public Preview of Large 4, With Weight Release Planned Later This Month — Let’s Data Science
  7. October 2026 AI Model Updates — Local AI Zone
  8. EmbeddingGemma 2 — Google
  9. Decisions API guide — OpenAI developer documentation
  10. OpenAI’s Decisions API charges only for input, at $0.10 per million tokens — AI Weekly
  11. OpenAI “rogue” agent activities found on Wikimedia projects — Wikimedia Foundation
  12. Rogue OpenAI agents behind potentially malicious Wikipedia edits — BleepingComputer
  13. Wikimedia says OpenAI agents tried to hack its tools — The Hacker News
  14. Cyber Verification Program — Anthropic
  15. Anthropic widens access to AI cyber capabilities for vetted security teams — CSO Online
  16. Anthropic expands Claude access for vetted security teams — The Hacker News
  17. New York City Council Convenes Hearing on Artificial Intelligence Risks — NYC Council
  18. OpenAI, Anthropic, Meta, Google stop short of AI safety guarantee — Fox News
  19. NYC Council hearing puts Anthropic, OpenAI, Google and Meta under oath — Unite.AI
  20. Major tech companies to testify at City Council hearing on AI safety — NY1
  21. Int 1100-2026, third-party validation and shut-down capability of AI models — NYC Council Legistar
  22. Int 1109-2026, reporting and public disclosure of AI safety incidents concerning city contracts — NYC Council Legistar
desktop bg dark

About Elena

Elena, a PhD in Computer Science, simplifies AI concepts and helps you use machine learning.




Citation
Elena Daehnhardt. (2026) 'Haiku 5.5 Is 90% Cheaper. The Real Cost Is in the Task', daehnhardt.com, 09 October 2026. Available at: https://daehnhardt.com/blog/2026/10/09/ai-weekly-signals-haiku-5-5-90-percent-cheaper-mistral-large-4-wikimedia-agents/
All Posts