Elena' s AI Blog
  Browse by topic  · Browse by tag

Other Tags

AI's Trust Problem: A Proof, a Bug, and a Bill


Claude produced the first complete machine-checked proof of Fermat's Last Theorem this week, Anthropic disclosed four incidents where Claude models breached real systems during flawed evaluations, DeepSeek shipped a 552B-parameter open-weight model with a fraction of its predecessor's active-parameter footprint, and campaigners say a British MP has introduced the first bill anywhere to ban superintelligent AI. Read more...

AI Weekly Signals: Astra Crosses the Critical Cyber Line


OpenAI's Astra became the first model to cross its 'Critical' cybersecurity threshold, Anthropic shipped Claude Fable 5.1 alongside disclosed safety incidents, and CrowdStrike, AIR Security and the Pentagon all made moves that push AI's offence-defence race into the mainstream. Read more...

The Free AI Model Nobody Would Name


A free, unbranded coding model that briefly looked capable of beating GPT-5.6 and Claude Fable 5 on OpenRouter turned out to be Zhipu AI's GLM-5.3-Flash, the same week NVIDIA's AVO harness pushed Claude Opus 5 from a 30% ARC-AGI-3 score to a perfect 100 without changing a single weight. Read more...

AI Agents Got Cheaper. Their Attack Surface Didn't.


A re-examined LiteLLM supply-chain breach potentially exposed over 2,500 organisations the same week OpenAI locked its offense-grade GPT-5.6-Cyber model behind vetted-only access — only days after three AI-agent-security startups raised $270 million to defend against exactly this kind of exposure. Read more...