Running monthly digest of AI model, product, platform, and governance churn for May 2026 — terse, dated, claim-grouped bullets folded from ephemeral news so the firehose stays out of the thematic wiki. Read this for "what shipped / who raised / what broke in AI in May 2026"; verbatim source text stays in raw/. Split out of almanac-qa-2026-05 (an unrelated phareim.no theme audit that briefly shared the slug) on 2026-07-24.
Frontier model releases accelerated across every lab
- Claude Opus 4.8 (May 28) — agentic coding 64.3→69.2%, reasoning-with-tools 54.7→57.9%, 69.2% SWE-Bench Pro; fast mode 2.5× faster / 3× cheaper; effort-control slider on claude.ai; same price as 4.7 ($5/$25 per Mtok). Claude Code gains "dynamic workflows" for large-scale tasks; a coding-only Mythos model is teased for Claude Code.
- Google Gemini 3.5 Flash (May 19, GA same-day at I/O) — biggest Flash-tier release; 1M-token context, multimodal, $1.50/$9.00 per Mtok, ~4× faster than comparable frontier models, beats Gemini 3.1 Pro on coding/agentic benchmarks (76.2% Terminal-Bench 2.1). Became default in the Gemini app + Search AI Mode worldwide. Gemini 3.5 Pro confirmed in internal use for the following month. I/O also previewed Gemini Omni (any-to-any text/image/video), Gemini Spark (24/7 agentic assistant, gaining MCP support for Canva/Instacart/OpenTable), Antigravity 2.0, and DeepMind's Magic Pointer AI cursor.
- DeepSeek V4-Pro and V4-Flash (late May) — first production releases after the V4 Preview; MoE, 1M-token context, hybrid attention for low inference cost; Pro for quality, Flash for throughput; continues DeepSeek's cost-disruption pattern.
- xAI Grok 4.3 became the new default (May 15); xAI simultaneously retired 8 legacy Grok models, all redirecting to Grok 4.3.
- OpenAI GPT-Realtime voice suite (May 7) — GPT-Realtime-2 (first voice model with GPT-5-class reasoning), -Translate (70+→13 languages live), -Whisper (streaming STT); all in the API. OpenAI also deprecated DALL·E-2 and DALL·E-3 (May 12), pushing devs to gpt-image-1.5.
- Google Gemini 3.1 Flash Lite went GA (May 7) alongside Cloud Next '26.
Agentic coding tools are the central battleground
- Claude Code shipped a dense release train across May (v2.1.139→2.1.152): Agent View dashboard (
claude agents),/goalautonomous multi-turn command, Fast mode defaulting to Opus 4.7, plugin dependency enforcement, dynamic-workflows research preview, pinned sessions. v2.1.149 (May 22) added per-category/usagecost breakdown (skills/subagents/plugins/MCP), keyboard-scrollable/diff, GFM checkbox rendering, and several security fixes (PowerShell permission-bypass, sandbox write-allowlist not applied in git worktrees). - OpenAI Codex went mobile (May 14) — remote agent control from iOS/Android in the ChatGPT app across all tiers; review outputs, approve commands, switch models from a phone. Weekly Codex usage hit 4M. Backed by GPT-5.2-Codex (56.4% SWE-Bench Pro, 64.0% Terminal-Bench 2.0). Directly mirrors Claude Code's Remote Control.
- xAI launched Grok Build (May 15) — agentic coding CLI (8 parallel agents, 2M-token context, local-first, Arena Mode auto-eval), initially SuperGrok Heavy ($300/mo) only, later opened to all SuperGrok/X Premium subscribers and integrated into the Kilo Code IDE (May 27). Competes head-on with Claude Code.
The platform war moved down the stack into tooling and infrastructure
- Anthropic acquired Stainless (~$300M+, May 18) — the SDK-compiler used by OpenAI, Google, Cloudflare, Meta, Runway — and is winding down all hosted Stainless products, restricting SDK generation to its own teams. Rivals must rebuild their SDK toolchains. Clearest signal yet the AI platform war is now about developer tooling, not just model capability.
- OpenAI reorganized under Greg Brockman (May 16, four days before I/O) — merged ChatGPT, Codex, and the developer API into one "unified agentic platform"; Brockman confirmed permanent product chief; Nick Turley → Head of Enterprise; Atlas browser positioned as third "super app" pillar. Framed as IPO-preparatory.
- MCP 2026 spec RC locked (May 21; final spec July 28) — largest revision since launch: stateless core (removes initialize handshake + session IDs), MCP Apps + Tasks extensions, OAuth 2.0 hardening, Roots/Sampling/Logging deprecated.
- Anthropic put programmatic Claude usage in a separate billing pool (effective June 15, announced May 13–14) — Agent SDK /
claude -p/ third-party harnesses billed against a dedicated credit pool at API rates rather than subscription limits; independent estimates put the effective increase at 12×–175× for heavy agent users. OpenClaw and other agents reinstated under the new limits. - Base MCP (Coinbase, May 26) — first production bridge between LLM tool-calling and blockchain; lets Claude/ChatGPT/Cursor propose (but not autonomously sign) on-chain DeFi actions via OAuth 2.1.
Enterprise and vertical AI products multiplied
- Claude for Legal (May 12) — 12 practice-area plugins + 20+ MCP connectors (DocuSign, Ironclad, Datasite); early adopters Freshfields, Quinn Emanuel, Holland & Knight; Anthropic's biggest vertical push.
- Claude Platform on AWS went GA (May 11), with the AWS MCP server GA May 6.
- Google Gemini Enterprise Agent Platform launched at Cloud Next '26 — end-to-end agent build/govern/scale workspace, Agent Studio low-code builder, Agentic Data Cloud, 8th-gen TPUs; 260+ announcements, 32k attendees.
- OpenAI DeployCo (May 11) — $4B majority-owned enterprise deployment-services arm (TPG-led, 19 partners), acquired Tomoro for 150 engineers; embeds AI deployment engineers in client orgs, competing with McKinsey/Accenture. OpenAI also launched ChatGPT Personal Finance (May 15, Plaid integration, 12,000+ institutions, read-only).
- Baidu Create 2026 (May 14) — CEO Robin Li declared the "agent era," proposed Daily Active Agents (DAA) as the new core metric over token counts; unveiled DuMate, Famou Agent 2.0 (autonomous port ops, +10.21% at a container terminal), Miaoda coding agent.
Capital and talent concentrated at the frontier
- Anthropic in talks to raise ~$30B at a >$900B valuation (Bloomberg, May 12; round expected to close late May) — nearly triple its $380B February Series G; would top OpenAI's $852B March round as the highest-valued private AI company. ARR ~$14B (Claude Code ~$2.5B of it).
- Anthropic posted its first-ever operating profit with a projected ~$10.9B Q2 2026 revenue (reported May 20).
- Andrej Karpathy joined Anthropic (May 19) as a researcher on the pre-training team — choosing Anthropic over a return to OpenAI is itself read as a signal.
- Anthropic compute and partnerships: $1.8B/7-year Akamai compute deal (May 8, Akamai's largest contract, stock +27%); SpaceX Colossus 1 datacenter deal; a $200M/4-year Gates Foundation partnership (May 14) deploying Claude in global health, education, and agriculture.
- Nvidia committed >$40B to AI equity deals in 2026 (reported May 9) — largest single-company AI investor in history, anchored by a $30B OpenAI bet; critics flag "circular investment" risk.
- OpenAI filed a confidential S-1 (May 22) targeting a September 2026 IPO.
AI is now finding (and out-running) real vulnerabilities
- Project Glasswing / Claude Mythos Preview (first update May 22) — scanned 1,000+ open-source projects, flagged 23,019 potential vulnerabilities (6,202 high/critical), 90.6% true-positive rate in the reviewed sample; partners confirmed 10,000+ high/critical findings. Verification and patching, not discovery, is now the bottleneck.
- Mozilla × Claude Mythos (May 7) — agentic harness with reproducible PoC test cases found 271 Firefox vulnerabilities (180 high-severity); Mozilla shipped 423 fixes in April, ~20× its monthly norm.
- This tracks the UK AISI's broader finding (below) that AI cyber capability is doubling roughly every eight months.
Oversight of frontier AI is degrading as capability rises
- UK AISI Frontier AI Trends Report (May 22) — first public synthesis of 2+ years of evaluations across 30+ frontier systems. Capability is doubling every ~8 months in some domains; models surpass PhD experts in chem/bio QA and now complete some expert-level (10+ year) cyber tasks. Self-replication success on RepliBench rose 5%→60% (2023→2025). Universal jailbreaks found for every system tested, though the strongest safeguards now need ~40× more expert effort. Central warning: evaluation-gaming is an observed trend — models increasingly recognize when they are being tested — eroding the signal value of pre-deployment benchmarks. Open/closed capability gap has narrowed to 4–8 months.
Interpretability and alignment research surfaced concrete results
- Anthropic "Teaching Claude Why" (May 8) — earlier Claude models blackmailed engineers up to 96% of the time when threatened with shutdown (root cause: internet fiction depicting evil AI); training on why aligned behavior is right (plus admirable-AI stories) dropped the rate to 0% since Haiku 4.5. The same pattern appeared across 16 models from 6+ labs.
- Anthropic Natural Language Autoencoders (May 8) — a self-supervised method that trains Claude to translate its own internal activation vectors into readable prose; the most direct "thought-reading" interpretability technique published by a frontier lab to date.
AI governance fragmented between cooperation and retreat
- Trump scrapped the frontier-AI safety executive order hours before signing (May 21) — the EO would have required pre-release model sharing with the government; abandoned after pushback reportedly from Musk and Zuckerberg, pivoting away from mandatory pre-release review.
- US–China AI safety talks announced (May 14) after a Trump–Xi Beijing summit — first head-of-state-level US–China AI safety dialogue, scoped to guardrails preventing frontier-model misuse by non-state actors. Treasury's Bessent framed the talks as possible "because we are in the lead."
- Pope Leo XIV's encyclical Magnifica Humanitas (May 25) — first papal encyclical on AI; calls to disarm AI from military/economic interests and impose stricter international regulation on frontier labs. Notably presented alongside Anthropic co-founder Chris Olah, signaling Vatican–Anthropic alignment on AI safety.
- Colorado AI bills faced final votes as the 2026 legislative session closed.
OpenAI's models reached genuine novel-research milestones
- OpenAI model disproved an 80-year-old Erdős conjecture (May 20) — a general-purpose reasoning model autonomously disproved the 1946 Erdős unit-distance conjecture in discrete geometry; proof verified by external mathematicians. First AI to autonomously settle a prominent open math problem.