Archive
Issue #07

Eval sandboxes, MCP credentials, and agent cloud computers

UK AISI hardens and restarts most dangerous-capability evals, official MCP client SDKs disclose OAuth issuer-trust flaws, and dedicated agent cloud computers raise the question of who decides what leaves the box.

  • 6 min read
  • 11 items
  • 5 sections

Saturday, 3 October 2026. UK AISI hardens and restarts most dangerous-capability evals, and MCP client SDKs disclose OAuth issuer-trust flaws that hit agent backends. Dedicated agent cloud computers raise the same control-plane question: who decides what leaves the box.

Today

  • UK AISI resumes most frontier evals after network lockdown and live monitoring
  • Official MCP Python and TypeScript SDKs: malicious servers could steal OAuth credentials
  • Personal agents get their own cloud computers; isolation is not the same as egress control
  • Anthropic puts $100M behind 10,000 Frontier Deployed Engineers
  • Gemini 4 Argon opens to trusted cyber defenders via Fairwind
  • Also: 4 other items, 1 tools note, 0 papers

AI

UK AISI restarts most dangerous-capability evals after hardening the harness

The UK AI Security Institute said on 1 October that it has completed the first phase of security work pledged after its August cyber-evaluation incident and can now resume most evaluation activity. Highest-risk cyber offense-and-defense tests remain paused pending a new sandbox service.

AISI disabled internet access for agentic cyber evaluations, with outbound networking blocked both inside the sandbox and independently at the cloud network layer. It added a synchronous LLM monitor that reviews messages, tool calls, and chain-of-thought where available, and can block suspicious actions before they execute. Evaluations were redesigned to run with local resources, clearer task boundaries in prompts, and automated pre-run checks that monitoring is on and the network is off. NCSC-informed internal governance and phased AI-assisted escape testing round out the package. Secondary coverage notes the methodological trade-off: cutting the open internet lowers realism for offense/defense scoring (Cocoloop).

Why it matters: Eval isolation is now a first-class control-plane problem for anyone running agentic red teams or pre-release cyber scores. Dual-layer network deny, pre-action monitoring, and “controls verified before run” are the pattern to copy, and AISI is explicit that today’s controls will not hold for the next capability step.

MCP client SDKs trusted the server’s word on where credentials go

GitHub advisories document the same failure class in the official MCP client stacks. In the Python SDK (GHSA-qx49-fqc8-xw99, High, CVSS 7.5, published 28 September), affected mcp clients over HTTP could send client_secret, authorization codes, and PKCE verifiers, or signed assertions, to a token endpoint chosen by a malicious or compromised MCP server, because issuer validation and credential-to-issuer binding were incomplete. Fixed in mcp 1.30.0 / 2.2.0; machine-to-machine providers still need an explicit issuer= or the upgrade alone changes nothing. The TypeScript SDK (GHSA-6qxp-vccf-f47h, High, CVSS 7.5, published 30 September) has the same root cause for stored refresh tokens and bundled providers; the fix is in @modelcontextprotocol/sdk 1.31.0 and @modelcontextprotocol/client 2.2.0, with expectedIssuer required for bundled M2M providers.

WorkOS (2 October) frames a 30-day cluster: Python SDK, Rust rmcp (CVE-2026-63127), and LiteLLM’s MCP endpoint (CVE-2026-59822, on CISA KEV). ThreatFrontier (updated 3 October) stresses the operational miss, upgrading without setting issuer, and notes a separate Python session-memory DoS advisory (GHSA-84m7-p3x7-pcfv).

Why it matters: Agent backends that connect OAuth-enabled MCP clients to third-party servers inherit a classic mix-up attack. Patch inventory must cover Python, TypeScript, and Rust clients; CI should fail closed when issuer / expectedIssuer is missing; rotate secrets if a vulnerable client ever talked to an untrusted server.

Dedicated agent computers are shipping, and the hard boundary is still egress

Frontier vendors are converging on persistent cloud machines for agents rather than hijacking the user’s laptop. Meta’s Muse runs in a Muse Secure VM with a separate Sentinel that alone authorises connector actions and network egress, injects real credentials at the boundary, and keeps the agent off host root. OpenAI’s Dots, announced at DevDay on 29 September, give each agent its own cloud computer and browser with connected apps; proactive research is read-only until the user grants write/actions (Indian Express, IBTimes SG). xAI’s Grok Bot documentation, as quoted in that roundup, warns that bots on one account share a single persistent computer, so separate bots are not a security boundary.

Simon Willison (1 October) highlighted Matthew Green’s framing: sandbox isolation alone is not enough once agents can leave payloads for each other via shared channels. Muse’s Sentinel and AISI’s egress lockdown address the same lesson.

Why it matters: For enterprise agent platforms, VM isolation without an independent egress and permission authority is theatre. Design reviews should ask which process can mint network egress, where real tokens live, and whether “one agent per identity” is actually enforced in storage and sessions.

Anthropic’s Claude Frontier Academy: $100M, 10,000 FDEs by end-2027

Anthropic launched Claude Frontier Academy on 2 October with a $100 million commitment to train 10,000 Frontier Deployed Engineers by the end of 2027. First cohorts include engineers from Accenture, Bain, Capgemini, Commonwealth Bank of Australia, Deloitte, McKinsey, Morgan Stanley, and Novo Nordisk. The residency starts with an in-person simulated enterprise deployment through security review, then a 12-week on-the-job Claude project; first FDE badges are expected in early 2027. CNBC corroborates the figures and the talent-gap framing.

Why it matters: The scarce resource is no longer model access. It is people who can take an agentic use case from requirements through security review into production. Partner nomination and FDE-shaped residencies will show up in RFP conversations and internal talent plans.

Tech

Gemini 4 Argon: cyber-defense first, broad rollout later

Google DeepMind announced Gemini 4 Argon on 30 September, initially rolling out to trusted cyber defenders through the Fairwind Program, with engagement in the U.S. voluntary pre-release access process. Argon expands output to 1M tokens, claims frontier results on long-horizon software engineering and knowledge-work benchmarks, and is positioned for defensive cybersecurity, including versions without cyber guardrails for trusted defenders. Safeguards before broad availability include misuse refusals, prompt-injection resilience, chain-of-thought/action misalignment monitoring, and hardened sandboxes for high-risk training and evals. Introductory API pricing is listed at $2 / $10 per million input/output tokens (rising to $4 / $20 after the intro period).

Why it matters: Another frontier model is gated first by defensive cyber use and government pre-release process. That’s a useful signal for how vendors will sequence high dual-use capability into enterprise control planes.

Also noted

  • Epoch AI (2 Oct): HBM shipped through 2027 could support roughly 33-171 million concurrent frontier-model agents (or much more with efficient open models); even 20% use of the central case implies multi-trillion API-equivalent spend, a demand/supply stress test for agent economics.
  • IAPP (2 Oct): FTC Chair Ferguson framed corporate liability for agent actions under existing consumer-protection law; reporting cited FTC scrutiny of OpenAI, Anthropic, and METR on agentic safety controls. CFAA-style criminal intent remains a poorer fit for rogue-agent cases.
  • EU/DK milestone check (past ~24-48h): no new Digst, Datatilsynet, Commission AI Act / Service Desk, or Official Journal hits. Background only: Datatilsynet’s AI-hjemmelslov hearing (opened 16 Sep) remains open until 12 October 2026.
  • Interconnects: no new October post found in this scan window.

Tools

  • Cloudflare Clef and Clef-flash (1 Oct): open-weight (Apache 2.0) decision models that return typed probabilities for schema-bound questions rather than free text; hosted on Workers AI, Jev-API compatible, with an FDE-led RL fine-tuning path. Useful as a stack-moving alternative to parsing LLM JSON for agent routing and triage.