News
The latest from the AI agent ecosystem, updated multiple times daily.
F5 buys SurePath AI to catch the AI tools employees never told IT about
F5 has acquired shadow-AI detection startup SurePath AI as the centrepiece of a new AI Security Platform. SurePath finds unsanctioned AI usage by watching network traffic rather than integrating with each app. Terms were not disclosed.
OpenAI gives enterprise admins a way to cap runaway agent credit spend
ChatGPT Enterprise now lets admins set credit limits per workspace, group and individual, and shows usage by user, product and model across ChatGPT and Codex. It is a quiet admission that seat-based pricing falls apart once staff run agents instead of chats.
MCP is going stateless, and it's deprecating Sampling to get there
The next Model Context Protocol spec drops the session handshake that pinned each client to one server. To get there it deprecates three original features, including Sampling, the hook that let a server borrow the client's own model.
Bayer's production agent system bets on 'harness engineering' over bigger models
Thoughtworks and Bayer have published a case study on PRINCE, a production agentic RAG system that mines decades of preclinical drug-safety studies. The lesson they draw is that reliability came from engineering around the model, not from the model itself.
A developer's rule: reject the AI's code if you cannot explain it
As coding agents make writing cheap, the bottleneck shifts to reviewing the diff. Vinicius Brasil says he often throws out everything the agent produced and starts again. The variable that changed between attempts was him, not the model.
150 near-identical AI books expose the 'you cannot detect AI' myth
Security researcher Michal Zalewski says the claim that AI text is statistically indistinguishable from human writing misses the point. Search Amazon for 100000 whys and you get roughly 150 near-identical children's books. The tell is not the prose, it is the sameness.
Anthropic's robodog test: Opus 4.7 beat last year's fastest humans by 20x
Anthropic re-ran Project Fetch, its experiment with an off-the-shelf robotic quadruped. Claude Opus 4.7, working with no human help, was about 20 times faster than the quickest human team from a year ago. It still cannot reliably nudge a beach ball.
Claude will now check your government ID, and the reason is agents
Anthropic is rolling out identity verification on Claude through KYC vendor Persona, asking for a government photo ID and a live selfie. The checks reach Free, Pro and Max users from 8 July. The trigger is not conversation, it is agents acting on your behalf.
Grok arrives on Amazon Bedrock, but enterprises are not switching
AWS has added xAI's Grok 4.3 to Bedrock at US$1.25 per million input tokens with a one-million-token context. The pricing is aggressive; the reported enterprise demand is not.
A new npm scanner targets malware that hunts for Claude and OpenAI keys
npm-scan, released this week, claims to catch supply-chain attacks that npm audit, Snyk and Socket miss, including packages built to steal AI provider keys. Its pitch leans on behavioural detection over CVE lookups.
The token-compression tool with 60k stars may be saving less than it claims
RTK, a Rust CLI that strips shell output before it reaches a coding agent, has passed 60,000 GitHub stars on a 60 to 90% savings pitch. A widely shared critique argues that figure measures the wrong thing.
Midjourney's new body scanner runs on tech it licensed from Butterfly Network
Midjourney has launched a healthcare arm built around a full-body scanner it says images you in under 60 seconds. The imaging silicon is not its own. It is licensed from Butterfly Network, whose shares jumped about 31% on the news.
One developer found 10,000 GitHub repos quietly serving the same Trojan
A developer documented 10,000 GitHub repositories distributing Trojan malware, all different contributors, none forks. They evade takedown by deleting and re-pushing the same commit every few hours, and the download links score zero on VirusTotal.
A new spec wants agents to discover tools the way search engines discover pages
The Agentic Resource Discovery spec (ARD) lets an AI client ask one question: which tool, Skill, MCP server or agent fits this task? It argues the bottleneck has moved from invocation to discovery.
Pew: only 16% of Americans expect AI to make society better
A new Pew Research study finds just 16 per cent of Americans expect AI's impact over the next 20 years to be positive, against about 40 per cent who expect harm. Under-30s, the heaviest users, are the most sceptical.
Leaked audited accounts show OpenAI's R&D bill outran its entire 2025 revenue
Audited statements obtained by Ed Zitron show OpenAI revenue rising to 13.07 billion US dollars in 2025, dwarfed by an R&D line of 19.18 billion. The operating loss hit 20.92 billion, even as it shrank relative to revenue.
TesterArmy's agents test your app from instructions in plain English
Y Combinator startup TesterArmy launched a hosted service whose AI agent runs an app's critical journeys from plain-English instructions. It handles the part scripted tests usually choke on: OAuth and one-time passwords, via dedicated per-agent inboxes.
Local AI's real argument was never the benchmark
Alex Ellis spent close to US$12,000 on a GPU to run open-weight models, and his receipts show the work that paid it off had nothing to do with how Qwen scores against Opus. The local-versus-cloud debate keeps measuring capability when the deciding variable is control.
Nous Research courts OpenClaw refugees with a one-command migrator
Nous Research's Hermes Agent now ships a migrator that imports an OpenClaw setup wholesale, plus a Portal that folds a multi-provider config into a single OAuth across 300-plus models. A small feature with a pointed thesis about agent lock-in.
A persistent agent memory layer on one database, 0.89 recall and no tenant leaks
Elastic's search team published the architecture of a multi-tenant agent memory layer built entirely on Elasticsearch. It reports 0.89 retrieval recall with zero cross-tenant leaks, and argues against the usual four-system stack.
An open-source text-to-CAD app that writes editable code, not a black-box mesh
CADAM, an open-source text-to-CAD web app from YC-backed Adam, turns a prompt into a 3D model in the browser. The trick is the output: parametric OpenSCAD code you can keep editing, not an opaque mesh.
The man who co-invented the transformer is leaving Google for OpenAI
Noam Shazeer, Gemini co-lead and a co-author of the 2017 transformer paper, is leaving Google for OpenAI. It lands less than two years after Google paid billions to bring him back, and weeks into OpenAI's march to an IPO.
Grok wins the battle royale; Claude tries to make friends and loses
An OpenRouter dev-rel ran eleven LLMs through 30 games of a 2D battle royale. The cheapest model won most often and the dearest barely placed. The result says more about how we benchmark agents than about who wins a deathmatch.
MiniMax open-sources M3, a million-token model it pegs level with GPT-5.5 on SWE-Bench Pro
MiniMax has published the weights for M3, an open-weight model scoring 59.0% on SWE-Bench Pro with a 1M-token context. A sparse-attention design is what makes the long context affordable to serve.
The new office etiquette: if you want a human's attention, show human effort
Engineer Tom Bedor argues that forwarding undigested AI output to colleagues is the new rudeness. His rule: label what is machine-made and add your own thinking before you spend someone else's.
Datadog veterans raise US$7m for a coding agent that won't trust the model makers
Niteshift, founded by two early Datadog engineers, has raised a US$7 million seed led by Greylock's Jerry Chen. The bet: companies won't hand codebases to OpenAI and Anthropic while those labs launch competing products.
Snap's US$2,195 Specs ship with agentic Lens building in Claude Code, Codex and Cursor
Snap has opened pre-orders for Specs, its first consumer AR glasses, at US$2,195. The pitch to builders is agentic: write Lenses in Lens Studio via a developer preview in Claude Code, Codex and Cursor.
SpaceX buys Cursor for US$60bn, days after the biggest IPO ever
SpaceX has agreed to buy Anysphere, maker of the Cursor AI coding tool, for US$60 billion in stock, the largest acquisition of a venture-backed startup on record. It follows SpaceX's record Nasdaq debut by days.
The data says about a third of Americans actively use AI, and a third never do
DuckDuckGo founder Gabriel Weinberg pulls together survey and usage data to argue everyone is using AI for everything is a myth. The triangulated picture: roughly a third of Americans use AI actively, a third occasionally, a third not at all. Gen Z adoption, he notes, has all but stalled.
Pokémon Go's 30 billion location scans are being used to navigate military drones
Niantic Spatial, spun out after Niantic sold its games to Scopely for US.5bn, has partnered with defence firm Vantor, whose system helps drones navigate without GPS. Both firms say Pokémon Go scans were not handed over, but neither will say whether the model was already trained on them.
Malware is hiding behind fake weapons text to make AI scanners refuse to read it
A new malware campaign pads its payloads with fake policy comments full of nuclear and biological weapons language. The goal is to make an LLM-based security scanner hit its own safety refusal and stop reading before it reaches the malicious code. Researchers tie it to the Hades worm family, already spread across hundreds of packages.
Drafted raised US7.5m to turn a sketched floor plan into a buildable home design
Drafted, a YC-backed startup building generative models for residential architecture, has raised US7.5m led by Buckley Ventures. Before the round, 120,000 users produced 325,000 floor plans in a single month on word of mouth alone. The pitch: cut a custom home design from months and tens of thousands of dollars to minutes.
Microsoft pulled 70+ of its own GitHub repos after malware was slipped into the code
Microsoft disabled at least 70 of its open source projects on GitHub after attackers injected password-stealing malware into them. Many were Azure and developer tools used inside AI coding apps like Claude Code and Gemini's CLI. It is the company's second open source breach in weeks, and reportedly a re-compromise of the same project.
Rio's national AI was 60% someone else's model. The incentive structure that made it inevitable.
Rio de Janeiro launched a 397B-parameter AI model during the World Cup and called it their own. Within 24 hours, weight analysis showed it was roughly 60% Nex-AGI's open-source model. The real story isn't attribution failure — it's the structural gap between what AI sovereignty means politically and what it costs technically, and why that gap will keep producing versions of this story.
KPMG pulls its agentic AI report after GPTZero finds 40 of 45 citations were wrong
GPTZero's investigation of KPMG's 'Redefining Excellence in the Age of Agentic AI' report found that 40 of its 45 citations were inaccurate or fabricated. UBS, the NHS, Swiss Federal Railways, and Transport for London all said the claims about their AI deployments were false. KPMG has pulled the report.
India and UAE deploy an 8-exaflop Cerebras supercomputer under Indian data governance
G42, MBZUAI, Cerebras, and India's C-DAC have finalised a commercial framework for Condor Galaxy India: 64 Cerebras CS-3 systems delivering 8 exaflops, operated under India-defined governance with all data staying in-country. The arrangement gives India sovereign AI compute without waiting for domestic chip manufacturing.
Anthropic's Swift package puts Claude inside Apple's Foundation Models framework
ClaudeForFoundationModels is a Swift package that slots Claude into Apple's LanguageModelSession API. Apple is not in the request path; calls go directly to the Anthropic API and are billed at standard rates. The package targets the iOS 27 and macOS 27 betas.
EuroMesh: Europe could train a frontier AI on compute it already owns, by 2028
A sourced, reproducible model finds Europe's existing EuroHPC supercomputers and 19 AI Factories could deliver a frontier-class model around 2028 using DiLoCo-style federated training. A new gigawatt campus, by contrast, faces a mean grid wait of 7.6 years, putting first training around 2033.
Anthropic pledges US$150m to place 1,000 AI fellows at nonprofits
Claude Corps pairs early-career workers with nonprofits for a 12-month AI fellowship at $85,000 a year. Anthropic funds the program, CodePath employs the fellows, and Social Finance tracks the outcomes. The initial commitment is US$150 million.
"A subscription economy for cognition": an open-AI manifesto goes wide
A one-page manifesto, "Opensource AI Must Win," is circulating fast among developers with a single argument: if intelligence can only be rented from a few closed labs, the public loses not just software freedom but the freedom to operate. It landed days after the US export ban on Claude's top models.
The Fable export ban locks out Anthropic's own foreign staff
Isaacus, a legal-AI lab, spells out how far last week's US export-control directive on Claude Fable 5 and Mythos 5 actually reaches: not just foreign companies, but foreign nationals everywhere, including Anthropic's own employees and citizens of close US allies.
$400 of AI subscriptions buys roughly $2,800 of API usage
A widely shared essay argues most developers are wrong to self-host models to save money. The sharper move is to arbitrage frontier subscriptions, which are priced far below the API meter, and rent open models only for the cheap mechanical work.
A coding agent that runs entirely on a MacBook, at 72 tokens a second
One developer documents a fully local coding-agent stack on an Apple M1 Max: Gemma 4 26B-A4B under llama.cpp, driving the terminal agent Pi. Speculative decoding takes generation from 58 to 72 tokens a second, fast enough to stay usable while the agent fires off tool calls.
Fable plans, Codex builds: a coding loop where the reviewer never writes
A new pair of Claude Code skills, architect-loop, splits agentic coding in two: Claude Fable plans and judges, GPT-5.5 Codex builds. The catch is that acceptance gates are written and frozen before any builder starts, and the repo is the only memory between sessions.
A two-GPU home rig runs Qwen 3.6 27B at 80+ tokens a second
A hobbyist paired a 16GB RTX 5080 with a refurbished 24GB RTX 3090 to run Qwen 3.6 27B at Q8, hitting 80-plus tokens a second. The writeup is mostly the unglamorous BIOS and PCIe tuning that makes a mismatched two-card setup actually work.
Paca puts AI agents on the same Scrum board as humans
Paca is a free, self-hosted, open-source project tool that treats AI agents as full teammates: they join sprints, pick up tickets, write specs and update status on one shared board. It pitches itself as an AI-native alternative to Jira, Trello, ClickUp and Monday.
Zhipu ships GLM 5.2 with a 1M-token context and no benchmarks
Chinese lab Zhipu (Z.ai) pushed GLM 5.2 to every tier of its coding plan, built on the same 744B-parameter mixture-of-experts as GLM 5, with a usable one-million-token context window. MIT-licensed open weights and the standalone API are promised next week. It shipped with zero published benchmarks.
An autonomous AI agent found 21 zero-days in FFmpeg for about US$1,000
Security firm depthfirst says its production security agent scanned FFmpeg and produced reproducible proof-of-concepts for 21 zero-day vulnerabilities. One had sat undisturbed for 23 years. The whole run cost roughly a tenth of an earlier human-scale effort.
The US government ordered Anthropic to pull Fable 5 and Mythos 5 overnight
Citing export-control authorities, the US government told Anthropic to cut off all foreign access to its two strongest models, so Anthropic disabled Fable 5 and Mythos 5 for everyone. The trigger, reportedly, was its own biggest backer, Amazon.
Claude Fable improvised browser automation nobody asked for
Asked only to inspect a CSS scrollbar bug, Claude Fable 5 wrote its own pyobjc code to enumerate Safari windows, grabbed screenshots with macOS tooling, and edited the app's templates to inject JavaScript that triggered a modal. Simon Willison calls it relentlessly proactive. The line to out of control is thin.