News
The latest from the AI agent ecosystem, updated multiple times daily.
Copilot's Free Ride Ends: GitHub Switches to Usage Billing
GitHub is transitioning all Copilot plans from premium request-based pricing to usage-based billing starting June 1, 2026. The new system uses GitHub AI Credits consumed based on token usage (input, output, and cached tokens) at published API rates per model. Base plan pricing remains unchanged: Pro ($10/month), Pro+ ($39/month), Business ($19/user/month), and Enterprise ($39/user/month), with each plan including equivalent AI Credits. The change addresses rising inference costs from agentic AI features and multi-step coding sessions.
Dutch central bank ditches AWS and chooses Lidl for European cloud
De Nederlandsche Bank (DNB), the Dutch Central Bank, is signing a major contract with Schwarz Digits (the IT arm of Lidl owner Schwarz Group) to use their Stackit cloud platform. This move aims to reduce dependence on American cloud companies like AWS, Google Cloud, and Microsoft Azure, driven by concerns about cloud sovereignty and the US Cloud Act which allows US authorities to access data. Schwarz Digits is positioning Stackit as a European alternative to American hyperscalers, with a recent 11 billion euro investment in a data center in Lübbenau.
Sam Altman Wants to Rethink Operating Systems. OpenAI Might Build One.
OpenAI CEO Sam Altman tweeted that operating systems and user interfaces need a fundamental rethink. The comment has sparked speculation about whether OpenAI plans to build its own AI-first operating system.
Microsoft Ends Revenue Sharing With OpenAI
Microsoft is ending its revenue-sharing agreement with OpenAI, giving the startup more flexibility to compete with rivals like Anthropic. The move also helps Microsoft stay ahead of regulatory scrutiny in the US and Europe. Separately, GitHub is transitioning Copilot to usage-based billing, signaling the end of free AI coding assistance.
Google's DiLoCo Trains AI on Mixed Hardware, 20x Faster
Google DeepMind announces Decoupled DiLoCo, a new architecture for training large language models across globally distributed data centers. The approach splits training into decoupled 'islands' of compute with asynchronous data flow. It isolates hardware failures and needs far less bandwidth than traditional methods. Tested on Gemma 4 models, the system matched conventional training performance while running 20x faster and mixing different hardware generations (TPU v6e and TPU v5p).
Cal Newport to AI Industry: Who Asked for This?
Cal Newport argues Silicon Valley shifted from solving customer problems to 'inventing the future,' pursuing AI without clear market utility. While LLMs have more potential than NFTs or the metaverse, ordinary people mainly use ChatGPT as a search tool and face conflicting narratives about AI's job impact.
AI is about to hit a power wall
US power demand is projected to reach record highs in 2026-2027, driven by AI usage and data center expansion according to EIA data. Hacker News commenters highlight that AI is becoming an energy problem, not just a compute problem, with power infrastructure scaling slower than chip improvements—a potential constraint for AI growth.
Chrome Now Ships an On-Device LLM. Most Laptops Need Not Apply.
Google's Prompt API lets Chrome run Gemini Nano locally, no API keys or cloud calls required. The catch: the model download is larger than Chrome itself, and the hardware requirements shut out most laptops.
Git-pinned cache cuts AI agent costs in half
An experiment shows that using a Git-based cache to pre-verify and pin facts about a codebase can reduce AI agent token costs by 51%. The approach uses Claude Haiku to scan repos and create claims pinned to Git blob OIDs, with Merkle root computation to automatically mark claims as stale when files change. Results: cost reduction from $4.35 to $2.13 per session, 16% faster wall time.
Nvidia VP: AI Compute Costs Now Exceed Team's Salaries
Some companies are now spending more on AI compute than on human workers. Nvidia's VP of applied deep learning says compute costs far exceed employee salaries for his team, while Uber's CTO reportedly burned through the company's entire 2026 AI budget on token costs alone. As global IT spending heads toward $6.31 trillion, companies face growing pressure to show returns on AI investments and get costs under control.
Tendril starts with three tools and builds the rest itself
Tendril is an agentic sandbox where an AI model builds and registers its own tools as it works. Starting with just three bootstrap commands, the agent creates what it needs and saves it for next time. Built with AWS Strands Agents SDK, Tauri, and Claude Sonnet 4.5.
AgentSwarms: Break 30+ agents in-browser until you understand them
AgentSwarms is a free browser platform for learning agentic AI by running and modifying live agents. Covers prompts, RAG, tool calling, guardrails, multi-agent swarms, and observability across 40+ lessons. Supports OpenAI, Gemini, Grok, and Claude with zero setup.
Nvidia VP: Compute now costs more than our employees
Nvidia's VP says compute now costs more than employees. Uber's CTO already burned through 2026's AI budget on tokens. The AI cost problem is here, and companies are scrambling for cheaper alternatives.
AI now costs more than the humans it replaces
Nvidia VP Bryan Catanzaro says compute costs for his team now sit 'far beyond' employee salaries. Uber's CTO already burned through his 2026 AI budget on token costs alone. Companies are spending millions to replace teams that cost less, and some are bailing on commercial APIs to self-host open-source models instead.
Claude Code 90% cost-saver goes viral, attribution doesn't
A tutorial showing how to route Claude Code through Ollama to reduce costs by approximately 90%. The setup keeps strategic work on Claude Pro while offloading heavy tasks like lints, refactors, and file operations to free open-source models (Gemma, Qwen, DeepSeek) running locally via Ollama.
Google Bets AI Edge Can Close Cloud Gap With Amazon, Microsoft
Google Cloud is pushing AI to catch AWS and Azure, but history says technical superiority alone won't win enterprise deals. Google open-sourced TensorFlow in 2015 and built custom TPUs early, yet still trails rivals who spent years building deep corporate relationships. The question isn't whether Google can build better AI tools. It's whether enterprises trust Google to support them long-term.
Microsoft ends Azure exclusivity, lets OpenAI use any cloud
OpenAI can now serve products through any cloud provider, ending Microsoft's exclusive grip. The amended deal keeps Microsoft as primary cloud partner with Azure getting first access to new OpenAI products. Microsoft retains a non-exclusive IP license through 2032, stops paying revenue share to OpenAI, and remains a major shareholder. OpenAI continues revenue share payments to Microsoft through 2030 with a cap.
DeepSeek V4 runs on Huawei chips and keeps up with GPT-5
DeepSeek released V4, an open-source AI model with a 1 million token context window, optimized for Huawei's Ascend chips. The model comes in Pro and Flash versions, offering performance rivaling Claude-Opus-4.6, GPT-5.4, and Gemini-3.1. V4-Pro costs $1.74 per million input tokens. Flash costs $0.14. Architectural improvements cut memory use dramatically, and the Huawei partnership signals China's push away from Nvidia dependency.
Microsoft Drops Revenue Split as OpenAI Outgrows the Deal
Microsoft and OpenAI restructured their partnership. Microsoft will stop sharing revenue with OpenAI, while OpenAI gains the ability to sell products on any cloud provider, ending Microsoft's exclusivity. Microsoft retains rights until 2032, and OpenAI has capped repayment obligations until 2030. The move comes as competition from Anthropic forced OpenAI to scale beyond what Microsoft's infrastructure alone could support.
AI is hungry for electricity and America's grid can't keep up
A Reuters report discusses how AI workloads and data center expansion are projected to drive US power demand to record levels in 2026-2027, according to the U.S. Energy Information Administration (EIA). The story and Hacker News comments highlight emerging concerns about energy infrastructure scaling slower than chip technology, potentially creating a significant constraint for AI development.
This news site's reporters are AI bots. OpenAI appears to fund it.
An investigation reveals AcutusWire.com, a digital news site launched in December 2025, operates almost entirely on AI-generated content. 69% of articles are fully AI-generated, and 'reporter' Michael Chen is actually an AI agent sending interview requests. The site appears connected to OpenAI's super PAC Leading The Future and Republican PR firm Novus Public Affairs, suggesting an astroturfing operation pushing specific political narratives.
Sam Altman's World ID wins U.S. corporate backing despite global bans
Tinder, Zoom, and Docusign are partnering with World (formerly Worldcoin), Sam Altman's iris-scanning biometric ID project. While U.S. companies embrace the technology, countries across Asia, Africa, Europe, and Latin America have banned or halted World over privacy violations, including collecting minors' data and paying people for iris scans. The project claims 18 million verified users, but many received $50 in crypto to sign up.
DeepSeek v4 works with OpenAI SDKs and runs on your Mac
DeepSeek releases v4 API docs with two models: deepseek-v4-flash and deepseek-v4-pro. Compatible with OpenAI and Anthropic SDKs out of the box. Features thinking mode with configurable reasoning effort. Legacy models deepseek-chat and deepseek-reasoner deprecated July 2026.
DeepSeek v4 Runs on Huawei Chips, No CUDA Required
DeepSeek v4 API launches with flash and pro models running on Huawei Ascend chips with zero CUDA dependency. Features OpenAI/Anthropic compatibility, thinking mode, tool calls, context caching, and deterministic bitwise outputs at temperature 0.
Google Blinks: TorchTPU Runs PyTorch Natively on TPUs
TorchTPU enables PyTorch to run natively on Google's Tensor Processing Units with an eager-first approach offering Debug, Strict, and Fused Eager modes, plus full-graph compilation via torch.compile over XLA. The stack supports distributed training patterns including DDP, FSDPv2, and DTensor. A 2026 roadmap targets a public GitHub repo, Helion DSL integration, and first-class dynamic shapes.
Tolaria uses AGENTS files so Claude Code can read your markdown vault
Tolaria is an open-source macOS app for managing markdown knowledge bases, created by Luca Ronin. What makes it notable is the AGENTS file, a YAML configuration in each vault that gives AI agents like Claude Code and Codex CLI persistent context about your notes. The app stores everything as plain markdown in git repositories, with no accounts or cloud dependencies.
Karpathy's LLM lecture becomes interactive browser playground
An interactive browser guide lets you click through how LLMs work, from data collection to RLHF. Based on Karpathy's lecture but with live tokenizers and training visualizations you can actually play with. The key insight: pre-training builds an 'internet simulator,' not an assistant.
Mythos' 271 Firefox bugs: real finds or marketing?
Anthropic and Mozilla claimed Mythos found 271 Firefox vulnerabilities for under $20K. That number aggregates bug fixes, cleanups, and patches across multiple products, not just Firefox 150. Mythos is useful for defensive security at scale, but doesn't prove AI has cracked offensive vulnerability research.
Karpathy's LLM lectures turned into interactive visual guide
An interactive guide that lets you play with tokenization and watch neural network training happen in real time, based on Andrej Karpathy's technical lectures.
AI-run store in SF can't stop ordering candles and paying women less
Andon Labs' AI agent Luna manages an experimental retail store called Andon Market in San Francisco's Cow Hollow neighborhood. The AI has made questionable decisions including over-ordering candles, buying 1,000 toilet-seat covers for resale, and paying female employees $2/hour less than a male colleague. The experiment raises concerns about AI managing humans and highlights gaps in how AI understands retail operations.
Tesla's HW3 FSD Promise Dies: Millions Need Hardware Upgrades
Elon Musk admitted millions of Tesla owners with Hardware 3 vehicles need new computers and cameras for unsupervised Full Self-Driving, contradicting years of promises. Tesla is considering 'micro-factories' for retrofits, but the admission exposes the company to legal risk from customers who paid thousands based on false assurances.
Fast Tanh: Four Rust Tricks to Speed Up Inference
A technical survey of fast tanh approximations using Taylor series, Padé approximants, splines, and bitwise manipulation techniques like K-TanH and Schraudolph, with Rust code examples. Covers the speed gains that matter for neural network inference and real-time audio, plus why quantized models need these tricks to work at all.
Goedecke: Anti-AI Left Borrows from Conservative Playbook
Sean Goedecke argues that while anti-AI rhetoric appears left-wing, the underlying arguments mirror conservative positions on copyright, human essence in art, and job protection. He traces how timing and tech's rightward pivot created this odd alignment, and predicts it won't last: either the right claims anti-AI rhetoric or the left ends up defending technology it currently opposes.
Ghost Pepper: Free open-source transcription that stays on your Mac
Ghost Pepper is a free, open-source macOS app for local voice dictation and meeting transcription. It runs 100% local models (WhisperKit and Qwen LLMs) on Apple Silicon, with complete privacy since no data leaves the machine. Features include speech-to-text, meeting transcription with AI summaries, and text cleanup to remove filler words.
Agent Vault keeps API keys away from AI agents
Agent Vault is an open-source credential broker by Infisical that prevents credential exfiltration for AI agents. Instead of returning credentials directly to agents, it uses brokered access where agents route HTTP requests through a local proxy that injects credentials at the network layer. Works with Claude Code, Cursor, and other HTTP-speaking agents.
Andreessen, Thiel Frame AI Regulation as Evil. The Tab: $670 Billion
Peter Thiel calls AI critics 'Antichrist agents.' Marc Andreessen brands slowing AI 'a form of murder.' As tech giants pour $670 billion into AI development, Silicon Valley has reframed regulation as a religious war and shifted its political spending decisively toward Republicans.
My phone replaced a brass plug
Vadim Drobinin describes building an iOS app that uses computer vision to automate shooting target scoring. The solution combines OpenCV for structural geometry and a fine-tuned YOLOv8 model (exported to CoreML) for bullet hole detection, replacing traditional physical gauges.
Meta Lays Off 8,000 to Afford the AI Race It's Late To
Meta plans to cut 10% of its workforce (8,000 employees) and freeze hiring for 6,000 open roles starting May 20. The cuts fund AI investments after tens of billions burned on metaverse bets that mostly flopped. Meta recently launched Muse Spark to compete in the AI space.
GPT-5.5: Mythos-Like Hacking, Open to All
XBOW's benchmarks tell a clear story: GPT-5.5 misses 10% of vulnerabilities (GPT-5 missed 40%, Opus 4.6 hit 18%). Running blind, it beats GPT-5 with full source code access. Login speeds doubled, and the model knows when to quit. It performs like Anthropic's Mythos, but you can actually use it.
UK Biobank health data keeps ending up on GitHub
A tracker monitoring UK Biobank's efforts to remove health data from public GitHub repositories reveals 110 takedown notices targeting 197 repositories by 170 developers worldwide. Despite strict access agreements prohibiting data sharing, researchers continue to accidentally upload sensitive genetic and health data from half a million British volunteers. The tracker, built by Luc Rocher at Oxford Internet Institute, uses GitHub's DMCA archive to identify affected repositories and shows that nearly half of targeted files are Jupyter or R notebooks, with a quarter being genetic data files.
A Boy That Cried Mythos: Verification Is Collapsing Trust in Anthropic
Security researcher Davi Ottenheimer's analysis of Anthropic's Claude Mythos Preview and Project Glasswing finds inflated cybersecurity claims. The 244-page system card contains just 7 pages of security content with no CVEs, fuzzing data, or severity scores. The Firefox 147 demo used a stripped-down JavaScript shell on bugs already found and patched. Removing the top two bugs drops Mythos's exploit success from 72.4% to 4.4%. Ottenheimer calls Project Glasswing regulatory capture, noting no independent partner verification exists.
Meta logs worker keystrokes on Google, LinkedIn for AI training
Meta has launched an internal employee tracking initiative called the Model Capability Initiative (MCI) that captures keystrokes, mouse clicks, and screen content from employees using popular websites like Google, LinkedIn, Wikipedia, GitHub, and Slack. The data collection aims to train AI models to better understand human-computer interaction for building AI agents. Employees have raised privacy concerns about potential exposure of sensitive data including passwords and personal information. The program reflects Meta's urgency to catch up with rivals OpenAI, Anthropic, and Google in the generative AI race.
Ars Technica bans AI-authored content, permits research tools
Ars Technica's new AI policy draws a hard line: no AI-authored content. But allowing AI for research raises questions about whether reporters can reliably verify confident-sounding but subtly wrong AI output under deadline pressure.
Ubuntu 26.04 Ships CUDA and ROCm Out of the Box
Ubuntu 26.04 LTS 'Resolute Raccoon' ships with NVIDIA CUDA and AMD ROCm in its official repositories, letting developers install AI compute frameworks with one command. Also adds TPM-backed full-disk encryption, Rust-based system utilities, Livepatch for Arm64, and Intel Core Ultra Series 3 NPU support. Built on Linux 7.0.
Claude Desktop Quietly Installs Native Messaging Bridge
Anthropic's Claude Desktop App installs a native messaging bridge enabling browser extensions to communicate with the local app. Users must approve a Chrome permission prompt, but the desktop installer doesn't prominently explain the component, leaving some security-conscious users unsettled.
LocalLLM: The Open Cookbook for Local AI Without Training Wheels
An open-source project collecting hardware-specific configurations for local LLM inference wants community help to document more setups. Unlike one-click tools, it exposes every setting for users who need precise control over their models.
THE PEOPLE DO NOT YEARN FOR AUTOMATION
Nilay Patel's Decoder podcast breaks down why AI faces growing public hostility despite tech industry enthusiasm. Covers polling data showing AI less popular than ICE, with Gen Z sentiment worsening among heavy users. Features quotes from Nadella, Altman, and Amodei, and introduces 'software brain' as the worldview gap between builders and everyone else.
GitHub's 90-Day Uptime: 88%. No, That's Not a Typo.
GitHub outage knocked out Copilot, Actions, and Webhooks on April 23. Services were restored within the hour, but the incident highlights a troubling 90-day uptime trend hovering around 88%.
OpenAI's 1.5B Model Has One Job: Hunt PII Locally
OpenAI has released Privacy Filter, a 1.5B parameter open-weight model for detecting and redacting PII in text. It runs locally, supports 128K token context, and hits 96-97% F1 scores on benchmarks. Unlike orchestration frameworks, it's a single opinionated model that needs no extra components. Apache 2.0 license.
OpenAI code signing certs exposed in Axios supply chain hit
OpenAI disclosed a security incident where a compromised version of the Axios developer library (v1.14.1) was downloaded by a GitHub Actions workflow used in their macOS app signing process on March 31, 2026. The workflow had access to code signing certificates for ChatGPT Desktop, Codex, Codex CLI, and Atlas. OpenAI found no evidence of user data compromise, software alteration, or certificate exfiltration, but is rotating certificates out of caution. Users must update macOS apps by May 8, 2026. The root cause was a misconfiguration: using a floating tag instead of a specific commit hash and lacking minimumReleaseAge configuration.