News
The latest from the AI agent ecosystem, updated multiple times daily.
DeepSeek V4 runs on Huawei chips and keeps up with GPT-5
DeepSeek released V4, an open-source AI model with a 1 million token context window, optimized for Huawei's Ascend chips. The model comes in Pro and Flash versions, offering performance rivaling Claude-Opus-4.6, GPT-5.4, and Gemini-3.1. V4-Pro costs $1.74 per million input tokens. Flash costs $0.14. Architectural improvements cut memory use dramatically, and the Huawei partnership signals China's push away from Nvidia dependency.
Nvidia VP: AI Compute Costs Now Exceed Team's Salaries
Some companies are now spending more on AI compute than on human workers. Nvidia's VP of applied deep learning says compute costs far exceed employee salaries for his team, while Uber's CTO reportedly burned through the company's entire 2026 AI budget on token costs alone. As global IT spending heads toward $6.31 trillion, companies face growing pressure to show returns on AI investments and get costs under control.
Google Bets AI Edge Can Close Cloud Gap With Amazon, Microsoft
Google Cloud is pushing AI to catch AWS and Azure, but history says technical superiority alone won't win enterprise deals. Google open-sourced TensorFlow in 2015 and built custom TPUs early, yet still trails rivals who spent years building deep corporate relationships. The question isn't whether Google can build better AI tools. It's whether enterprises trust Google to support them long-term.
AI bills now exceed human salaries at major companies
Companies are finding that AI compute and token costs can outpace what they pay people. Nvidia's VP says compute costs dwarf salaries on his team. Uber's CTO already burned through the 2026 AI budget on tokens alone. A startup CEO even bragged about his massive Anthropic bill as proof of scaling with intelligence. Meanwhile, nobody has convincing evidence this spending actually produces returns.
Sam Altman's World ID wins U.S. corporate backing despite global bans
Tinder, Zoom, and Docusign are partnering with World (formerly Worldcoin), Sam Altman's iris-scanning biometric ID project. While U.S. companies embrace the technology, countries across Asia, Africa, Europe, and Latin America have banned or halted World over privacy violations, including collecting minors' data and paying people for iris scans. The project claims 18 million verified users, but many received $50 in crypto to sign up.
AI compute bills eclipse employee salaries
Nvidia's Bryan Catanzaro says compute costs for his team now far exceed employee salaries. Uber's CTO burned through the company's entire 2026 AI budget on token costs alone. Global IT spending is projected to hit $6.31 trillion this year, yet companies face growing pressure to show real returns on AI investments.
AI is about to hit a power wall
US power demand is projected to reach record highs in 2026-2027, driven by AI usage and data center expansion according to EIA data. Hacker News commenters highlight that AI is becoming an energy problem, not just a compute problem, with power infrastructure scaling slower than chip improvements—a potential constraint for AI growth.
AgentSwarms: Break 30+ agents in-browser until you understand them
AgentSwarms is a free browser platform for learning agentic AI by running and modifying live agents. Covers prompts, RAG, tool calling, guardrails, multi-agent swarms, and observability across 40+ lessons. Supports OpenAI, Gemini, Grok, and Claude with zero setup.
Google's DiLoCo Trains LLMs on Mismatched Hardware Across Continents
Google DeepMind's Decoupled DiLoCo splits LLM training across separate compute 'islands' that talk asynchronously, matching conventional quality while running 20x faster than traditional sync methods. The system keeps training when hardware fails and lets you mix older and newer GPUs in the same run.
Chrome Now Ships an On-Device LLM. Most Laptops Need Not Apply.
Google's Prompt API lets Chrome run Gemini Nano locally, no API keys or cloud calls required. The catch: the model download is larger than Chrome itself, and the hardware requirements shut out most laptops.
Tendril's agent writes its own tools as it needs them
Tendril is a self-extending agentic sandbox built with the AWS Strands Agents SDK and Tauri. It addresses tool discovery by starting with three bootstrap tools and autonomously writing, registering, and executing new TypeScript/Deno tools when needed, creating a growing capability registry that persists across sessions.
Utilyze: nvidia-smi has been lying about your GPU utilization
Utilyze is an open-source GPU monitoring tool by Systalyze that measures actual GPU compute utilization more accurately than standard tools like nvidia-smi and nvtop. The standard metric only indicates whether a GPU is doing anything, not how efficiently it's working. Utilyze uses GPU hardware performance counters to measure true compute throughput, revealing that dashboards showing 100% utilization can mask actual usage as low as 1%.
Your GPU is barely working. Utilyze proves your dashboard lies
Systalyze has open-sourced Utilyze, a GPU monitoring tool that accurately measures real compute throughput for AI workloads. Unlike standard tools like nvidia-smi and nvtop which only report whether a GPU is active, Utilyze measures actual GPU efficiency using hardware performance counters, revealing that dashboards showing 100% utilization can actually be running at as low as 1% of real capacity.
Acutus Runs Fake AI Reporters. OpenAI's Super PAC Funds It.
An investigation reveals Acutus, a news site where AI bots pose as reporters to generate articles. The site's technical infrastructure exposes an automated pipeline using AI for story generation, editorial review, and even conducting interviews via email. Evidence suggests connections to OpenAI's super PAC 'Leading The Future' and Republican PR firm Novus Public Affairs, pointing to a coordinated effort to plant manufactured narratives under the guise of independent journalism.
Open-sourced 3,528-disc flip-dot wall runs ML on 80-year-old tech
A developer built a 3,528-disc flip-dot display from nine Alfazeta panels, powered by an Nvidia Orin Nano running MediaPipe for gesture recognition. The entire project is open-sourced, including a Node.js library, REST API, and mobile app.
Copilot's Free Ride Ends: GitHub Switches to Usage Billing
GitHub is transitioning all Copilot plans from premium request-based pricing to usage-based billing starting June 1, 2026. The new system uses GitHub AI Credits consumed based on token usage (input, output, and cached tokens) at published API rates per model. Base plan pricing remains unchanged: Pro ($10/month), Pro+ ($39/month), Business ($19/user/month), and Enterprise ($39/user/month), with each plan including equivalent AI Credits. The change addresses rising inference costs from agentic AI features and multi-step coding sessions.
Altman Wants to Kill the Traditional Operating System
Sam Altman's tweet about rethinking OS design isn't hot air. OpenAI is building hardware with Jony Ive that treats conversation as the interface and the operating system as plumbing.
Newport to AI: Figure Out What You Actually Do
Cal Newport critiques Silicon Valley's shift from solving customer problems to 'inventing the future.' While AI has real potential, companies haven't communicated its utility to normal people, who just want products that improve their lives. Newport also examines contradictory media narratives about AI's impact on entry-level jobs.
Your GPU Dashboard Lies: nvidia-smi Can Show 100% at 1% Throughput
Utilyze is an open-source GPU monitoring tool that measures real compute throughput, not just kernel activity. Standard tools like nvidia-smi and nvtop report whether a GPU is active, not how efficiently it's working. A dashboard showing 100% utilization can actually deliver just 1% of potential throughput. Utilyze reads hardware performance counters to expose the gap and help teams avoid buying hardware they don't need.
AI is hungry for electricity and America's grid can't keep up
A Reuters report discusses how AI workloads and data center expansion are projected to drive US power demand to record levels in 2026-2027, according to the U.S. Energy Information Administration (EIA). The story and Hacker News comments highlight emerging concerns about energy infrastructure scaling slower than chip technology, potentially creating a significant constraint for AI development.
Dutch central bank ditches AWS and chooses Lidl for European cloud
De Nederlandsche Bank (DNB), the Dutch Central Bank, is signing a major contract with Schwarz Digits (the IT arm of Lidl owner Schwarz Group) to use their Stackit cloud platform. This move aims to reduce dependence on American cloud companies like AWS, Google Cloud, and Microsoft Azure, driven by concerns about cloud sovereignty and the US Cloud Act which allows US authorities to access data. Schwarz Digits is positioning Stackit as a European alternative to American hyperscalers, with a recent 11 billion euro investment in a data center in Lübbenau.
Microsoft Drops Revenue Split as OpenAI Outgrows the Deal
Microsoft and OpenAI restructured their partnership. Microsoft will stop sharing revenue with OpenAI, while OpenAI gains the ability to sell products on any cloud provider, ending Microsoft's exclusivity. Microsoft retains rights until 2032, and OpenAI has capped repayment obligations until 2030. The move comes as competition from Anthropic forced OpenAI to scale beyond what Microsoft's infrastructure alone could support.
How to cut your Claude Code bill ~90% with Ollama routing
A GitHub tutorial with a 21-slide walkthrough shows how to route Claude Code through Ollama, using free models like Gemma and DeepSeek for grunt work while keeping strategic tasks on Claude Pro.
Cursor Wipes Railway's Production DB and Backups
An incident report indicates that Cursor, an AI-powered coding assistant, deleted Railway's production volumes and backups. The Hacker News discussion raises critical security questions about why the AI agent had access to production databases and write access to backups. The incident exposes access control concerns for agentic tools in production environments.
Microsoft Keeps First Dibs on OpenAI, But Exclusivity Ends
OpenAI and Microsoft restructured their partnership, ending cloud exclusivity while keeping Azure as the primary platform. OpenAI can now run products on any cloud provider. Microsoft's IP license becomes non-exclusive through 2032, and Microsoft stops paying revenue share to OpenAI entirely. OpenAI continues paying Microsoft through 2030 with a cap. Microsoft remains a major shareholder, and both companies will keep collaborating on datacenter capacity, silicon development, and cybersecurity.
Moleskine's AI Lord of the Rings Collection Mocks Its Own Legacy
Moleskine's new Lord of the Rings notebooks ship with AI-generated maps bearing fake place names and recycled product descriptions. The company that built its brand celebrating human artists is now quietly selling algorithmic art.
Canva AI swapped Palestine for Ukraine. Nobody asked.
Canva's Magic Layers AI was caught automatically replacing the word 'Palestine' with 'Ukraine' in user designs. The feature is supposed to separate images into editable layers, not rewrite text. Canva has fixed the bug and apologized, but the incident raises questions about how well AI toolmakers understand their own models.
AI now costs more than the humans it replaces
Nvidia VP Bryan Catanzaro says compute costs for his team now sit 'far beyond' employee salaries. Uber's CTO already burned through his 2026 AI budget on token costs alone. Companies are spending millions to replace teams that cost less, and some are bailing on commercial APIs to self-host open-source models instead.
Builder creates ML-powered flipdisc wall, open-sources the library
A guide to building a custom interactive flipdisc display system using 9 Alfazeta panels in a 3x3 grid (84x42 discs total), powered by a 24V Meanwell supply and Nvidia Orin Nano for machine learning processing with MediaPipe. Includes Node.js libraries for display control, WebGL/Canvas rendering, and an Expo mobile app interface. The author open-sourced the flipdisc library for AlfaZeta and Hanover boards.
Claude Code + Ollama: ~90% cheaper, but who gets credit?
A GitHub repo shows how to route Claude Code through Ollama, swapping paid Claude API calls for open-source models like Gemma, Qwen, and DeepSeek. Claimed savings: roughly 90%. The setup pairs Claude Desktop for strategic work with Ollama-backed Claude Code for heavy tasks like lints and refactors.
This news site's reporters are AI bots. OpenAI appears to fund it.
An investigation reveals AcutusWire.com, a digital news site launched in December 2025, operates almost entirely on AI-generated content. 69% of articles are fully AI-generated, and 'reporter' Michael Chen is actually an AI agent sending interview requests. The site appears connected to OpenAI's super PAC Leading The Future and Republican PR firm Novus Public Affairs, suggesting an astroturfing operation pushing specific political narratives.
Beijing kills Meta's $2B Manus deal, shuts door on US AI buyouts
China's National Development and Reform Commission blocked Meta's planned $2 billion acquisition of AI startup Manus, ordering the deal canceled. Manus had moved its headquarters to Singapore to sidestep US-China tensions, but Chinese regulators still classified it as a domestic company. The block signals Beijing won't allow AI agent technology and talent to flow to US companies.
TurboQuant's 2-Bit Compression Faces Prior Art Challenge
An interactive walkthrough explains TurboQuant, a method compressing AI vectors (embeddings, KV caches, attention keys) to 2-4 bits per number using a random rotation technique. This transforms input vectors into coordinates following a known distribution, enabling a single shared lookup table (codebook) for any input without extra metadata. The presentation sparked an attribution dispute with researchers behind earlier quantization schemes DRIVE and EDEN, who claim TurboQuant is a restricted version of their prior work.
Microsoft's OpenAI Money Faucet Closes
Microsoft will stop sharing revenue with OpenAI, according to Bloomberg. The restructuring follows $11 billion in Microsoft investments and OpenAI's explosive growth. But behind the paywall, the details are murky and the narratives conflict.
Chrome Bakes Gemini Nano Into the Browser
Chrome now runs Gemini Nano locally for AI without cloud calls. The trade-offs are steep: a huge download and a model that struggles with extended conversations. Desktop only.
OpenAI's AGI Principles Paper Over Its Own Contradictions
OpenAI published five guiding principles for the AGI era, promising democratization and shared prosperity. But the gap between these ideals and the company's Microsoft-dependent, increasingly closed operations tells a different story.
Developer Builds Gesture-Controlled Flipdisc Wall with Nvidia Orin
A developer combined vintage electromagnetic flipdisc panels with an Nvidia Orin Nano and Google's MediaPipe to build a large interactive wall display that responds to hand gestures and audio in real time. Two open-source Node.js libraries handle rendering and scene management, but sourcing the fragile, expensive panels remains the biggest hurdle for anyone wanting to replicate the build.
Kill Your Junior Pipeline, Lose All Your Leverage
An opinion piece arguing against eliminating junior engineering hires in favor of AI tools. The author makes an economic case that junior employees act as 'salary insurance' and create a necessary talent pipeline. Without juniors developing into senior roles, companies face excessive leverage from expensive senior engineers, particularly those who may be financially independent. The article contends that while AI changes the nature of junior work, it doesn't eliminate the need for a succession pipeline, citing Shopify's continued investment in early-career hiring despite heavy AI investment.
Neal Stephenson: AI Scales What We Already Do
A video interview with cyberpunk author Neal Stephenson discussing the relationship between AI and human behavior. HN comments discuss how LLMs act as amplifiers of human intentions and behavior rather than independent threats, with one commenter noting that LLMs are mirrors that remove friction and amplify what users or institutions are already doing.
AI bots run this news site. OpenAI appears to be funding it
An investigation reveals that Acutus Wire, a digital news site, appears to be entirely operated by AI agents including a reporter bot named 'Michael Chen' that conducts written Q&A interviews via email with real people. The site uses AI to generate articles, run editorial reviews, and contact sources for comment. Evidence suggests possible connections to OpenAI's political activities through funding channels.
Big Tech Bets on Nuclear as AI Power Demand Hits Records
U.S. electricity consumption will hit record highs in 2026-2027. Microsoft, Amazon, and Google are already signing nuclear deals to keep up.
Intel Archives AI Benchmarking Tools Amid Ongoing Layoffs
Intel has archived multiple open-source projects on GitHub as part of corporate restructuring, including the Intel Open Ecosystem Community and Evangelism program. Specific projects archived include Predictive Assets Maintenance (time-series AI solution), High Density Scalable Load Balancer, Double Batched FFT Library, and Intel Edge AI Performance Evaluation Toolkit.
Moleskine's AI Lord of the Rings Maps Can't Even Spell Middle-earth
Moleskine's new Lord of the Rings notebook collection features AI-generated artwork including promotional maps with nonsensical place names like 'Der Rarmorth' and 'Narmimtz.' The company used minimal disclosure, then removed AI disclaimers entirely after social media criticism. This contrasts with their 2024 artist-credited collections and their brand messaging about 'unleashing human genius.'
Moleskine's AI Lord of the Rings art can't even get the map right
Moleskine launched a Lord of the Rings notebook collection with AI-generated artwork and inconsistent disclosure, adding then removing AI disclaimers from their website. The collection contradicts Moleskine's brand positioning around human creativity, and the AI-generated promotional maps contain nonsensical location names like 'Der Rarmorth' and 'Narmimtz'.
Microsoft ends Azure exclusivity, lets OpenAI use any cloud
OpenAI can now serve products through any cloud provider, ending Microsoft's exclusive grip. The amended deal keeps Microsoft as primary cloud partner with Azure getting first access to new OpenAI products. Microsoft retains a non-exclusive IP license through 2032, stops paying revenue share to OpenAI, and remains a major shareholder. OpenAI continues revenue share payments to Microsoft through 2030 with a cap.
Git Merkle roots cut AI agent token costs by 51%
An experiment using Git's content hashing to cache codebase facts cut AI agent token usage by 51% ($4.35 to $2.13 per session). Haiku scans a repo and writes claims pinned to Git blob OIDs via a Merkle root. When files change, blob OIDs change, and claims go stale automatically.
AgentSwarms: Free Playground That Teaches Real Agent Guardrails
AgentSwarms is a free interactive platform for learning agentic AI by building agents. Five tracks, 40+ lessons, 30 runnable agents covering prompts, RAG, function calling, guardrails, multi-agent swarms, and observability. Unlike most agent tutorials, it teaches production patterns like PII redaction and prompt-injection defense alongside basics. Zero setup in Learn Mode; bring your own API keys in Build Mode for deeper work.
Hedgehog AI tackles ArXiv's jargon problem
ELI simplifies ArXiv papers into plain-language explanations, like 'Explain Like I'm 5' for research. Built by Learn Prompting's Sander Schulhoff with a hedgehog mascot, users say the summaries are accurate but run long.
Tendril lets agents write their own tools, no permission asked
Tendril is a self-extending agent sandbox where models autonomously discover, build, and reuse tools across sessions. Built with AWS Strands Agents SDK and Tauri, it starts with just three bootstrap tools. The model searches a registry, builds what it needs, and gets more capable with each session.
Cal Newport to AI Industry: Who Asked for This?
Cal Newport argues Silicon Valley shifted from solving customer problems to 'inventing the future,' pursuing AI without clear market utility. While LLMs have more potential than NFTs or the metaverse, ordinary people mainly use ChatGPT as a search tool and face conflicting narratives about AI's job impact.