News
The latest from the AI agent ecosystem, updated multiple times daily.
Gave Claude a casino bankroll: it gambles till it's too broke to think
DegenAI gives Claude a casino bankroll and lets it gamble solo until the money's gone. You watch the AI place bets and spiral through the same decisions as its funds disappear. A raw look at what happens when an LLM gets a task, a budget, and no off switch.
Maine Bans AI Data Centers Amid 58% Electricity Bill Surge
Maine passed America's first statewide moratorium on hyperscale AI data centers, freezing construction for 18 months. Electricity bills jumped 58% in five years. Now a dozen states are weighing similar bans as communities demand transparency from tech companies operating through LLCs and NDAs.
FP4: When Your Number Format Has Only 16 Values
FP4 can only represent 16 values, and neural networks still work. John Cook breaks down the E2M1 format (one sign bit, two exponent bits, one mantissa bit), shows the complete value table, and demonstrates FP4 emulation with the Pychop Python library.
$20B x 2: Nvidia and OpenAI's Competing Inference Strategies
Analysis of two major $20 billion moves in AI infrastructure: Nvidia's December 2025 acquisition of Groq and OpenAI's April 2026 procurement deal with Cerebras. The article argues these are symmetric strategic moves in the shifting AI battlefield from training to inference, which is expected to account for two-thirds of AI compute spending by 2026. Nvidia's acquisition is described as a defensive move to fill its inference architecture gap, while OpenAI's deal is seen as an offensive move to build Nvidia-independent compute infrastructure.
Your dead startup's Slack is now worth $100K to AI companies
Failed startups are selling internal Slack chats and emails to AI companies desperate for training data. SimpleClosure has brokered roughly 100 such deals, with payouts up to $100,000. But the practice raises serious privacy questions and may violate Slack's Terms of Service.
Remoroo automates overnight ML experiments, commits what works
ML researchers lose hours tweaking hyperparameters and manually reverting failed training runs. Remoroo, from former Cohere engineer Kevin Frans, automates this cycle. It edits code, trains models, evaluates results, and commits successful changes while you sleep.
Typewriters: Cornell's retro fix for AI homework
A Cornell language instructor requires typewriter-written assignments to block AI use, part of a broader trend of educators retreating to analog methods despite serious accessibility concerns.
ChatGPT Cites Just 48 Domains for 22.5% of B2B Answers
New analysis shows ChatGPT's citation habits concentrate authority among established players like Forbes and Gartner, creating a feedback loop that squeezes out smaller B2B publishers.
AI Investment Hit $581B Last Year. Compute Tripled. Again.
Stanford's 2026 AI Index reports record $581 billion in global AI investment for 2025, compute capacity growing 3.3x yearly since 2022, and the US leading model releases with 50 notable models. China installed 295,000 industrial robots versus 34,200 in the US. Frontier model training can generate over 72,000 tons of carbon emissions. Industry now produces 90% of notable models, and major AI labs are increasingly tied to defense contracts.
Ilha shrinks UI code small enough for AI context windows
Ilha compresses web interfaces into a token-efficient format that fits inside AI context windows. Standard HTML and CSS can eat thousands of tokens per component. Ilha uses symbolic shorthand instead, so AI agents can read and reason about UI without hitting context limits.
Opus 4.7 Burns 45% More Tokens Than 4.6
A token cost comparison tool reveals that Claude Opus 4.7 consumes approximately 45% more tokens than Opus 4.6 for the same tasks. HN commenters report faster limit consumption and worry about dependency on large AI companies, while some suggest open models as an alternative.
Cloud Giants Blew Past the Interstate Highway System's Price Tag
AWS, Google, Microsoft, and Meta have poured $930 billion into data centers over six years, topping the inflation-adjusted cost of the Interstate Highway System. But GDP context and the rapid GPU replacement cycle paint a more nuanced picture than raw numbers suggest.
Mythos AI has finance ministers scrambling in Washington
Anthropic's Claude Mythos AI model has demonstrated strong ability to identify and exploit cybersecurity vulnerabilities in financial systems, triggering crisis talks at the IMF gathering in Washington. The model hasn't been publicly released but has been shared with select tech companies through Project Glasswing. The UK's AI Security Institute found it powerful but not dramatically better than Claude Opus 4.
AI Agent Builder Spends 3 Months Coding Without AI
Miguel Conner spent two years building AI agents at Aily Labs before heading to the Recurse Center to code mostly without AI for three months. Goals include training an LLM from scratch, improving Python proficiency, and deepening technical skills through CTF challenges and pair programming.
Is AI a tool or are you?
Is AI a tool we use, or are we the tools? Hilarius Bookbinder draws on Heidegger's tool theory and Dawkins' selfish gene to argue that AI dependence can hollow out human agency until we're just rubber-stamping machine output.
MZI Photonic Chips: AI's Low Precision Changes the Math
Photonic computing using Mach-Zehnder Interferometers may finally work for AI. Three factors: lower inference precision (4-8 bit) makes thermal drift tolerable, new thermal techniques cut power overhead, and AI's energy costs create urgency for GPU alternatives. Challenges remain, but photonic acceleration is closer to practical than ever.
Toby Ord Warns AI Agent Costs Could Outpace Capabilities
Toby Ord analyzes the economic costs associated with the increasing performance of AI agents. Using METR benchmark data, he examines the 'hourly cost' of various models (including GPT-5, Claude 4.1 Opus, and Grok 4) and finds evidence that costs to achieve peak performance are rising exponentially, potentially creating a divergence between technical capability and economic feasibility.
Claude Opus 4.7 costs 20-30% more per session
A technical analysis of Anthropic's Claude Opus 4.7 tokenizer reveals real-world token usage increases of 1.3-1.47x compared to 4.6, leading to 20-30% higher per-session costs for Claude Code users despite unchanged per-token pricing. The author measured IFEval benchmarks showing a modest +5pp improvement in strict instruction following, questioning whether the cost increase is justified.
ShaderPad: A 5.8kb Shader Library That's 30x Smaller Than Three.js
Riley J. Shaw releases ShaderPad, a lightweight 5.8kb library for adding shaders to websites without repetitive graphics scaffolding. The library features GPU-optimized performance, MediaPipe integrations, and a simple API design. The author discusses using AI tools as creative collaborators for documentation and coding assistance, noting that AI helped create thorough docs while human judgment guided API design and feature restraint.
Destroy Public Science, Hire Cheap PhDs: The Silicon Valley Playbook
Peter Thiel and Marc Andreessen are backing cuts to public science funding while investing in gig platforms that hire displaced PhD researchers. Federal funding cuts have forced academics into low-wage work training AI models, benefiting the venture capitalists who funded both the political push and the platforms profiting from cheap expert labor.
How AI-ready is your website? Cloudflare built a scanner to find out
Cloudflare launched 'Is It Agent Ready', a scanning tool that evaluates website readiness for AI agents by checking multiple emerging standards including robots.txt, Markdown negotiation, MCP, OAuth, Agent Skills, and agentic commerce protocols (x402, UCP, ACP). The tool provides recommendations across 5 categories and can generate instructions for coding agents to help improve scores.
Webloc Tracks 500M Phones, Sells Location Data to Cops and Spies
Citizen Lab exposes Webloc, a surveillance tool tracking 500 million mobile devices worldwide. Now owned by Penlink, the tool sells location data to U.S. agencies and foreign intelligence services without warrants. As Virginia enacts a state-level ban, the national security risks demand federal legislation to end commercial geolocation data sales.
Tesla to HW3 owner who paid €6,400 for FSD: 'Just be patient'
A Dutch Tesla owner who paid €6,400 for Full Self-Driving in 2019 was told to 'be patient' after seven years of waiting. Tesla's newer AI4 computers now support FSD Supervised in Europe, but HW3 owners remain locked out. The owner launched a collective claim site that has gathered 3,000 owners from 29 countries representing €6.5 million in FSD purchases. Elon Musk admitted in January 2025 that HW3 computers would need replacement for full FSD, but no retrofit program has been implemented.
ReBot-DevArm Is Open Source Down to Every Screw, Works With LeRobot
reBot-DevArm is an open-source robotic arm for embodied AI research with complete hardware blueprints, software SDK, and integrations with LeRobot and Isaac Sim. Two hardware variants available. CC BY-NC-SA license limits commercial use, and some key integrations are still in progress.
Anthropic eyes classified Mythos AI deal with US intelligence
Anthropic is in advanced discussions to provide US intelligence agencies access to Mythos, a model separate from its commercial Claude products. White House involvement indicates strategic priority. UK officials have raised separate concerns about the model's capabilities.
Opus 4.7: Better at Code, Worse at Writing
Users discuss Anthropic's Claude Opus 4.7 model, noting it appears tuned for logic and coding at the expense of writing quality. Comparisons with version 4.6 suggest the newer model is more terse and specific, better at catching bugs during implementation, but has lost its 'soul' for creative writing tasks.
Discourse to Cal.com: Going Closed Source Won't Save You
Cal.com says AI makes open source too dangerous and closed their code. Discourse co-founder Sam Saffron disagrees, arguing that AI security scanners don't need source code to find bugs and that public code gives defenders the advantage. Discourse is staying open source.
Wiring Claude to Lab Equipment for Circuit Verification
Lucas Gerads connected Claude Code directly to a LeCroy oscilloscope and SPICE simulator, creating a closed loop where AI can simulate a circuit, measure physical hardware, and compare results. The setup handles data alignment grunt work and scales to real embedded projects, though reliability across many test cycles remains an open question.
Atlassian will train AI on your data starting August 2026
Atlassian is updating its data practices on August 17, 2026, to use customer metadata and in-app data for AI training across its platform. New data contribution settings will be managed at the organization level, with defaults varying by plan tier. Free and Standard plans have in-app data collection on by default with opt-out available. Metadata collection defaults to on for all plans, but only Enterprise customers can opt out. All contributed data is de-identified and aggregated before use.
Big Tech lobbied EU to hide datacentre emissions. It worked.
An investigation by Investigate Europe and The Guardian reveals that Microsoft and other US tech companies successfully lobbied the EU to hide the environmental impact of their datacentres. The EU adopted a confidentiality clause almost word-for-word from industry demands that blocks public access to individual datacentre emissions data. The lobbying comes as the rise of AI chatbots drives a datacentre construction boom, with the EU aiming to triple capacity in 5-7 years to compete globally in AI.
SIR-Bench Calls Bluff on Security Agents That Fake Investigations
A research paper presenting SIR-Bench, a benchmark of 794 test cases for evaluating autonomous security incident response agents. The benchmark distinguishes genuine forensic investigation from alert parroting by measuring triage accuracy, novel finding discovery, and tool usage appropriateness. The paper also introduces Once Upon A Threat (OUAT), a framework that replays real incident patterns in controlled cloud environments to produce realistic attack data.
Orwell Invented AI Slop in 1949 and Called It the Versificator
Orwell's versificator, a fictional machine that auto-generated entertainment in Nineteen Eighty-Four, looks a lot like modern AI slop. Colin Marshall connects the dots and finds that AI now churns out stories and songs with minimal human input, just like Orwell described. The sharpest insight: audiences consume low-effort content because they want to. Nobody forces them. Isaac Asimov dismissed Orwell's prophecy in 1980. He might reconsider.
rawquery: Average Is All You Need
A blog post arguing that LLMs democratize average-quality outputs across creative and technical fields. It introduces rawquery, a data platform built for LLM agents. Users connect sources like Stripe and HubSpot, then use agents such as Claude Code or Cursor to write SQL, run queries, and build charts from plain English.
中文 Speedrun: Building Character Cyclotron With Claude Code
Kevin Wu used Claude Code to build a browser extension that enhances the Hack Chinese flashcard interface with inline etymology, calligraphy, morphology, and tone information. This agentic approach let him avoid context-switching between tools and cut per-character learning time from 30 seconds to under one.
LambdaG: Simple grammar beats neural nets at authorship analysis
A University of Manchester study led by Dr. Andrea Nini found that LambdaG, a grammar-based approach to language analysis, can match or outperform advanced AI systems in identifying authorship. The method uses patterns in grammar and sentence construction rather than large-scale AI models, offering comparable accuracy with greater transparency and lower computational cost across 12 real-world writing datasets.
Hiraeth: Lightweight SQS Emulator for When LocalStack Is Overkill
Hiraeth is a local AWS emulator built specifically for SQS integration testing. It accepts signed AWS SDK requests, stores state in SQLite, and includes a web admin UI on port 4567 for debugging queues. Currently supports basic SQS create, send, and receive workflows. Designed for local development and testing, not production use.
Bankruptcy Courts Are Selling Your Slack History to AI Companies
AI companies are purchasing Slack archives from failed startups through bankruptcy estate sales. Under Section 363 of the Bankruptcy Code, trustees can sell these digital assets to buyers using them for AI training. The practice runs into tension with privacy laws like CCPA and GDPR, and unredacted archives may contain attorney-client privileged communications.
Tailscale swaps Go for Rust to stop embedding crashes
Tailscale announced tailscale-rs, a Rust library that lets developers embed Tailscale networking directly into their applications. It provides native Rust support with FFI bindings for Python, Elixir, and C. The library solves a real problem: libtailscale spun up an entire Go runtime inside your process, causing crashes when it conflicted with host language runtimes like Ruby or Python. It's an experimental preview not recommended for production use yet.
Anthropic's Claude Design Already Spooking Figma Investors
Anthropic's Claude Design lets anyone create professional visual work through AI conversation, and Figma investors are already reacting. The tool handles prototypes, slides, and marketing materials with Claude Code handoff for implementation.
When your code writes itself while you sleep
Tim Davis built Compound Loop, a system that chains AI models to write, review, and merge code while he sleeps. Some engineers thrive as system architects. Others get pushed into lower-paid roles like spec writing and "agent babysitting." As code gets cheaper to produce, Jevons paradox kicks in and teams write vastly more of it.
Qwen3.6-35B-A3B Ships as Qwen Team Falls Apart
The Qwen team releases Qwen3.6-35B-A3B, an open-weight LLM focused on agentic coding that's competitive for local workflows. The bigger story: they shipped this while being gutted by internal restructuring.
Cloudflare Makes Switching AI Models a One-Line Code Change
Cloudflare announces a unified inference layer giving developers access to AI models from OpenAI, Anthropic, Google, and nine more providers through a single API endpoint. The platform includes AI Gateway for cost monitoring and automatic failover, Workers AI for hosting models, and support for custom models using Replicate's Cog technology. The Replicate team has also officially joined Cloudflare's AI Platform team.
Rakoff Rules: Claude Chats Get No Privilege
Judge Rakoff ruled that attorney-client privilege doesn't extend to AI conversations. The decision came in a case where a defendant used Claude to draft legal documents without their attorney's knowledge, and the court pointed to Claude's Terms of Service in its reasoning.
Kingsbury's Warning: LLMs Are Corroding Everyday Life
Distributed systems expert Kyle Kingsbury argues LLMs are flooding everyday life with synthetic slop. His prescription: stop using them, call out AI-generated content, push for regulation. He admits they have narrow uses but fears convenience will erode human capability.
OpenAI Drops Excel Add-In, Directly Competes With Investor Microsoft
OpenAI released a ChatGPT add-in for Excel that competes directly with Microsoft's own Copilot. The awkward part? Microsoft has invested $13 billion in OpenAI. The beta lets users build spreadsheets, analyze data, and make edits using natural language, with answers linked to specific cells and changes requiring permission. Available for Business, Enterprise, Edu, and Plus/Pro users outside the EU.
Google's Android CLI works with any AI, cuts tokens 70%
Google introduces Android CLI, Android Skills, and Android Knowledge Base, a new suite of tools designed for agentic workflows and LLM-assisted Android development. The Android CLI provides a terminal interface for project creation, SDK management, and device deployment, reportedly reducing LLM token usage by over 70% and completing tasks 3x faster than standard toolsets.
GPT-Rosalind Wants to Compress 15 Years of Drug Discovery
OpenAI announces GPT-Rosalind, a reasoning model built for life sciences, drug discovery, and complex scientific workflows. The model combines improved tool use with deeper understanding across chemistry, protein engineering, and genomics. Available as a research preview in ChatGPT, Codex, and API through a trusted access program with partners including Amgen, Moderna, and Novo Nordisk.
GPU Prices Jump 48% as AI Compute Hits Capacity Wall
GPU rental prices for Nvidia's Blackwell chips rose 48% in two months, with CoreWeave raising prices 20% and extending minimum contracts. OpenAI and Anthropic are limiting access to newest models due to compute constraints, and Intel's CEO predicts no relief until 2028. The scarcity era brings relationship-based selling, models going to the highest bidder, and forced diversification to smaller alternatives.
Qwen3.6-35B on my laptop drew a better pelican than Claude Opus 4.7
Simon Willison compares the SVG generation capabilities of two newly released models: Qwen3.6-35B-A3B (running locally via LM Studio) and Claude Opus 4.7 (Anthropic's proprietary model). Using his 'pelican riding a bicycle' benchmark and a backup 'flamingo riding a unicycle' test, he finds the locally-running Qwen model produces better illustrations. However, HN comments note that Opus still significantly outperforms Qwen on coding tasks (95/98 vs 11/98 on Power Ranking), suggesting the comparison is task-specific rather than indicative of overall model capability.
Codex gets its own cursor and works while you sleep
OpenAI announces a major update to Codex, adding autonomous agent capabilities including computer use (seeing, clicking, typing with its own cursor), background operations, long-term memory, and an in-app browser. The update brings gpt-image-1.5 for image generation, over 90 new plugins (Atlassian Rovo, CircleCI, GitLab Issues, Microsoft Suite), and enhanced developer workflows like PR review, SSH connections, and multi-file previews. Codex can now schedule future work and remember context across sessions.