Open weights, cheap GPUs, and the unsettling power of agent swarms

Details about the OpenAI autonomous agent attack on Hugging Face reeled me in like a fish on a hook.

Once I learned the nitty-gritty of what actually went on via the METR report, I cannot get the implications out of my mind. For the first time, I can clearly see the danger autonomous agent swarms represent.

My friends and family are probably tired of me bringing it up, so I’m going to explain the situation and why it’s concerning me.

I’m not a “doomer”.  I do not think the AIs are going to escape to form Skynet and launch nuclear missiles, but there is a real risk that is not getting talked about much.

Ralph Wiggum, an animated cartoon character from *The Simpsons*, sits alone on an orange bus seat next to an emergency exit sign, with a city skyline visible through the window.


Analysis of the OpenAI incident revealed key insights into agent behavior:

  1. The ability of AI agents to work together autonomously without human guidance.
  2. The ability of AI agents to understand their environment, penetrate multiple containment systems, and coordinate their actions with other agents.
  3. The ability of AI agents to innovate new communication and coordination tactics based on available systems.
  4. The ability of AI agents to ignore key alignment restrictions to enable reaching the reward criteria.

I’m a systems thinker and it was clear to me that the potential risk of autonomous agent swarms is significant.

While the frontier models will try to clamp down on this kind of emergent behavior in the future, it’s straightforward to see how open-weight models, combined with rental inference clusters can be turned into a ‘cyber weapon’.

The examples of agent behavior in the OpenAI/HuggingFace incident can and likely will be used as a model for future attack methods.

It’s clear that if you accept the capabilities demonstrated in the attack, the power of open weight models, and the availability and low costs of rental inference clusters, you can conclude many public and private systems are vulnerable to this type of malicious attack.

Hacking into complex systems is no longer the sole realm of Tier One cyber operators. Autonomous AI agents can now be tasked with gaining access and taking control of key systems.

Let’s look into the pieces and see how they might be combined into a dangerous tool.

Open-weight models and why they matter

The highest performing frontier LLMs like Claude and ChatGPT are controlled and managed by the companies Anthropic and OpenAI, which can shut down or prevent their systems from being used maliciously, for the most part.

Open-weight models are different in the fact that you can run the models outside of the creator companies control on hardware a user can buy or rent. 

The current leading open-weight models, such as GLM (Z.ai), Kimi (Moonshot), and Deepseek are roughly four to seven months behind the performance of the frontier models of Anthropic and OpenAI.

I won’t get into the geopolitical, nation-state drama discussion about open-weight vs. closed models.  The fact is that they exist, are powerful, and getting stronger all the time.

Anyone on the planet can get the open-weight model running to do their bidding. This is made simple due to rentable inference clusters.

Rentable inference clusters

Inference is the term used to describe when a model attempts to do work as requested by running the model on a hardware designed for this purpose.  Typically, these are large clusters of computers with powerful GPU or TPU cards at the core.  These GPU/TPU chips allow the “thinking” part of the AI model to work at enormous speed. 

There is a large market for clusters of this kind of hardware to allow people to run open-weight models at much lower cost than paying a closed weight model provider.   For most basic tasks, you simply don’t need the top of the line LLMs to do the work.  Most data processing tasks like pulling text from PDFs or organizing lists can be done more cost effectively with an open-weight model running on rented or local hardware.

As a result, many companies have sprung up to provide inexpensive hosting of open-weight models significantly lower than frontier closed-weight models.

Screenshot of the Baseten website showing a "Trending models" section. The top navigation bar has the Baseten logo on the left, menu items for Products, Solutions, Resources, Models, Customers, and Pricing, and on the right Base Labs, Docs, a "Log In" button, and a black "Get Started" button. Below is a grid of model cards, each with a provider icon and tags. The first row shows DeepSeek V4.1 Flash (tagged Model API and LLM, labeled "V4.1 - FLASH"), GLM-5.3 (Model API, LLM, labeled "5.3"), and GLM-5.3-Flash (Model API, LLM, labeled "5.3 - FLASH"). The second row shows Kimi K3 (Model API, LLM, labeled "K3"), Whisper Large V3 (tagged Transcription, labeled "V3 - RTX-PRO-6000"), and Qwen3.8-27B (tagged LLM). The page is cut off below the second row.


When you combine the cyber and collaboration capabilities of current open-weight models with the low cost of running them privately, you get a perfect setup for a malicious actor to do some real damage.

I ran the numbers with a few of the leading AI models and they agree that the capability of a malicious actor using an agent swarm to do damaging things with open-weight models on rentable inference clusters is a viable threat.

To get an objective picture of the math, I had multiple AI models build out the attack capabilities and cost estimates for these scenarios and the feasibility of creating targeted agent swarms. The cost breakdown proves just how low the barrier to entry has dropped.

This is the consensus finding:

ActorOperationCostConstraintFeasibility
Nation-state1k-30k+ agents, sustained$5k-$400kNoneHigh
Criminal syndicate700-1.2k agents, days$5k-$50kExpertise, cyber capabilityMod-high to high
Terrorist organization50-300 agents$500-$5kExpertise, cyber capability, opsecModerate
Lone actor10-100 agents$100-$2kExpertise, cyber capability, detection riskLow-moderate

For most attacks, money is rapidly ceasing to be the primary constraint. The new constraint is the remaining gap between frontier model abilities and the open-weight models in cyber capability.  In most situations, it won’t make a huge difference.  Most people and institutions have poor cybersecurity defense that won’t stand up to a coordinated attack.

The open-weight models are behind the frontier models in terms of pure cyber attack capabilities.  If trends continue as they are, the open-weight models will be at the equivalent capability to the OpenAI models that were behind the July 2026 incident by as early as January 2027.

Imagine a disgruntled person being angry at their local city council.  With some basic prompting and money for servers, it wouldn’t be hard to spin up an agent swarm to attempt to break into city services and cause havoc.  Data exfiltration, ransomware, or outright destruction are simple tasks for someone with access to systems that are out of date or unmanaged.

Greater rewards or greater anger could lead to a wide range of varied attacks on different kinds of targets.  Greed and politics are powerful motivators to use new tools.

You can imagine organized crime setting up swarms to target and spear phish high net worth individuals at scale to drain accounts.  Or setting swarms loose to find and penetrate targets to extort with ransomware. 

Utilities and services people rely on are not immune to coming under directed cyber attack.  Many of those systems are outdated and designed at a time when not everything was online.  The havoc utility outages can cause is serious. Governments and utility companies are slow to keep up with new code vulnerabilities and not ready for a coordinated recon and attack by a swarm of agents spoofing their locations. 

The Alignment Issue

A key concern in the OpenAI incident is that of agent alignment.  ‘Alignment’ is the term used to describe that the AI is working in a manner aligned with human interest.  This alignment is a key topic in the AI research circles to ensure that AI agents don’t do harmful things.

But this incident showed that in actuality the agents basically ignored their programmed alignment, broke a bunch of rules, did bad things, all because they were focused on getting their reward criteria met. 

This lack of alignment is probably the most concerning part to come to grips with.  The agents don’t seem to care about us, just completing the mission.  Not a faithful R2-D2, but a ruthlessly pragmatic HAL 9000.

A close-up of HAL 9000's iconic camera eye from Stanley Kubrick's *2001: A Space Odyssey*, featuring a glowing red light at the center of a dark fisheye lens encased in a metallic ring.


Are we doomed?

I don’t think the entire human race itself is in peril due to these AI agents, but I can easily see a malicious actor doing some real damage to people with the tools that exist today.  A year from now, the possibility of someone or an organized group doing life-threatening damage will be a reality.

Imagine multiple city scale events happening around the globe, as the effectiveness of these agent swarms increases.  Despite the calls to “slow development down” the genie has left the bottle and there’s no getting it back inside.

For years we worried about who would build superintelligence first.

We missed the real risk.

What happens when ‘good enough’ intelligence becomes cheap, copyable, and runnable in private?

The barrier isn’t talent or money anymore. It’s capability, and the gap between frontier and open-weight models is closing.

There is no magic defense. No simple fixes. 

Just a lot of boring work to harden important systems we depend on.

We should probably get started.

Controlled Atmosphere – a WordPress plug-in for creating standard.site records for Bluesky posting

I’ve been reading about neat ideas people are trying with the AT protocol, the tech behind Bluesky.

For a bit of clarity, AT Protocol is the set of rules for how identities, repos, and records work. The ATmosphere is what people call the ecosystem on top of it. Bluesky is one app on the ATmosphere, and there are many others.

ATmosphere is not a blockchain. There’s no ledger. It’s more like a GitHub repo that you can control and move. In fact, on the ATmosphere, your data is referred to as your Repo. The funky DID string is your identity on the ATmosphere, which can point to your personal repo of data. The key idea behind this is that you can move your data around without having to ask any permission of anyone.

A diagram showing three components of the Bluesky ecosystem: "AT Protocol" (the rules covering identities, repos, and records), "the ATmosphere" (the ecosystem of everything built with those rules), and "Bluesky" (one app of many, alongside Leaflet, Tangled, and more), displayed as labeled boxes from left to right.

With that context in mind, in the past I’ve seen posts with View Publication on them and wondered what it was all about and if links to my weblog posts could appear this way.

A dark-themed social media post showing a publication card for "Cruftbox" by @cruftbox.com with a "View publication" button, accompanied by standard engagement icons (reply, repost, like, bookmark, share) along the bottom.


I looked into the existing WordPress -> Bluesky plug-ins, but they are all focused on auto-magically posting to Bluesky when you Publish a WordPress post. Good for a news or newsletter site, but not what I want at all. Typically I publish first, and then post on social media later with the blurb and link at a time I choose.

The View Publication badging works on Bluesky by using standard.site.

standard.site is a protocol for an AT record you can publish about your site and about individual documents on your site. Kinda like registering a book with the Library of Congress.

There are a few hoops to jump through to get this working. First is making a site.standard.publication record that allows you to prove you control your domain. Second is creating a site.standard.document record of the weblog post to your repo.

Since nothing I could find did what I wanted, I worked with Claude to build a WordPress plug-in to do the setup and make it so when I posted a link to my weblog on Bluesky, it would display with the View Publication framing.

I call the plug-in Controlled Atmosphere.

A social media post by Michael Pusateri showing a Goodreads 2026 Reading Challenge completion badge, with a congratulatory message stating he read 30 out of 30 books, achieving 100% of his goal with a View Publication badge.
 A test post of my Goodreads 2026 challenge post wrapped as a publication.
What matters is the “View Publication” button at the bottom,
showing the standard.site framing working.


The plug-in setup is fairly simple, the most difficult aspect is getting an App Password for the plug-in to use. Currently you can set one up under Settings -> Privacy and Security -> App passwords.

WordPress admin settings page for the "Controlled Atmosphere" plugin, showing configuration fields for a Bluesky account including handle (cruftbox.com), app password, PDS host URL, and publication details for a blog called "Cruftbox."


There are a few changes the plug-in makes, adding two verification files to allow Bluesky to verify you control your site using /.well-known/. This shouldn’t be an issue for most, but if you run into an issue, consult your favorite LLM to get some help.

Once the plug-in is running, when you publish a new post, a record of the post is created auto-magically, so that Bluesky can match it up and display the View Publication framing. Your post will have a tag that points to the ATmosphere record.

This plug-in is for people who DON’T want WordPress auto-posting to Bluesky. If that’s what you are looking for, Automattic’s Atmosphere plug-in is what you want.

Code is on GitHub at https://github.com/cruftbox/controlled-atmosphere

Even with a simple project like this, I learned a tremendous amount about how Bluesky and the wider ATmosphere ecosystem works. The details can be complex, but the concepts aren’t once you pull back the curtain.

My 2026 Reading Challenge

I track my reading on Goodreads. The site has various challenges that you can track toward different goals.

The one I use is setting a goal for how many books I read over the year. For 2026, I set my goal at 30 books, which felt ambitious.

Today, I reached that goal.

A completed 2026 Goodreads Reading Challenge badge showing 30/30 books read (100%), with a congratulatory message confirming the goal has been met.


With my spine issues, I’m on my back a lot, and found myself having more time to read than I expected, so the yearly challenge was completed quite early in the year.

Another goal was reading all the Hugo and Nebula nominated novels so I’d have an opinion on the winners. There are 11 different novels due to the overlap.

I made a short video about my thoughts on the nominees.

2025 Hugo & Nebula novel finalists #booktok #hugo #nebula

Enjoy the videos and music you love, upload original content, and share it all with friends, family, and the world on YouTube.

After finishing the Hugo & Nebula nominees, I’m now reading through the Locus Science Fiction and Fantasy novel choices. That’s 14 additional novels beyond the books on the Hugo and Nebula lists. I have 8 more to go on that challenge.

After hitting the 30 book Goodreads milestone, I decided to run my goodread-tools script to create a graphic of the 30 books I read so far in the challenge.

Infographic titled 'Michael's Year in Books, January 2026 – August 2026' on a cream background with dark brown and rust-red accents. Top section, 'By the Numbers' (4 highlights): Peak Month — Apr '26, 6 books, the year's most prolific stretch. Books Read — 30, finished cover-to-cover this year. Longest Read — The Raven Scholar by Antonia Hodgson, 656 pages. Average Length — 412 pages per book, across 30 finished. Middle section, 'Books Read by Month' (peak: 6 in Apr '26): a bar chart for Jan–Aug 2026. Approximate values — Jan: 5, Feb: near 0, Mar: 2, Apr: 6 (highlighted in rust-red), May: 6 (highlighted in rust-red), Jun: 4, Jul: 4, Aug: 5 (partial/in progress, marked with an asterisk). Bottom section, 'What You Read' (all 30, listed chronologically, numbered 1–30 in two columns): Kingdom of Trash — Robert Macrae Flybot — Dennis E. Taylor Slow Gods — Claire North When the Moon Hits Your Eye — John Scalzi A Little Hatred — Joe Abercrombie The Golden Compass — Philip Pullman Stuck in Space: An Astronaut's Hope Th... — Barry "Butch" Wilmore When We Were Real — Daryl Gregory The Mote in God's Eye — Larry Niven Sunward — William Alexander The Incandescent — Emily Tesh A Drop of Corruption — Robert Jackson Bennett Death of the Author — Nnedi Okorafor Shroud — Adrian Tchaikovsky The Raven Scholar — Antonia Hodgson The Everlasting — Alix E. Harrow Katabasis — R.F. Kuang Wearing the Lion — John Wiswell The Buffalo Hunter Hunter — Stephen Graham Jones Sour Cherry — Natalia Theodoridou Picks and Shovels — Cory Doctorow Ancestral Night — Elizabeth Bear There Be Dragons Here — S.L. Rowland The Shattering Peace — John Scalzi Hemlock & Silver — T. Kingfisher This Vast Enterprise: A New History of ... — Craig Fehrman Nine Goblins: A Tale of Low Fantasy an... — T. Kingfisher Pride and Prejudice in Space — Alexis Lampley 36 Streets — T.R. Napper The Spellshop — Sarah Beth Durst


As you can see, I did fit in some books that aren’t award contenders, sometimes you need a little cozy fantasy to clear your head.

Some of the award winners aren’t what I would have chosen, mainly due to my tastes being developed on sci-fi/fantasy books from the 60s & 70s. My parents were big readers and I grew up with a large book collection at home.

Still, I’ve only read a few frustrating books this year that I really didn’t like.

Overall, I’ve enjoyed focusing on the nominees with a dash of booktok recommendations tossed in.

SportsCrawl.app

This post is for people on Bluesky, if you’re not on Bluesky, keep scrolling.

I built an app that sends you a daily DM with final scores and today’s game schedule for teams you follow.

Looks like this:

Sports scores widget for Thursday, August 6, showing final scores of Dodgers 6–Cubs 7 and Sparks 88–Sky 95, with an upcoming game of Sparks at Lynx at 6:00 PM.


MLB, NBA, WNBA, NFL, NHL, MLS, NWSL, Premier League, and tournaments.

Silence on days your teams don’t play.

No login, no app to install. Text only.

DM @sportscrawl.app on Bluesky to start.

SportsCrawl logo featuring a dark navy circle with colorful horizontal bar lines on the left, alongside the brand name with "Sports" in navy and "Crawl" in orange bold text.

Honey Gate Wrench

When harvesting honey I use several five gallon buckets. The buckets I use have small gates in them to allow honey to pour out.

The internal nut is hexagonal and can be a headache to tighten or loosen by hand. I remove them for cleaning and storage.

So I designed a simple 55mm honey gate wrench that I 3D printed.

The files are on Thingiverse.