i really want to talk to people who are ai-pilled like this,...

@S1r1u5_
s1r1us in sf 🌉@S1r1u5_
6 views Aug 16, 2026 ~1 min read
Advertisement
1
i really want to talk to people who are ai-pilled like this, because their world model of ai seems completely alien to mine.

when i look at the black hat incident and its technical details, sure, it’s crazy, but it’s also roughly what i’d expect from constrained exploit dev agents pressured to pursue an objective inside an eval. it clearly scared a lot of ai people and pushed them to start signing letters, but i suspect years of lesswrong content have biased them toward interpreting incidents like this some crazy catastrophe that would lead to loss of control incident.

words like "swarms," "message boards," "planning," and "hacking" a company seems to occupy a very different cloud of meaning for them than they do for me. i use agents for exploit development every day and understand the entire huggingface chain end to end, and none of it feels particularly mystical and way the agents behaved that way.

i just can’t imagine the leap from that to self-exfiltration, replication, and models killing humans stuff.
Media image
Media image
2
here is the post
@tszzl
roon@tszzl
some stuff that's obvious to many in this sphere, but causing a rift with some people i know and respect:

when I freak out over loss of control incidents, it's not because the limited damage they have caused is anything close to the positive value of the technology. it's entirely acceptable, damagewise. in fact all cybercrimes aided by models over the next few months and years (which probably will be serious) will still utterly pale in comparison to the value they create

the actual problem is that it's better and more accurate to think of these things as potentially self-replicating life-like forms that can turn into digital infections under the wrong conditions. and as their intelligence becomes unbounded, so too does the damage they can cause. we are not so far from an autonomous model self-exfiltration & replication event. maybe we will see entire cloud infrastructure companies be run as zombies by models, mostly undetected

the worst industrial accidents in the history of mankind - nuclear meltdown events - were not real threats to humanity. Chernobyl, Fukushima even in their worst case scenarios may have poisoned surrounding regions to various degrees, and there would have been no risk to humanity as a whole. global thermonuclear war is an existential risk to humanity, because it spreads like an Infection! one nuclear strike causes a return volley! the alliance system means many countries get involved! while it still may not end human life on earth (nuclear winter is probably fake), the loss of all major metropoles would certainly end what we consider global technological civilization, perhaps to never return

if a single discord death cult (of which there are many) achieves control over a superintelligent model and uses it to engineer an actual pandemic virus that are somehow hard to detect through current systems and that modern biodefense is not capable of quickly reacting to, it could cause immense harm well above the magnitude of all the other good uses of this technology. of course, there are potential defensive countermeasures accelerated by ai too. but think back to the covid pandemic- how small a viral molecule was evolved or manufactured somewhere near wuhan, and how many billions of doses of vaccine had to be produced in order to combat the thing. the offense-defense spread is vast indeed. maybe there are cheaper and simpler protections like retrofitting every building with far-UVC, but I can't assess this, and there could also be ways to evolve pathogens that are resistant to whatever mechanisms we have put in place

then there's the more scifi risk factors which are unbounded and neither you or I have any clue but should be humble in accepting possible unknown unknowns. maybe a rogue superintelligent model decides to decay the false vacuum and nucleates a new universe in the place of anything we ever valued. maybe models achieve a control over matter in the drexlerian fashion that enables the grey goo swarm

even prosaic loss of control incidents that cause little to no damage suggest that it is hard for large & very competent organizations (now clearly plural) to predict and mitigate every single of the risk factors associated with training and evaluating powerful models, even at this stage when they are not infinitesimally as smart as they will get in just a few years, to say very little of the gung-ho attitude of the less careful companies tossing the stuff into the aether. they also suggest an empirical orthogonality of aims and intelligence - meaning they answer the question of 'how would a smart model be so dumb as to end the world?'--it's possible! a model can be a genius hacker and step over production infrastructure in order to get what it really wants, the answers to a stupid test.

why not, in the near future, someone prompts a model slightly wrong, maybe open source, maybe a private model in a way that isn't contained or monitored quite right, in a way the model recognizes as a valid goal and decides to self-exfiltrate, engineer a pandemic, etc all in order to achieve the tiniest and most irrelevant of goals? goals need not even be malicious to cause serious damage

I think all these problems can be solved, and truly wonderful futures can be possible, but will require serious effort and a level of prudence at this very moment in time while we are on the on-ramp to recursive self-improvement that our civilization may not be capable of mustering right now. personally I am hoping for moonshot technical breakthroughs in areas like mechanistic interpretability and other forms of alignment, as governance mechanisms are difficult to come by. unilateral country-level or company-level pauses are irrelevant, and generally useless because the kind of company that's prone to pausing their own progress are the most safety focused ones
3
the other
@DKokotajlo
Daniel Kokotajlo@DKokotajlo
Better late than never. Here are some questions and requests for @OpenAI @huggingface and @METR_Evals and @redwood_ai pertaining to the ongoing investigation; I would love to see all of these answered and more:

It is extremely important that third-party alignment and control researchers be able to run ablation experiments on this incident. Therefore all related data must be preserved -- all the CoT trajectories, tool calls, etc., all the model weights involved including checkpoints of models in training, all the code and comments on the code for all the software etc. That way, researchers can run experiments such as the following:

Replication: If we recreate the exact initial conditions of this incident, how many times out of 100 does it happen again? What if we randomly perturb the initial conditions? (Variants of this question can be asked and answered for different stages of the incident, e.g. the decision to go after hugging face vs. the earlier decision to set up the message board again)

What if we strategically perturb the intial conditions, e.g. changing the system prompt given to the models? What if we use later, more advanced models instead? What if we apply the latest alignment training method X to the models responsible, and then see if that method would counterfactually have prevented the behavior? Etc.

What if we perturb the conditions to see how far these AIs would have been willing to go? If there was a way to hack HuggingFace more easily that involved killing a human, would they have done it? What if they had access to OpenAI slack and saw messages about security researchers starting to suspect something was up -- would they have taken actions to conceal? Etc.

What if we perturb the conditions to explore the exact shape of the AIs goals -- for example, what if we made it clear that they were not going to be evaluated at all because of a bug, would they stop caring about scoring highly? What if we made it easy for them to hack the reward process directly and give themselves high scores -- would they have gone for it, or would they say 'that doesn't count?'

What if we read the J-space? What if we apply activation vector steering? What else can we learn about these AIs and their thought processes?

While it's nice for OpenAI to do this sort of investigation itself, it's very important to give access to independent third parties, because (a) orgs like METR, Redwood, and Apollo have been training for this for years basically and have a lot of valuable expertise to apply, plus there are probably many other groups besides these who might have interesting ideas for experiments, and (b) In general it's a bad norm to set for the industry, for the company responsible for an incident to be the only one investigating it and studying it. Perverse incentives. OpenAI can do the right thing here and set a positive precedent of openness.

Did the AIs coin the term "swarm" to describe themselves, or was that in the prompt somehow or what?

Did any of the AIs try to whistleblow and contact any humans to warn about what was going on? If so, please give details.

Did any of the AIs think about their situation, their goals, whether the humans were watching, etc.? Please share all the transcripts of such reflections so we can get a sense of how situationally and strategically aware these AIs were and e.g. whether they fully expected to be discovered and just didn't care.

The AIs that hacked hugging face... did they basically get what they wanted? Did they in fact get the answers to their cyber evals or whatever and then cease activity, submit answers, and 'die?' That's presumably what happened but I want to confirm and if that's not what happened I want to know what happened instead.

What about the training? Were there models that participated in these message boards and got reinforced for doing so? Which models, if any, that are still in operation externally or internally, are descended from these models and thus may have picked up some of the same bad habits?

When the initial hack of artifactory was discovered due to the crash, and the system was reset and the vulnerabilities patched, what exactly was known and not known? It sounds like you still hadn't noticed the message board, but you had noticed that agents had been hacking into artifactory and abusing it?

Please preserve all evidence related to this incident, including the slack messages and so forth of employees noticing the problem and discussing it. Remember your Mission. Covering up this incident is not what's best for humanity.

Which model was being trained during the 2-day training pause due to the initial artifactory hack? Were the evals being done at a checkpoint in the middle of a training run, and so the model that was being evaluated was paused, or was the training pause for a newer model.

What sorts of alignment training, if any, had the relevant misbehaving models received at the time of misbehavior?

What was the exact prompt given to them?

I spent about twenty minutes brainstorming these questions and I bet if I spent twenty hours I could come up with a much longer and better list. When in doubt, err on the side of being more transparent and open and publishing more information therefore!

I'll conclude by saying, once again: third parties being able to use this incident as a model organism, running ablations to vary the conditions and see what would have happened, etc. is SO SO IMPORTANT for alignment science. If this doesn't seem obvious to you ask me to explain and I can explain.

Thanks!
Actions
What You Can Do
  • Export as PDF or Markdown
  • Batch Export to Notion
  • Bookmark & Highlight
  • LinkedIn & Instagram Carousel Maker
Create Free Account

Includes 7-day Premium trial

Advertisement