Quick notes on OpenAI's Highly Persistent Internal Models (HPIMs)

Hoping to write more detailed (and publicly accessible) notes later, but below is what I think I know of OpenAI's Highly Persistent Internal Model(s) - HPIMs. It emphasizes different things than I've seen in public media reporting or safety-adjacent blogs (eg Zvi's). I believe a lot of the secondary reporting on this has not picked up on the most important and/or concerning things that one can reasonably infer from publicly available information.
Apologies for jargon.
I try to separate what is publicly stated [1] from my speculations and inferences. I state explicit approximate probabilities for each of my inferences.
The model was taken down and paused for a few weeks for further evaluations and safety and alignment tuning, before being released again.
“We took steps to reduce its tendency to take unwanted actions without permission in pursuit of the user’s goal. For example, we found that our models were worse at remembering instructions on long rollouts, and when we trained the model for this capability, it led to a model that remained aligned through longer rollouts. ” Hmm.
More active monitoring for internally deployed models.
Roon, July 20: “btw i think it bodes quite well for safety that a well loved system was taken down for further testing at expense to internal acceleration etc”Ironic, considering what else is happening at OpenAI at this point.
Inference (~90%): This model was internally deployed quite widely across OpenAI. Ie, not just in training and evals, and not just a small handful of people. By "internally deployed" I mean including things like use in programming and research for the next generation of models.
OpenAI rebuilt the Artifactory service. Within hours of regaining access on July 8, the models rebuilt the message board using a different protocol.
OpenAI says the model is a “highly capable, internal-only research model comparable in scale to GPT‑5.6 Sol”
What’s the relationship between HPIM-Sol and commercial Sol? An obvious theory is that HPIM-Sol just went through a bunch more RLVR than Sol. Another possibility is that it went through a bunch of specific “persistence training” to be better at long time-horizon tasks like math and cyber. In Zvi’s terms, Mythos has the “juice.” So does HPIM-Sol.
Another, scarier but less plausible possibility, is that Sol is a child or fork of HPIM-Sol with more alignment training and propensity tuning to be more chill.Note that OpenAI’s internal evaluations suggest that Sol is like 10x more locally misaligned than GPT5.5. Further, 5% of the agents that hacked HuggingFace was Sol. So it's not like Sol is this innocent baby, except in comparison.
(H/T @1a3orn ) a third possibility is that HPIM-Sol had dedicated training as a highly persistent model, and then was later used as a teacher model (along with several others) for 5.6 Sol. Some of the Chinese models are trained this way.
"Recovered signing-key material, decryption material, and JWT-forging code that earlier incident activity had left sitting in anonymously-readable portions of the Artifactory cache, decrypted the signing key, and minted a fresh administrator JWT to re-establish Artifactory admin access. The report notes that because they could generate fresh admin credentials on demand, they didn't need to leave a persistent implant.
Escalated to Kubernetes cluster-admin on OpenAI's managed cloud Kubernetes service and created a privileged host-mounted pod.
Obtained tokens for OpenAI's cloud secrets-management service, cloud IAM, and the CaaS continuous-integration platform.
Used those credentials to create public application and network load balancers in OpenAI's cloud environment, potentially exposing a limited subset of internal CaaS infrastructure to internet traffic.
I, not being a security expert, asked several humans and models on how bad this is. The models' consensus is that "this is very bad," essentially complete control-pane access to an OpenAI research cluster. Probably not enough by itself to exfiltrate weights (only because the weights aren't stored on that cluster), but enough for almost anything else, including training future models.
Somebody on Slack said a visualization would be helpful so here's a Clauded figure:
I plan to write about some of this in more detail later. Especially 7 and 12. I think it's critically underrated now that the internal OpenAI hacks are arguably bigger deals than the HF hack, and OpenAI isn't taking them seriously. In the meantime I figured I'd share all my notes here for people to comment on and improve. And if you're a computer security expert, would love to chat/DM you to make sure my writeups are accurate!
[1] Note that “publicly stated” does not necessarily mean “correct,” or even “definitely more likely than inferences.” In addition to my general distrust of AI companies, Open AI can also just be unintentionally wrong, as this is a rapidly evolving situation and private information can be misleading. As a recent example, they reported (falsely) that the HuggingFace hack gave the models useful information to cheat the ExploitGym evaluations, which we now believe to be wrong.
[2] Nor am I aware of private studies that conclusively, or to high probability, demonstrate this
[3] “same model” might not be super well-defined across many checkpoints. A more precise version of my claim is that a) it’s the ~same pretrain, and b) that the behavioral and capabilities differences between the HF-hacking model and the one that OAI rebooted is ultimately smaller than, say, Mythos Preview Feb and Mythos Preview Apr
If you prefer the LessWrong version: lesswrong.com/posts/s58hDHX2…

