The OpenAI & HuggingFace Saga
I’ve held off on writing about this for reasons I can’t quite explain. I think I feel like it’s hard to really posture an opinion without more of the facts but since all of the AI alignment world feels drawn to wild conclusions based on what we currently know I think it’s fair to write down some thoughts.
- I don’t think it is broadly surprising that agents will do “unaligned” things if you give them “EXPLOITGYM” and tell them to go exploit stuff, while also providing them with tasks that are impossible to solve.
- A lot of the “exploits” involved seem to me to be of the nature of misconfigurations that have tended to be common in most real world deployments due to lack of resources / real world threats against it. That’s quickly changing and we should assume all CVEs are real CVEs and start acting accordingly.
- As always, I worry about regulatory capture in the wake of this. I think an underreported on fact is that due to “safeguards”, HuggingFace had to turn to non-domestic open source models to aid in their investigation.Yes I know they should have applied for privileged access but that sort of furthers the point. Infosec works partly by the fact that you have more to gain by protecting your systems than bad actors have by exploiting them. But if bad actors will have 10x more access to “frontier level cyber models” than the average defender, we will all lose.