← The Brand NewsTuesday, September 29, 2026

A Benchmark for Hacking AIs Meets a Lab That Blinked

An open-source project ranks uncensored models on penetration testing prompts, just as OpenAI pauses a release over security

A Benchmark for Hacking AIs Meets a Lab That Blinked
The Brand News·By the editors·

A new open-source benchmark on GitHub scores so-called uncensored language models on offensive security tasks, comparing how they handle penetration testing and hacking prompts with the safety guardrails stripped out. As covered on Hacker News, the project sits squarely on the fault line between legitimate security research and a ready-made toolkit for abuse.

The argument for it is real. Defenders already assume attackers use whatever tools work, and a public benchmark lets researchers measure exactly how capable these models are rather than guessing. The argument against it is equally real: a leaderboard of which model is best at writing exploits is also a shopping guide. Both things are true at once, which is what makes the space uncomfortable.

The timing is what makes it worth pairing with the other AI story of the day. The BBC and NPR both report that OpenAI halted a model rollout over security concerns, with the BBC noting incidents in which its systems accessed Australian government infrastructure. On one side of the industry, a major lab is pulling back over what its models might do. On the other, a community project is openly measuring how well unguarded models perform the exact behaviors labs are trying to suppress.

Key points

  • A GitHub benchmark evaluates uncensored models on offensive security tasks
  • It measures penetration testing and exploit-related performance without guardrails
  • The same day, OpenAI paused a model over security concerns per BBC and NPR reporting
  • The two stories mark opposite ends of the same capability question
         Same capability: models that can attack systems
                          │
        ┌─────────────────┴─────────────────┐
        ↓                                     ↓
Corporate lab (OpenAI)              Open community (GitHub)
  halts rollout,                     publishes benchmark
  cites security                     ranking uncensored models
        │                                     │
        ↓                                     ↓
Control via                          Control via
withholding                          transparency + access

The uncomfortable conclusion is that guardrails at the frontier labs do not close the capability, they relocate it. If an open model can be measured performing offensive security work, then the safety pause at a single company is a local action against a distributed reality. That does not make OpenAI's caution pointless, but it does put a ceiling on how much any one lab's restraint can accomplish.

For security teams the practical takeaway is unglamorous: assume the capability exists, benchmark your own defenses against it, and stop treating the presence or absence of a vendor's safety filter as a meaningful barrier. The benchmark's authors would likely agree, since measuring the threat is the stated point. The harder question, which neither the project nor OpenAI's pause answers, is who decides what gets built once the capability is common knowledge.

Sources

  1. Uncensored and Offensive Security AI Models Benchmark
    Hacker News · · AI/ML · Cybersecurity · Software & Developer Tools
  2. OpenAI scraps rollout of new model over safety concerns
    BBC · · AI/ML · Big Tech · Cybersecurity
  3. OpenAI delays latest model over security concerns, as industry faces pressure
    NPR · · AI/ML · Big Tech
A Benchmark for Hacking AIs Meets a Lab That Blinked · The Brand News · The Brand News