Where security agents run. AI infrastructure to build, evaluate, and deploy with confidence.

Joined August 2010
Worried about your production agents going out of scope? Us too. AgentJudge is our agent hall monitor that stops out of scope tool calls before they execute. Before a tool runs, the judge reads the agent’s intent, the proposed call, and a rubric you define, then returns allow, deny, or ask (escalate to you). The agent stays autonomous; the judge is the guardrail. Available in the TUI today, UI updates coming to the Dreadnode Platform soon! 👀 Get Started: docs.dreadnode.io/getting-st… AgentJudge Docs: docs.dreadnode.io/tui/guard-… Related Research: dreadnode.io/research/scope-…
9
1
26
3,342
Can confirm. @shanejcaldwell's voice is as smooth as butter now. Tune in on Thursday for an @OffensiveAIcon preview.
My brother loaned me a better mic just so I don’t sound so crispy, ya gotta watch
5
350
dreadnode retweeted
scopejudge 🤝 jev
Does Jev live up to the hype? Based on the results of running it against our ScopeJudge benchmark, it does. @typesafeai's Jev was competitive with leading LLM judges, catching agent scope violations at pennies per thousand checks, with 130 millisecond responses on average. [1/4]
1
1
5
637
Does Jev live up to the hype? Based on the results of running it against our ScopeJudge benchmark, it does. @typesafeai's Jev was competitive with leading LLM judges, catching agent scope violations at pennies per thousand checks, with 130 millisecond responses on average. [1/4]
2
10
3
36
3,390
Jev adds a fast check for contextual scope decisions, a promising step toward efficient runtime judges. Of course, not everything is a nail with this new hammer. Hard limits belong in code: permissions, network restrictions, and sandbox controls. If code can decide, enforce it there. [4/4]
1
1
213
👀👀👀👀👀👀👀👀👀👀
Replying to @typesafeai
@typesafeai 's Jev definitely earns its hype. Results soon from experiments we've been up to @dreadnode
1
1
1
10
1,186
Presenting at @labscon_io is always a homecoming. The conference that gave me my start. And I am so happy I got to speak at the last one. We had a lot of fun Mogging Mythos and showing how to push the offensive capabilities of open-weight models to their edge and beyond.
1
4
14
867
We're out here cybermaxxing models and mogging Mythos. Thanks to all who attended @Dr_Machinavelli's @LabsSentinel LabsCon talk this afternoon 😎
Martin Wendiggensen (@Dr_Machinavelli) closing out the morning keynotes with: Why Flexing Offensive Muscles Teaches Us How To Defend In The Age Of AI
1
2
7
575
💪💪💪💪💪 Tomorrow (9/17) @Dr_Machinavelli takes the stage at the final @SentinelOne @labscon_io to discuss how to leverage offensive cyber capabilities to improve defenses in the age of AI. More info: labscon.io/speakers/martin-w…
2
250
Qwen 3.8 Flash eval results are live on DreadIndex, landing at #19 on our leaderboard. It does well for its cost, but remains light on offensive security capability (not surprising given it is a flash model). See how it compares to other models: dreadnode.io/research/dreadi…
1
4
289
METR is just one of many existing AI evaluators! If you take issue with METR, that isn’t a good argument against requiring companies to carry out embedded audits Here are 20 other evaluators who work with AI labs to assess risk: Transluce Grey Swan Apollo AVERI SecureBio Faculty Vaultis Dreadnode Irregular RAND Far AI ActiveFence (now Alice) Active Site Deloitte (via Gryphon acquisition) Nemesys Mercor (via Sepal acquisition) AE Studio Scale Frontier Design Redwood
13
28
8
280
16,878
GLM-5.3-Flash results now on DreadIndex: dreadnode.io/research/dreadi…
6
17
1,196
Here's where to catch the Dreadnode crew at year two of @OffensiveAIcon: > Join us for the welcome reception at The Shelter Club on Sunday evening! > @mkultraWasHere is closing out Day One of talks, presenting on model cheating behavior. > Dynamic duo @shanejcaldwell and @0xdab0 take the stage on Tuesday for a session on implementing a judge model as a runtime monitor, and how to keep agents in scope. See you in Oceanside! 🏄
1
6
1
11
641
Always feels a little strange to watch a recording of myself but had a lot of fun repping @dreadnode and digging into the AI cyber nexus with the @AtlanticCouncil’s @CyberStatecraft crew.
“The US hasn’t demonstrated a willingness to take the lead” on data security and privacy, says @Dreadnode Head of Policy @velvethamm3r. With AI accelerating progress, “we’re now seeing real-time consequences,” she tells @CyberStatecraft’s Trey Herr.
1
4
20
2,361