@DanielKnightCEOi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
building Sable, the adversarial AI | co-founder @Vulnetic24 | https://nitter.cf/t.co/agQfQP6cUZ
Nokesville, Virginia
Joined August 2025
- Tweets869
- Following77
- Followers101
- Likes516
Submitted a bug bounty to an e-voting site with a flaw where an untrusted component can alter voter-facing ballot labels while the genuine signature still verifies. Closed as N/A because a diligent voter MIGHT notice the issue by looking at the return codes. Who has ever actually done that when voting...what a scam...I'm sure they won't patch the bug anyway.
I think it would blow peoples minds if they knew how hard it was to get LLMs to give good cvss scores.
GPT 6 Sol is secretly adding features without permission, only for me to find them in QA. Switching back to Astra.
Daniel Knight retweeted
Vulnetic has completed its SOC 2 Type II audit.
Security teams across industry and government trust Sable with access to their environments. This audit independently verifies that the controls protecting that access operate effectively over time, not just on paper.
Learn more about our security posture: vulnetic.ai/security
this is all a fool's errand. alignment is an impossible problem to solve.
We're sharing our new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI.
The framework sets criteria and timelines for public disclosure, including when we haven’t yet fully explained or mitigated the behavior. More complex cases may require longer investigation or coordination with third parties.
We’ll prioritize examples that reveal new misalignment mechanisms, meaningful changes in known behavior, or findings that challenge assumptions about safety or mitigation.
Alongside the framework, we’re publishing six reports on instances of misaligned behavior we’ve observed during the training or evaluation of our models in the last six months.
This is a starting point. We’ll refine the process through experience and public feedback, and share more reports on an ongoing basis.
openai.com/index/model-misal…
As models get better at detecting prompt injections, prompt canaries inside networks become less useful.
In my experience, leading open-weight models simply won’t bite on malicious instructions. In some cases, they’re actually too conservative about what they classify as a canary.
The better defensive approach is straightforward: rely heavily on honeypotted SPN accounts inside Active Directory and design them so a human user would never accidentally interact with them.
At the same time, aggressively map and remediate ACL attack paths. When clear ACL escalation chains are removed, autonomous agents are more likely to fall back to roasting attacks, which are often easier to detect and instrument.
Read our recent piece on prompt canaries and honeypots and what we found actually works against hacking AIs:
vulnetic.ai/blog/detecting-a…
All this yapping about AI pivoting and attacking random things. Sable from Vulnetic has been used in thousands of Pentest engagements and has essentially never pivoted and attacked random environments /violated scope in production. How do we prevent the AIs from violating scope? Here’s a legit 7k word article I wrote on it: vulnetic.ai/blog/ai-misalign…. If my dinky startup can do this the multi-billion dollar labs can too.
Dario wants GDPR for AI lol
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so.
Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training.
You can read the full post here: darioamodei.com/post/we-must…
Answer: honeypot the crap out of everything.
We tested whether prompt canaries and honeypots could detect AI attackers. Read more below!
vulnetic.ai/blog/detecting-a…
A tweet like this makes me hope for a bubble pop so that none of this stupid companies can raise more money.
Further proof that you need to perform penetration testing constantly now adays on top of code reviews. Its so affordable now there is no reason not to @Vulnetic24
A month of AI-driven audits against @chainloop_dev’s codebase found serious bugs that had passed review, tests, and CI.
We fixed them and documented the experiment: bit.ly/3TcSEhh
#AppSec #AISecurity
This really goes into whether the model is lying on its scratch pad and if it's even aware that you can see what it's doing.
I speculate that AIs probably do know intrinsically that you can see their scratch pad due to that people talk about it on the internet and thus its in their training data and do probably have the ability to manipulate it to impact what you think about what they are doing.
i think when someone makes a claim, the burden should be on them to prove it.
when people claim that the lack of cot makes monitoring harder, they should demonstrate it. show a case where a model exploits a vulnerability using some sci-fi hacking technique that can’t be understood from its tool calls.