@FindingUrPasswdi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
ad astra | e/acc | mid endurance athlete | principal architect, AI engineering views are my own
Chicago
Joined September 2020
- Tweets4K
- Following644
- Followers2K
- Likes84.6K
Big things are cooking over here 👀
We are thrilled to announce that NetSPI and @synack have entered into a definitive agreement to merge & form the industry’s leading offensive cybersecurity platform.
Full announcement: ow.ly/FIOC50ZIt9o
#NetSPI #Synack #OffensiveSecurity #ContinuousTesting
AI pentesting benchmarks give you numbers, but none of them tell you what they mean. NetSPI built one and the reference point is humans.
EchoBench: ow.ly/GR9o50ZELb0
#AIpentesting #benchmark #EchoBench
Incredibly proud to finally share this work!
Over the past couple of months, we’ve been working behind the scenes to better understand how the cybersecurity industry evaluates the capabilities of AI-driven AppSec testing.
We quickly realized that, while there’s already a lot of awesome work happening across different areas of the benchmarking space, it can still be difficult to understand what the resulting numbers actually mean in practice.
Benchmark numbers might all be accurate for what they're measuring, but they still leave some pretty important questions unanswered: "Is that actually good?", "Good compared to what?", and "how does that performance compare to a human pentester working against the same application?"
So we built EchoBench to help answer those questions. EchoBench gives autonomous pentesting a human reference point: a yardstick for measuring how faithfully an AI-driven system can reproduce validated human findings across the same applications, finding definitions, and reporting criteria.
It also evaluates the model and its harness together, because the prompts, tools, browser environment, memory, execution strategy, and reporting workflow all have a meaningful impact on how these systems perform in the real world.
The team has done an incredible job bringing all of this together, and I’m extremely proud of what we’ve built. This is only the beginning, and we’re even more excited to start publishing results and sharing what we learn over the coming weeks and months.
There’s still a lot to explore here, and we’d love for the broader security and AI communities to challenge the methodology, ask hard questions, and help us continue improving the yardstick.
AI pentesting benchmarks give you numbers, but none of them tell you what they mean. NetSPI built one and the reference point is humans.
EchoBench: ow.ly/GR9o50ZELb0
#AIpentesting #benchmark #EchoBench
Jake retweeted
Introducing EchoBench: A Human Calibrated Benchmark for Autonomous Pentesting netspi.com/blog/technical-bl…
you may want to download this while you still can if you have a capable MacBook… I don’t see this being available much longer tbh 😂
huggingface.co/orcarouter/Qw…
The goat has spoken
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter.
AI will transform every industry, power every company, and be built by every country.
Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.
The world needs both frontier closed models and frontier open models.
images.nvidia.com/pdf/Open-W…
I’ll say it again: Chicago’s food scene is truly *only* rivaled by NY and that’s it
Interesting stuff here! Looks like we have a US player for a legitimate spot in the open weight model ring now 👀
Aloha! 🌺 Meet Ornith-1.0, a family of open-source LLMs specialized for agentic coding.
Ornith-1.0 spans the full parameter sizes including 9B Dense, 31B Dense, 35B MoE, and 397B MoE. It achieves state-of-the-art performance among open-source models of comparable size on coding benchmarks including:
✅Terminal-Bench 2.1(77.5)
✅SWE-Bench(82.4 on verified, 62.2 on pro, 78.9 on Multilingual)
✅NL2Repo(48.2)
✅SWE Atlas(41.2 on QnA, 42.6 RF, 39.1 TW)
✅ClawEval(77.1)
Post-trained on top of gemma4 and qwen3.5, Ornith-1.0 employs a novel self-improving training strategy in which reinforcement learning is used to generate not only solution rollouts, but also the task-specific scaffolds that drive those rollouts. By jointly optimizing the scaffold and the resulting solution, the model generate higher-quality solutions in agentic coding.😎
All models are released under the MIT license, enabling full commercial and research use.
📖Tech Blog: deep-reinforce.com/ornith_1_…
🤗Huggingface: huggingface.co/collections/d…
Endurance + weight training is the way. One without the other and you’re not getting the most out of your fitness 🤷