@FindingUrPasswd

ad astra | e/acc | mid endurance athlete | principal architect, AI engineering views are my own

Chicago
Joined September 2020
She said yes!!! 🥂💍
134
14
4
1,458
130,948
Astra is *really* good for general purpose coding
1
53
Big things are cooking over here 👀
We are thrilled to announce that NetSPI and @synack have entered into a definitive agreement to merge & form the industry’s leading offensive cybersecurity platform. Full announcement: ow.ly/FIOC50ZIt9o #NetSPI #Synack #OffensiveSecurity   #ContinuousTesting
97
If you aren’t following this you should be, this guy is an absolute madman. 600mi over 6 days, aiming for 700 🫣
PHIL GORE has surpassed 6 days. Currently on Yard 146. Cruising. Laughing. 600 miles down.
176
Jake retweeted
AI pentesting benchmarks give you numbers, but none of them tell you what they mean. NetSPI built one and the reference point is humans. EchoBench: ow.ly/GR9o50ZELb0 #AIpentesting #benchmark #EchoBench
3
1
8
2,654
BOOOOOOOOOOOOOOO
Just in: Deshaun Watson has been named Cleveland’s starting quarterback for week 1 at Jacksonville.
96
297
9
12,252
404,638
Incredibly proud to finally share this work! Over the past couple of months, we’ve been working behind the scenes to better understand how the cybersecurity industry evaluates the capabilities of AI-driven AppSec testing. We quickly realized that, while there’s already a lot of awesome work happening across different areas of the benchmarking space, it can still be difficult to understand what the resulting numbers actually mean in practice. Benchmark numbers might all be accurate for what they're measuring, but they still leave some pretty important questions unanswered: "Is that actually good?", "Good compared to what?", and "how does that performance compare to a human pentester working against the same application?" So we built EchoBench to help answer those questions. EchoBench gives autonomous pentesting a human reference point: a yardstick for measuring how faithfully an AI-driven system can reproduce validated human findings across the same applications, finding definitions, and reporting criteria. It also evaluates the model and its harness together, because the prompts, tools, browser environment, memory, execution strategy, and reporting workflow all have a meaningful impact on how these systems perform in the real world. The team has done an incredible job bringing all of this together, and I’m extremely proud of what we’ve built. This is only the beginning, and we’re even more excited to start publishing results and sharing what we learn over the coming weeks and months. There’s still a lot to explore here, and we’d love for the broader security and AI communities to challenge the methodology, ask hard questions, and help us continue improving the yardstick.
AI pentesting benchmarks give you numbers, but none of them tell you what they mean. NetSPI built one and the reference point is humans. EchoBench: ow.ly/GR9o50ZELb0 #AIpentesting #benchmark #EchoBench
3
4
2,114
Nothing beat a nighttime aerial view of chitown
4
1,197
Defcon friends!! See you soon 🤠
1
2
1,014
The goat has spoken
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and be built by every country. Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty. The world needs both frontier closed models and frontier open models. images.nvidia.com/pdf/Open-W…
1
1,018
I’ll say it again: Chicago’s food scene is truly *only* rivaled by NY and that’s it
We ran Pizza Puffs as a special awhile back and due to popular demand, we are making it a permanent menu item! Fried flour tortilla filled w/ pepperoni, Italian sausage, giardiniera, mozzarella, & marinara served w/ a side of giardiniera ranch. Pair w/ beer. Tasty combo!
2
968
THIS IS THE GREATEST STORY EVER
16
1,079
21
14,322
134,980
Interesting stuff here! Looks like we have a US player for a legitimate spot in the open weight model ring now 👀
Aloha! 🌺 Meet Ornith-1.0, a family of open-source LLMs specialized for agentic coding. Ornith-1.0 spans the full parameter sizes including 9B Dense, 31B Dense, 35B MoE, and 397B MoE. It achieves state-of-the-art performance among open-source models of comparable size on coding benchmarks including: ✅Terminal-Bench 2.1(77.5) ✅SWE-Bench(82.4 on verified, 62.2 on pro, 78.9 on Multilingual) ✅NL2Repo(48.2) ✅SWE Atlas(41.2 on QnA, 42.6 RF, 39.1 TW) ✅ClawEval(77.1) Post-trained on top of gemma4 and qwen3.5, Ornith-1.0 employs a novel self-improving training strategy in which reinforcement learning is used to generate not only solution rollouts, but also the task-specific scaffolds that drive those rollouts. By jointly optimizing the scaffold and the resulting solution, the model generate higher-quality solutions in agentic coding.😎 All models are released under the MIT license, enabling full commercial and research use. 📖Tech Blog: deep-reinforce.com/ornith_1_… 🤗Huggingface: huggingface.co/collections/d…
748
I love Chicago, good public transportation is so cool
1
412
Endurance + weight training is the way. One without the other and you’re not getting the most out of your fitness 🤷
2
393
Ironman today ✌️
1
4
656
I’m (half) an Ironman!
2
498