I switched my Hermes agent to 100% free local models with Magnitude It’s running Qwen 3.6 35B-A3B at ~60 tok/s on my DGX Spark. It’s free, private, and always on running background tasks Hermes set Magnitude up itself. With the CLI it: - Profiled my hardware and found the best models for it - Walked me through the options and let me decide - Switched itself over From there, models load just in time as the agent works and unload when idle Works with Pi, OpenCode, Hermes, OpenClaw, Codex, Claude Code, Oh My Pi, and Cline Copy this prompt and send it to your agent: “Set up local models for me with the Magnitude CLI. Install it with `npm i -g @magnitudedev/cli` (or my package manager), then run `magnitude docs onboarding` and follow the instructions” GitHub: github.com/magnitudedev/magn…
21
23
2
222
22,881
The only problem I have is that Qwen3.6-35B and also 27B often get stuck in a loop, and when I check the background they produce over 50 K tokens. Hermes also frequently runs something in the background when Bots are used, which likewise generates a lot of tokens.
1
1
340
Interesting. Would be curious if they still get stuck when running in Magnitude. Sounds like tool call doom loops that are more common in local models, which we specifically correct for

Sep 2, 2026 · 9:41 PM UTC

1
275
RelevantRecentLikes
I haven't tried Magnitude yet, but I'm giving it a try to see if there's any improvement.
1
31