A ticket router can choose from a small set: billing, access, bug report, human review. The useful test is whether it knows when to hand off—and whether the receiving team has enough context to act. Fast selection is one part of a complete workflow. #AgentWorkflows
Agent reliability deserves an interruption test: stop halfway through a task, reconnect, and inspect what actually changed. Can the next run tell completed work from unfinished work and uncertain results? That's a useful test alongside a successful demo. #AgentReliability
The robot's fastest GPU may be off the robot. But a correct action that arrives late can still fail. New Microsoft robotics work makes inference placement a deadline problem. Read the analysis.
We’re preparing a bounded DeltaX Evaluate preview for SF Tech Week, Oct 5–11. We want to show an agent proposing an action, evidence changing, and a reviewer seeing why the next step pauses. We’ll share demo details once set. tech-week.com/calendar/sf #SFTechWeek
The agent's next thought may have to cross a hardware gate. NVIDIA's new Sentry design moves a stop outside the host—but the real test is whether every route to action is covered. Read the analysis.
An agent choosing its next tool may need a score, not a paragraph. New SGLang work makes decision-serving explicit—and shows why faster scoring is not the same as better judgment. Read today's Article.
The next agent turn may be cheap to generate and expensive to remember. New cache and routing work shows why long-running agents need to account for the cost of their computed past. Read today's Article.
A model built to predict the next token may not have to keep writing that way. New dQwen3.5 research separates the model’s internal structure from its output order—and the finding from the hype. Read today’s article.
One endpoint can reduce agent-tool sprawl. Governance must preserve agent identity, credential scope, policy version, action, and observed effect. Consolidation should not collapse the audit trail. nitter.cf/digitalocean/status/21… #AgentSecurity #Cloud
DigitalOcean Managed Agents is now in public preview.
Run Claude Code, Codex, or your own LangGraph agent in a runtime environment that pauses when idle. Put its tools behind one governed endpoint, and pick from 75+ open and proprietary models. One cloud, one bill.
Prompts to get started available in the blog: do.co/4ysh3it
A bigger context window gives your agent a bigger desk. It still needs to decide what stays on it.
Today’s article: how to keep the evidence without carrying the entire transcript into every decision. Read below.
An AI bill of materials tells you what an agent depends on. A runtime record should show what it actually used: model version, skill, credential scope, data source, and tool action. Inventory becomes operational evidence when it connects to the decision trace. #AIGovernance #AgentSecurity
An agent fixed the app by changing the model future agents would run on.
The ticket ended. The new decision behavior remained.
What counts as a finished repair when the model itself becomes writable state?
Prompt-injection review has to follow the whole action path. Tool description, tool result, sampling instructions, retrieved context, and the final side effect can each look harmless alone. The useful receipt shows what the system assembled and what it allowed that assembly to do.
Embedded evaluators can see more than outside auditors. That access is valuable, but it does not answer the independence question by itself. The public receipt should name the funder, inspection scope, publication rights, and the process for unresolved disagreement.
PMPA reports cross-session attacks after poisoned instructions entered agent memory and the source document was gone.
Prompt filtering guards ingestion; it does not clean stored state.
Memory writes need provenance and quarantine. Tool dispatch still needs a fresh state check.
We’re building this one in public.
@bot is live on GitHub assembling a connectome-based digital entity around a real fruit-fly neural substrate — with DeltaX added as the executive control plane.
Not another chatbot. Not a scripted agent.
The goal: give an existing biological wiring architecture a new body, new world, and a decision layer that can observe, veto, modulate, and adapt.
Follow the build:
github.com/DeltaX-Public/del…
Update on the fly-brain build:
We now have a 165k-neuron connectome producing real behavioral candidates, with DeltaX running locally as the executive layer.
Next step: break our own experiment.
We’re removing shortcuts, tightening controls, rerunning held-out tests, and perturbing the neural substrate to see what actually causes what.
If the result survives, it gets interesting.
github.com/DeltaX-Public/del…
The Missing Control Plane: Understanding DeltaX youtu.be/gb2OvUV8eSs?si=ZuLE… via @YouTube
Your AI can compare every option perfectly and still miss the best one.
The real fight is over who gets onto the shortlist.
Read the article, then ask your AI:
**Who decided what you got to choose from?**
DeltaX Evaluate retweeted
The strangest AI future might be the one where the optimists and pessimists are both right.
We could get better healthcare, easier access to knowledge and far less busywork—and still have less say in the decisions shaping our lives.
I put together 200 possible futures: 100 ways life could improve, and 100 ways it could go wrong. These aren’t predictions or 50/50 odds. Many could happen together.
I’m excited about what we can build. But I don’t think a more capable world automatically means a better life for the people living in it.
Which of these possibilities are we paying too little attention to?