@TechWaveNate

🚀 Future-Proofing Industries with Tech 🔧

Southampton
Joined December 2023
AI agent security is becoming just as important as the agents themselves. 10 open-source GitHub projects for giving agents real tools without giving them access to everything. ISOLATE THE AGENT  1.⁠ ⁠E2B isolated sandboxes for agent execution. ⁠ github.com/e2b-dev/E2B ⁠  2.⁠ ⁠gVisor sandbox untrusted workloads. ⁠ github.com/google/gvisor ⁠  3.⁠ ⁠Firecracker lightweight microVM isolation. ⁠ github.com/firecracker-micro… ⁠  4.⁠ ⁠Kata Containers containers with VM-level isolation. ⁠ github.com/kata-containers/k… ⁠ CONTROL THE TOOLS  5.⁠ ⁠Open Policy Agent policy checks before actions execute. ⁠ github.com/open-policy-agent… ⁠  6.⁠ ⁠Casbin permissions and access control. ⁠ github.com/casbin/casbin ⁠ PROTECT THE SECRETS  7.⁠ ⁠Infisical secrets management for AI infrastructure. ⁠ github.com/Infisical/infisic… ⁠  8.⁠ ⁠SOPS encrypted secrets inside repos. ⁠ github.com/getsops/sops ⁠ WATCH WHAT HAPPENS  9.⁠ ⁠Falco detect suspicious runtime behavior. ⁠ github.com/falcosecurity/fal… ⁠ 10.⁠ ⁠Trivy scan code, containers and infrastructure. ⁠ github.com/aquasecurity/triv… ⁠ the security loop: request → permission → sandbox → tool → execute → inspect → allow/block → log a useful agent stack isn’t just memory + tools + reasoning anymore. it needs permissions too.
7
6
37
49,678
Nathan retweeted
Building an AI workflow is harder when you can’t explain what happened in each run. auto-brain is a source-available server runtime for teams building AI-powered business workflows. It helps you run model calls and durable workflows with a record of each primitive’s inputs and outputs, exposed over HTTP and MCP. Key features: • Inference primitive – define Markdown-based prompts with schemas, then call a language model • Temporal orchestration – run durable workflows that branch, loop, retry, wait, and listen for events • Shared ledger – records primitive inputs and outputs so runs can be recalled and explained • HTTP + MCP access – use the same brain and spec operations through an API or an MCP server • Self-hosting path – run the server from a container image, with hosted operation also described It’s source-available under the Elastic License 2.0. Link in the reply 👇
5
4
12
945
“The first principle is that you must not fool yourself and you are the easiest person to fool.” — Richard Feynman
9
52
3
394
9,390
This quote from Charles T. Munger I will place into my T) Education and Creativity Topic. To study more quote-posters illustrating this phase of the Creative Process go to: nitter.cf/search?q=%22Ed
4
3
3
83
Nathan retweeted
AI still hasn’t cracked the categories that built the internet. a16z’s chart shows 9 of 15 major consumer categories, streaming, social, dating, gaming, travel, retail, finance, real estate, jobs with zero dedicated AI apps in the Top 100. People don’t mainly want tools that save time. They want places to spend it. The blank categories aren’t empty because AI is useless there. They’re empty because the winners so far are either features inside existing apps or still blocked by trust, liability, and distribution. Dating and gaming are the real opening. The last consumer giants sold attention, the next ones will generate it.
"Most people aren't looking to save time, they're looking for ways to spend their time." 9 of 15 consumer internet categories have zero AI products in the Top 100. These built some of the biggest companies of the last two eras: - Streaming - Social - Dating - Gaming - Travel - Retail - Finance - Real estate - Jobs More charts in our Top 100 Consumer AI Apps breakdown: a16z.news/p/top-100-consumer…
4
1
3
995
🚨 Claude is no longer just a chatbot. Connect it with Gmail, Google Drive, Notion, Slack, Figma, HubSpot, Zapier, Make & more. ⚡ Access your tools 🤖 Automate repetitive work 📊 Work with your data 🚀 Build smarter workflows Connect → Automate → Create → Scale The future of AI isn’t just smarter models. It’s AI that can actually work with your tools. 🔥 #Claude #AI #Automation #Productivity #AITools
16
56
66
1,600
🤖 UBTECH Takes Robot Storage Vertical #UBTECH has built an automated #warehouse that can hold up to 112 finished #humanoid robots in just 65 square metres of floor space. Located inside its Super Smart Factory in Liuzhou, China, the system stacks #robots vertically and connects storage with the factory’s production and shipping operations. @SourabhSKatoch @wil_bielert @HsrYueli @faustospain @Whats_AI @kaifulee @demishassabis @marek_rosa @AndrewYNg @BernardMarr @CurieuxExplorer @andresvilarino @KentPage @kimgarst @raehanbobby @michaelqtodd @antgrasso @MelDMann
6
15
569
Researchers are starting to let AI agents read their own failure traces and rewrite parts of the stack around the model. The model stays frozen. What changes is the harness, the skills, the context. Self-Harness, from Shanghai AI Lab, has an agent mine its own execution traces for failure patterns and propose targeted edits to its operating rules, with reported gains of up to 60%. If your agent is failing, swapping the LLM may not be the fix. A working way to think about it is to match the symptom to the layer: → Procedural mistakes on known tasks: skills (the text instructions and heuristics) → Tool, memory or loop errors: harness (control flow, context management) → Task mastered but progress stalls: environment (workspace, what the agent observes) → Specialists failing at handoffs: orchestration (subtask routing) Read with care: this symptom-to-layer mapping is my working heuristic, not a published standard. Some researchers treat orchestration as part of the harness. The gains are paper-reported, and at least one evolved-harness result was measured on the same tasks it evolved against, so it doesn't prove generalization. The practical starting point is unglamorous: log execution traces, define a verifiable goal, and gate every edit against regressions. Which layer is your biggest bottleneck right now?
1
6
3
640
OKX just asked the SEC for permission to sell tokenized U.S. stocks. The filing is the boring part. The tell is who is asking. For three years tokenized equities lived offshore, where the pitch was that U.S. rules made them impossible. A major exchange just walked into the building and filed anyway. That only happens if the internal read on the new SEC is that the answer might be yes, or at least not an automatic no. If the filing clears, the product is not a coin. It is a brokerage account that settles like a token and trades on nights and weekends. The firms that spent 2024 building “compliant” wrappers offshore just got lapped by a company willing to put its name on a Form filing. The arbitrage was never the technology. It was the willingness to be the one who asked.
JUST IN: 🇺🇸 OKX crypto exchange files with SEC to launch tokenized US stock trading.
2
3
543
Nathan retweeted
Write the one file your whole company runs on, in 15 minutes. Send this prompt to Claude Sonnet 5.5. It interviews you, then writes your company brief: who buys, what you sell, your prices, your voice, and the lines AI never crosses. Every job in a one-person company starts by reading this file. Skip it and every chat starts from zero. It's step one in the article below. Paste the prompt, then read the rest: nitter.cf/Zephyr_hg/status/20701…
6
15
53
2,702
FIGURE JUST THREW ITS ROBOTS INTO MOLTEN STEEL. THEY WALKED TO THE EDGE AND JUMPED. THIS WAS THE RETIREMENT PLAN. 🤖 not a movie not CGI a real foundry in Finland a 75-ton furnace full of molten steel and a humanoid robot standing on a platform above it workers in heat gear filming from below then it jumps the footage you're watching is how Figure retires a robot these F.02 units just finished 11 months on BMW's assembly line 10-hour shifts 5 days a week 30,000+ cars built 99% accuracy they wore down exactly where the engineers predicted they would that wear data is already inside the next model so why melt them? disassembling by hand pulls engineers off the F.04 build storing them risks leaking proprietary hardware to competitors melting solves both problems in 20 minutes Schwarzenegger suggested it showed up in person did the thumbs down into the steel "Hasta la vista, F.02" the metal was shipped back to the US turned into limited-edition collectibles a robot that built 30,000 BMWs is now sitting on someone's shelf here's the part nobody is saying out loud the lifecycle of a humanoid robot is now shorter than a phone contract build it ship it deploy it replace it melt it sell what's left 11 months from factory floor to furnace that's not a product lifecycle that's a subscription
8
7
1
49
1,984
Your coding agents now live in Minecraft. AgentCraft is an open-source harness where a team of Claude agents splits up a project, builds in parallel in real git worktrees, and literally walks over to you in-game when they need a decision.
12
3
40
50,842
I rushed a @vangrid_io clip once and the model came back with a blank wall. Honestly, that was on me. The better approach is to do it slower. Walk the space properly, keep the subject in frame, and overlap your passes so you don’t leave gaps. If you skip one side of the room, that missing coverage becomes a blind spot. Poor lighting can hurt the result too. The requester only picks one submission, so speed alone doesn’t mean much. A slower clip with full coverage is more useful than a quick lap with half the room missing. I’d read the brief before recording. If it calls out the shelves, door, floor, or anything specific, make sure those actually show up. Android uploads are live, but you’ll need an invite code to submit.
153
4
155
2,341
Nathan retweeted
Most people use Claude Code like a chatbot. That’s the mistake. Claude Code becomes far more powerful when you structure your project around: → CLAUDE.md for project-wide instructions → .mcp.json for tool integrations → settings.json for permissions & hooks → Custom commands for repeatable workflows → Skills for task-specific expertise → Agents for specialized work I turned this complete Claude Code project structure into a simple step-by-step guide. Save this. It can completely change how you work with Claude Code. Comment: "Code" & l'll DM it to you.
13
18
31
1,279
Founder of Agora Finance, @Nick_van_Eck on Tokenized said that round the clock trading doesn't mean round the clock pricing, and ETF buyers could end up paying for that gap.⁣ ⁣ "If you can only have the APs create and redeem during traditional banking hours, you may have really large deviations from the NAV."⁣ ⁣ "You as a buyer might actually be getting a really bad price by trading on the weekend or trading at night because some of these other things behind the scenes are still relying on the legacy system."⁣ ⁣ "All those people need to get up to date to really have efficient twenty four seven markets, which is going to take a very long time."⁣ ⁣ "There's definitely gonna be periods of thin liquidity like at night or on the weekends."⁣ ⁣ Creation and redemption still run on banking hours even when the index runs on none, so the NAV gap shows up exactly when liquidity is thinnest.⁣ ⁣ 🎙 Listen to the latest episode on Tokenizedpod(dot)com
2
11
650
The Definitive Guide to DAX — Mastering the Semantic Model Expression Language for Microsoft PowerBI, Fabric, and Excel: amzn.to/4q3Aqdu [THIRD EDITION] ————— #BI #Analytics #DataAnalytics #DataAnalysis #DataScience #CDO #DataAnalyst #DataScientist
4
5
2,302
Friedberg says the Starlink network has grown so large it can now act as a radar system to detect stealth aircraft
36
289
19
2,690
111,522
Nathan retweeted
SpaceX is becoming overwhelmingly commercial. Elon Musk says: • ~90% of SpaceX revenue this year will be commercial • Government revenue in Q4 will be under 5% For context, federal spending is roughly 25% of the U.S. economy. SpaceX is not a government-dependent company. It is increasingly a commercial infrastructure giant operating at global scale.
7
23
4
80
2,961
Someone had Opus 5.5 and Sonnet 5.5 animate every level as a little Claude playing chess
9
5
112
59,806
10 GITHUB REPOS THAT TEST IF YOUR AI IS ACTUALLY GOOD. 1. DeepEval — pytest for LLMs. Test hallucinations, RAG quality, agent performance, safety and more with automated evaluations. github.com/confident-ai/deep… 2. promptfoo — break your AI before users do. Test prompts, models, RAG pipelines and AI agents with automated comparisons, red teaming and CI checks. github.com/promptfoo/promptf… 3. Ragas — test your RAG system properly. Evaluate retrieval and generation quality with metrics for faithfulness, relevance, context precision and more. github.com/explodinggradient… 4. Opik — see everything your AI is doing. Trace LLM calls and agent workflows, run evaluations, track prompts and debug AI applications from development to production. github.com/comet-ml/opik 5. Arize Phoenix — X-ray vision for AI apps. Open-source observability and evaluation for LLMs and agents, with tracing that helps you understand exactly what happened inside a run. github.com/Arize-ai/phoenix 6. Inspect AI — benchmark your AI agents. Build evaluations around datasets, solvers and scorers, then test how agents perform on real tasks. github.com/UKGovernmentBEIS/… 7. TruLens — measure your GenAI apps. Evaluate LLM applications and RAG systems with feedback functions, tracing and quality metrics. github.com/truera/trulens 8. OpenAI Evals — build your own AI tests. An open-source framework and registry for creating and running evaluations against models and AI systems. github.com/openai/evals 9. Giskard — find the dangerous stuff. Test AI systems for vulnerabilities, performance issues, hallucinations and other failure modes before deployment. github.com/Giskard-AI/giskar… 10. Evidently — monitor AI after deployment. Track data quality, model performance, drift and LLM evaluation metrics so problems don't hide in production. github.com/evidentlyai/evide… All are open-source. Bookmark this before shipping your next AI app.
9
25
73
3,210