blessdyb retweeted
Cloudflare の Public DNS 1.1.1.1 に脆弱性があることを示す PoC
blog.alice.poc.debiru.net/
1.1.1.1 を使っている人は FAKE サーバーに辿り着くはずです。
(参考)
検証用URL blog.alice.poc.debiru.net/
FAKEサーバー debiru.static.jp/
TRUEサーバー blog.bob.poc.debiru.net/
blessdyb retweeted
Your job may be surviving on inertia. Your resentment won’t extend the runway. What do you actually want to do with the intelligence now at your disposal?
See your apps' network activity. Understand your AI agents.
Flowlight is an open-source macOS network monitor for apps and AI agents: destinations, protocols, bytes, agent-launched tools/MCP, alerts, rules, and optional HTTPS tool-call inspection.
github.com/xinbetween/flowli…
blessdyb retweeted
第三期 小路软件工程系统架构分享
Cloudflare 如何省下了 100TB RAM?!
答案是给一致性哈希环“减肥”。
Pingora Backend Router 原本给每台服务器生成大量 hash,用来降低负载不均。但 Cloudflare 发现:hash 数量增加到一定程度后,收益已经非常小,RAM 却还在疯狂消耗。
于是他们重新推了一遍数学,发现可以把每台服务器的 hash 数量减少约 90%,而几乎不影响负载均衡精度。
更有意思的是 Rust:原本 hash: u32 + index: u32 占 8 bytes,通过紧凑存储压到 6 bytes,又省掉 25%。
最终,全球直接回收超过 100TB RAM。
这就是我喜欢的大规模基础设施优化:没有新模型,没有新硬件,只是重新审视一个“用了很多年”的设计,然后用数学 + Rust 把浪费挤出来。
我把一致性哈希在这里的用法画在第一张图里,不了解的可以去看看。顺便这个consistent hashing就是aws dynamo db论文里提出的那个。可以去看那篇论文
[原文:Cloudflare — Saving another 100TB of RAM with math (and Rust)](blog.cloudflare.com/saving-1…)
第二期 小路软件工程系统架构分享。
OpenAI 的 Habitat,最开始只是 2023 DevDay 的一个 Python 客户端库,后来一路长成了支撑 ChatGPT、API、Codex 的在线存储平台。
现在每周服务 10 亿+ 用户,覆盖近 40 个区域,存储超过 500 PB,在线请求峰值超过 7000 万 QPS。
它采用简单、稳定的 NoSQL API,复杂查询则通过 CDC 送到 Rockset。底层用了 Cosmos DB、Nanobase、Valkey 和 Blob。
有意思的是性能优化:2026 Q2,OpenAI 用 Codex + GPT-5.5 把核心服务从 Python 重写成 Rust,已经承载约 95% 生产流量,CPU 约降 6 倍,内存约降 15 倍。
前面再用 Envoy 做连接池和 HTTP/2 复用。
openai.com/index/scaling-sto…
blessdyb retweeted
周末适合沉下心来系统性的学习基础知识
CMU 秋季最新课程 11-768: AI Agents
主讲:@dan_fried @gneubig
课程:cmu-agents.com/
视频:youtube.com/watch?v=UwfjzyLn…
blessdyb retweeted
1B requests a day in Weixin. Now open source.
WeMM-Embedding is built for real traffic and verified in real use. It reads text, image and video in the order they show up.
The 9B ranks #1 on MMEB-v2 (80.6); the 2B keeps 98.7% at just 256 dims.
Truncate to 64 dims for fast recall or 2048 for fine ranking, no retraining.
Try it: github.com/Tencent/WeMM-Embe…
blessdyb retweeted
Follow along here if you are interested 😊 We will keep this site updated with assignments and lectures!
cs329z.stanford.edu/
blessdyb retweeted
We just raised $5.7M for @PolarBrowser, the AI browser that beats Anthropic and OpenAI on every major web agent benchmark.
- 4,500,000+ actions taken for users, automating sales, recruiting, and ops
- One company cancelled Clay and saves 25+ hrs/wk per person
- Team is from MIT, YC, Prod, Citadel, Jane Street, Perplexity, Modal, and Apple
100 hours of work. From 30 seconds of typing.
Download the world's most powerful AI browser: polarbrowser.com
blessdyb retweeted
Vector Database by hand ✍️ ~ 10 steps walkthrough below
Vector databases are the backbone of Retrieval Augmented Generation (RAG).
How do they actually work?
Goal: index three sentences, then answer a query by finding the nearest one, filling in every cell yourself.
= 1. Given =
A dataset of three sentences, three words each. In practice it is millions of them.
= 2. Word embeddings =
Let us look up each word in an embedding table. Here the vocabulary is 22 words; in practice it is tens of thousands, and the vectors have thousands of dimensions rather than four.
= 3. Encoding =
We feed the sequence to an encoder, one linear layer and a ReLU, and get one feature vector per word. In practice the encoder is a transformer.
= 4. Mean pooling =
Let us average across the columns. Three word vectors collapse into one, which is what people mean by a text embedding or a sentence embedding.
= 5. Indexing =
We multiply by a projection matrix and the four dimensions become two. It is doing the job of a hash: a short representation that is faster to compare, and it is what gets saved in the vector storage.
= 6. Process "who are you" =
Let us repeat steps 2 to 5 on the second sentence.
= 7. Process "who am I" =
We do it a third time. The database is now indexed.
= 8. Query "am I you" =
Let us push the query through the very same pipeline: lookup, encoder, mean pooling, projection, and it lands as a 2D vector in the same space.
= 9. Dot products =
We transpose the query and multiply, which takes the dot product against every stored vector at once. The dot product is the estimate of similarity.
= 10. Nearest neighbour =
Let us scan for the largest: 60/9 beats 44/9 and 40/9, so the answer is "who am I". Scanning billions of vectors one at a time is what makes this the slow step in practice, which is why real databases use an approximate nearest neighbour index like HNSW.
The outputs:
Stored index vectors = [5/3, 2/3], [5/3, 0], [7/3, 2/3]
Query vector = [8/3, 2/3]
Dot products = 44/9, 40/9, 60/9
Nearest neighbour = "who am I"
The takeaway: a vector database is an embedding pipeline, a projection, and a dot product. Every step here is arithmetic you can do in pen, which is worth remembering when the word "database" makes it sound like something else.
💾 Save this post!
blessdyb retweeted
defending-code-reference-harness
Skills for threat modeling, scanning, triage, patching, plus an autonomous scanning harness you can customize.
github.com/anthropics/defend…
LLMs have a limited context window. When conversations grow too long this affects output quality, performance and cost.
New blog post from Earendil engineer @vegardstikbakke on how compaction addresses this and how we’ve implemented it in Pi.
Read the full post below
blessdyb retweeted
#inference, #llm Everyone can call an inference API. Few can explain why vllm serve beats their PyTorch loop.
So I wrote the course: 22 chapters from a while-loop to PagedAttention, continuous batching and speculative decoding.
No GPU, no PyTorch. llminference.xinbetween.com
blessdyb retweeted
"How to Write an Effective Software Design Document".
Writing a design doc forces you to think through important decisions before you waste time on the wrong implementation or paint yourself into a corner.
by Michael Lynch
refactoringenglish.com/excer…
Good materials to practice your prompt skills.
gandalf.lakera.ai/word-black…
red.giskard.ai/register
promptairlines.com/