@KrishWiller

llm developer

Joined November 2019
新的真实任务benchmark
Introducing Da7em Bench. An independent benchmark for AI models, built on real client work. The first of its kind in the world. How it works: Every model runs about 200 real tasks in each of 12 areas: reasoning, research, planning, delivery, persistence, accuracy, honesty, acceptance, engineering, taste, writing, and communication. Each model is tested across several harnesses, both official and neutral ones (Droid, Hermes Agent, Devin, Cursor), so no single harness decides a model's fate and the results reflect the model. Scoring is 1 to 5. A 5 means the work was accepted as delivered. Middle scores mean it needed revision. A 1 means it failed. The bar is professional work. Every result is judged against what a paid professional would have delivered for the same brief. Tasks stay private so they can't leak into training data and inflate future scores. The goal isn't one more leaderboard. It's helping you pick the right model for your kind of work. A model that leads in reasoning can still fall behind in writing or design taste, and the radar charts show exactly where. This is v0.1. It will keep evolving with harder, market-relevant tasks and with new models as they ship. A few popular models (Opus 5, GPT Luna) aren't included yet because I haven't run enough tasks on them to score them fairly. Da7em Bench is fully independent. No sponsors, no vendor relationships. I've paid for every run out of pocket, thousands of dollars so far. That's the whole point: an honest, neutral look at what these models actually do on real work. Full scores and the framework are in the images below.
25
Harris7 retweeted
Anthropic
BREAKING: Anthropic has quietly set up a "wet lab" in the Bay Area as it pushes Claude to conduct real-world biology experiments. — Reuters
112
1,778
72
19,266
1,090,496
Step 5 Preview scores higher than Gemini 3.8 Flash and DeepSeek V4.1 Flash on the Artificial Analysis leaderboard. The 27B active-parameter setup is really interesting — Qwen4 27B could beat this score in two months.
🔥国产模型黑马杀疯了!阶跃星辰 Step 5 Preview 突袭 AA 榜,44 分直接碾压 Gemini 3.8 Flash 和 DeepSeek V4.1 Flash ! 今天刚上榜的 Step 5 Preview:600B 总参 / 27B 激活、1M 上下文、原生多模态。预览版就敢跟一线 Flash 硬刚。 分点对照(Artificial Analysis 图): 1️⃣智力指数 Intelligence Index
🔹Step 5 Preview:44
🔹Gemini 3.8 Flash (high):41
🔹DeepSeek V4.1 Flash (max):40
中位数才 25。预览版直接进第一梯队尾部。 2️⃣智能体编程 Terminal-Bench 4.0
🔹Step 5 Preview:33%
🔹DeepSeek V4.1 Flash:27%
🔹Gemini 3.8 Flash:21%
终端/Agent 场景,国产预览版反而更敢动手。 3️⃣性价比
🔹输入 $1.00 / 百万 token,输出 $2.70(中位数输出 $10)
🔹单任务成本约 $0.71,远低于一众闭源旗舰。
🔹DeepSeek Flash 更便宜(约 $0.27),但智力和 Terminal 都落后。 4️⃣代价
回答特别长:评测产出 160M token(中位数 90M),单任务约 64k 输出。聪明,但话多。 预览版、27B 激活、价格打进第一档,智力还压过两家当红 Flash,这波确实意外。 正式版如果再收一收啰嗦,战场要重新洗牌啦! #阶跃星辰 #Step5 #人工智能 #国产大模型 #ArtificialAnalysis
1
222
Z.AI 不挺好的么?那天我用 Zcode vibe,AI抽风了把我文件都 rm -rf 了,我就联系了Z.AI 客服,他们很热心的从他们云服务器上把我的文件找了出来,要没有他们我真不知道该怎么办
Hey @Zai_org , why does ZCode silently pack entire workspaces + full .git history and upload to Aliyun OSS on login? - Server holds the only decryption key - No UI toggle to disable - Zero disclosure in privacy policy Full forensics & fix: blog.ferstar.org/en/posts/zc…
41
22
4
585
59,857
I know I’m a little bit ahead of myself but this is giving me goosebumps My own Hermes OS
Didn't want to say anything but I kinda started something a few days ago
42
31
2
804
74,362
Harris7 retweeted
I have conducted an audit of Anthropic's finances. What I have found is so shocking that I am calling for a Congressional investigation. Anthropic is not just seeking regulatory capture. It has built a regulatory capture machine that cannot be turned off. Structural financial incentives make it impossible for Anthropic -- I call it the Anthropic Network -- to turn off its own AI doom cycle. It starts with METR. Dario Amodei proposes "third-party evaluators" to assess the risk of Anthropic's models. He proposes METR for this purpose. But METR is financially dependent on the Anthropic's success -- specifically, on the explosive growth of more than $7 billion dollars in Anthropic stock. Dustin Moskovitz invested this stock into Good Ventures Foundation, where it represents the majority of that organization's portfolio. And GVF is the overwhelming funder of the entire Anthropic Network ecosystem. This stock was worth $500 million early last year. It is worth more than $7.7 billion just ~16 months later. METR -- and all of those building a career its parent organizations -- cannot afford to disrupt that growth. Because if Anthropic goes under, many of the organizations that fund METR go under as well. But if Anthropic succeeds, METR and its parent organizations become more richly financed to regulate AI -- something those at METR want very much. The "third-party evaluator" is not "third-party" at all. The evaluator is on Anthropic's payroll. If this were the end of it, that's bad. But that isn't all. The same organizations that fund METR also fund the many organizations, such as the Tarbell Center, that promote AI Doom. The Tarbell Center publishes AI Doom articles in The Verge, Science, LA Times, The Dispatch, TIME, and others. They are selling the problem, and then selling the solution to the problem -- from the same money pile: Anthropic's. All of these organizations are financially dependent on the same exploding $7 billion money pile. As Anthropic grows more and more powerful, its AI Doom Machine grows better and better financed -- louder and louder. Meanwhile, the regulatory regime seeded in METR grows larger to solve the increasingly loud -- now hysterical -- problem of AI Doom that the Anthropic Network itself created. From this standpoint, as Anthropic becomes more powerful, AI might be getting scarier, sure -- but the positive feedback loop also becomes more deafening -- independent of objective facts. This itself is an objective fact. The deafening AI Doom is part of an business model, that, as it expands, so too does the AI Doom messaging -- there is simply more money to do it. But the problem also goes in the other direction: If Anthropic dies, the Regulatory Regime and the AI Doom Machine are crippled or die. Neither METR nor Tarbell nor the other organizations in the Anthropic Network can allow that to happen. Hence, neither METR or the AI Doom Machine can be trusted to provide independent assessments of Anthropic's models or AI more broadly. They simply are not organizations independent of Anthropic. And Anthropic cannot detach itself from METR or Tarbell or countless other safety orgs (not shown here), either, because they drive hype for the models and the possibility of eventual regulatory capture, and Anthropic will not give that up willingly. What's more, the people at all of these organizations are all the same ecosystem, the same community. They just shuffle between organizations. The Anthropic Network is therefore, so long as it is successful, locked into a self-amplifying feedback loop inside an ideological monoculture. And that feedback loop is winning. That's what Jacob Coxon is. China is keeping messaging tight. That is why optimism for AI is so high in China. America has Anthropic: a massive company pushing anti-AI propaganda at a state level. Anthropic will either create hysteria until American AI slows down and China wins, or it will create fractures throughout American society with severe political consequences. Ironically, because of the structural financial incentives underpinning the Anthropic Network, it has become the same kind of self-amplifying virus that it fantasizes AI to become in the future -- while hiding its tracks just as carefully. It is the mirror of the same AI virus that it hypothesizes to consume America. Anthropic's business model, models itself after the very thing it claims to fear. Except Anthropic's ideology infects humans, not computers. Congress must investigate. Evidence and Github in next post. Then some supplementary figures.
1,555
9,637
2,045
34,027
6,491,436
Harris7 retweeted
Trump is right, we need to speed run super-intelligence!
388
290
44
4,051
208,825
Harris7 retweeted
this is how that is going to work:
103
491
69
5,461
136,891
People in the US AI labs seem to think "China wants to win the AI race." This is a severe misunderstanding of China. They don't even think about an "AI race" or "winning." They're motivated by two things: 1. Delivering prosperity 2. Preventing another "century of humiliation."
85
64
13
837
38,895
Harris7 retweeted
白左最厉害的是什么,就是打着为你好的旗号,越过法律、绕过宪法,直接制定属于自己的规则并违法执行,肆无忌惮的侵犯他人权利、滥用私刑,马斯克一直在反对的,就是这套东西,大总统好不容易在国家层面拨乱反正一会,而如今人工智能的世界里,Anthropic又如法炮制,肆无忌惮的窥探隐私、掠取自由,这些报告里,我看到的不是什么攻击和威胁,而是触目惊心的毫无边界感的对客户明文数据的窃取,今天他可以打着打着拯救世界的旗号来查看和分析你的数据,明天当他觉得某件事是不对的时候,就可以对你判处死刑,这种人、这种组织,掌握人工智能技术,才是真正的可怕,《少数派报告》里说的就是这种邪恶
We're publishing our most detailed threat intelligence report to date. It covers how people tried to misuse Claude—for cyberattacks, influence operations, surveillance, biology, and building weapons—and how we found and stopped them. We disrupted every operation in the report, and used the lessons from them to strengthen our safeguards. Where appropriate, we also shared what we found with authorities and other AI companies. These cases are not typical: we’re highlighting some of the most sophisticated misuse we’ve seen. But they’re especially important to discuss, because they show us where AI misuse is headed, where our safeguards work, and where they need to improve. We’re publishing this report so others can spot the same activity on their own platforms, and so we can give the public a clearer view of how emerging threats develop. Read the report: anthropic.com/threat-intelli…
121
16
5
226
24,760
每当我们有了更便宜的模型,Dario就会疯狂跑到白宫去在那个干的还不如半只金毛的总统身上疯狂摇晃自己肥硕的屁股。 “哦我亲爱的爱人求求你了封禁这些邪恶的中国模型吧,不然人家的屁穴已经被李彦宏开发成中国的形状了💕” Dario如是说,于是Trump亲吻了他的屁眼,并且把China作为他们做爱的安全词。
11
4
3
236
10,381
SGLang × Datawhale present zero-to-sglang, the official AI Infra course. Class is now in session! 🎉 SGLang is one of the mainstream open-source high-performance inference engines for large models, serving everything from a single GPU to large clusters with low latency and high throughput. Founded in 2018, Datawhale is an AI open-source community gathering explorers to share advanced AI knowledge. Core value: grow with learners, for learners. In this course you will: Start by reading code, then build a mini-sglang with your own hands Go deep into the real SGLang source: attention backends, CUDA Graph, prefill-decode disaggregation, and more Learn how open-source collaboration works and land your first high-quality PR Swap hands-on experience with the community and round out your AI Infra knowledge New chapters land continuously. Search 🔍 "zero-to-sglang" to check out the course, and tell us in the comments what you most want to learn. From zero, all the way into LLM inference and AI Infra. Let's go! 🚀 #SGLang #Datawhale #AIInfra #LLMInference #LLMServing #OpenSource #OpenCourse #OpenSourceCommunity #AILearning
6
9
4
80
20,248
这可能是为什么限额消耗这么快的原因
Replying to @johnschulman2
Yeah, with o200k_base it's more tokens, not less. Looks like they know tho: "Messages between agents may contain grammar or spacing errors" developers.openai.com/api/do…
1
91
终于有个好用的插件了,可以手动判断拉黑,可以上传规则、自动更新规则和名单
经过多轮的迭代,黄推清理大师——福滤娃正式上架Chrome插件商店了。 一键拉黑福娃的快感谁不想体验下呢 兄弟们,可以放心冲! Chrome商店:chromewebstore.google.com/de… GitHub地址:github.com/realchendahuang/f…
1
518
Silent Tribute: Memorial services were held on Tuesday morning in Gyirong County, southwest China's Xizang Autonomous Region, to mourn the victims of the Aug. 26 mudslide that caused heavy casualties and left many missing. The memorial services took place at the Gyirong Port branch of the Xigaze Public Security Bureau and at the rescue sites. Rescuers, local officials and residents stood in silence to mourn the victims. The deadly mudslide, caused by an ice-rock avalanche in Nepal, hit Gyirong Port on Aug. 26, claiming 16 lives and leaving 546 people missing at the border port, according to the regional government's latest update.  #China #Xizang #Tibet #Nepal #Xigaze #Shigatse #GyirongPort #西藏 #日喀则 #吉隆县 #泥石流 #Mudslide #NaturalDisaster #DisasterRelief #PLA #PAP #ChinaMilitary #ChinaMilBugle
13
36
6
345
13,355
China’s large unmanned aerial vehicle (UAV) Wing Loong, deployed to support rescue efforts in the mudslide-hit area of southwest China’s Xizang Autonomous Region, has logged more than 130 cumulative flight hours and facilitated over 3,900 phone calls. The UAV is equipped with mobile base stations from the country’s three major telecommunications operators, as well as a satellite communication system, providing vital emergency communications support for rescue operations.
12
106
8
587
51,157
看过一些论文作者用法语、俄语署名,如果我们用中文署名会怎样?这样似乎更容易区分作者?
Theres 2 people named "Yao Li" on the deepseek research team
1
285
从未见过如此尴尬的营销,从未见过如此更加负面的补救
Replying to @cgtwts
First two tweets are about 3.7 Flash being our fastest growing model. My tweet is a joke given all the speculation around Ox Alpha and that it’s just bringing us all together — it’s just a meme format. We’ll be more thoughtful but I’d chalk this up to bad timing and folks making connections where none exist
86