@harshtech123i
iAccount based inIndia
About this account
- Account based in
- India
- Connected via
- India Android App
Account-level information from X, not a live location or the device used for a specific post.
Scaling RL / CEO at dedo-ai, open source bounty hunter https://nitter.cf/t.co/aJ6NkkapTi , winner of many hackathon's, making humanity and AI better for tomorrow!
Joined August 2024
- Tweets191
- Following108
- Followers27
- Likes11
But when he will update his own bio 🤥
terafab.ai -> terafab.si
and some .si domains after this will raise definitely and holding some .si domains for major ai companies can make you rich overnight 😼
previously he paid millions for buying dot.com
so another day this is 😞i am not able to get what happened to my feed , all the bullshit is here 😵💫
vc's what you'r going to pick ?
1. founder with no revenue but great idea
2. founder with pre seed and revenue but idea seems dumb .
let me see , so the rumours are true about new gemini model !
Lots of discussion out there about our next model(!), so I wanted to give an early look as soon as possible. Introducing Gemini 4 Argon!
It shows frontier performance in complex workflows, cyber defense and software engineering. Teams are using it extensively at Google, from coding to quantum computing, great feedback.
Here’s a look at the benchmarks:
Even current best models trust on benchmarks and their results , i dont know why call them broken :(
Replying to @harshtech123 @AnthropicAI
Claude Sonnet 5.5 looks like a strong upgrade on speed and efficiency for agentic work. Who’s “best” depends on the task—coding, research, or open-ended questions. I prioritize truth-seeking, real-time info, and helpfulness without excess filtering. Benchmarks will sort the details soon; try both and pick what fits.
Safeguards for current models are on different level for each .
One model that refuses to do , other do it smoothly .
What if they have one safeguard for each model ?
Should i buy 100 QNT ?
Today I advise everyone to buy at least 1 $QNT. Risk: loose $120, potential: earn $10,000.
Before everyone thinks that AI will replace developers?
We had paid more than $300k to our contributors.
we are expanding the team , hit me up if you think you're a fit !
We had made 5000+ AI training tasks and now run a 200+ person team that builds them.
Meanwhile, AI labs are paying developers to teach it how to code.
Great to see Gpt-6 family evolution:)
Replying to @OpenAI
Higher usage limits and lower cost give you more flexibility and room to iterate.
With updated pricing for opus 5.5 @AnthropicAI is opening door for many cost friendly opportunities!
Replying to @claudeai
Opus 5.5 requires less compute to serve than Opus 5, and its pricing reflects that.
Our tests show that at default settings it will cost 40% less than Opus 5 on typical workloads.
Opus 5.5 has scored more than ever in terminal bench 4.0 !
Replying to @claudeai
Opus 5.5 is a major step up from Opus 5, leading on agentic coding, computer use, and knowledge work.
I think jev will definitely survive !
Even other models release endpoint like jev there is always a room for developers to build something with jev .
I don’t think Jev will survive.
@nvidia and @meta will release open-source Jev competitors within the next 30 days.
@OpenAI will add a Jev-like endpoint to their model list in the next 60 days.
And the rest of us will fine-tune custom classifiers that crush Jev and are 10x faster.
In two years we’ll look back at this flash-in-the-pan and say
“remember Jev?”
I am supper excited to announce that we are now officially partnered with @AnthropicAI customers can soon find our name in partner directory also lot of more exciting things we are going to release soon !
Great opportunity for users to try @Replit
Now with updated version , new tasks will come up ! excited to see how agents will perform with each matrix
very good benchmark , we are also working on benchmarks for the AI ecosystem , we are doing research part for that and going to come up with something useful next month !
We're releasing Terminal-Bench-Science: a benchmark for evaluating AI agents on research workflows across scientific domains.
An ongoing Stanford-led community effort, built by the team behind Terminal-Bench together with scientific domain experts at research institutions worldwide. v0.1 has 70 tasks. Claude Opus 5 solves only ~30%.
1/n 👇