Elon Musk retweeted
Grok 4.6 just took #1 on CursorBench. Not only the highest score. It did it at a fraction of the cost of the models sitting right behind it. That combination is the real signal. Top-tier results are one thing. Top-tier results that stay cheap enough to run for long agentic coding sessions are something else. I see this as the practical edge that matters for real work. Benchmarks are useful. Sustained performance at low cost is what actually gets used.
Grok 4.6 on extra high thinking mode now achieves #1 score on CursorBench!
82
185
1,224
755,421
It’s that easy to use Grok @Bot
I launched Grok Bot for 24 hours, gave it 3 jobs, and stopped treating AI like a chatbot. I had it scan live X + web data for emerging AI products, research what people were actually saying about them, filter the obvious engagement bait, compare sources and turn the strongest signals into a structured brief. At the same time, I could have separate research paths looking at competitors, code or market opportunities instead of forcing everything through one conversation. That’s the difference I underestimated. I don’t want to sit inside a chat writing another prompt every 5 minutes. I want to define the objective, give the agent access to the right tools and data, let it execute, then step back in when judgment or approval is actually needed. The old workflow was prompt → answer → copy → repeat. The workflow I’m moving toward is assign → execute → verify. Once AI stops waiting for your next message, it stops feeling like software you use and starts feeling like work you delegated.
1,428
1,498
8,169
4,023,641