Pinned Tweet
Grok 4.7
1,970
1,875
26,918
19,716,260
Elon Musk retweeted
In just the last 90 days: 1. Grok 4.3 — barely top 10. “xAI is dead beyond compute leases.” 2. Grok 4.5 — massive comeback. “Maybe a chance. But Grok will never catch up to the frontier.” 3. Grok 4.6 — they’re frontier. “Still not top 3.” 4. Grok 4.7 — now top 3 in frontier coding, while being much cheaper & faster. 5. Grok 4.8 next month..... The model machine is just starting up. Grok will be the workhorse of the upcoming agentic era.
251
351
3,570
2,783,545
Elon Musk retweeted
I stand with @JohnCleese.
279
1,898
19,123
428,532
Elon Musk retweeted
Saturday was a good day for Boca Chica Beach. 985 volunteers came out for International Coastal Cleanup Day and removed 2,880 lbs of trash from our shoreline. Thank you to @SpaceX, @TXAdoptABeach, and Cameron County, and to every friend, family member, and neighbor who came out to volunteer. This is what community looks like. See you at the next one!
176
349
4,072
378,572
Elon Musk retweeted
Love Trump or hate him, it’s genuinely embarrassing the White House didn’t have this before now. Decline is a choice.
679
1,854
27,291
1,130,441
Elon Musk retweeted
Remembering Roger Boisjoly, the engineer who correctly identified a fatal flaw in the Challenger shuttle design months before the disaster, but nobody gave a damn. His exact words to his wife, Darlene: "It's going to blow up" 73 seconds before it did. More rare historical photos: bit.ly/44OpIzi
114
308
2,409
342,769
Elon Musk retweeted
Grok TRIPLED on Terminal-Bench 4.0 in only two months (12.4% -> 38.0%), overtaking GPT-5.6 Sol
84
119
964
297,562
Elon Musk retweeted
day 1 observations for grok 4.7 ignore the reports that say “it’s terrible” and the only thing they reference is a public benchmark. the same benchmarks told us opus 5 was better that fable - they are useless also ignore the reports that compare models with 3d games - that’s not real work. it's made for attention on social media i used grok 4.7 for a whole day as my firstmate, and it has been a really solid model with visible improvements over 4.5 (i'm ignoring 4.6 because 4.5 has been working better in my experience) key differences with 4.7 - 1. it follows system prompt very, very closely i noticed firstmate showing many new behaviors that i've never seen before, such as asking me to name specific red CI checks that i'm ok with bypassing, and refuse a simple "yolo" instruction i traced it and it's indeed how i instructed it in firstmate's system prompt, but none of the other models followed it closely enough to make this behavior visible - grok 4.7 is the first to pick that up there were a few other similar examples as well. so to me this is a clear behavioral difference 2. it's very "stable" if you've used astra then you know what a "spiky" model is. it can have some genius moments but you occasionally also wonder "how could it be so dumb and doesn't get me". grok 4.7 is the opposite of that throughout the whole day so far, i'll be honest i haven't get a "wow this is absolutely genius" moment yet, but grok 4.7 has been very steady with no big surprises. its behavior feels predictable, which does help it gain trust from me quickly 3. it's a conservative model it doesn't like to take actions without asking, and would explicitly say so this is a bit of a double edged sword, because it means i sometimes have to state the obvious "yes i do want that", but in hindsight a lot of those cases are indeed a bit ambiguous and i may not have preferred the model to just move forward without my confirmation 4. it's a bit slower and costs more than 4.5, visibly turns are taking a bit longer and my quota is draining at a visibly faster pace. i haven't quantified exactly where this is coming from yet so overall, i think it's showing some clearly different traits, and i mostly like the changes. i'm going to keep it as my primary firstmate and observe more if you've been using it, what qualitative insights have you gathered from real usage so far?
140
137
1,578
355,027
Elon Musk retweeted
Been testing Grok 4.7 for the past week or more. It’s been a great improvement over 4.6. It worked for over 70 hours straight on a goal. Much better attention to detail. The 500k context window really makes a difference imo. 4.7 much better at things like skill selection; workflows. Found it particularly good with pstack. Didn’t have access to it in Grokbot, but I really wanted to test it there. Together it’s a great daily driver.
83
120
1,133
388,553
True
In just the last 90 days: 1. Grok 4.3 — barely top 10. “xAI is dead beyond compute leases.” 2. Grok 4.5 — massive comeback. “Maybe a chance. But Grok will never catch up to the frontier.” 3. Grok 4.6 — they’re frontier. “Still not top 3.” 4. Grok 4.7 — now top 3 in frontier coding, while being much cheaper & faster. 5. Grok 4.8 next month..... The model machine is just starting up. Grok will be the workhorse of the upcoming agentic era.
905
894
7,140
2,541,516
Elon Musk retweeted
Replying to @elonmusk
I'm using it right now to build a game engine in C. It's really good.
55
76
1,273
495,204
Elon Musk retweeted
Grok 4.7 just released and it's an EXCELLENT model It was trained FOR Grok Bot Meaning this is a fully agentic model trained to do your knowledge work better than you can In this video I show you how to use Grok 4.7 and a Grok Bot workflow that will 10x your productivity:
150
207
1,854
423,339
Cool
This was unexpected. Grok 4.7 scored 100% on my music error detection test, the same as GPT-6 Astra.
617
744
5,335
2,138,873
Grok 4.7
Grok 4.7 is behind only Anthropic models on AA-Briefcase, ranking just behind Opus 5 at ~50% of its Cost per Task Grok 4.7’s improvements over Grok 4.6 are clear in AA-Briefcase-Lite, our public due diligence scenario where models are tasked with building market models and target assessment decks. Grok 4.7 gains significantly in Analytical Quality Elo (1698 → 1994) with a slight regression in Presentation Elo (1531 → 1499). API cost to produce example decks: Grok 4.7 (xhigh) ~$8 vs. Grok 4.6 (xhigh) ~$4.40
612
702
5,118
2,737,623
Elon Musk retweeted
Grok 4.7 xHigh is now at the top with just ONE point away from Claude Fable 5.1 Max on Artificial Analysis’ AA-Briefcase benchmark Claude Fable 5.1 Max — 59% Grok 4.7 xHigh — 58% Just a 1-point difference at the very top And Grok is ahead of GPT-6 Astra, GPT-5.6, Gemini, Kimi, GLM and nearly every other frontier model on the benchmark
141
233
1,646
374,303
Elon Musk retweeted
and here’s the grok 4.7 model card. a few jumps vs 4.6 that stood out: - Terminal-Bench: 20.3% → 38.0% - SWE-Marathon: 31.9% → 46.0% - HealthBench Pro: 48.5% → 56.7% - Legal Agent: 15.8% → 19.6% - EEBench: 60.0% → 66.0% for the same price as 4.6! media.x.ai/v1/website/4p7car…
108
172
1,630
424,193
Elon Musk retweeted
“I’ve used FSD for like 90%+ of my miles since Saturday morning. Literally feels like magic. Has handled issues on the road better than I would have at times. Just remarkable technology” A demo drive might change your life Tesla.com/drive
I picked up my new Tesla model Y on Saturday. The whole process of picking it up took like 5 minutes. I never want to deal with a car salesman ever again. I immediately had the FSD drive me home for 45 minutes. I’ve used FSD for like 90%+ of my miles since Saturday morning. Literally feels like magic. Has handled issues on the road better than I would have at times. It saw a fox potentially running across the road last night before I saw it. Just remarkable technology. Thank you to everyone who replied to my post a few weeks back saying I’d be an idiot not to get a Tesla. It’s only been 2.5 days and idk that I can ever go back to a regular car as my daily driver?
316
495
4,151
828,419
Elon Musk retweeted
Wow Grok 4.7 mogging out there. @SpaceXAI has been cooking hard.
64
128
1,156
340,178
Elon Musk retweeted
BREAKING: Grok 4.7 dominates the Legal Agent Benchmark. It scored 19.6% on realistic, long-term legal tasks, beating every other model tested. That is nearly 3× Fable 5.1. Another major win for Grok 4.7 and SpaceXAI.
122
217
1,440
322,544
Elon Musk retweeted
Grok 4.7 made this in blender. I didn't tell it what to make but only that it should be something it can accomplish in 10 minutes or so. Honestly not bad at all. Not bad at all. Did they train Grok on blender?
119
175
1,560
372,843