Pinned Tweet
Grok 4.7
925
671
8,704
4,823,177
Elon Musk retweeted
Saturday was a good day for Boca Chica Beach. 985 volunteers came out for International Coastal Cleanup Day and removed 2,880 lbs of trash from our shoreline. Thank you to @SpaceX, @TXAdoptABeach, and Cameron County, and to every friend, family member, and neighbor who came out to volunteer. This is what community looks like. See you at the next one!
60
145
1,626
47,133
Elon Musk retweeted
Love Trump or hate him, it’s genuinely embarrassing the White House didn’t have this before now. Decline is a choice.
142
379
7,372
133,651
Elon Musk retweeted
Remembering Roger Boisjoly, the engineer who correctly identified a fatal flaw in the Challenger shuttle design months before the disaster, but nobody gave a damn. His exact words to his wife, Darlene: "It's going to blow up" 73 seconds before it did. More rare historical photos: bit.ly/44OpIzi
41
70
393
42,315
Elon Musk retweeted
Grok TRIPLED on Terminal-Bench 4.0 in only two months (12.4% -> 38.0%), overtaking GPT-5.6 Sol
63
67
449
72,550
Elon Musk retweeted
day 1 observations for grok 4.7 ignore the reports that say “it’s terrible” and the only thing they reference is a public benchmark. the same benchmarks told us opus 5 was better that fable - they are useless also ignore the reports that compare models with 3d games - that’s not real work. it's made for attention on social media i used grok 4.7 for a whole day as my firstmate, and it has been a really solid model with visible improvements over 4.5 (i'm ignoring 4.6 because 4.5 has been working better in my experience) key differences with 4.7 - 1. it follows system prompt very, very closely i noticed firstmate showing many new behaviors that i've never seen before, such as asking me to name specific red CI checks that i'm ok with bypassing, and refuse a simple "yolo" instruction i traced it and it's indeed how i instructed it in firstmate's system prompt, but none of the other models followed it closely enough to make this behavior visible - grok 4.7 is the first to pick that up there were a few other similar examples as well. so to me this is a clear behavioral difference 2. it's very "stable" if you've used astra then you know what a "spiky" model is. it can have some genius moments but you occasionally also wonder "how could it be so dumb and doesn't get me". grok 4.7 is the opposite of that throughout the whole day so far, i'll be honest i haven't get a "wow this is absolutely genius" moment yet, but grok 4.7 has been very steady with no big surprises. its behavior feels predictable, which does help it gain trust from me quickly 3. it's a conservative model it doesn't like to take actions without asking, and would explicitly say so this is a bit of a double edged sword, because it means i sometimes have to state the obvious "yes i do want that", but in hindsight a lot of those cases are indeed a bit ambiguous and i may not have preferred the model to just move forward without my confirmation 4. it's a bit slower and costs more than 4.5, visibly turns are taking a bit longer and my quota is draining at a visibly faster pace. i haven't quantified exactly where this is coming from yet so overall, i think it's showing some clearly different traits, and i mostly like the changes. i'm going to keep it as my primary firstmate and observe more if you've been using it, what qualitative insights have you gathered from real usage so far?
77
70
644
101,981
Elon Musk retweeted
Been testing Grok 4.7 for the past week or more. It’s been a great improvement over 4.6. It worked for over 70 hours straight on a goal. Much better attention to detail. The 500k context window really makes a difference imo. 4.7 much better at things like skill selection; workflows. Found it particularly good with pstack. Didn’t have access to it in Grokbot, but I really wanted to test it there. Together it’s a great daily driver.
74
92
917
284,452
True
In just the last 90 days: 1. Grok 4.3 — barely top 10. “xAI is dead beyond compute leases.” 2. Grok 4.5 — massive comeback. “Maybe a chance. But Grok will never catch up to the frontier.” 3. Grok 4.6 — they’re frontier. “Still not top 3.” 4. Grok 4.7 — now top 3 in frontier coding, while being much cheaper & faster. 5. Grok 4.8 next month..... The model machine is just starting up. Grok will be the workhorse of the upcoming agentic era.
604
550
4,595
1,429,957
Elon Musk retweeted
Replying to @elonmusk
I'm using it right now to build a game engine in C. It's really good.
48
56
898
320,473
Elon Musk retweeted
Grok 4.7 just released and it's an EXCELLENT model It was trained FOR Grok Bot Meaning this is a fully agentic model trained to do your knowledge work better than you can In this video I show you how to use Grok 4.7 and a Grok Bot workflow that will 10x your productivity:
115
150
1,449
317,072
Cool
This was unexpected. Grok 4.7 scored 100% on my music error detection test, the same as GPT-6 Astra.
468
482
3,568
1,344,076
Grok 4.7
Grok 4.7 is behind only Anthropic models on AA-Briefcase, ranking just behind Opus 5 at ~50% of its Cost per Task Grok 4.7’s improvements over Grok 4.6 are clear in AA-Briefcase-Lite, our public due diligence scenario where models are tasked with building market models and target assessment decks. Grok 4.7 gains significantly in Analytical Quality Elo (1698 → 1994) with a slight regression in Presentation Elo (1531 → 1499). API cost to produce example decks: Grok 4.7 (xhigh) ~$8 vs. Grok 4.6 (xhigh) ~$4.40
509
491
3,593
1,644,242
Elon Musk retweeted
Grok 4.7 xHigh is now at the top with just ONE point away from Claude Fable 5.1 Max on Artificial Analysis’ AA-Briefcase benchmark Claude Fable 5.1 Max — 59% Grok 4.7 xHigh — 58% Just a 1-point difference at the very top And Grok is ahead of GPT-6 Astra, GPT-5.6, Gemini, Kimi, GLM and nearly every other frontier model on the benchmark
128
195
1,455
307,666
Elon Musk retweeted
and here’s the grok 4.7 model card. a few jumps vs 4.6 that stood out: - Terminal-Bench: 20.3% → 38.0% - SWE-Marathon: 31.9% → 46.0% - HealthBench Pro: 48.5% → 56.7% - Legal Agent: 15.8% → 19.6% - EEBench: 60.0% → 66.0% for the same price as 4.6! media.x.ai/v1/website/4p7car…
101
158
1,522
384,851
Elon Musk retweeted
“I’ve used FSD for like 90%+ of my miles since Saturday morning. Literally feels like magic. Has handled issues on the road better than I would have at times. Just remarkable technology” A demo drive might change your life Tesla.com/drive
I picked up my new Tesla model Y on Saturday. The whole process of picking it up took like 5 minutes. I never want to deal with a car salesman ever again. I immediately had the FSD drive me home for 45 minutes. I’ve used FSD for like 90%+ of my miles since Saturday morning. Literally feels like magic. Has handled issues on the road better than I would have at times. It saw a fox potentially running across the road last night before I saw it. Just remarkable technology. Thank you to everyone who replied to my post a few weeks back saying I’d be an idiot not to get a Tesla. It’s only been 2.5 days and idk that I can ever go back to a regular car as my daily driver?
278
419
3,594
738,717
Elon Musk retweeted
Wow Grok 4.7 mogging out there. @SpaceXAI has been cooking hard.
61
117
1,091
310,845
Elon Musk retweeted
BREAKING: Grok 4.7 dominates the Legal Agent Benchmark. It scored 19.6% on realistic, long-term legal tasks, beating every other model tested. That is nearly 3× Fable 5.1. Another major win for Grok 4.7 and SpaceXAI.
113
197
1,313
294,956
Elon Musk retweeted
Grok 4.7 made this in blender. I didn't tell it what to make but only that it should be something it can accomplish in 10 minutes or so. Honestly not bad at all. Not bad at all. Did they train Grok on blender?
109
156
1,420
332,681
Elon Musk retweeted
Compare Grok 4.7 (first) and 4.6 (second) building an open world city game.
126
248
3,166
482,718
Elon Musk retweeted
Electrical engineering score is a pretty big deal. SpaceX has the engineering data that software companies don't.
Grok 4.7 is here. It's a notable improvement over Grok 4.6 at the same price and speed.
59
129
1,416
381,136