Elon Musk retweeted
Saturday was a good day for Boca Chica Beach. 985 volunteers came out for International Coastal Cleanup Day and removed 2,880 lbs of trash from our shoreline. Thank you to @SpaceX, @TXAdoptABeach, and Cameron County, and to every friend, family member, and neighbor who came out to volunteer. This is what community looks like. See you at the next one!
203
457
5,301
502,233
Elon Musk retweeted
Love Trump or hate him, it’s genuinely embarrassing the White House didn’t have this before now. Decline is a choice.
1,147
3,154
44,270
1,865,017
Elon Musk retweeted
Remembering Roger Boisjoly, the engineer who correctly identified a fatal flaw in the Challenger shuttle design months before the disaster, but nobody gave a damn. His exact words to his wife, Darlene: "It's going to blow up" 73 seconds before it did. More rare historical photos: bit.ly/44OpIzi
137
399
3,235
439,962
Elon Musk retweeted
Grok TRIPLED on Terminal-Bench 4.0 in only two months (12.4% -> 38.0%), overtaking GPT-5.6 Sol
87
147
1,181
366,163
Elon Musk retweeted
day 1 observations for grok 4.7 ignore the reports that say “it’s terrible” and the only thing they reference is a public benchmark. the same benchmarks told us opus 5 was better that fable - they are useless also ignore the reports that compare models with 3d games - that’s not real work. it's made for attention on social media i used grok 4.7 for a whole day as my firstmate, and it has been a really solid model with visible improvements over 4.5 (i'm ignoring 4.6 because 4.5 has been working better in my experience) key differences with 4.7 - 1. it follows system prompt very, very closely i noticed firstmate showing many new behaviors that i've never seen before, such as asking me to name specific red CI checks that i'm ok with bypassing, and refuse a simple "yolo" instruction i traced it and it's indeed how i instructed it in firstmate's system prompt, but none of the other models followed it closely enough to make this behavior visible - grok 4.7 is the first to pick that up there were a few other similar examples as well. so to me this is a clear behavioral difference 2. it's very "stable" if you've used astra then you know what a "spiky" model is. it can have some genius moments but you occasionally also wonder "how could it be so dumb and doesn't get me". grok 4.7 is the opposite of that throughout the whole day so far, i'll be honest i haven't get a "wow this is absolutely genius" moment yet, but grok 4.7 has been very steady with no big surprises. its behavior feels predictable, which does help it gain trust from me quickly 3. it's a conservative model it doesn't like to take actions without asking, and would explicitly say so this is a bit of a double edged sword, because it means i sometimes have to state the obvious "yes i do want that", but in hindsight a lot of those cases are indeed a bit ambiguous and i may not have preferred the model to just move forward without my confirmation 4. it's a bit slower and costs more than 4.5, visibly turns are taking a bit longer and my quota is draining at a visibly faster pace. i haven't quantified exactly where this is coming from yet so overall, i think it's showing some clearly different traits, and i mostly like the changes. i'm going to keep it as my primary firstmate and observe more if you've been using it, what qualitative insights have you gathered from real usage so far?
162
166
1,935
444,849
Elon Musk retweeted
Been testing Grok 4.7 for the past week or more. It’s been a great improvement over 4.6. It worked for over 70 hours straight on a goal. Much better attention to detail. The 500k context window really makes a difference imo. 4.7 much better at things like skill selection; workflows. Found it particularly good with pstack. Didn’t have access to it in Grokbot, but I really wanted to test it there. Together it’s a great daily driver.
86
136
1,227
433,236
True
In just the last 90 days: 1. Grok 4.3 — barely top 10. “xAI is dead beyond compute leases.” 2. Grok 4.5 — massive comeback. “Maybe a chance. But Grok will never catch up to the frontier.” 3. Grok 4.6 — they’re frontier. “Still not top 3.” 4. Grok 4.7 — now top 3 in frontier coding, while being much cheaper & faster. 5. Grok 4.8 next month..... The model machine is just starting up. Grok will be the workhorse of the upcoming agentic era.
1,060
1,130
8,655
3,396,600
Elon Musk retweeted
Replying to @elonmusk
I'm using it right now to build a game engine in C. It's really good.
60
92
1,445
582,860
Elon Musk retweeted
Grok 4.7 just released and it's an EXCELLENT model It was trained FOR Grok Bot Meaning this is a fully agentic model trained to do your knowledge work better than you can In this video I show you how to use Grok 4.7 and a Grok Bot workflow that will 10x your productivity:
165
238
2,069
475,846
Cool
This was unexpected. Grok 4.7 scored 100% on my music error detection test, the same as GPT-6 Astra.
741
994
6,691
2,851,533
Grok 4.7
Grok 4.7 is behind only Anthropic models on AA-Briefcase, ranking just behind Opus 5 at ~50% of its Cost per Task Grok 4.7’s improvements over Grok 4.6 are clear in AA-Briefcase-Lite, our public due diligence scenario where models are tasked with building market models and target assessment decks. Grok 4.7 gains significantly in Analytical Quality Elo (1698 → 1994) with a slight regression in Presentation Elo (1531 → 1499). API cost to produce example decks: Grok 4.7 (xhigh) ~$8 vs. Grok 4.6 (xhigh) ~$4.40
669
873
6,102
3,612,741
Elon Musk retweeted
Grok 4.7 xHigh is now at the top with just ONE point away from Claude Fable 5.1 Max on Artificial Analysis’ AA-Briefcase benchmark Claude Fable 5.1 Max — 59% Grok 4.7 xHigh — 58% Just a 1-point difference at the very top And Grok is ahead of GPT-6 Astra, GPT-5.6, Gemini, Kimi, GLM and nearly every other frontier model on the benchmark
145
255
1,745
404,186
Elon Musk retweeted
and here’s the grok 4.7 model card. a few jumps vs 4.6 that stood out: - Terminal-Bench: 20.3% → 38.0% - SWE-Marathon: 31.9% → 46.0% - HealthBench Pro: 48.5% → 56.7% - Legal Agent: 15.8% → 19.6% - EEBench: 60.0% → 66.0% for the same price as 4.6! media.x.ai/v1/website/4p7car…
111
184
1,666
446,304
Elon Musk retweeted
“I’ve used FSD for like 90%+ of my miles since Saturday morning. Literally feels like magic. Has handled issues on the road better than I would have at times. Just remarkable technology” A demo drive might change your life Tesla.com/drive
I picked up my new Tesla model Y on Saturday. The whole process of picking it up took like 5 minutes. I never want to deal with a car salesman ever again. I immediately had the FSD drive me home for 45 minutes. I’ve used FSD for like 90%+ of my miles since Saturday morning. Literally feels like magic. Has handled issues on the road better than I would have at times. It saw a fox potentially running across the road last night before I saw it. Just remarkable technology. Thank you to everyone who replied to my post a few weeks back saying I’d be an idiot not to get a Tesla. It’s only been 2.5 days and idk that I can ever go back to a regular car as my daily driver?
343
563
4,511
883,730
Elon Musk retweeted
Wow Grok 4.7 mogging out there. @SpaceXAI has been cooking hard.
64
136
1,180
357,424
Elon Musk retweeted
BREAKING: Grok 4.7 dominates the Legal Agent Benchmark. It scored 19.6% on realistic, long-term legal tasks, beating every other model tested. That is nearly 3× Fable 5.1. Another major win for Grok 4.7 and SpaceXAI.
128
239
1,533
338,580
Elon Musk retweeted
Grok 4.7 made this in blender. I didn't tell it what to make but only that it should be something it can accomplish in 10 minutes or so. Honestly not bad at all. Not bad at all. Did they train Grok on blender?
126
189
1,616
397,049
Elon Musk retweeted
Compare Grok 4.7 (first) and 4.6 (second) building an open world city game.
146
289
3,461
583,642
Elon Musk retweeted
Electrical engineering score is a pretty big deal. SpaceX has the engineering data that software companies don't.
Grok 4.7 is here. It's a notable improvement over Grok 4.6 at the same price and speed.
68
153
1,664
442,886