Grok 4.7 Kept Its $2/$6 Price and Can Still Double Cost per Task

Grok 4.7's per-token price didn't move. Its per-task cost did: up 47% at matched effort on Artificial Analysis's index, down 10% on Cursor's own benchmark. Price it on your workload.

By Rajesh Beri·September 22, 2026·11 min read
Share:
A developer's laptop open on a desk showing a code editor, beside a small thermal receipt printer pushing out a paper receipt so long it curls off the desk edge onto the floor.

Illustration generated using AI

Grok 4.7 has exactly the same token price as Grok 4.6 — $2 per million input tokens, $6 per million output — and that tells you almost nothing about what a task will cost. On Artificial Analysis's independent index, a task at the default "high" reasoning effort costs $2.73, against $1.86 for Grok 4.6 at the same effort: 47% more. The xhigh setting behind most of SpaceXAI's launch numbers costs $3.74 a task, roughly double 4.6 at high. On Cursor's own coding benchmark it goes the other way, and 4.7 is slightly cheaper. The price per token didn't change. What you pay per task now depends on your workload and your effort setting, so measure both before a default changes under you.

It is already changing. SpaceXAI released Grok 4.7 on September 21. According to its model card, it is the default model in Grok Build and in the Grok add-ins for Word, PowerPoint and Excel. It is available on every Cursor plan tier, and you can reach it through OpenRouter, Vercel, Cloudflare, Snowflake and Databricks Mosaic. A router, a gateway or a developer that picks models from the list-price column will see "same price, better scores" and switch.

What Does "Same Price" Actually Cover?

It covers the price per token and nothing else. xAI's price list gives Grok 4.7 and 4.6 identical rows. Both cost $2.00 input, $0.50 cached input and $6.00 output per million tokens below 200k tokens of context, and $4.00 / $1.00 / $12.00 at or above it. The change is in how many tokens the model spends to finish a job.

Reasoning effort is a per-request setting that controls how long a model thinks before it answers, and you pay for the thinking. xAI's reasoning guide lists four levels (low, medium, high and xhigh), makes high the default, and says reasoning tokens "are billed as part of your total consumption." Cursor's Grok 4.7 page adds the line that matters: "effort levels are more separated than in Grok 4.6, so harder tasks spend more time thinking."

Artificial Analysis measured what that separation costs across its Intelligence Index. These are its figures as published on September 21-22, 2026:

Configuration AA Intelligence Index Cost per task
Grok 4.6, high 44 $1.86
Grok 4.6, xhigh 44 $2.32
Grok 4.7, high (default) 46 $2.73
Grok 4.7, xhigh 46 $3.74

Per-row sources: 4.6 high, 4.6 xhigh, 4.7 high, 4.7 xhigh.

Compare the rows. At the same effort, a task costs 47% more at high and 61% more at xhigh. SpaceXAI's launch table pairs 4.7 at xhigh with 4.6 at high, and on that pairing the cost per task roughly doubles for a two-point gain on the index. Inside 4.7 itself, xhigh shows the same score of 46 as high and costs 37% more.

The token counts explain why. Artificial Analysis's write-up says Grok 4.7 at xhigh "uses approximately 81k output tokens per Intelligence Index task." Grok 4.6 used 38k at xhigh and 36k at high, and 4.7 averaged about 7.1 minutes per task. Even at the default setting, AA's model page calls Grok 4.7 "slower than average and very verbose". It calls Grok 4.6 at high "fairly concise". The same write-up concludes that "outside of agentic knowledge work, Grok 4.7 broadly matches Grok 4.6 (high) on the other Intelligence Index tasks."

Why Does Cursor's Benchmark Point the Other Way?

It measures a different workload, and on that workload the extra tokens don't raise the bill. CursorBench 4.0 is Cursor's leaderboard of "ambiguous, multi-file tasks from real Cursor sessions." It publishes cost, output tokens and agent steps for every effort level:

Configuration CursorBench 4.0 Cost per task Output tokens per task Steps
Grok 4.6, high 40.4% $5.20 41,387 48
Grok 4.7, high 43.9% $4.69 56,382 71
Grok 4.6, xhigh 41.4% $6.10 49,814 56
Grok 4.7, xhigh 46.3% $6.01 70,141 88
Fable 5.1, low 45.1% $5.44 34,795 51

At high, Grok 4.7 produces 36% more output tokens and takes 71 steps instead of 48, yet the task costs 10% less. Cursor says its cost column applies "each model's published per-million-token pricing (input, cache read, cache write, and output) to the tokens it used on each task." Output tokens went up at the same output price, so the saving has to be on the input and cache side. The leaderboard doesn't show which of those lines moved.

This is SpaceXAI's strongest case, and it holds up. The launch pitch is "twice as fast, at half the price of comparable models", and the model card claims Grok 4.7 gets results "with fewer steps and fewer output tokens than other frontier models." On CursorBench the price and token claims hold against the premium tier at maximum effort. At xhigh, Grok 4.7 costs $6.01 a task and uses 70,141 output tokens. Opus 5 at max costs $11.95 and uses 85,384. Fable 5.1 at max costs $17.28 and uses 117,236. The token claim fails against GPT-5.6 Sol at max, which used 42,944. It also fails against Grok 4.6, and 4.6 is the model your budget line was built on.

Then there's the row that didn't make any launch post. Fable 5.1 at its lowest effort scores 45.1% for $5.44 a task. That beats Grok 4.7 at its default high setting, and it comes within 1.2 points of Grok 4.7 at xhigh for 57 cents less per task. Yet according to SpaceXAI's own comparison table, Fable 5.1 lists at $10 input and $50 output per million tokens: five times Grok's input rate and more than eight times its output rate. The table leaves out the rate that dominates a long agent session: Anthropic prices Fable 5.1 cache reads at $0.25 per million tokens, half Grok's $0.50. On this benchmark, the headline input and output prices didn't predict the cost per task in either direction.

Which Launch Numbers Compare Like With Like?

One of the seven. The benchmark table in SpaceXAI's announcement heads its columns "Grok 4.7 (xHigh)" and "Grok 4.6 (High)". In six of its seven rows, the new model runs at its most expensive setting and the old one at its default. The exception is DeepSWE v1.1, which carries an asterisk marking it as a high-effort score: 71.0% against 65.2%. The model card says Datacurve ran that one on the mini-SWE-agent harness.

Where same-effort numbers exist, the gap shrinks:

  • CursorBench 4.0: the launch table shows 46.3% against 40.4%, a 5.9-point gain. At the same effort, Cursor's leaderboard shows a 4.9-point gain at xhigh and a 3.5-point gain at high.
  • EEBench: the launch table shows 64.0% against 53.0%. The model card reports 66.0% for 4.7 at xhigh against 60.0% for 4.6 at xhigh, a six-point gain at the same effort. The two documents also disagree on the new model's own score.
  • Terminal-Bench 4.0: the launch's biggest jump, 38.0% against 20.3%, comes from Harbor running Grok 4.7 at xhigh in the Grok Build harness. The launch reports 4.6 only at high. When Artificial Analysis ran Terminal-Bench 4.0 itself, it found a 4.5-point gain over Grok 4.6 at high. The model card admits that "absolute scores remain sensitive to the agent harness," and harness choice can reorder a model ranking outright.

The model did improve. Artificial Analysis's Coding Agent Index scores each model in its own harness. It puts Grok Build with Grok 4.7 (xhigh) at 56, up from 47 with Grok 4.6 (xhigh), which is fourth among native harnesses, behind Fable 5.1, GPT-6 Astra and Opus 5. That nine-point gain is real, and so is its price. On the same index, Artificial Analysis lists a Grok Build task at $8.82 and 39.2 minutes with Grok 4.7 (xhigh), against $3.57 and 19.5 minutes with Grok 4.6 (xhigh). The launch table doesn't show that cost.

Who Built the Benchmark Where Grok 4.7 Looks Cheapest?

Its parent company. SpaceX completed its $60 billion acquisition of Anysphere, the company behind Cursor, on August 14, 2026. CursorBench is built from "real Cursor sessions". A footnote on page 3 of the model card says: "Grok 4.7 received supplemental training on anonymized Cursor workflow data to improve coding and agentic performance."

Start with the case for trusting it. CursorBench scores competitors too, and Fable 5.1 tops it. Grok 4.7's same-effort gain there, 3.5 to 4.9 points, is close to the independent DeepSWE result of 5.8 points at high. Nothing suggests the numbers are wrong.

But a benchmark built from the same kind of sessions the model trained on is an in-distribution test: it checks the model on familiar material. CursorBench is the only published cost-per-task measurement we found where 4.7 comes out cheaper. Artificial Analysis's broader index has it 47% more expensive at high, and AA's own coding-agent run, in SpaceXAI's Grok Build harness, has it about two and a half times as expensive at xhigh. If your developers' work looks like typical Cursor sessions, CursorBench may predict your bill better. If you are routing document work, analysis, or anything through the Office add-ins, it won't. The footnote also doesn't say whose sessions were used or which plans they came from, and that question is worth putting to Cursor in writing. When the deal was announced, we covered what it meant for Cursor's neutrality. This footnote shows that risk playing out.


What Should You Change Before the Default Moves?

Treat Grok 4.7 as a price change, not a free upgrade. Price it on your own work at the same effort before anything switches automatically.

This Week:

  1. Find where 4.7 already arrived without anyone deciding. The model card makes it the default in Grok Build and in the Word, PowerPoint and Excel add-ins. On the API, xAI's model docs say a <modelname> or <modelname>-latest alias updates automatically, and only <modelname>-<date> stays fixed. Search your gateway and router configs for Grok aliases without a date, and pin them until you've re-priced.
  2. Cap reasoning effort at high in the gateway. On Artificial Analysis's index, xhigh scored no higher than high and cost 37% more. Make xhigh an explicit exception for named workflows, each with an owner.
  3. Check Cursor's speed tier. Per Cursor's docs, "Fast is the default speed tier on Pro and higher plans," at $4 / $1 / $12 per million tokens, twice the standard rate. That was already true for Grok 4.6, so it isn't new. But for a developer who never changed the setting, every per-task cost above doubles. Decide team by team whether Fast is worth it.

This Month:

  1. Re-price on 30 to 50 of your own tasks at the same effort. Run 4.6 against 4.7 at high, then at xhigh, in the harness you actually use. Record the cost per accepted result, not per token and not per attempt. Runs on private codebases have already produced cost-per-resolved-task rankings that differ from list-price rankings. And small eval sets rarely separate close models, so size the sample before you trust a three-point gap.
  2. Add a third option to the test: a more expensive model at low effort. CursorBench's Fable 5.1 low-effort row is the reason. A premium model that thinks briefly can beat a cheap model that thinks hard on both score and cost.
  3. Alert on output tokens per task, not just monthly spend. The increases measured here run from 36% on CursorBench to 113% on Artificial Analysis's index (81k against 38k, both at xhigh). A spend alert fires a month late. A tokens-per-task alert on your gateway fires on day one.

Before Renewal:

  1. Put the long-context price tier into your cost model. xAI charges $4 / $12 per million tokens once a prompt reaches 200k tokens. Cursor charges twice the standard rate for requests above 256k, and three times for Fast ones. Agents that run longer build up longer contexts, so either cap context in the harness or budget for the higher tier.
  2. Ask Cursor which sessions trained Grok 4.7. Get a written answer on three points: whether your organisation's sessions were eligible for the "anonymized Cursor workflow data" the model card cites, under which plan settings, and how to opt out for future models.

The Bottom Line

Model pricing used to be one number you could compare across vendors. With configurable reasoning, it is two numbers. The first is the price per token, which vendors publish and compete on. The second is the number of tokens per task, which the price list never shows and which now moves with every release. Grok 4.7 held the first number steady and moved the second: up 47% on one independent index, down 10% on its parent company's coding benchmark. Both results can be true, because tokens per task depend on your workload and harness, not just the model. This publication drew the same lesson from Grok 4.5's per-task pricing in July and from comparing per-million-token prices this month. The rate card is where the calculation starts, not the answer.

SpaceXAI held the price steady. You have to hold the effort setting.

Continue Reading

Share:

Frequently Asked Questions

Is Grok 4.7 more expensive than Grok 4.6?

Per token, no: both list at $2 input, $0.50 cached and $6 output per million tokens below 200k context. Per task, it depends on the workload. Artificial Analysis measured $2.73 vs $1.86 at high effort (+47%) and $3.74 vs $2.32 at xhigh (+61%). Cursor's CursorBench 4.0 shows $4.69 vs $5.20 at high, about 10% cheaper, but Artificial Analysis's coding-agent run in Grok Build lists $8.82 vs $3.57 per task at xhigh.

What reasoning effort does Grok 4.7 use by default?

High. xAI's API supports low, medium, high and xhigh and defaults to high, and Cursor also defaults Grok 4.7 to high. On Artificial Analysis's Intelligence Index, Grok 4.7 at xhigh scored the same 46 as at high while costing 37% more per task.

Were Grok 4.7's launch benchmarks compared at the same reasoning effort?

Mostly not. SpaceXAI's launch table compares Grok 4.7 at xhigh with Grok 4.6 at high; only DeepSWE v1.1 is high versus high (71.0% vs 65.2%). At matched effort, the CursorBench 4.0 gain is 3.5 points at high and 4.9 at xhigh, versus 5.9 in the launch table.

Was Grok 4.7 trained on Cursor data?

Yes. The model card says Grok 4.7 received supplemental training on anonymized Cursor workflow data. SpaceX completed its acquisition of Cursor's maker, Anysphere, on August 14, 2026, and CursorBench is built from real Cursor sessions. The card does not say which sessions or plans the data came from.

How much does Grok 4.7 cost in Cursor?

Standard usage is $2 input, $0.50 cached and $6 output per million tokens; the Fast variant is $4, $1 and $12, and Fast is the default speed tier on Pro and higher plans. Long-context requests above 256k tokens bill at twice the standard rate, or three times on Fast.

Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe

Latest Articles

View All →