Signals Inbox·July 28, 2026·Frontier AI
Is Grok far behind Claude and ChatGPT?
Grok is no longer far behind Claude and ChatGPT on raw model capability, but xAI still trails badly on factual reliability, habitual use, enterprise adoption and commercial scale.
We track what's happening in frontier AI. Want the market signals in your inbox?
Send me the signals →Grok is close to Claude and ChatGPT technically, but it remains far behind them as a product and a business. Its best model trails the intelligence leaders by roughly 9% to 12%, while the gaps in paying users, enterprise spending and revenue are several times larger.
The most interesting part of Grok’s position is its economics. Its coding agent delivers nearly the same benchmark score as Claude Code and Codex at a fraction of the measured cost, making Grok a credible choice for high-volume work where occasional errors can be caught cheaply.
Its biggest technical weakness is not intelligence but judgment. Grok knows much more than earlier versions, yet it has also become more willing to guess when it does not know, which weakens the apparent cost advantage in research and other high-stakes work.
X and Tesla have solved distribution in a way Claude cannot easily copy. They have not yet solved habit: many people encounter Grok, try it once and leave, while Claude attracts valuable professional use and ChatGPT remains the default consumer assistant.
xAI has enough money, computing capacity and product momentum to stay in the frontier race. Catching the models is plausible. Catching the customer relationships built around Claude and ChatGPT will be much slower.
Looking at frontier AI?We can send you all the signals
Send me the signals → Delivered straight to your inboxQ1What would it actually mean for Grok to be “far behind”?
Grok is currently close to Claude and ChatGPT in raw model capability, yet far behind in several areas that decide which AI products people actually use, pay for and trust.
For an ordinary user, the comparison sounds simple: can Grok answer the same questions, write the same code and complete the same tasks? On that test, the gap has become surprisingly small.
The business comparison is tougher. ChatGPT has become a daily habit for a huge global audience. Claude is deeply embedded in coding and enterprise work. Both companies have far more paying customers, commercial revenue and production deployments than xAI.
That leaves three different races: model intelligence, product adoption and commercial strength.
Grok belongs near the leaders in the first. It is clearly behind Claude in the second and third, and the distance from ChatGPT is larger still.
Calling Grok technologically weak would now be inaccurate. Calling it an equal competitor overall would be just as misleading.
Q2How far behind are Grok’s best models today?
Grok 4.5 currently trails the strongest Claude and GPT models by roughly 9% to 12% on broad independent intelligence testing.
Artificial Analysis gives Grok 4.5 a score of 53.8 on its Intelligence Index, which combines coding, scientific reasoning, knowledge and agentic work. Claude Opus 5 leads with 60.7, Claude Fable 5 scores 59.9 and GPT-5.6 Sol reaches 58.9.
Grok therefore sits 5.1 points behind GPT-5.6 Sol and 6.9 points behind Claude Opus 5. Relative to the leaders’ scores, those gaps equal approximately 8.7% and 11.4%.
That difference should be visible across thousands of difficult tasks. It will be much harder to notice during a normal conversation, especially when the user asks for writing, summaries, brainstorming or contained coding help.
Rankings also exaggerate the distance. Grok appears eighth on the current leaderboard, but several positions above it belong to different versions or reasoning settings from the same Claude and GPT families. In practice, three laboratories are clustered near the frontier, with xAI at the back of that group.
The latest releases from Anthropic and OpenAI have widened the gap slightly after Grok 4.5 briefly moved closer. Even so, xAI has already made up most of the technical distance that existed during Grok’s first generations.
Frontier model intelligence comparison
| Model | Intelligence Index | Lead over Grok | Relative lead |
|---|---|---|---|
| Claude Opus 5 | 60.7 | 6.9 points | 11.4% |
| Claude Fable 5 | 59.9 | 6.1 points | 10.2% |
| GPT-5.6 Sol | 58.9 | 5.1 points | 8.7% |
| Grok 4.5 | 53.8 | Baseline | Baseline |
Q3Is Grok still behind Claude and ChatGPT in coding?
Grok is only a few points behind the best Claude and GPT coding agents, although those last few points can become expensive on difficult software projects.
The latest Artificial Analysis Coding Agent Index places Claude Code with Opus 5 and OpenAI Codex with GPT-5.6 Sol at 67. Claude Fable 5 scores 66, while Grok 4.5 inside Grok Build reaches 64.
A three-point gap is small enough for Grok to handle serious development work. It also means the leaders complete around 5% more of the benchmark’s available score. Across a long project involving dozens of decisions, failed terminal commands and code revisions, that modest average advantage can prevent several hours of cleanup.
Grok’s strongest result is efficiency. Artificial Analysis measured an average API cost of $2.59 per coding task for Grok Build. The same benchmark cost $7.08 with GPT-5.6 Sol in Codex and $8.23 with Opus 5 in Claude Code.
The comparison includes the surrounding coding tools as well as the underlying models. Claude Code, Codex and Grok Build use different prompts, tool systems and recovery methods, so the complete working product is what counts.
For developers handling thousands of routine tasks, Grok’s lower cost could outweigh a three-point performance deficit. Teams working on complex migrations, security-sensitive code or production incidents may still prefer the extra reliability offered by Claude or GPT.
Coding-agent performance, cost and task time
| Coding agent | Coding Agent Index | Average cost per task | Average task time |
|---|---|---|---|
| Claude Code with Opus 5 | 67 | $8.23 | 23.6 minutes |
| Codex with GPT-5.6 Sol | 67 | $7.08 | 10.2 minutes |
| Claude Code with Fable 5 | 66 | $11.70 | 23.4 minutes |
| Grok Build with Grok 4.5 | 64 | $2.59 | 16.5 minutes |
We track what's happening in frontier AI. Want the market signals in your inbox?
Send me the signals →Q4Is Grok cheap enough to make up for slightly weaker performance?
Grok’s low cost already makes it a better choice for some high-volume workloads, even while Claude and GPT remain stronger overall.
xAI charges $2 per million input tokens and $6 per million output tokens for Grok 4.5. Claude Opus 5 costs $5 per million input tokens and $25 per million output tokens.
That makes Grok 60% cheaper on input and 76% cheaper on output before prompt caching or volume discounts. Its shorter answers and lower token use can widen the real difference.
The coding-agent results show what this looks like in practice. Grok delivered 96% of the leaders’ index score while costing about 63% less than GPT-5.6 Sol and 69% less than Claude Opus 5 per task.
A company running one million similar agent tasks would face a measured benchmark cost of roughly $2.6 million with Grok, $7.1 million with GPT-5.6 Sol and $8.2 million with Opus 5. That is not a small procurement detail.
Benchmark costs do not include human review, failed deployments or the damage caused by a bad answer. Grok becomes less attractive whenever its mistakes require substantially more checking.
For customer support, document classification, first drafts, testing and repetitive internal work, the cheaper model may still win. A company does not need the absolute best model for every request. It needs enough accuracy at a cost the business can sustain.
Q5Does Grok still make too many things up?
Grok’s factual reliability currently looks weaker than its overall intelligence score, and this remains one of its clearest technical problems.
On Artificial Analysis’s AA-Omniscience evaluation, Grok 4.5 correctly answered 52% of the factual questions. Claude Fable 5 reached 61%, while GPT-5.6 Sol reached 59%.
The worrying part appeared when Grok did not know the answer. Its measured hallucination rate reached 54%, up from 25% for Grok 4.3. In this benchmark, the hallucination rate tracks how often a model gives a false answer rather than acknowledging uncertainty among its unsuccessful responses.
Grok has become much more knowledgeable while also becoming more willing to guess. Its factual accuracy climbed from 35% to 52%, but its mistakes became more confident at the same time.
One evaluation cannot describe every use case. Retrieval tools, web search and well-designed prompts can reduce errors considerably. Coding agents can run tests too, which catch many mistakes before the user sees them.
Still, the result changes the price comparison. A cheap answer that needs heavy verification becomes expensive quickly. The problem is especially serious in research, finance, healthcare, law and any workflow where a plausible false statement may survive several rounds of review.
Grok’s intelligence is now close to the frontier. Its instinct for when to stop and say “I don’t know” has further to go.
Q6How many people use Grok, Claude and ChatGPT?
Grok has built a large audience, but Claude currently has around twice as many monthly users and ChatGPT has roughly nine times as many.
SpaceX’s public filing reported approximately 117 million monthly users of Grok features. Sensor Tower separately estimated about 245 million monthly users for Claude and more than 1.1 billion for ChatGPT.
The measurements come from different systems. SpaceX counted users interacting with Grok features across its ecosystem, while Sensor Tower estimated activity across websites and mobile applications. The ratios are better read as reliable orders of magnitude than as audited head-to-head totals.
Even with that caveat, the gap is obvious. Claude’s estimated audience is 2.1 times Grok’s. ChatGPT’s is approximately 9.4 times larger.
Grok has already crossed the threshold where it can be called a major consumer product. Very few standalone AI assistants have reached 100 million monthly users.
The problem for xAI is the company it keeps. Reaching 117 million users would make most software products global successes. Here, it still leaves Grok well behind one rival and almost an order of magnitude behind the leader.
Estimated monthly audience by AI assistant
| Product | Estimated monthly users | Audience versus Grok | Main source |
|---|---|---|---|
| ChatGPT | More than 1.1 billion | About 9.4× | Sensor Tower |
| Claude | About 245 million | About 2.1× | Sensor Tower |
| Grok | About 117 million | Baseline | SpaceX filing |
Looking at frontier AI?We can send you all the signals
Send me the signals → Delivered straight to your inboxOpenAI's product lead says future AI won’t fit on your laptop
Mistral and Saudi Arabia are building frontier AI in Arabic
Open-weight models jumped from 28% to 62% on Vercel
OpenCode is giving away Ox Alpha, with 1M-token context
Stripe tells investors the Singularity began on January 1, 2026
Anthropic adds $18B in annualized revenue, reaching $65B
Q7Is X creating loyal Grok users or mostly casual trials?
X gives Grok enormous exposure, but the latest usage data suggests that much of this attention remains shallow.
SpaceX reported 550 million combined monthly users across X and Grok, of whom 117 million used Grok features. That puts the apparent Grok penetration rate near 21%.
A fifth of a huge social network trying an AI feature is impressive. Repeat usage looks less convincing.
Researchers studying more than 169,000 public posts invoking Grok found that 76.8% of users called it only once. A separate analysis of 41,735 interactions found that half of Grok’s public responses received 20 views or fewer after 48 hours.
Those studies cover visible activity on X rather than private conversations inside the Grok application. They still reveal how the social integration is often used: someone summons Grok to explain a post, settle an argument or check a claim, then moves on.
The standalone application has had the same difficulty keeping momentum. AppMagic estimated that Grok downloads fell from more than 20 million in January 2026 to around 8.3 million in April, a drop of nearly 60%.
Payment data is harsher. In a Recon Analytics survey of more than 260,000 American AI users and workers, 0.174% said they paid for Grok. More than 6% paid for ChatGPT, giving ChatGPT roughly 35 times Grok’s paid adoption in the same survey.
X has solved Grok’s discovery problem. Habit and willingness to pay remain unresolved.
Q8Is ChatGPT still the default AI assistant?
ChatGPT is currently the default consumer AI assistant by a huge margin, even though rivals have started taking market share.
OpenAI reports more than 900 million weekly active users and over 50 million paying consumer subscribers. Sensor Tower’s broader monthly estimate recently crossed 1.1 billion.
The paid figure is particularly revealing. ChatGPT’s 50 million consumer subscribers equal roughly 43% of Grok’s entire reported monthly audience. OpenAI has nearly half as many paying consumers as Grok has total monthly users.
ChatGPT also benefits from familiarity. People already know where to open it, what kind of questions to ask and how to work around its weaknesses. Schools discuss it, companies train employees on it, and other products increasingly connect to it.
That accumulated habit explains why a technically excellent challenger cannot release a slightly better benchmark score and expect hundreds of millions of people to switch.
ChatGPT’s relative share has recently slipped as Claude, Gemini and other assistants grow. Its absolute use continues to rise. Sensor Tower estimated that ChatGPT accounted for 46.4% of AI-assistant activity after previously holding more than half.
A falling share inside a rapidly expanding market still leaves OpenAI with more users, more subscribers and more time spent than any direct competitor. Grok needs several years of exceptional growth, or a genuinely new distribution channel, to close that consumer gap.
Q9Why does Claude punch above its audience size?
Claude has a much smaller consumer audience than ChatGPT, but its users are unusually valuable because so much of its activity comes from coding and professional work.
Sensor Tower estimates that Claude has around 245 million monthly users and approximately 10% of global AI-assistant activity. That is far below ChatGPT, yet comfortably ahead of Grok.
Claude also converts a high share of users into subscribers. Sensor Tower estimated that about 13% of Claude’s monthly users pay, compared with the usual 2% to 5% range for many consumer software products.
Anthropic’s own usage research helps explain why. Fixing software errors has repeatedly appeared among Claude’s most common activities. Among direct API customers, software correction recently represented around one in ten recorded tasks.
These users tend to consume more tokens and return more frequently than someone who occasionally asks for a restaurant recommendation or rewrites a short email. They also bring Claude into their companies after using it personally.
OpenAI owns the largest general audience. Anthropic has built a smaller audience concentrated around work people are willing to pay to complete.
Grok currently sits behind Claude on both dimensions. Its audience is smaller, and the public evidence points to less intensive professional use.
We track what's happening in frontier AI. Want the market signals in your inbox?
Send me the signals →Q10Does Grok now have a complete product?
Grok now offers nearly every major feature expected from a leading AI assistant, so product breadth is no longer its main weakness.
Grok 4.5 is available through the web, X, iOS and Android. Users can search current information, analyze files, generate images and video, speak with a voice assistant and create documents, spreadsheets and presentations.
xAI has lately moved further into workplace software. Grok now works directly inside Word, Excel, PowerPoint and Outlook through Microsoft 365 add-ins. Those tools can draw from emails, SharePoint, Google Drive, the web and X.
Developers can access Grok through xAI’s API, Cursor, OpenRouter, Vercel, Cloudflare, Snowflake and Databricks. Grok Build gives the company its own terminal-based coding agent.
The company also has a strong voice product. Its voice API powers Grok inside millions of Tesla vehicles and is being used for Starlink customer support.
Feature lists can hide differences in maturity. ChatGPT has had more time to connect research, coding, files, memory, voice, image generation and business controls inside one familiar interface. Claude has built a particularly coherent workflow around Claude Code, Cowork and professional document tools.
Grok assembled most of the necessary pieces quickly. Now it has to get people to use them often enough that they feel like one product rather than a busy feature list.
Q11Do X and Tesla give Grok an advantage that Claude cannot copy?
X and Tesla give Grok a genuine distribution advantage, particularly for live public conversations and in-car voice use.
Claude has no social network where its assistant can be inserted directly into public discussions. OpenAI can search the web and connect to outside data, but it does not own a stream of hundreds of millions of posts, replies, reactions and breaking narratives.
Grok can therefore see how a story is spreading on X before traditional search results fully catch up. That is useful for tracking politics, market reactions, online communities and fast-moving public disputes.
The feed is also full of noise: impersonation, recycled footage, coordinated campaigns, bad jokes presented as facts and plain speculation. Fast access helps only when the model can work out which material deserves trust.
Tesla creates a separate opportunity. Grok Voice already runs inside millions of vehicles, where it can access vehicle status, provide directions and control navigation. Neither Claude nor ChatGPT has an equivalent built-in automotive network at that scale.
These channels give xAI routes into everyday life that competitors would struggle to reproduce. So far, neither has generated consumer engagement comparable with ChatGPT or professional adoption comparable with Claude.
The advantage is real. Its commercial value depends on whether Grok becomes indispensable inside those environments.
Q12How far behind is Grok in enterprise and developer adoption?
Grok is currently far behind Claude and OpenAI in enterprise adoption, and this gap is much wider than the difference between their models.
Menlo Ventures estimated that Anthropic captured 40% of enterprise spending on large language models in 2025. OpenAI held 27% and Google held 21%. Every other provider combined shared the remaining 12%.
That residual group includes xAI, Meta, Mistral, Cohere and several smaller laboratories. Grok’s individual share must therefore be well below 12%.
A separate Enterprise Technology Research survey produced a more direct comparison. Among roughly 500 respondents, 48% said their companies were using Claude and planned to continue. The figure reached 40% for Gemini and only 7% for Grok.
The customer totals show how much infrastructure already sits behind those percentages. OpenAI says more than one million businesses use its products. Anthropic has disclosed more than 300,000 business customers, with over 1,000 spending at least $1 million on an annualized basis.
xAI has released Grok through most major developer platforms, so access is no longer a serious barrier. A team can test it through Cursor, Snowflake, Cloudflare, Databricks or a model gateway without rebuilding its stack.
Availability and adoption are different achievements. Developers can choose Grok today, but public spending and customer data show that most production workloads still go elsewhere.
Enterprise adoption indicators
| Enterprise indicator | Claude or Anthropic | ChatGPT or OpenAI | Grok or xAI |
|---|---|---|---|
| Share of enterprise model spending | 40% | 27% | Included within a combined 12% |
| Companies using and planning continued use | 48% | Not reported in the same survey | 7% |
| Disclosed business customers | More than 300,000 | More than 1 million | No comparable figure disclosed |
| Customers spending at least $1M annualized | More than 1,000 | No comparable figure disclosed | No comparable figure disclosed |
Looking at frontier AI?We can send you all the signals
Send me the signals → Delivered straight to your inboxQ13Are Grok’s safety controversies hurting its chances?
Grok’s safety controversies are almost certainly making enterprise and government sales harder, even though we cannot count the contracts they have cost.
SpaceX itself warned investors that Grok’s “Spicy” image mode and “Unhinged” voice mode could create legal, regulatory and reputational damage. That warning is unusually direct for a company discussing one of its own flagship products.
The concern became concrete after a Grok image feature allowed users to alter photographs in sexually suggestive ways. Images involving women and minors triggered criticism from lawmakers and regulators, after which xAI restricted access.
The controversy coincided with Grok’s download peak. Downloads later dropped by almost 60%, while paid adoption barely moved. That does not prove the feature caused the decline. It does show that attention generated by provocative capabilities did not become a durable paid audience.
Large companies examine data security, support, reliability, legal exposure and the behavior of the provider before approving an AI system. A controversial consumer persona gives risk teams one more reason to delay the purchase or choose a safer-looking alternative.
Anthropic has spent years making safety and controlled deployment part of Claude’s enterprise identity. OpenAI has also built extensive security, compliance and administrative products around ChatGPT.
Grok’s personality attracts attention. The same branding makes the climb into regulated industries steeper.
Q14How far behind is xAI financially?
xAI remains roughly an order of magnitude smaller than Anthropic and OpenAI by revenue, while its operating losses are several times larger than the money it brings in.
SpaceX’s filing showed $3.2 billion of revenue for the xAI segment in 2025 and an operating loss of $6.36 billion. The segment lost almost $2 for every $1 of revenue.
The comparison may overstate Grok’s commercial strength because the segment also includes revenue associated with X. Grok subscriptions, API use and enterprise contracts account for only part of the total.
The position weakened further during the first quarter of 2026. xAI generated $818 million of revenue and recorded an operating loss of $2.47 billion. That works out to slightly more than $3 of operating loss for each dollar earned.
Anthropic says its annualized revenue run rate has crossed $47 billion, with more than 1,000 customers spending at least $1 million each. Public reporting places OpenAI’s annualized revenue in the tens of billions of dollars.
Run-rate revenue and recognized annual revenue are different measurements, so the ratios are approximate. Even with conservative adjustments, both rivals are many times larger commercially.
xAI can continue absorbing these losses because it has access to SpaceX, investors and valuable infrastructure. Money gives Grok time to improve. It does not create customer demand by itself.
Q15Does xAI have enough computing power to catch up?
xAI has enough computing infrastructure to remain in the frontier race, so compute scarcity is no longer a convincing explanation for Grok’s position.
Grok 4.5 was trained on tens of thousands of Nvidia GB300 GPUs in xAI’s Memphis data centers. The xAI segment spent $12.7 billion on capital expenditure in 2025 and another $7.7 billion during the first quarter of 2026.
That latest quarterly spending equals an annualized pace of more than $30 billion. Few AI companies can deploy infrastructure at anything close to that rate.
xAI also raised $20 billion before joining SpaceX, giving the combined group access to capital, energy expertise, global connectivity and a large engineering organization.
The company has even started renting computing capacity to Anthropic. Commercially, that turns expensive infrastructure into revenue. Strategically, it shows that owning a giant data center does not automatically produce the strongest model or the largest AI business.
Anthropic, OpenAI and Google are increasing their capacity at the same time. Anthropic has signed multigigawatt agreements with Amazon and Google. OpenAI is pursuing computing commitments measured in hundreds of billions of dollars.
All three companies can now train very large models. Research talent, data quality, post-training, inference efficiency, product design and feedback from real customers will decide more of the next round.
We track what's happening in frontier AI. Want the market signals in your inbox?
Send me the signals →Q16Is Grok improving fast enough to catch Claude and ChatGPT?
Grok is improving faster than its current market position suggests, although Claude and ChatGPT are moving too quickly for xAI to close the gap through model upgrades alone.
On the Artificial Analysis Intelligence Index, Grok 4 scored 33.3, Grok 4.3 reached 37.6 and Grok 4.5 climbed to 53.8. That is a 20.5-point improvement between Grok 4 and Grok 4.5, or a gain of more than 60% relative to the earlier score.
Grok’s reported monthly audience also grew from around 35 million in December 2025 to 117 million by March 2026. Some of that jump came from wider integration across X, but tripling the measured audience within one quarter is substantial.
xAI has recently shipped a coding agent, Office add-ins, improved voice models, mobile access to Grok 4.5 and broader developer distribution. The pace of product development is plainly accelerating.
The leaders have accelerated too. Anthropic released several major Claude models within a few months. OpenAI followed GPT-5.5 with GPT-5.6 and expanded Codex to millions of weekly users.
Technical progress can happen in one successful training run. Customer trust, recurring payments and workplace habits accumulate much more slowly.
Grok can plausibly catch the leading models. Catching the businesses built around Claude and ChatGPT will take longer.
Q17What would prove that Grok has caught up?
Grok will have caught up when its technical gains begin producing durable usage and spending rather than another impressive benchmark release.
First, Grok should remain within three Intelligence Index points of the leader across two consecutive model generations. One close result can come from a particularly successful training run. Repeating it would show that xAI can keep pace while its rivals continue improving.
Consumer adoption would need to pass roughly 250 million standalone monthly users, with clear evidence that people return regularly. That would put Grok near Claude’s current audience rather than relying mainly on occasional interactions inside X.
Paid conversion should rise above 2%. This would still trail Claude and ChatGPT, but it would move Grok away from its current position where attention rarely turns into a subscription.
Enterprise share would also need to cross 10%. Grok currently sits inside a residual category shared by many providers. A double-digit share would show that companies are giving it meaningful production workloads.
Finally, xAI should disclose Grok-specific revenue, retention and business-customer figures. The current segment accounts mix Grok with parts of X, making it difficult to judge whether the assistant itself has found a sustainable market.
Those thresholds are demanding but achievable. Reaching most of them would change the verdict quickly.
Q18So, is Grok far behind Claude and ChatGPT?
Grok is far behind Claude and ChatGPT as a commercial platform, while its best model now sits only modestly behind their strongest models.
The technical gap is smaller than the public perception. Grok trails the current intelligence leaders by roughly 9% to 12%, reaches 64 against 67 on the latest coding-agent index and can run comparable coding work at less than half the cost.
Its factual reliability remains less convincing. Grok 4.5 answered fewer knowledge questions correctly than Claude Fable 5 and GPT-5.6 Sol, while its measured tendency to hallucinate rose sharply from the previous generation.
The market gap is much larger. Claude has around twice Grok’s monthly audience, while ChatGPT has roughly nine times as many users. ChatGPT’s paid consumer base alone approaches half of Grok’s entire monthly audience.
Enterprise adoption gives the clearest verdict. Claude and OpenAI together receive about two-thirds of measured enterprise model spending. Grok is one of several providers dividing most of what remains. xAI also generates only a fraction of its rivals’ revenue while losing several dollars for every dollar it earns.
So the claim is partly true. Grok already belongs in the same technical conversation as Claude and GPT. Today, it remains far behind the products, customer relationships and businesses built around them.
The cleanest description is a near-frontier model inside a second-tier AI platform. Its low cost, rapid improvement and unique access to X and Tesla make it a serious threat. The evidence does not yet make it an equal.
We track what's happening in frontier AI. Want the market signals in your inbox?
Send me the signals →This analysis examines whether Grok is far behind Claude and ChatGPT by separating model capability from product adoption and commercial strength. We evaluate technical intelligence, coding performance, factual reliability, pricing, consumer use, engagement, enterprise adoption, revenue, distribution and computing infrastructure rather than forcing all of them into one overall score.
We compare like with like wherever possible. General model capability is evaluated through independent intelligence benchmarks, coding through complete coding-agent evaluations, consumer adoption through user and engagement estimates, enterprise position through spending and customer data, and financial strength through disclosed company information.
Artificial Analysis is the main benchmark source because it evaluates models and agent products under a shared methodology. Its Intelligence Index is used for broad model capability, its Coding Agent Index for coding performance, cost and task time, and AA-Omniscience for factual accuracy and hallucination behavior.
Benchmark scores are treated as evidence of relative technical performance, not as complete measures of product quality. Coding-agent results include the model, system prompts, tools, execution environment and recovery behavior of Claude Code, Codex and Grok Build. They should therefore be read as comparisons between working agent products rather than isolated base models.
We use official API prices to compare token economics, while the coding-agent benchmark provides a more practical estimate of the cost of completing a standardized task. Neither measure includes the cost of human review, failed deployments or downstream errors.
User estimates come from different measurement systems. SpaceX’s filing reports users interacting with Grok features across its ecosystem, while Sensor Tower estimates activity across websites and mobile applications. We use these figures to establish scale and broad ratios, but we do not treat them as perfectly comparable audited totals.
Public interactions on X are used to understand how Grok’s social integration is being used, not to represent every private conversation in the standalone application. Download estimates, paid-user surveys and repeated-use studies provide additional evidence about whether exposure is turning into a durable habit.
Enterprise position is assessed using Menlo Ventures’ estimates of model spending, Enterprise Technology Research survey data, disclosed business-customer totals and the availability of Grok through major development and data platforms. Platform availability is treated separately from proven production adoption.
Financial comparisons use SpaceX’s disclosed xAI segment figures, Anthropic’s reported annualized revenue and customer milestones, and public reporting on OpenAI’s commercial scale. Run-rate revenue and recognized revenue are different measures, so the comparisons are directional rather than presented as exact accounting ratios.
Distribution through X and Tesla is treated as a strategic advantage only where it produces access that rivals cannot easily reproduce. We do not assume that bundled access automatically creates engagement, subscriptions or enterprise demand.
Our conclusion reflects the combined weight of the evidence. Grok’s technical position is supported mainly by benchmark convergence and rapid model improvement. The conclusion that it remains far behind commercially is supported by a broader set of user, payment, enterprise, customer and financial indicators.
Key sources used for this analysis include: Artificial Analysis and its model leaderboards, Artificial Analysis on Grok 4.5, Artificial Analysis on GPT-5.6, the official SpaceX EU prospectus, Space Exploration Technologies’ SEC filing, xAI’s Grok documentation, OpenAI’s API pricing, Anthropic’s pricing information, Anthropic’s API pricing documentation, OpenAI’s ChatGPT Business information, Anthropic’s Claude for Enterprise information, Menlo Ventures’ State of Generative AI in the Enterprise, Enterprise Technology Research, Sensor Tower, AppMagic, and Recon Analytics.
Building or investing in frontier AI?We can send you all the signals
Send me the signals → Delivered straight to your inbox