What AI PR Actually Costs: 11,160 Provider Attempts and One Expensive Surprise
We measured 74 production AI Visibility runs and 11,160 stored provider attempts over a fixed 30-day window. One engine allocated roughly 57% of the modeled cost of a deep scan — because a single answer fanned out into about 5 upstream requests and 14 billable web searches.
One engine accounted for roughly 57% of the modeled cost of a deep AI visibility scan — not because it wrote more, but because a single answer quietly fanned out into about 5 upstream API requests and about 14 billable web searches.
That is the finding. Here is the measurement behind it.
The population
We took a fixed 30-day window: 2026-07-18T15:04:07Z inclusive to 2026-08-17T15:04:07Z exclusive. Inside it, RunPR ran 74 production AI Visibility runs — 49 standard and 25 deep — across six engines.
Those runs stored 11,160 provider attempts. Of those, 10,448 returned a stored response and 712 stored a provider error. Attempts, not answers, is the honest denominator: an attempt that fails still costs upstream work and still tells you something about the workflow.
By engine, the attempt counts were:
- ChatGPT: 2,763
- Perplexity: 2,763
- Exa: 2,763
- Claude: 957
- Gemini: 957
- Grok: 957
Standard runs accounted for 5,418 attempts (5,392 stored responses, 26 stored errors) at an average of 36.857 prompts per engine per run, ranging from 24 to 45. Deep runs accounted for 5,742 attempts (5,056 stored responses, 686 stored errors) at an average of 38.28 prompts per engine per run, over the same 24-to-45 range.
The cost
At the workload shape we actually observed, the modeled direct-variable API cost was:
- about $1.15 for a standard scan — 37 rounded prompts across ChatGPT ($0.0152 per prompt, invoice-reconciled), Perplexity ($0.011, modeled), and Exa ($0.005, list-rate allocation)
- about $10.06 for a deep scan — 38 rounded prompts across those three engines plus Grok ($0.151, invoice-reconciled), Claude ($0.080, ledger/model reconciled within 0.7% of the Anthropic bill), and Gemini ($0.0025, invoice-reconciled free-grounding allocation)
Read those as historical modeled estimates at the observed workload shape, not as invoice-grade per-scan bills. Nobody receives a receipt for $10.06. The number is what the direct variable API usage of a typical run in this window costs when you allocate each engine's reconciled or modeled rate across the prompts that engine actually ran.
Where the money went
Inside that modeled $10.06 deep scan, about $5.75 — roughly 57% — allocates to a single engine, Grok. That share is a derived modeled allocation across a typical 38-prompt deep run, not a line item on an invoice for one scan.
The reconciliation is where it gets interesting. We re-ran the xAI management usage API read-only for the exact window. It reported $144.87377 of aggregate cost for grok-4.5. Divided by the 957 Grok attempts we stored, that is $0.151383 per logical attempt — invoice-reconciled.
The same verified window shows 4,875 upstream API requests and 13,400 billable web searches behind those 957 stored attempts. That is about 5 upstream requests and about 14 billable web searches per single logical attempt we made.
We asked for one answer, 957 times. The provider did roughly five times that many requests and fourteen times that many searches to produce them.
That is not a bug and it is not overcharging. Agentic search is the product. But it does break the mental model most teams use to budget AI work.
A model call is not an economic unit
An AI PR workflow is a chain:
- Form the research question.
- Search or retrieve evidence.
- Ask one or more engines to reason over it.
- Classify and normalize the result.
- Compare it against other engines or prior scans.
- Generate the useful output.
- Store enough detail to explain where the answer came from.
Pricing the generation step alone misses most of the system. Two providers can each return exactly one answer while doing wildly different amounts of metered work behind it. Two features can both claim the same model and differ tenfold in fan-out, search volume, retries, prompt count, and classification calls.
The useful operational question is not "what does this model cost per token?" It is "what can this provider cause to happen after we ask for one answer?"
What we changed
Meter workflows, not prompts. A scan receipt has to include engine attempts, search activity, and model usage — and mark each value as measured, reconciled, or modeled.
Budget the expensive behavior, not the cheap text. Deep-scan cadence, engine selection, and per-account limits move this number far more than prompt trimming does. Internal and complimentary accounts get caps too.
Keep confidence visible. Invoice-reconciled figures, modeled rates, and known undercounts should never be blended into one authoritative-looking total. This article labels every rate for that reason.
Fail loudly when metering is incomplete. A cost dashboard that silently omits a workflow is worse than one that admits the number is partial.
Limitations
We are publishing these because they are part of the result.
Historical `model_usage_events` coverage for this window was incomplete — 2,411 visibility rows against 11,160 stored attempts, with token columns null. That is why the per-engine rates above are modeled and invoice-reconciled allocations rather than invoice-grade per-run measurements. Proof scans run after the window closed showed real run-to-run variance, and future studies will use the improved exact metering rather than back-filled allocation.
The Perplexity rate is modeled and the actual monthly bill suggested a lower effective rate; we kept the higher modeled figure rather than restate it downward without reconciliation. The Exa figure is a list-rate allocation. The OpenAI rate comes from $42.12 aggregate across 2,763 attempts, or $0.015244 per attempt. The Gemini allocation comes from a $2.43 monthly GCP bill against free-tier grounding.
These figures cover direct variable API usage for RunPR's AI Visibility scan workload over one 30-day window. They exclude engineering time, human review, fixed infrastructure, coverage-workflow costs whose metering remains incomplete, and every other RunPR workflow. They are not a universal price list, and provider pricing and behavior change.
For the provider-side accounting we used, xAI documents its cost tracking and management billing API publicly.
The takeaway
AI PR does not get expensive because a machine writes a paragraph.
It gets expensive when one useful output fans out across dozens of prompts, six engines, live searches, retries, classifications, and evidence checks — and nobody measures the chain end to end.
If you are building or buying an AI workflow, ask for more than the model name. Ask what one user action causes the system to do.
That is where the cost actually lives.
RunPR's AI Visibility scans run this workflow in production across ChatGPT, Claude, Gemini, Perplexity, Exa, and Grok. If you want to see where your brand currently stands, run a free visibility check.
Want a tighter PR workflow?
See how RunPR helps teams move from signal to drafted, approved, and sent outreach.
Start free trial