AI Observability Pricing: Same 10 GB, $49 or $930

Eleven AI observability options priced against one workload — 500,000 agent runs a month, 5 million spans, 10 GB of traces. The same telemetry costs $49 at SigNoz Cloud and about $930 at W&B Weave, and retention is a bigger multiplier than traffic.

By Rajesh Beri·August 19, 2026·17 min read
Share:
A long steel shelving rack in a records room holding identical grey archive boxes, the first two shelves packed tight and every shelf beyond them standing empty, a rolling ladder parked at the end under a single overhead

Illustration generated using AI

Put $100 to $300 a month in the budget for each production agent doing half a million runs — but only if you split the stream. Send every span to a cheap OpenTelemetry backend and forward a sampled slice to whichever LLM-native tool your evaluation workflow lives in. Send 100% of agent traffic to an LLM-native platform at list price and the same number is $340 to $2,900. Add the six-month log retention the EU AI Act will require of high-risk systems from December 2027 and it clears $2,800 on published rates. The telemetry is identical in all three cases. The bill is not.

Here is the same workload — 500,000 agent runs a month, 10 steps per run, 5 million spans, roughly 10 GB of trace payload, 10 engineers with access — priced against every vendor's own published page on 19 August 2026.

Product Billed on Cost at this workload Retention included Overage rate published?
SigNoz Cloud Teams GB of traces $49/mo (plan floor, unlimited seats) 15 days Yes — $0.30/GB
Grafana Cloud Pro GB of traces + seats $75/mo ($19 platform fee, 50 GB included, 3 of 10 users included) 30 days Yes — $0.400/GB write
Pydantic Logfire Team records (spans) $174/mo (10 seats) 30 days Yes — $2 per million
Braintrust Pro GB processed + scores ~$339/mo 30 days Yes — $3/GB, $1.50/1k scores
Langfuse Core (cloud) traces + observations + scores ~$416/mo 90 days Yes — graduated $8 to $6/100k
Langfuse (self-hosted) your own infrastructure $0 licence your call MIT, no meter
LangSmith Plus traces + seats ~$635/mo at 14 days 14 days Yes — $0.50/1k base
LangSmith Plus (long retention) traces + seats ~$2,840/mo 400 days Yes — $5.00/1k extended
W&B Weave Pro MB ingested ~$930/mo not published Yes — $0.10/MB
Datadog Agent Observability LLM spans only $160 buys 5% of it 15 days No
Arize AX Pro spans + GB $50 buys 1% of it 30 days No

The verdict: default to self-hosted Langfuse for the trace store, or Langfuse Core at $416 for this workload if you would rather not run ClickHouse. If you already operate Grafana or SigNoz, put the spans there and do not buy a second product for the operational half at all. Reserve the LLM-native premium for the fraction of traffic you actually evaluate.


What We Priced: 500,000 Agent Runs, 10 Steps Each

Every vendor in this category quotes a different unit — traces, spans, records, scores, gigabytes, requests — so a list price is not comparable to another list price. The only honest method is to fix one workload and push it through all of them.

Ours is a single production agent at 500,000 runs a month, averaging 10 steps per run. That decomposes into about 4 model calls and 6 tool, retrieval or routing steps per run: 500,000 traces, 5 million spans, 2 million of them LLM calls. At an average serialized payload of 2 KB per span — which assumes you have turned prompt and completion capture on, and you have, or the tool is a latency chart — that is roughly 10 GB of trace data a month. Ten engineers need read access.

Three of those assumptions are yours to change and they move the bill more than any vendor choice: steps per run, bytes per span, and how many people get a login. The third is the one nobody sizes: seats are 61% of the LangSmith bill below and they turn Grafana's $19 platform fee into $75.

The Same Run Is 1 Unit, 4 Units, or 11

One agent run is a different quantity of billable thing depending on who sends the invoice. That single fact, not the headline rate, is what makes these bills incomparable.

  • LangSmith bills the whole trace. One run is 1 billable unit, whether the agent took 3 steps or 30. LangChain's billing documentation puts the base charge at ".05¢ per trace" — $0.50 per 1,000.
  • Datadog bills only model calls. Its Agent Observability page states that "an LLM span is a single call to an LLM provider" and that "tool, workflow, agent, embedding, and retrieval spans are all free." One run is 4 billable units.
  • Langfuse bills everything it ingests. Its pricing page defines a unit as any tracing data point — "traces (complete application interactions), observations (individual steps: spans, events, and generations), and scores (evaluations)." One run is 11 billable units, plus one more for every score you attach.
  • The GB-priced tools bill bytes. Step count is irrelevant; payload size is everything.

An eleven-to-one spread on the same request is not a rounding error, and it inverts as your agents get deeper. Move from 10 steps to 20 and LangSmith's bill does not change at all, Datadog's and Langfuse's roughly double, and the per-GB bills double with them. Move to shorter runs at higher volume and LangSmith becomes the expensive one. Pick the meter that runs in the opposite direction to your architecture's growth.

Retention Is a Cliff, and Article 19 Sits on the Wrong Side

Retention is where a modest bill becomes an unrecognisable one, because every vendor here prices it as a step function rather than a slider.

LangSmith is the cleanest illustration and the most expensive lesson. Base traces are retained 14 days at $0.50 per 1,000. Upgrading a trace to 400-day retention costs an extra ".45¢" per the same billing docs — $4.50 per 1,000 on top of the base charge. An extended-retention trace therefore costs ten times a base one — $5.00 per 1,000 against $0.50, which is how LangChain's own docs describe it. At our workload that is the difference between $635 and $2,840 a month for identical telemetry.

Now put the compliance floor next to it. Article 19 of the EU AI Act requires providers of high-risk AI systems to keep the logs referred to in Article 12(1) — those their systems generate automatically, and only "to the extent such logs are under their control" — "for a period appropriate to the intended purpose of the high-risk AI system, of at least six months."

That obligation no longer bites on 2 August 2026, and any vendor or consultant still telling you it does is working from a superseded calendar. Regulation (EU) 2026/1744, the Digital Omnibus on AI, was adopted on 8 July 2026, published in the Official Journal on 24 July and entered into force on 27 July. It moves the application date of Chapter III, Sections 1 to 3 — which is where Article 19 lives — to 2 December 2027 for standalone high-risk systems under Annex III, and 2 August 2028 for high-risk AI embedded in Annex I products. Six months is still 180 days, and the multi-year contract you sign this quarter will still be running when the clock finally starts.

Read the retention column again. Datadog's base plan is 15 days. Arize AX Pro is 30. LangSmith base is 14. Braintrust Pro is 30 then $0.50/GB/mo to keep it. Grafana Cloud Pro is 30. SigNoz is 15. Logfire Team is 30, Growth up to 90. Not one entry tier in this category reaches the six-month floor. Only two clear it on a published rate, and neither does so at its sticker price: Langfuse Pro, whose three-year data access is advertised at $199 a month but costs $586 at this workload once the same graduated units are added, and LangSmith's extended tier at ten times the base charge, or $2,840.

If your agent touches recruitment, credit, education, essential services or law enforcement, the retention line is not an optimisation — it becomes the requirement on 2 December 2027. The deferral is eighteen months of relief on the compliance date and none at all on the procurement date, because the contract that has to satisfy it is the one in front of you now.

The Two Vendors Whose Rate You Cannot Look Up

Datadog publishes a free tier of "up to 40k LLM spans" and a Pro plan at "$160 per month" for "up to 100K LLM spans," with base retention of 15 days and add-ons extending traces to 30, 60 or 90 days. It then says additional on-demand usage "is billed after the first 100K LLM spans" — and does not publish the rate. At 2 million LLM spans you are twenty times past the included volume, so $160 prices 5% of your workload and a salesperson prices the other 95%. The retention add-ons are the same story.

Arize AX is tighter still: Free is "25k spans per month" with "1 GB per month" and 15-day retention, Pro is "$50 per month" for "50k spans per month" and "10 GB per month" at 30 days. Fifty thousand spans is 1% of our workload. Everything above it is Enterprise, and no overage rate appears anywhere on the page.

Neither is a bad product. Datadog's meter is genuinely the friendliest one in the category for deep agents — free tool and retrieval spans mean a 12-step agent costs no more than a 4-step one. But you cannot put a number in a budget from either page, and that is a procurement fact, not a feature complaint.

Arize carries a second consideration this quarter. Dynatrace announced on 13 August 2026 that it will acquire Arize for $915 million, expected to close "later this quarter or early in Dynatrace's third quarter." Signing a multi-year Enterprise agreement with an unpublished rate, weeks before the counterparty changes owner, is a decision to make deliberately rather than by default. Ask for the renewal terms in writing now.

Weave Is the Loser on Price, and It Is Not Close

Weights & Biases prices Weave trace ingestion at "$0.10/MB" beyond the 1.5 GB included in its Pro plan, which starts at $60 a month. It defines ingested bytes as "trace metadata, LLM inputs/outputs, and any other information you explicitly log to Weave."

Ten cents per megabyte is roughly $102 per gigabyte. SigNoz charges "$0.3/GB ingested" for traces. Grafana Cloud charges "$0.400/GB Write" past a Pro plan whose "$19 per month" platform fee already includes 50 GB. The same bytes cost about 340 times more at Weave's list rate than at SigNoz's. Our 10 GB is $1,024 at Weave's raw rate, about $930 after the Pro allowance and base fee — against $49 at SigNoz and $75 at Grafana, both of which absorb the entire volume inside their plan minimum and charge only for seats on top.

The steel-man is real: Weave is one surface of a machine-learning platform that also does experiment tracking, model registry and sweeps, and if you already run W&B for training then a second contract has its own cost. And 1.5 GB is generous if you never capture prompt content. But agent tracing without prompt content is a dashboard, not an audit trail — and the moment you turn content on, this is the most expensive published rate in the category by two orders of magnitude. Do not route production agent traffic here on a per-MB meter.

The Sleeper: General-Purpose OTel Backends Are 20x Cheaper

The pricing gap between LLM-native platforms and general-purpose OpenTelemetry backends is not a quality gap. It is a market gap. The general-purpose vendors set their rates against ordinary application telemetry, where a customer ships terabytes; the LLM-native vendors set theirs against a new category where a customer ships gigabytes and pays for the analysis.

Pydantic Logfire is the clearest example. Its Team plan is $49 a month, includes 10 million records, and charges "$2 per additional million records above the included credits." That is $0.20 per 100,000 units. Langfuse's rate for the same 100,000 units starts at $8.00 and only falls to $6.00 above 50 million. Logfire absorbs our entire 5.5-million-unit workload inside its included allowance — the bill is $49 plus $25 for each seat past the first five, so $174 for a team of ten. New Relic gives away "100 GB of free data ingest/month" and then charges "$0.40/GB ingested beyond"; the cost there is seats, at $349 per full-platform user per month on an annual commitment, which is why its agentic platform makes sense only where the observability team is already licensed.

What you give up is real and specific: LLM-as-judge evaluators, dataset and experiment workflows, prompt versioning, annotation queues, trajectory scoring. Those are the reasons to pay the premium, and they are worth paying for — on the traffic you evaluate. They are not worth paying for on the 95% of production spans nobody will ever open.

The Second Bill Nobody Budgets: Evaluation

Tracing is the line item people size. Evaluation is the line item that surprises them, because it is metered twice.

Braintrust makes this legible, which is to its credit. Its Pro plan is "$249 / month" and includes "5 GB processed data" then "$3/GB", plus "50k scores" then "$1.50/1k", with "30-day retention" and "$0.50/GB/mo" after that. At our workload — 10 GB of payload, and scoring 10% of runs with two evaluators — the arithmetic lands near $339 a month. Then there is a separate model-credit line, because every LLM-as-judge call is an inference call somebody pays for. Braintrust bundles "$249 credits" into the Pro plan and bills token rates beyond it.

Langfuse folds the same cost into its unit meter: a score is a billable unit exactly like a span. Two evaluators on every run at our volume adds a million units, about $70 a month at the 1M-to-10M tier — plus the judge tokens, on your own model contract.

Budget evaluation as a percentage of runs, not as a feature. Scoring 100% of production traffic with two LLM judges roughly doubles the observability bill and adds an inference bill on top. Scoring 5% and every error costs almost nothing and catches almost everything.

What Changes the Answer

Four variables predict whether you regret this decision. None of them is a feature.

  1. Steps per run, and where it is going. Agents get deeper, not shallower. If your trajectory length is climbing quarter over quarter, a per-span meter compounds against you and a per-trace meter does not. Measure the trend before you sign, not the current average.
  2. Whether you capture content. The OpenTelemetry GenAI conventions define gen_ai.input.messages, gen_ai.output.messages and gen_ai.system_instructions for exactly this, and the OpenTelemetry project's own GenAI observability writeup notes that by default "no prompt content or tool arguments are captured with GenAI telemetry" because "these can contain sensitive data." Turning it on is a 10x payload event. On a per-span meter it costs nothing; on a per-GB meter it is the whole bill. Note also that the registry now marks those attributes as moved out to a separate GenAI semantic conventions repository, so instrumentation churn is a live risk regardless of vendor.
  3. Your retention floor. Six months for high-risk systems under Article 19 once it applies on 2 December 2027, longer if a financial-services or data-protection rule reaches you first — and those rules are in force today. Price the floor, not the default.
  4. Whether you will run infrastructure. Langfuse is MIT licensed "except for the ee folders," so the licence is genuinely free — but self-hosting it means running Postgres for transactional data, ClickHouse for traces and observations, Redis or Valkey for queues, S3-compatible object storage for raw events, and separate web and worker containers. That is a real platform team commitment, roughly $200 to $400 a month of managed infrastructure at this volume, and it is the correct trade only if you already run a ClickHouse or have a data-residency requirement that removes the choice.

Who Should Not Buy Each of These

  • Not Langfuse Cloud if you cannot predict your observation count. The unit meter counts every span and every score, so a framework upgrade that adds internal spans raises your bill without you shipping a line of code.
  • Not LangSmith if you need long retention on high volume. Ten times the base rate for 400-day traces is the single most expensive retention decision in this comparison — and the seat charge of "$39 / seat per month" from LangChain's pricing page is on top. It is the right buy if your team lives in LangGraph and your traces are short.
  • Not Datadog or Arize AX if you have to produce a defensible three-year budget number this quarter. Both are strong products with unpublished overage rates.
  • Not Weave for production agent tracing at any meaningful volume, unless you are already a W&B shop and the alternative is a second vendor contract.
  • Not Braintrust if evaluation is not actually your bottleneck. You are paying a premium for the experiment and scoring workflow; if you are only watching latency and errors, that money is idle.
  • Not Grafana, SigNoz or Logfire if you need LLM-as-judge evaluators, prompt versioning or trajectory scoring built in. They are trace stores. Excellent, cheap trace stores.
  • Not Helicone — Pro at "$79/month" with "1 month" retention, Team at "$799/month" with "3 months" — if you route through multiple frameworks. It is a proxy-first design, strongest when your traffic already goes through one gateway.

How to Size the Budget Line This Quarter

This week: instrument one agent and count. Runs per month, spans per run, bytes per span, and how many people actually need a login. Every number in this article is a function of those four, and you almost certainly have them wrong — teams consistently underestimate span count and seat count, and overestimate payload.

This month: put a sampling decision in the OpenTelemetry Collector, not in the vendor SDK. Head sampling at 100% to the cheap backend, tail sampling of errors plus 1–5% of successes to the LLM-native tool. Doing it in the Collector means switching vendors later is a config change, not a re-instrumentation project.

Before you sign: get the overage rate in writing, get the retention step function in writing, and get the change-of-control language in writing. Two of the eleven options here will not give you the first, one of them is being acquired, and every one of them prices retention as a cliff you will walk off in month seven.

The Bottom Line

This category is repeating the log-management story of 2015, beat for beat. A new telemetry type arrives, the tools that understand it charge a large premium for that understanding, everyone routes 100% of traffic to them because sampling feels like giving up visibility, and eighteen months later a finance business partner asks why the monitoring bill grew faster than the thing it monitors. The teams that came out of that cycle well were the ones that separated cheap storage of everything from expensive analysis of a sample — and they did the separation in the pipeline, where it was reversible.

Do the same thing here, and do it before the traffic arrives. The bill is not set by which vendor you pick. It is set by how many spans you agree to pay someone to think about.

Continue Reading

Share:

Frequently Asked Questions

How much does AI observability cost per month?

For one production agent at 500,000 runs a month (5 million spans, about 10 GB of trace payload, 10 engineers with access), published rates read on 19 August 2026 range from $49/month at SigNoz Cloud Teams and $75 at Grafana Cloud Pro — whose $19 platform fee covers only 3 of the 10 seats — to roughly $416 at Langfuse Core, $635 at LangSmith Plus and about $930 at W&B Weave. Datadog and Arize do not publish an overage rate at that volume.

Is AI observability billed per trace or per span?

It depends on the vendor, and the difference is large. LangSmith bills whole traces, so a 10-step agent run is one billable unit. Datadog bills only LLM spans — tool, retrieval, agent and embedding spans are free — so the same run is four. Langfuse bills every trace, observation and score, so the same run is eleven. Grafana, SigNoz, Braintrust and W&B Weave bill by gigabyte instead, which makes payload size the driver rather than step count.

Why does trace retention cost so much more than tracing?

Every vendor in this category prices retention as a step function rather than a slider. LangSmith charges $0.50 per 1,000 traces at 14-day retention and an extra $4.50 per 1,000 to upgrade to 400 days, which makes an extended trace $5.00 per 1,000 — ten times the base charge. Braintrust keeps data 30 days on Pro then charges $0.50/GB/month. Datadog's base retention is 15 days with paid add-ons for 30, 60 or 90.

Should we self-host Langfuse instead of paying for cloud?

Only if you already run the infrastructure or have a data-residency requirement. Langfuse is MIT licensed except for its enterprise folders, so the software is free, but a self-hosted deployment needs Postgres, ClickHouse, Redis or Valkey, S3-compatible object storage, and separate web and worker containers. At this volume that is roughly $200 to $400 a month of managed infrastructure plus a platform team's attention — against $416 for Langfuse Core cloud.

How do we stop the observability bill scaling with agent traffic?

Split the stream in the OpenTelemetry Collector rather than the vendor SDK. Send 100% of spans to a cheap OTel-native backend for the operational view, and tail-sample errors plus 1 to 5 percent of successful runs into the LLM-native tool where your evaluators and datasets live. Doing the split in the Collector also makes a later vendor change a config edit instead of a re-instrumentation project.

Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe

Latest Articles

View All →