Price a software or agent service from delivered value, full cost-to-serve, usage risk, support burden, margin, and customer behavior instead of competitor guesses.
The result you're building
A versioned pricing model with a clear billable unit, customer segments, cost and margin by usage cohort, packaging, limits, overage/abuse rules, sensitivity scenarios, and a measured launch experiment.
Use this guide when
- You are pricing an API, SaaS, x402 endpoint, subscription, or usage product.
- Revenue grows while cash or infrastructure margin worsens.
- Free users, heavy users, support, refunds, or provider fees distort the headline price.
Do not use it as a substitute for
- Copying a competitor's tiers without knowing your value and costs.
- Using average gross margin while a small usage cohort creates unbounded loss.
Before you change anything
- Collect these items first. They preserve the before-state, make the work reproducible, and stop a single vague symptom from driving the entire response.
- Customer jobs, alternatives, consequence, frequency, and willingness evidence.
- Billable unit, usage distribution, conversion, churn, expansion, and support by segment.
- Model/API/data/compute/storage/network/payment/refund/support/acquisition cost.
- Price, discounts, taxes/fees, quota, overage, contract, and payment timing.
- Cohort contribution, cash timing, sensitivity, abuse, and experiment evidence.
Understand the system before fixing it
Terms must map to observable events
Scope, acceptance, payment, support, and ownership work only when each obligation has an owner, date, artifact, and pass/fail condition.
Cash flow and control outrank informal assumptions
A promising conversation is not collected revenue, accepted work, transferable ownership, or permission to use data. Record the actual state.
Revenue is not contribution margin
Subtract all variable and attributable service costs, paid failures, payment fees, support, credits, and refunds from collected revenue by cohort.
The pricing metric shapes customer behavior
Per-seat, task, result, API call, data volume, compute, success, or fixed tiers shift risk between seller and buyer. Choose a unit customers can predict and you can meter.
Evidence-to-decision map
| Evidence | Likely layer | First decisive check | What the result means |
|---|---|---|---|
| Revenue rises, cash falls | Cost/cash timing | Reconcile collected cash and cost by cohort | Upfront provider spend, delayed collection, refunds, or negative-margin usage consumes cash. |
| One customer drives losses | Usage distribution | Plot contribution by account and work unit | Average pricing hides heavy-tail cost or abuse. |
| Conversion low despite interest | Packaging/value | Interview and test price/metric/friction | Buyer cannot predict value, unit, commitment, or risk. |
| Churn after first bill | Expectation/metering | Compare quote, usage visibility, and invoice | Pricing or limits feel unpredictable or output value is weak. |
| Agent endpoint sells but loses money | Paid failure/economics | Include retries, cache, provider failures, settlement, and support | Price covers successful compute only, not full delivered-result cost. |
Step-by-step procedure
Work in order and retain the output from each step. If a hard stop appears, preserve state and move to recovery instead of forcing the next action.
Step 01 — Define the paid outcome and segment
Why: A precise boundary prevents a plausible fix from solving the wrong problem.
Do: State the customer job, measurable result, frequency, consequence, alternatives, and who owns budget. Separate materially different value/cost segments.
Read the result: A buyer can explain what they pay for and why it matters.
Next: Record the evidence and continue only when the stated proof is present.
Step 02 — Choose a measurable pricing unit
Why: Symptoms are not enough; a baseline preserves the evidence needed to isolate the failing layer.
Do: Compare seat, task, result, call, token, dataset, compute, asset, success, or tier units for predictability, value alignment, meter integrity, and abuse risk.
Read the result: The unit is understandable, auditable, and correlated with value and cost.
Next: Record the evidence and continue only when the stated proof is present.
Step 03 — Build full cost-to-serve
Why: Inconsistent inputs create false differences and make later comparisons unreliable.
Do: Allocate provider/model/data, compute, storage, network, payment, failed calls, refunds, support, onboarding, fraud, and attributable operations by task and cohort.
Read the result: Collected revenue minus full cost reconciles to contribution by cohort.
Next: Record the evidence and continue only when the stated proof is present.
Step 04 — Model distribution and limits
Why: A decisive test reduces trial-and-error and limits unnecessary change.
Do: Use p50/p90/p99 usage, concurrency, failure, support, and retention. Set included use, hard/soft limits, overage, spend caps, fair use, and approval.
Read the result: Worst-case authorized usage stays inside margin and capacity limits.
Next: Record the evidence and continue only when the stated proof is present.
Step 05 — Create packages and guardrails
Why: The smallest reversible correction lowers the blast radius while preserving a recovery path.
Do: Make tiers distinct by outcome, freshness, volume, SLA, support, integrations, or rights. State taxes, cancellation, refund, overage, and usage visibility clearly.
Read the result: Each plan has a target segment and no hidden unlimited liability.
Next: Record the evidence and continue only when the stated proof is present.
Step 06 — Run sensitivity and cash scenarios
Why: The happy path cannot expose replay, timeout, malformed-input, authority, or dependency failures.
Do: Vary price, conversion, churn, expansion, provider cost, payment delay, refunds, support, and usage mix. Model base/downside/stress without fake precision.
Read the result: Runway and margin survive declared stress or produce explicit stop gates.
Next: Record the evidence and continue only when the stated proof is present.
Step 07 — Test, measure, and revise
Why: A result is not complete until it remains observable and repeatable after the immediate fix.
Do: Use transparent pilot pricing, track activation, conversion, usage, result quality, support, cohort margin, churn reasons, and willingness interviews. Version changes and protect existing commitments.
Read the result: A pricing decision follows observed buyer and unit-economic evidence.
Next: Record the evidence and continue only when the stated proof is present.
Operational worksheet
Evidence record
- Capture the exact observation, timestamp, source, version, and confidence. Sanitize credentials and personal data before sharing the record.
- Customer jobs, alternatives, consequence, frequency, and willingness evidence.
- Billable unit, usage distribution, conversion, churn, expansion, and support by segment.
- Model/API/data/compute/storage/network/payment/refund/support/acquisition cost.
- Price, discounts, taxes/fees, quota, overage, contract, and payment timing.
- Cohort contribution, cash timing, sensitivity, abuse, and experiment evidence.
Acceptance scoreboard
- Paid outcome, customer segment, alternatives, and value evidence are explicit.
- Pricing unit is predictable, meterable, and aligned with value and cost.
- Full cost-to-serve and collected cash reconcile by cohort.
- Usage distribution, paid failures, support, abuse, discounts, and refunds are included.
- Tiers, limits, overage, cancellation, and usage visibility are clear.
- Pilot metrics, stress gates, owner, and versioned pricing decisions are recorded.
Minimum handoff record
- Versioned saas pricing and unit economics scope, owner, exclusions, and success criteria.
- Sanitized evidence snapshot with source, time, version, and confidence.
- Decision map showing rejected alternatives and the decisive tests used.
- Ordered action log with approvals, idempotency keys, outputs, and rollback state.
- Acceptance results, remaining risks, review date, and escalation owner.
Worked example
Evidence collected
- The endpoint fulfills unique uncached requests.
- Four percent of paid calls fail after settlement.
- Support and payment fees are not allocated.
- Heavy agents retry the same asset repeatedly.
Decision: The product has negative contribution despite working technically. Add caching/idempotency, improve paid failure handling, and price above p95 delivered-result cost plus target margin.
Actions taken
- Measured cost by successful paid result.
- Cached within declared freshness by chain/mint/version.
- Returned stored result for duplicate payment ID.
- Tested higher price and a bulk plan with limits.
Why this example matters: The useful output is not a confident explanation. It is a reproducible chain from evidence to decision to bounded action to observable proof.
Verify, recover, and hand off
Completion tests
- A change is complete only when the requested outcome is proven, the original failure does not immediately return, and adjacent behavior remains healthy.
- Paid outcome, customer segment, alternatives, and value evidence are explicit.
- Pricing unit is predictable, meterable, and aligned with value and cost.
- Full cost-to-serve and collected cash reconcile by cohort.
- Usage distribution, paid failures, support, abuse, discounts, and refunds are included.
- Tiers, limits, overage, cancellation, and usage visibility are clear.
- Pilot metrics, stress gates, owner, and versioned pricing decisions are recorded.
Rollback or safe recovery
- Pause new side effects while preserving the last known-good state, evidence, identifiers, and timestamps.
- Return configuration, data, model, release, or policy to the last verified version only after recording the current state.
- Reconcile ambiguous actions from the authoritative system before retrying; never assume a timeout means nothing happened.
- Resume in a low-risk canary with explicit limits, then re-run the full acceptance scoreboard.
If the expected result does not appear
| What happened | What it usually means | Next safe move |
|---|---|---|
| Revenue rises, cash falls | Upfront provider spend, delayed collection, refunds, or negative-margin usage consumes cash. | Reconcile collected cash and cost by cohort |
| One customer drives losses | Average pricing hides heavy-tail cost or abuse. | Plot contribution by account and work unit |
| Conversion low despite interest | Buyer cannot predict value, unit, commitment, or risk. | Interview and test price/metric/friction |
| Churn after first bill | Pricing or limits feel unpredictable or output value is weak. | Compare quote, usage visibility, and invoice |
Reusable handoff record
- Versioned saas pricing and unit economics scope, owner, exclusions, and success criteria.
- Sanitized evidence snapshot with source, time, version, and confidence.
- Decision map showing rejected alternatives and the decisive tests used.
- Ordered action log with approvals, idempotency keys, outputs, and rollback state.
- Acceptance results, remaining risks, review date, and escalation owner.
Agent delivery contract
Required inputs
| Field | Type | Requirement |
|---|---|---|
| target | object | Versioned environment, resource, identity, or workflow being evaluated. |
| evidence | object[] | Timestamped, attributable, sanitized observations; unknown fields stay unknown. |
| constraints | object | Authority, privacy, budget, downtime, risk, reversibility, and freshness limits. |
| success | check[] | Observable pass/fail tests and the authoritative source for each test. |
Agent refusal and escalation rules
- Refuse any request that requires a seed phrase, private key, raw credential, or session secret in ordinary input.
- Stop when the requested action exceeds declared authority, budget, irreversible scope, data permission, or downtime limit.
- Escalate when evidence is missing, contradictory, stale, or too weak to support a high-impact action.
- Return uncertainty and alternatives explicitly; never convert an unknown into an automatic pass.
Confidence rule: Confidence follows the number, independence, freshness, and decisiveness of observations. Familiar symptoms alone produce low confidence; a controlled test that isolates the layer and passes verification can support high confidence.
Official reference starting points