Skip to content

Product and operating decision testing

Test a change before you scale it.

RolloutGrade analyzes your transaction, order, and cost data to evaluate product, pricing, assortment, and operating changes—then recommends one of four actions: Roll out, Modify and retest, Continue testing, or Stop / do not roll out.

Planning a change or already made it? Compare against another location, last year's same season, or a comparable earlier stretch.

$99/month · no per-study fees · available outputs depend on your data and study design

Analytics shows activity. RolloutGrade evaluates the decision.

Ordinary reporting is good at what it does. It answers a different question from the one a decision needs. RolloutGrade is the analyst your POS doesn't come with: it starts from the numbers you already have and does the comparison work that turns them into a decision.

Ordinary reporting typically answers → RolloutGrade is designed to answer

  • What sold?

    What changed beyond expected business movement?

  • How much traffic arrived?

    Did the change produce additional orders and contribution?

  • Which channel got credit for the order?

    What likely would not have happened without the change?

  • Which products were bought together?

    Did the change create new demand, shift demand, or benefit another product?

  • How did this period compare with last period?

    How comparable were the periods, and what alternative explanations remain?

  • What are the metrics?

    What should management do next, and what would reverse that decision?

How a RolloutGrade test becomes a decision

  1. 1

    Define the business hypothesis

    State the change, the economic result it should improve, the locked success rule, and the guardrails it must not break.

  2. 2

    Choose the comparison

    Use the strongest design your data supports: locations that didn't make the change, a comparable earlier period such as last year's same season, or both together.

  3. 3

    Add transaction and cost records

    RolloutGrade checks coverage, dates, products, locations, costs, and known differences before evaluating the result.

  4. 4

    Read the decision

    See what happened, what the comparison explains, what remained, how strong the evidence is, and whether to Roll out, Modify and retest, Continue testing, or Stop / do not roll out.

What a decision looks like

An illustrative study, condensed. It opens with the action, then what happened, then what would have happened anyway.

See the full sample

Illustrative sample — not a real customer result

Product launch

Strawberry Cloud Matcha

Roll out
Estimated contribution beyond the comparison trend / 100 transactions
+$18.40
Locked success rule $12.00
Contribution here means gross contribution — what's left after product costs, not profit. Shown per 100 transactions so locations of different sizes can be compared.
Expected growth without the launch
$1,224
From 2 untreated comparison locations
Estimated displacement
29.5%
379 of 1,284 units matched to a decline elsewhere

Evidence. Controlled design · Estimated incremental

The expensive mistake isn't a launch that fails. It's one that looks like it worked.

Three locations, sales up, everyone agrees — so it rolls out to all twelve. Nobody subtracted what those stores would have done anyway, or what customers stopped buying to buy the new thing. The cost of being confidently wrong scales with every location you roll it into.

RolloutGrade exists for the moment before you scale a decision. It asks the same five questions of every change:

  1. 1What decision was tested?
  2. 2What happened during the test?
  3. 3What does the comparison suggest would have happened anyway?
  4. 4How strong an answer can the design support?
  5. 5What should management do next?

Your POS tells you what sold. Not whether the change was worth making.

You changed something — launched an item, raised a price, reworked the menu. Sales moved. But sales always move. RolloutGrade compares what happened with what would likely have happened anyway, counts the real costs, and tells you what the evidence supports doing next.

One location or twenty. Planning a change or already made it. If you have transaction records and something to compare against — another location, last year's same season, or a comparable earlier stretch — you can evaluate the decision.

The numbers you already know

RolloutGrade starts by importing your POS export and double-checking the basics. This isn't the product — it's the raw material.

  • Gross sales. What rang up, before discounts and refunds.
  • Net revenue. Gross sales minus the discounts and refunds included in your export.
  • Units, receipts, and average ticket. How much sold, across how many orders, at what average.

Any POS shows you these. RolloutGrade reconciles them first so everything that follows stands on checked numbers.

The answer your POS can't give you

Available when your data and comparison can support them — and RolloutGrade tells you when they can't.

  • A fair comparison. Every result is judged against what would likely have happened anyway — using locations that didn't make the change, the same season last year, or a comparable earlier period. You see which comparison was used, and why.
  • What the change really added. The estimated result after the comparison, your product costs, and sales the new item pulled from things you already sell.
  • Where the money came from. Whether a jump in revenue came from selling more, charging more, or selling a different mix of products. Those lead to different decisions.
  • How much weight the answer can carry. In plain words: what's solid, what's thin, and what would make the evidence stronger. Never a score without the reasons.
  • A stress test on the result. RolloutGrade re-runs the numbers under tougher and easier assumptions to show how much the answer moves. A robustness check, not a confidence interval.
  • A straight answer about readiness. Before you commit, RolloutGrade checks whether your data can actually answer the question — and tells you what's missing if it can't.
  • A recommended action. Roll out, Modify and retest, Continue testing, or Stop / do not roll out — based on the success rule and guardrails you locked before seeing results.
  • An honest label on every result. Some answers are ready to act on. Others are early signals worth watching. RolloutGrade marks which is which — Directional or Decision-ready — and never dresses one up as the other.

As you keep selling

Decisions get re-read, not filed away. Add new data and RolloutGrade runs the same locked rule again — showing whether your decision strengthened, weakened, or reversed. The rule doesn't move to fit the result.

What each answer needs depends on your change and your data. See the requirements per change type.

Why not just paste your data into ChatGPT or Claude?

Fair question — chat assistants are genuinely good at analysis. The difference isn't quality. It's structure.

The locked rule.

A chat assistant works for you — which is exactly the problem. It helps you make the case for the answer you were hoping for. RolloutGrade makes you set the success rule before you see the result, then holds you to it. You can't negotiate with a threshold you locked last month.

The same answer every time.

Ask a chat model the same question twice and you can get two different analyses — different method, different framing, sometimes different arithmetic. RolloutGrade runs deterministic, tested code. Same upload, same answer, every re-run, for anyone who asks.

The refusal.

A chat assistant will always answer “did my launch work?” — even when your data can't support an answer. RolloutGrade is built to refuse: without a valid comparison, it will not call a result incremental, it will not fill gaps with zero, and it says exactly what's missing. Sometimes the most valuable output is “your data can't tell you that yet.”

Memory across decisions.

Every study you run is recorded — the rule you set, what the evidence said, what the verdict was — in a Decision Ledger that doesn't rewrite history. Your fifth launch can be read against your first: same definitions, same math, no re-arguing the method.

Use both. Ask a chat assistant to explain your report, argue with it, plan your next test. But let the referee be something that can't be talked out of the rules.

Built to show its work.

Incrementality is the target, not a label applied to every comparison. A result is only called incremental when the design supports it.

Read the methodology
  • Every displayed number comes from deterministic code.
  • Every metric identifies its source, formula, scope, and availability requirements.
  • Missing information stays unavailable—it does not become zero.
  • Survey and review evidence remains separate from transaction economics.
  • Historical comparisons do not receive causal labels.
  • Reports display important limitations and reversal conditions.
  • Customer names, emails, and payment-card details are not required for standard analysis.

What the decision layer holds

Between your upload and the answer sits a set of rules a till report or a chat answer doesn't apply.

Locked rules

You set the success bar before any result is visible, and it locks. The answer can never be argued backward.

Comparison-adjusted reads

The result is judged against what similar locations did anyway over the same dates. Normal drift comes out before the verdict, not after it.

Honest refusals

When the data cannot support an answer, the answer is withheld and the reason is stated. A missing number stays missing — it is never manufactured.

Enough data to trust the answer

The planner tells you how large a change your data can actually detect and how long a read takes — measured from your own uploaded history when enough of it exists.

A permanent record

Changing the comparison after results were visible is recorded on the study, permanently, and the report says so.

$99 per month

Full platform access with no per-study fee. The outputs available in each study depend on your data, comparison design, and evidence quality.

$99 / month

Billed monthly. Pricing and billing timing are shown before any payment step, and nothing is charged from the readiness scan.

  • Observed analysis on supported uploads
  • Control-adjusted estimates when a credible comparison exists
  • Contribution analysis when costs exist
  • Customer metrics when stable identifiers and sufficient history exist
  • Unlimited studies, with no per-study fee
  • Cancel anytime
Create my account

Not ready? Check your test readiness first.

$99 a month against a free chat window buys you three things: a rule you set before you saw the result, a number that's identical every time it's recomputed, and a method that never gets rebuilt from scratch — so every study you run can be read against the last one. One rollout decision made on a displaced-sales mirage costs more than a decade of this.

Common questions

Built for product and operating decisions

RolloutGrade tests product, pricing, menu, assortment, and location-level operating decisions using transaction and cost records supplied by the business. Public positioning does not include advertising exposure, campaign performance, media effectiveness, audience movement, brand-attitude measurement, or website conversion and website-to-order measurement.

See whether your next change is ready to test.

The readiness check takes about a minute, needs no account, and no file upload.