Skip to content

How RolloutGrade evaluates an operating decision.

RolloutGrade begins with the business records you provide, separates observed movement from the best available comparison, and states only the conclusion the design can support. This page explains the calculations, evidence standards and limitations behind each recommendation.

Observe first

Start with what happened in the transaction, cost, and operating data.

Compare

Use baseline and comparison evidence to estimate what likely would have happened without the decision.

Qualify the answer

Clearly distinguish measured results, modeled estimates, and conclusions that cannot yet be supported.

The five questions RolloutGrade asks

  1. 1What decision was tested?
  2. 2What happened during the test?
  3. 3What does the comparison suggest would have happened anyway?
  4. 4How strong an answer can the design support?
  5. 5What should management do next?

Everything below explains how each of those five questions is answered, and where the answer stops being supportable.

RolloutGrade is designed to tell you when the evidence supports a decision — and when it doesn’t.

We would rather return “Continue testing” than turn weak evidence into a confident answer.

Built for operating decisions

The method is designed for product launches, pricing changes, menu and assortment decisions and location-level operating changes where transaction and cost data can support a comparison.

No sensitive data needed

Standard analysis only needs sales or transaction exports. Customer names, emails, and payment-card details are not required.

Delete it, it is gone

When you delete a test, its sales files and transaction data are deleted from RolloutGrade as well.

How RolloutGrade thinks

Sales going up does not automatically mean a decision worked.

A new product, price change, promotion, or operational decision can coincide with higher sales for many reasons. The store may have been busier. Prices may have increased. Seasonality may have helped. Customers may have switched away from another product.

RolloutGrade is built to separate those effects before recommending whether to roll out, modify, stop, or continue testing.

  1. 1What happened
  2. 2What would likely have happened anyway
  3. 3What changed elsewhere in the business
  4. 4What remained
  5. 5How strong is the evidence?
  6. 6Roll out / Modify / Do not roll out / Continue testing

RolloutGrade does not require every test to produce an answer. If the available evidence cannot support a credible conclusion, the correct output is to continue testing.

Change and comparison groups

What is a change group?

The change group — also called the test or treatment group in statistical analysis — is the part of the business that received the product, pricing or operating change being evaluated.

New product
A location where the item was launched.
Price test
A location where the price changed.
Recipe or menu change
A location selling the revised product.
Product removal
A location where the item was removed.
Operating change
A location using the new hours, staffing model or process.

Being in the change group does not mean the change succeeded. It simply identifies where the change occurred.

What is a comparison group?

A comparison group is a similar part of the business that did not receive the change being evaluated. It helps estimate what was happening under ordinary business conditions during the same period.

Imagine Location 1 raises the price of a drink while Location 2 leaves the price unchanged. If Location 1 sales rise 8%, that initially looks positive. But if Location 2 sales also rise 7% during the same period, much of Location 1’s growth may reflect normal business conditions rather than the price decision.

Location 1 — change group

The price changed here.

Location 2 — comparison group

The price remained unchanged here.

RolloutGrade compares how both changed over the same period rather than looking only at the location where the change happened.

Not every location is a good comparison.

A useful comparison location should behave similarly to the change location before the change. Factors that may be considered when evaluating comparison quality:

  • Transaction volume
  • Historical sales patterns
  • Product mix
  • Daypart mix
  • Location behaviour
  • Seasonality
  • Channel mix
  • Baseline trend
  • Available data coverage

A nearby location is not automatically the best comparison. The goal is similarity in business behaviour, not simply geographic distance. RolloutGrade lowers confidence when available comparison locations are materially different from the change location.

Why RolloutGrade does not rely only on before vs after.

Before-and-after comparisons are useful for showing what changed. They are weaker for determining why it changed. A store can improve after a decision because traffic increased, seasonality changed, a holiday occurred, weather changed, another offer ran, pricing changed elsewhere, or local conditions changed.

Baseline

What was happening before?

Comparison

What happened elsewhere without the change?

Change group

What happened where the change occurred?

Together, these provide a stronger basis for estimating decision impact.

What would have happened without the decision?

This is the central question RolloutGrade is designed to answer. The alternative result cannot be directly observed because the same location cannot simultaneously receive and not receive the same change. RolloutGrade therefore uses the strongest available comparison evidence to estimate what likely would have happened without the change — the counterfactual. Counterfactual simply means: what likely would have happened if you had not made the change.

RolloutGrade looks at how the change location was behaving before the change, how comparable unaffected locations or groups behaved during the same period, and whether the observed result differs materially from that expected pattern.

This is an estimate, not an alternate reality we can directly observe. The quality of the estimate therefore affects the confidence of the verdict.

Time-period comparison

Comparing this period with a comparable prior period.

Some businesses have only one location, and some decisions are seasonal by nature. In those cases the most useful comparison is not another store — it is the equivalent period before: the same weeks last year, or the same holiday.

RolloutGrade can judge a decision against comparable locations, against a comparable prior period, or against both at once. Both together is the strongest observational design available, because it separates the decision from seasonality and from general trading conditions at the same time. It is not the strongest design that exists: a randomized assignment — a genuine holdout — is stronger still, because it needs no assumption that the two groups were comparable in the first place.

Comparison locations

What happened at similar locations that did not receive the change?

Comparable prior period

How does this period compare with the equivalent period before?

Both

Consistent movement across both is much harder to explain away.

What makes two periods comparable?

  • Similar length, so a longer window doesn't read as growth
  • Similar mix of weekdays and weekends
  • The same holiday or event, aligned by week rather than by date
  • Both periods actually trading, at the same locations
  • No major known disruption such as a closure or an outage
  • Comparable operating conditions — hours, channels, menu or range

Holidays move around the calendar, so aligning by exact dates is usually wrong. RolloutGrade aligns the prior window by whole weeks so weekday and weekend patterns line up, and lets you name the event — “Thanksgiving week”, for example — so the two windows describe the same trading occasion.

Same-store only, and traffic-adjusted.

Locations that were not trading in both periods are excluded from the comparison, so opening or closing a site cannot masquerade as growth. Where transaction counts are available, results are also shown per 100 transactions: a busier season should not be mistaken for product success.

Known differences are recorded, not hidden.

When you set up a comparison you can confirm differences between the two periods — a another offer, a price change, an operating-hours change, a closure, a stockout, a local event or a concurrent business initiative. Anything you note is carried into the report as a stated limitation and reduces confidence rather than being quietly ignored.

What a year-over-year change is, and is not.

A prior-period comparison tells you how this period differs from a comparable one. On its own that is observed change, not proof of cause: a year is long enough for pricing, competition, staffing and demand to all have moved. RolloutGrade reports it as observed change, and only uses decision-impact language when the comparison evidence is strong enough to support it.

No comparison is presented as a perfect match. Every comparison in RolloutGrade is rated — strong, good, weak, or not comparable — with the specific reasons and limitations shown alongside the result.

Two ways to compare a decision.

Historical comparison. Compare the test period with an equivalent earlier period, such as this year's holiday season versus last year's holiday season.

Comparison locations. Compare a location that received the change with a similar location that did not.

Using both. When available, RolloutGrade can combine historical and store comparisons to better separate the decision from broader seasonal or business changes.

A historical file is not required to contain the same dates as the current file. Its purpose is to represent a comparable earlier period. A comparison-location file, by contrast, must cover the same calendar window as the change period.

POS data

How RolloutGrade reads POS data.

Different POS systems export data differently. RolloutGrade normalizes uploaded transaction data into a consistent structure before analysis. Depending on the available export, it may identify information such as:

  • Location
  • Transaction / check
  • Date and time
  • Item
  • Category
  • Quantity
  • Sales
  • Discounts
  • Channel
  • Customer identifier when available
  • Product cost when supplied

Missing information remains missing. RolloutGrade does not silently invent fields that were not present in the source data. If a required field cannot be confidently interpreted, you are asked to confirm the mapping.

Before analysis, RolloutGrade checks whether the data can support the question.

  • Missing store-days
  • Incomplete periods
  • Duplicate records
  • Unidentified products
  • Unavailable costs
  • Missing comparison data
  • Suspicious gaps
  • Unexpected date ranges

Data quality problems appear as visible limitations before they affect the verdict.

Economics

Not every number means the same thing.

Observed

Calculated directly from uploaded data.

  • Sales
  • Units
  • Transactions
  • Average realized price
  • Basket size
  • Channel mix
  • Daypart mix

Estimated

Derived using available evidence and modeling assumptions.

  • Estimated cannibalization
  • Expected performance
  • Modeled displacement

Decision impact

Contribution remaining after direct costs, displacement and the movement estimated from the comparison.

  • Comparison-adjusted contribution
  • Estimated contribution beyond the comparison trend
  • Contribution per 100 transactions attributable to the decision

RolloutGrade never presents an estimated metric as directly observed.

Higher sales do not always mean better economics.

Price

Did customers pay more or less per unit?

Volume

Were more or fewer units purchased?

Cost

Did the direct cost of delivering each unit change?

Contribution

After direct product cost, did the economics improve?

A price reduction may generate more orders while reducing contribution. A price increase may generate higher revenue while unit demand falls. A recipe change may produce the same sales with lower cost. RolloutGrade evaluates these outcomes separately rather than treating revenue growth as automatic success.

The economic verdict prioritizes contribution, not sales alone, when valid cost information is available.

Why RolloutGrade uses realized price.

The price printed on a menu is not always the amount customers actually pay. Discounts, promotions, modifiers, bundles, and other transaction behaviour can change the effective selling price. When supported by the uploaded data, RolloutGrade uses the average realized price from actual transactions rather than assuming the listed price tells the full story.

Displacement

A new product can sell well without growing the business.

Suppose a new matcha drink generates $20,000 in sales. That does not necessarily mean the business gained $20,000 in new demand. Some customers may simply have purchased the matcha instead of another existing drink.

RolloutGrade calls that estimated displacement. It is estimated by matching declines elsewhere in the product mix against the units the new item sold — it is not an observed customer switching from one product to another, and no report describes it as one. The older word cannibalization appears only where the calculation engine’s own labels use it, which is a comparison-adjusted design with a valid counterfactual behind the estimate.

Comparison-adjusted displacement

Item-level declines measured against a credible counterfactual.

Estimated displacement

Item-level declines matched against the new units, without a counterfactual strong enough to size the effect.

Directional substitution signal

Item data too incomplete for either. Reported as a direction, never as a figure.

There is no default displacement rate. When displacement cannot be estimated, it stays unavailable rather than becoming zero.

What basket size can — and cannot — tell you.

RolloutGrade can compare receipts containing a product with receipts that do not contain it. This can reveal useful purchasing patterns. However, a larger basket among customers buying a product does not automatically mean the product caused the larger basket. RolloutGrade therefore reports this as an observed association unless stronger evidence supports an incremental conclusion.

Evidence

Every verdict should tell you how much to trust it.

RolloutGrade does not treat every analysis as equally reliable, and it keeps four separate things separate so none of them can be mistaken for another:

Design
How the comparison was built: randomized, controlled, matched, historical, or descriptive.
Evidence grade
The engine’s own grade for the study, A through D. It is issued once by the calculation engine and displayed verbatim — nothing in the interface recalculates or upgrades it.
Claim type
The strongest wording the design supports: a comparison-adjusted decision estimate, observed change, directional signal, descriptive association, or historical comparison.
Limitations
What weakens this specific result, stated in the report rather than left for the reader to infer.

A comparison-adjusted estimate is available only when the study includes a credible comparison group, comparable prior period or randomized holdout. Without that support, RolloutGrade reports observed change, a directional signal or a descriptive association.

None of these is a statistical significance claim, and none of them is a p-value or a confidence interval. What raises or lowers evidence strength:

  • Baseline coverage
  • Length of the test
  • Comparison quality
  • Cost completeness
  • Transaction detail
  • Item-level coverage
  • Known confounders
  • Stockouts
  • Promotions
  • Closures
  • Customer launch feedback when relevant

Missing evidence lowers the score. It does not get silently filled in.

Reports also show a store sensitivity range: the result recalculated with each changed location removed in turn. It is a robustness range, not a confidence interval. If dropping one location moves the answer across your bar, the decision is downgraded rather than presented as a rollout.

Weak

An early signal. Not enough to act on a high-stakes decision.

Directional

It points one way, but the comparison cannot size the effect reliably.

Comparison-adjusted

A credible comparison group carries the counterfactual.

Comparison-adjusted and stable

A credible comparison that also survives the pre-trend and leave-one-location-out checks.

A weak-evidence result is not a failed test. It means more evidence is needed before making a high-stakes decision.

Sometimes the correct answer is: we do not know yet.

RolloutGrade is intentionally designed to withhold claims when the available evidence cannot support them.

No cost data
Gross contribution cannot be calculated.
No valid comparison group
Observed performance can be reported, but causal decision impact cannot yet be estimated credibly.
Incomplete item data
Displacement may be directional rather than measured.
Short test
Early performance may not represent durable demand.

RolloutGrade never turns missing evidence into a zero, an assumption presented as fact, or an artificially confident verdict. Survey responses and location reviews provide explanatory context. They never replace transaction economics or independently upgrade the evidence grade.

Auditability

Can I audit my result?

Yes. RolloutGrade is designed so the evidence behind a verdict can be inspected without exposing proprietary modeling. Every report lets you inspect:

Source data
What files and periods were analyzed?
Test definition
What decision was evaluated?
Baseline
What period represented normal behaviour before the decision?
Comparison
What location or group was used for comparison?
Inputs
What prices, costs, thresholds, and assumptions were provided?
Observations
What directly changed?
Estimates
Which results required modeling?
Limitations
What data was missing or potentially confounded?
Verdict
What evidence caused RolloutGrade to recommend Roll out, Modify, Do not roll out, or Continue testing?

Each report carries a “Why you can trust this result” panel summarising the data used, what was observed, what was estimated, what is missing, and how strong the evidence is — populated from that analysis, never from an example.

Data & privacy

Your operating data is sensitive.

What data is needed

Sales or transaction exports covering periods before and after the decision, the locations involved, the decision date, and unit costs when you want contribution conclusions.

What data is not needed

Customer names and payment-card details are not required for standard product-decision analysis.

How data is used

Data is used to analyse your own business decisions and generate your reports.

Access

Uploaded data is scoped to your organisation account. Members of your organisation can access it; access is enforced at the database level by row-level security rules tied to your account.

Deletion

You can delete a test and its uploaded data from the workspace, and request deletion of your entire account and its data from Settings. When a test is deleted, the sales files and transaction data uploaded for that test are deleted with it.

Security

Data is transmitted over encrypted connections and stored in a managed cloud database with per-account access rules. We do not claim certifications we have not completed.

See the full privacy policy.

About the methodology

Built for operating decisions.

The method is designed for product launches, pricing changes, menu and assortment decisions and location-level operating changes where transaction and cost data can support a comparison.

It applies familiar measurement discipline to those decisions: clearly defining the change, separating the change and comparison groups, establishing a baseline, distinguishing observation from inference, identifying confounding factors, and communicating uncertainty.

Method boundary

RolloutGrade evaluates changes made inside a business using first-party operational records supplied by that business. It does not observe advertising exposure, track people across platforms, evaluate media channels or measure brand attitudes. Customer surveys and location reviews provide explanatory context and never replace transaction economics.

Why this approach?

Businesses constantly make decisions with incomplete information. A disciplined comparison is what separates signal from noise well enough to make a better decision. RolloutGrade applies that discipline to transaction and cost data.

Customer launch feedback

Where a study collects it, customer launch feedback records new-versus-returning status, self-reported visit influence, intended substitution, repeat interest, product satisfaction, price perception and reasons for purchasing. Survey responses and location reviews provide explanatory context. They never replace transaction economics or independently upgrade the evidence grade.

Define the question
What exactly changed?
Define the change group
Which part of the business received the change?
Define the comparison
What can help represent what would have happened otherwise?
Measure the outcome
What actually changed?
Identify alternative explanations
What else could have caused the result?
Grade the evidence
How much confidence should be placed in the conclusion?
Make the decision
Is the evidence strong enough to act?

Plain English

Every word we use, explained once

You don't need a statistics background to read a RolloutGrade report. These are the only terms we use, and what each one means in ordinary language.

Comparison location
Another one of your locations that did NOT get the change, watched over the same weeks.
If both moved the same way, the change probably wasn't the cause. If only the launch location moved, it probably was.
Independent control
A comparison location that carried on as normal while the change ran somewhere else.
This is the strongest evidence we can build without a lab.
Same weeks last year
The same calendar weeks from a previous year at the same location.
Good for stripping out season, but it can't rule out things that only happened this year.
Internal benchmark
Other parts of the same store that the change shouldn't have touched, used as a stand-in for 'business as usual'.
Directional only — it shows a gap, not proof the change caused it.
What would have happened anyway
Our estimate of the sales you'd have made if you had never made the change.
Everything above that line is what we credit to the change.
Extra sales
Sales above what would have happened anyway.
Not the same as total sales of the new item.
Estimated displacement
Sales the new thing appears to have pulled away from items you were already selling, estimated by matching declines elsewhere.
A launch can look great on its own and still leave the store flat once this is counted. It is an estimate, not an observed customer switching products.
Gross contribution
Sales minus the direct cost of the goods you sold — the money left over.
This is the number that decides whether the change was worth it. It is not profit: it carries no rent, labour or overhead.
Evidence grade
A to D on how good the evidence is — how much data, how clean the comparison, how steady the pattern. The engine issues it; nothing else recalculates it.
A high result on weak evidence is still a weak result.
Evidence strength
How much weight the result can carry, from the design, the coverage and the stability of the pattern.
Weak evidence means run it longer or add a comparison. It is not a statistical significance test and carries no p-value.
Pre-trend check
We check the two locations already moved together BEFORE the change.
If they didn't, the comparison can't be trusted.
Lift
How much higher (or lower) things were than the comparison.
Baseline
The weeks before the change, used as your starting point.
Measurement window
The days from the change onwards that we score.
Checks
Individual customer transactions — one receipt is one check.

FAQ

Common questions

Where to go next

Method boundary

RolloutGrade evaluates changes made inside a business using first-party operational records supplied by that business. It does not observe advertising exposure, track people across platforms, evaluate media channels or measure brand attitudes.

Create account