How RolloutGrade evaluates an operating decision.
RolloutGrade begins with the business records you provide, separates observed movement from the best available comparison, and states only the conclusion the design can support. This page explains the calculations, evidence standards and limitations behind each recommendation.
Observe first
Start with what happened in the transaction, cost, and operating data.
Compare
Use baseline and comparison evidence to estimate what likely would have happened without the decision.
Qualify the answer
Clearly distinguish measured results, modeled estimates, and conclusions that cannot yet be supported.
The five questions RolloutGrade asks
- 1What decision was tested?
- 2What happened during the test?
- 3What does the comparison suggest would have happened anyway?
- 4How strong an answer can the design support?
- 5What should management do next?
Everything below explains how each of those five questions is answered, and where the answer stops being supportable.
RolloutGrade is designed to tell you when the evidence supports a decision — and when it doesn’t.
We would rather return “Continue testing” than turn weak evidence into a confident answer.
Built for operating decisions
The method is designed for product launches, pricing changes, menu and assortment decisions and location-level operating changes where transaction and cost data can support a comparison.
No sensitive data needed
Standard analysis only needs sales or transaction exports. Customer names, emails, and payment-card details are not required.
Delete it, it is gone
When you delete a test, its sales files and transaction data are deleted from RolloutGrade as well.
How RolloutGrade thinks
Sales going up does not automatically mean a decision worked.
A new product, price change, promotion, or operational decision can coincide with higher sales for many reasons. The store may have been busier. Prices may have increased. Seasonality may have helped. Customers may have switched away from another product.
RolloutGrade is built to separate those effects before recommending whether to roll out, modify, stop, or continue testing.
- 1What happened
- 2What would likely have happened anyway
- 3What changed elsewhere in the business
- 4What remained
- 5How strong is the evidence?
- 6Roll out / Modify / Do not roll out / Continue testing
RolloutGrade does not require every test to produce an answer. If the available evidence cannot support a credible conclusion, the correct output is to continue testing.
Change and comparison groups
What is a change group?
The change group — also called the test or treatment group in statistical analysis — is the part of the business that received the product, pricing or operating change being evaluated.
- New product
- A location where the item was launched.
- Price test
- A location where the price changed.
- Recipe or menu change
- A location selling the revised product.
- Product removal
- A location where the item was removed.
- Operating change
- A location using the new hours, staffing model or process.
Being in the change group does not mean the change succeeded. It simply identifies where the change occurred.
What is a comparison group?
A comparison group is a similar part of the business that did not receive the change being evaluated. It helps estimate what was happening under ordinary business conditions during the same period.
Imagine Location 1 raises the price of a drink while Location 2 leaves the price unchanged. If Location 1 sales rise 8%, that initially looks positive. But if Location 2 sales also rise 7% during the same period, much of Location 1’s growth may reflect normal business conditions rather than the price decision.
Location 1 — change group
The price changed here.
Location 2 — comparison group
The price remained unchanged here.
RolloutGrade compares how both changed over the same period rather than looking only at the location where the change happened.
Not every location is a good comparison.
A useful comparison location should behave similarly to the change location before the change. Factors that may be considered when evaluating comparison quality:
- Transaction volume
- Historical sales patterns
- Product mix
- Daypart mix
- Location behaviour
- Seasonality
- Channel mix
- Baseline trend
- Available data coverage
A nearby location is not automatically the best comparison. The goal is similarity in business behaviour, not simply geographic distance. RolloutGrade lowers confidence when available comparison locations are materially different from the change location.
Why RolloutGrade does not rely only on before vs after.
Before-and-after comparisons are useful for showing what changed. They are weaker for determining why it changed. A store can improve after a decision because traffic increased, seasonality changed, a holiday occurred, weather changed, another offer ran, pricing changed elsewhere, or local conditions changed.
Baseline
Comparison
Change group
Together, these provide a stronger basis for estimating decision impact.
What would have happened without the decision?
This is the central question RolloutGrade is designed to answer. The alternative result cannot be directly observed because the same location cannot simultaneously receive and not receive the same change. RolloutGrade therefore uses the strongest available comparison evidence to estimate what likely would have happened without the change — the counterfactual. Counterfactual simply means: what likely would have happened if you had not made the change.
RolloutGrade looks at how the change location was behaving before the change, how comparable unaffected locations or groups behaved during the same period, and whether the observed result differs materially from that expected pattern.
This is an estimate, not an alternate reality we can directly observe. The quality of the estimate therefore affects the confidence of the verdict.
Time-period comparison
Comparing this period with a comparable prior period.
Some businesses have only one location, and some decisions are seasonal by nature. In those cases the most useful comparison is not another store — it is the equivalent period before: the same weeks last year, or the same holiday.
RolloutGrade can judge a decision against comparable locations, against a comparable prior period, or against both at once. Both together is the strongest observational design available, because it separates the decision from seasonality and from general trading conditions at the same time. It is not the strongest design that exists: a randomized assignment — a genuine holdout — is stronger still, because it needs no assumption that the two groups were comparable in the first place.
Comparison locations
Comparable prior period
Both
What makes two periods comparable?
- Similar length, so a longer window doesn't read as growth
- Similar mix of weekdays and weekends
- The same holiday or event, aligned by week rather than by date
- Both periods actually trading, at the same locations
- No major known disruption such as a closure or an outage
- Comparable operating conditions — hours, channels, menu or range
Holidays move around the calendar, so aligning by exact dates is usually wrong. RolloutGrade aligns the prior window by whole weeks so weekday and weekend patterns line up, and lets you name the event — “Thanksgiving week”, for example — so the two windows describe the same trading occasion.
Same-store only, and traffic-adjusted.
Locations that were not trading in both periods are excluded from the comparison, so opening or closing a site cannot masquerade as growth. Where transaction counts are available, results are also shown per 100 transactions: a busier season should not be mistaken for product success.
Known differences are recorded, not hidden.
When you set up a comparison you can confirm differences between the two periods — a another offer, a price change, an operating-hours change, a closure, a stockout, a local event or a concurrent business initiative. Anything you note is carried into the report as a stated limitation and reduces confidence rather than being quietly ignored.
What a year-over-year change is, and is not.
A prior-period comparison tells you how this period differs from a comparable one. On its own that is observed change, not proof of cause: a year is long enough for pricing, competition, staffing and demand to all have moved. RolloutGrade reports it as observed change, and only uses decision-impact language when the comparison evidence is strong enough to support it.
No comparison is presented as a perfect match. Every comparison in RolloutGrade is rated — strong, good, weak, or not comparable — with the specific reasons and limitations shown alongside the result.
Two ways to compare a decision.
Historical comparison. Compare the test period with an equivalent earlier period, such as this year's holiday season versus last year's holiday season.
Comparison locations. Compare a location that received the change with a similar location that did not.
Using both. When available, RolloutGrade can combine historical and store comparisons to better separate the decision from broader seasonal or business changes.
A historical file is not required to contain the same dates as the current file. Its purpose is to represent a comparable earlier period. A comparison-location file, by contrast, must cover the same calendar window as the change period.
POS data
How RolloutGrade reads POS data.
Different POS systems export data differently. RolloutGrade normalizes uploaded transaction data into a consistent structure before analysis. Depending on the available export, it may identify information such as:
- Location
- Transaction / check
- Date and time
- Item
- Category
- Quantity
- Sales
- Discounts
- Channel
- Customer identifier when available
- Product cost when supplied
Missing information remains missing. RolloutGrade does not silently invent fields that were not present in the source data. If a required field cannot be confidently interpreted, you are asked to confirm the mapping.
Before analysis, RolloutGrade checks whether the data can support the question.
- Missing store-days
- Incomplete periods
- Duplicate records
- Unidentified products
- Unavailable costs
- Missing comparison data
- Suspicious gaps
- Unexpected date ranges
Data quality problems appear as visible limitations before they affect the verdict.
Economics
Not every number means the same thing.
Observed
Calculated directly from uploaded data.
- Sales
- Units
- Transactions
- Average realized price
- Basket size
- Channel mix
- Daypart mix
Estimated
Derived using available evidence and modeling assumptions.
- Estimated cannibalization
- Expected performance
- Modeled displacement
Decision impact
Contribution remaining after direct costs, displacement and the movement estimated from the comparison.
- Comparison-adjusted contribution
- Estimated contribution beyond the comparison trend
- Contribution per 100 transactions attributable to the decision
RolloutGrade never presents an estimated metric as directly observed.
Higher sales do not always mean better economics.
Price
Did customers pay more or less per unit?
Volume
Were more or fewer units purchased?
Cost
Did the direct cost of delivering each unit change?
Contribution
After direct product cost, did the economics improve?
A price reduction may generate more orders while reducing contribution. A price increase may generate higher revenue while unit demand falls. A recipe change may produce the same sales with lower cost. RolloutGrade evaluates these outcomes separately rather than treating revenue growth as automatic success.
The economic verdict prioritizes contribution, not sales alone, when valid cost information is available.
Why RolloutGrade uses realized price.
The price printed on a menu is not always the amount customers actually pay. Discounts, promotions, modifiers, bundles, and other transaction behaviour can change the effective selling price. When supported by the uploaded data, RolloutGrade uses the average realized price from actual transactions rather than assuming the listed price tells the full story.
Displacement
A new product can sell well without growing the business.
Suppose a new matcha drink generates $20,000 in sales. That does not necessarily mean the business gained $20,000 in new demand. Some customers may simply have purchased the matcha instead of another existing drink.
RolloutGrade calls that estimated displacement. It is estimated by matching declines elsewhere in the product mix against the units the new item sold — it is not an observed customer switching from one product to another, and no report describes it as one. The older word cannibalization appears only where the calculation engine’s own labels use it, which is a comparison-adjusted design with a valid counterfactual behind the estimate.
Comparison-adjusted displacement
Estimated displacement
Directional substitution signal
There is no default displacement rate. When displacement cannot be estimated, it stays unavailable rather than becoming zero.
What basket size can — and cannot — tell you.
RolloutGrade can compare receipts containing a product with receipts that do not contain it. This can reveal useful purchasing patterns. However, a larger basket among customers buying a product does not automatically mean the product caused the larger basket. RolloutGrade therefore reports this as an observed association unless stronger evidence supports an incremental conclusion.
Evidence
Every verdict should tell you how much to trust it.
RolloutGrade does not treat every analysis as equally reliable, and it keeps four separate things separate so none of them can be mistaken for another:
- Design
- How the comparison was built: randomized, controlled, matched, historical, or descriptive.
- Evidence grade
- The engine’s own grade for the study, A through D. It is issued once by the calculation engine and displayed verbatim — nothing in the interface recalculates or upgrades it.
- Claim type
- The strongest wording the design supports: a comparison-adjusted decision estimate, observed change, directional signal, descriptive association, or historical comparison.
- Limitations
- What weakens this specific result, stated in the report rather than left for the reader to infer.
A comparison-adjusted estimate is available only when the study includes a credible comparison group, comparable prior period or randomized holdout. Without that support, RolloutGrade reports observed change, a directional signal or a descriptive association.
None of these is a statistical significance claim, and none of them is a p-value or a confidence interval. What raises or lowers evidence strength:
- Baseline coverage
- Length of the test
- Comparison quality
- Cost completeness
- Transaction detail
- Item-level coverage
- Known confounders
- Stockouts
- Promotions
- Closures
- Customer launch feedback when relevant
Missing evidence lowers the score. It does not get silently filled in.
Reports also show a store sensitivity range: the result recalculated with each changed location removed in turn. It is a robustness range, not a confidence interval. If dropping one location moves the answer across your bar, the decision is downgraded rather than presented as a rollout.
Weak
An early signal. Not enough to act on a high-stakes decision.
Directional
It points one way, but the comparison cannot size the effect reliably.
Comparison-adjusted
A credible comparison group carries the counterfactual.
Comparison-adjusted and stable
A credible comparison that also survives the pre-trend and leave-one-location-out checks.
A weak-evidence result is not a failed test. It means more evidence is needed before making a high-stakes decision.
Sometimes the correct answer is: we do not know yet.
RolloutGrade is intentionally designed to withhold claims when the available evidence cannot support them.
- No cost data
- Gross contribution cannot be calculated.
- No valid comparison group
- Observed performance can be reported, but causal decision impact cannot yet be estimated credibly.
- Incomplete item data
- Displacement may be directional rather than measured.
- Short test
- Early performance may not represent durable demand.
RolloutGrade never turns missing evidence into a zero, an assumption presented as fact, or an artificially confident verdict. Survey responses and location reviews provide explanatory context. They never replace transaction economics or independently upgrade the evidence grade.
Auditability
Can I audit my result?
Yes. RolloutGrade is designed so the evidence behind a verdict can be inspected without exposing proprietary modeling. Every report lets you inspect:
- Source data
- What files and periods were analyzed?
- Test definition
- What decision was evaluated?
- Baseline
- What period represented normal behaviour before the decision?
- Comparison
- What location or group was used for comparison?
- Inputs
- What prices, costs, thresholds, and assumptions were provided?
- Observations
- What directly changed?
- Estimates
- Which results required modeling?
- Limitations
- What data was missing or potentially confounded?
- Verdict
- What evidence caused RolloutGrade to recommend Roll out, Modify, Do not roll out, or Continue testing?
Each report carries a “Why you can trust this result” panel summarising the data used, what was observed, what was estimated, what is missing, and how strong the evidence is — populated from that analysis, never from an example.
Data & privacy
Your operating data is sensitive.
What data is needed
What data is not needed
How data is used
Access
Deletion
Security
See the full privacy policy.
About the methodology
Built for operating decisions.
The method is designed for product launches, pricing changes, menu and assortment decisions and location-level operating changes where transaction and cost data can support a comparison.
It applies familiar measurement discipline to those decisions: clearly defining the change, separating the change and comparison groups, establishing a baseline, distinguishing observation from inference, identifying confounding factors, and communicating uncertainty.
Method boundary
RolloutGrade evaluates changes made inside a business using first-party operational records supplied by that business. It does not observe advertising exposure, track people across platforms, evaluate media channels or measure brand attitudes. Customer surveys and location reviews provide explanatory context and never replace transaction economics.
Why this approach?
Businesses constantly make decisions with incomplete information. A disciplined comparison is what separates signal from noise well enough to make a better decision. RolloutGrade applies that discipline to transaction and cost data.
Customer launch feedback
Where a study collects it, customer launch feedback records new-versus-returning status, self-reported visit influence, intended substitution, repeat interest, product satisfaction, price perception and reasons for purchasing. Survey responses and location reviews provide explanatory context. They never replace transaction economics or independently upgrade the evidence grade.
- Define the question
- What exactly changed?
- Define the change group
- Which part of the business received the change?
- Define the comparison
- What can help represent what would have happened otherwise?
- Measure the outcome
- What actually changed?
- Identify alternative explanations
- What else could have caused the result?
- Grade the evidence
- How much confidence should be placed in the conclusion?
- Make the decision
- Is the evidence strong enough to act?
Plain English
Every word we use, explained once
You don't need a statistics background to read a RolloutGrade report. These are the only terms we use, and what each one means in ordinary language.
- Comparison location
- Another one of your locations that did NOT get the change, watched over the same weeks.
- If both moved the same way, the change probably wasn't the cause. If only the launch location moved, it probably was.
- Independent control
- A comparison location that carried on as normal while the change ran somewhere else.
- This is the strongest evidence we can build without a lab.
- Same weeks last year
- The same calendar weeks from a previous year at the same location.
- Good for stripping out season, but it can't rule out things that only happened this year.
- Internal benchmark
- Other parts of the same store that the change shouldn't have touched, used as a stand-in for 'business as usual'.
- Directional only — it shows a gap, not proof the change caused it.
- What would have happened anyway
- Our estimate of the sales you'd have made if you had never made the change.
- Everything above that line is what we credit to the change.
- Extra sales
- Sales above what would have happened anyway.
- Not the same as total sales of the new item.
- Estimated displacement
- Sales the new thing appears to have pulled away from items you were already selling, estimated by matching declines elsewhere.
- A launch can look great on its own and still leave the store flat once this is counted. It is an estimate, not an observed customer switching products.
- Gross contribution
- Sales minus the direct cost of the goods you sold — the money left over.
- This is the number that decides whether the change was worth it. It is not profit: it carries no rent, labour or overhead.
- Evidence grade
- A to D on how good the evidence is — how much data, how clean the comparison, how steady the pattern. The engine issues it; nothing else recalculates it.
- A high result on weak evidence is still a weak result.
- Evidence strength
- How much weight the result can carry, from the design, the coverage and the stability of the pattern.
- Weak evidence means run it longer or add a comparison. It is not a statistical significance test and carries no p-value.
- Pre-trend check
- We check the two locations already moved together BEFORE the change.
- If they didn't, the comparison can't be trusted.
- Lift
- How much higher (or lower) things were than the comparison.
- Baseline
- The weeks before the change, used as your starting point.
- Measurement window
- The days from the change onwards that we score.
- Checks
- Individual customer transactions — one receipt is one check.
FAQ
Common questions
Where to go next
Method boundary
RolloutGrade evaluates changes made inside a business using first-party operational records supplied by that business. It does not observe advertising exposure, track people across platforms, evaluate media channels or measure brand attitudes.
- Test Readiness ScanSix questions. See what your setup can support before you upload anything.
- Sample decisionA product launch, read end to end.
- Data requirementsExactly which fields a transaction export needs.
- Privacy policyWhat is stored, for how long, and what is never required.