Product launch
Strawberry Cloud Matcha
1
The decision
Roll out
Roll out. Take Strawberry Cloud Matcha to 9 further locations — the 9 locations whose category mix most closely matches the test group, in three waves of three.
Why this matters
A study that ends in a number ends in an argument. A study that ends in an action ends in a decision you can hold someone to.How this was calculated
The engine picks the action by strict precedence: reconciliation first, then evidence strength, then guardrails, then economics. Nothing later can override something earlier.Required data
A transaction export covering both periods, unit costs, and the success rule locked before the result was visible.Important limitation
The action follows from the locked success rule. Change the rule and the action can change with it.2
The business hypothesis
Adding Strawberry Cloud Matcha will add gross contribution in the launch locations without displacing too much of the existing range or creating too much waste.
- Locked success rule: at least $12.00 of contribution beyond the comparison trend per 100 transactions
- Guardrail: displacement of existing sales stays below 60%
- Guardrail: waste stays below 8% of units made
Why this matters
A success rule written after the result is visible is not a rule. Locking it first is what makes the decision reviewable later.How this was calculated
Nothing calculates the hypothesis. It is stated at setup; the engine only checks the result against it.Required data
A stated economic result, a locked success rule, and the guardrails.Important limitation
The guardrails and the success rule are business judgements. RolloutGrade tests the change against them; it does not set them.3
What happened during the test
+$4,547
The launch locations sold 1,284 units of Strawberry Cloud Matcha across 6 weeks in 3 locations, leaving +$4,547 of gross contribution after waste and displaced sales.
- 1,284 units sold, on 7.1% of eligible transactions
- +$32.99 gross contribution per 100 transactions before deductions
- $267 of retail value wasted, at $76 of recipe cost
Why this matters
This is the number an ordinary sales report would stop at. It is real, and on its own it cannot tell you whether the business is better off.How this was calculated
Units sold times contribution per unit, less waste at recipe cost, less the contribution of the existing lines that lost volume in the same locations over the same weeks.Required data
Item-level transaction rows and a unit cost for every product in scope.Important limitation
This is gross contribution, not profit. It carries no rent, labour or overhead.4
What the comparison indicates
$1,224
The 2 comparison locations, which never received the launch, grew over the same weeks. Applied to the launch locations, that growth accounts for $1,224 of the movement.
- 2 untreated comparison locations, matched on category mix and trading pattern
- Measured over the same 6 weeks, on the same category
Why this matters
Without this line, every seasonal upswing looks like a successful launch. This is the line that separates the two.How this was calculated
The comparison locations' contribution per transaction is compared between the baseline and the change period; that movement is applied to the launch locations' transaction volume.Required data
Untreated comparison locations with complete coverage in both the baseline and the change period.Important limitation
A comparison location is only as good as how closely it tracked the launch locations beforehand. This study has only two of them, below the three preferred.5
What remained after the comparison
+$18.40
Subtracting the movement seen at comparison locations leaves an estimated +$3,323 of contribution beyond the comparison trend — +$18.40 per 100 transactions against the locked $12.00 success rule.
- $4,547 observed − $1,224 expected anyway = $3,323
- 53.3% above the locked $12.00 threshold
- 13.6% unit movement relative to comparison locations
Why this matters
This is the figure the rollout decision turns on: how much contribution remained after costs, displaced sales and the movement seen at comparison locations.How this was calculated
Difference-in-differences on contribution per transaction: the launch locations' change, less the comparison locations' change over the same dates.Required data
Both periods, both location groups, complete cost coverage, and a comparison that already tracked the launch locations.Important limitation
This is an estimate, not a directly observed result. It depends on the untreated comparison locations providing a credible estimate of ordinary business movement.6
Why the result changed
The item earns $4.64 per unit, but an estimated 29.5% of its units displaced something the customer already bought, and the comparison locations were growing anyway. Both are subtracted before the result stands.
- +$32.99 gross contribution per 100 transactions
- −$0.42 waste at recipe cost
- −$7.39 estimated displacement (379 of 1,284 units)
- −$6.78 growth the comparison locations show would have happened anyway
- = +$18.40 estimated contribution beyond comparison trend per 100 transactions
Why this matters
A single headline figure hides which lever moved. These four lines are the levers, and each one is separately actionable.How this was calculated
Each line is the same calculation as the headline, expressed per 100 eligible transactions, so the four steps sum to the headline exactly.Required data
Item-level rows for the products that could have lost volume, plus unit costs for each of them.Important limitation
Displacement is estimated by matching declines in other lines. It is not an observed customer switching from one product to another.7
How strong the decision evidence is
controlled
A controlled design with 2 untreated comparison locations. The engine grades this study "controlled". The comparison design is strong enough to support a comparison-adjusted estimate rather than only an observed before-and-after change.
- Grade "controlled" is the engine's own word for a design strong enough to carry a comparison-adjusted read. It is a design-quality grade, not a significance test.
- only 2 comparison locations (3+ preferred)
- 6 weeks observed (8+ preferred for a full cycle)
- 2 stockout day(s) in the test window
- 2 noted confounder(s) mitigated but not eliminated
Why this matters
A strong number on weak evidence is still a weak result. The grade is what stops one being mistaken for the other.How this was calculated
The engine penalises missing comparison locations, missing costs, incomplete coverage, missing trading days and unresolved confounders, then reports the design's own weaknesses separately as cautions.Required data
Nothing extra — the grade is computed from the same upload as the result.Important limitation
This is an evidence grade, not a statistical significance claim. It carries no p-value and no confidence interval.8
What would reverse the decision
Illustrative example written for this sample—not calculated by RolloutGrade.
Comparison-adjusted contribution per 100 transactions falls below the $12 threshold for two consecutive weeks.
- Cannibalization rises above the 60% guardrail.
- Waste rises above the 8% guardrail.
- Displaced items fail to recover once novelty demand settles.
Why this matters
A decision with no stated reversal condition cannot be revisited honestly. This is what makes it falsifiable.Where this came from
Illustrative example written for this sample—not calculated by RolloutGrade.Each condition restates a guardrail this study locked before the result was visible — the success rule, the displacement ceiling and the waste ceiling. RolloutGrade does not compute a break-even or reversal threshold, so nothing here is a calculated or recommended threshold.Nothing calculated these. Each condition is written prose restating a guardrail this study locked before the result was visible — the success rule, the displacement ceiling and the waste ceiling. RolloutGrade computes no break-even or reversal threshold, so none of this is a calculated break-even point or a recommended threshold.Required data
Continued transaction and cost data from the rolled-out locations.Important limitation
These conditions were locked against this study's design, and they are authored rather than system-generated. A wider rollout changes the population, and the guardrails should be re-locked with it.9
What to test after rollout
Re-read at 30, 60 and 90 days against the same $12.00 per 100 transactions rule, using the same comparison locations.
- Day 30 — confirm comparison-adjusted contribution per 100 transactions is still at or above $12.00, and that novelty demand has settled rather than collapsed.
- Day 60 — re-check estimated displacement against the 60% guardrail and confirm the displaced lines have recovered.
- Day 90 — re-check waste against the 8% guardrail and re-run the full comparison-adjusted read across the rolled-out locations.
Why this matters
Most launch effects decay. A decision that is never re-read is a decision that quietly stops being true.How this was calculated
Each checkpoint re-runs the same calculation on new dates. Nothing about the method changes between reads.Required data
Continued uploads covering both the rolled-out locations and the comparison locations.Important limitation
Once the change reaches every location, the comparison group disappears and later reads fall back to a historical comparison.
Important limitations
- Two stockout days at Harbor St excluded from the comparison calculation.
- One comparison location ran a loyalty promotion in week 4; that week is weighted down.
- Cannibalization is measured on category units, not individual customer switching.
- Six weeks captures novelty decay but not a full seasonal cycle.
- only 2 comparison locations (3+ preferred)
- 6 weeks observed (8+ preferred for a full cycle)
- 2 stockout day(s) in the test window
- 2 noted confounder(s) mitigated but not eliminated
Why this design. A comparable group did not receive the change over the same dates. Strength depends on the two groups already moving together beforehand.