Skip to content

What a product-launch decision looks like

Start with the operating action, then work backward through the observed result, comparison-location movement, contribution economics and the limits of the evidence.

Illustrative sample — not a real customer result

Product launch

Strawberry Cloud Matcha

Roll outControlled designEstimated incrementalEngine grade controlled
  1. 1

    The decision

    Roll out

    Roll out. Take Strawberry Cloud Matcha to 9 further locations — the 9 locations whose category mix most closely matches the test group, in three waves of three.

    Why this matters
    A study that ends in a number ends in an argument. A study that ends in an action ends in a decision you can hold someone to.
    How this was calculated
    The engine picks the action by strict precedence: reconciliation first, then evidence strength, then guardrails, then economics. Nothing later can override something earlier.
    Required data
    A transaction export covering both periods, unit costs, and the success rule locked before the result was visible.
    Important limitation
    The action follows from the locked success rule. Change the rule and the action can change with it.
  2. 2

    The business hypothesis

    Adding Strawberry Cloud Matcha will add gross contribution in the launch locations without displacing too much of the existing range or creating too much waste.

    • Locked success rule: at least $12.00 of contribution beyond the comparison trend per 100 transactions
    • Guardrail: displacement of existing sales stays below 60%
    • Guardrail: waste stays below 8% of units made
    Why this matters
    A success rule written after the result is visible is not a rule. Locking it first is what makes the decision reviewable later.
    How this was calculated
    Nothing calculates the hypothesis. It is stated at setup; the engine only checks the result against it.
    Required data
    A stated economic result, a locked success rule, and the guardrails.
    Important limitation
    The guardrails and the success rule are business judgements. RolloutGrade tests the change against them; it does not set them.
  3. 3

    What happened during the test

    +$4,547

    The launch locations sold 1,284 units of Strawberry Cloud Matcha across 6 weeks in 3 locations, leaving +$4,547 of gross contribution after waste and displaced sales.

    • 1,284 units sold, on 7.1% of eligible transactions
    • +$32.99 gross contribution per 100 transactions before deductions
    • $267 of retail value wasted, at $76 of recipe cost
    Why this matters
    This is the number an ordinary sales report would stop at. It is real, and on its own it cannot tell you whether the business is better off.
    How this was calculated
    Units sold times contribution per unit, less waste at recipe cost, less the contribution of the existing lines that lost volume in the same locations over the same weeks.
    Required data
    Item-level transaction rows and a unit cost for every product in scope.
    Important limitation
    This is gross contribution, not profit. It carries no rent, labour or overhead.
  4. 4

    What the comparison indicates

    $1,224

    The 2 comparison locations, which never received the launch, grew over the same weeks. Applied to the launch locations, that growth accounts for $1,224 of the movement.

    • 2 untreated comparison locations, matched on category mix and trading pattern
    • Measured over the same 6 weeks, on the same category
    Why this matters
    Without this line, every seasonal upswing looks like a successful launch. This is the line that separates the two.
    How this was calculated
    The comparison locations' contribution per transaction is compared between the baseline and the change period; that movement is applied to the launch locations' transaction volume.
    Required data
    Untreated comparison locations with complete coverage in both the baseline and the change period.
    Important limitation
    A comparison location is only as good as how closely it tracked the launch locations beforehand. This study has only two of them, below the three preferred.
  5. 5

    What remained after the comparison

    +$18.40

    Subtracting the movement seen at comparison locations leaves an estimated +$3,323 of contribution beyond the comparison trend — +$18.40 per 100 transactions against the locked $12.00 success rule.

    • $4,547 observed − $1,224 expected anyway = $3,323
    • 53.3% above the locked $12.00 threshold
    • 13.6% unit movement relative to comparison locations
    Why this matters
    This is the figure the rollout decision turns on: how much contribution remained after costs, displaced sales and the movement seen at comparison locations.
    How this was calculated
    Difference-in-differences on contribution per transaction: the launch locations' change, less the comparison locations' change over the same dates.
    Required data
    Both periods, both location groups, complete cost coverage, and a comparison that already tracked the launch locations.
    Important limitation
    This is an estimate, not a directly observed result. It depends on the untreated comparison locations providing a credible estimate of ordinary business movement.
  6. 6

    Why the result changed

    The item earns $4.64 per unit, but an estimated 29.5% of its units displaced something the customer already bought, and the comparison locations were growing anyway. Both are subtracted before the result stands.

    • +$32.99 gross contribution per 100 transactions
    • −$0.42 waste at recipe cost
    • −$7.39 estimated displacement (379 of 1,284 units)
    • −$6.78 growth the comparison locations show would have happened anyway
    • = +$18.40 estimated contribution beyond comparison trend per 100 transactions
    Why this matters
    A single headline figure hides which lever moved. These four lines are the levers, and each one is separately actionable.
    How this was calculated
    Each line is the same calculation as the headline, expressed per 100 eligible transactions, so the four steps sum to the headline exactly.
    Required data
    Item-level rows for the products that could have lost volume, plus unit costs for each of them.
    Important limitation
    Displacement is estimated by matching declines in other lines. It is not an observed customer switching from one product to another.
  7. 7

    How strong the decision evidence is

    controlled

    A controlled design with 2 untreated comparison locations. The engine grades this study "controlled". The comparison design is strong enough to support a comparison-adjusted estimate rather than only an observed before-and-after change.

    • Grade "controlled" is the engine's own word for a design strong enough to carry a comparison-adjusted read. It is a design-quality grade, not a significance test.
    • only 2 comparison locations (3+ preferred)
    • 6 weeks observed (8+ preferred for a full cycle)
    • 2 stockout day(s) in the test window
    • 2 noted confounder(s) mitigated but not eliminated
    Why this matters
    A strong number on weak evidence is still a weak result. The grade is what stops one being mistaken for the other.
    How this was calculated
    The engine penalises missing comparison locations, missing costs, incomplete coverage, missing trading days and unresolved confounders, then reports the design's own weaknesses separately as cautions.
    Required data
    Nothing extra — the grade is computed from the same upload as the result.
    Important limitation
    This is an evidence grade, not a statistical significance claim. It carries no p-value and no confidence interval.
  8. 8

    What would reverse the decision

    Illustrative example written for this sample—not calculated by RolloutGrade.

    Comparison-adjusted contribution per 100 transactions falls below the $12 threshold for two consecutive weeks.

    • Cannibalization rises above the 60% guardrail.
    • Waste rises above the 8% guardrail.
    • Displaced items fail to recover once novelty demand settles.
    Why this matters
    A decision with no stated reversal condition cannot be revisited honestly. This is what makes it falsifiable.
    Where this came from
    Illustrative example written for this sample—not calculated by RolloutGrade.Each condition restates a guardrail this study locked before the result was visible — the success rule, the displacement ceiling and the waste ceiling. RolloutGrade does not compute a break-even or reversal threshold, so nothing here is a calculated or recommended threshold.Nothing calculated these. Each condition is written prose restating a guardrail this study locked before the result was visible — the success rule, the displacement ceiling and the waste ceiling. RolloutGrade computes no break-even or reversal threshold, so none of this is a calculated break-even point or a recommended threshold.
    Required data
    Continued transaction and cost data from the rolled-out locations.
    Important limitation
    These conditions were locked against this study's design, and they are authored rather than system-generated. A wider rollout changes the population, and the guardrails should be re-locked with it.
  9. 9

    What to test after rollout

    Re-read at 30, 60 and 90 days against the same $12.00 per 100 transactions rule, using the same comparison locations.

    • Day 30 — confirm comparison-adjusted contribution per 100 transactions is still at or above $12.00, and that novelty demand has settled rather than collapsed.
    • Day 60 — re-check estimated displacement against the 60% guardrail and confirm the displaced lines have recovered.
    • Day 90 — re-check waste against the 8% guardrail and re-run the full comparison-adjusted read across the rolled-out locations.
    Why this matters
    Most launch effects decay. A decision that is never re-read is a decision that quietly stops being true.
    How this was calculated
    Each checkpoint re-runs the same calculation on new dates. Nothing about the method changes between reads.
    Required data
    Continued uploads covering both the rolled-out locations and the comparison locations.
    Important limitation
    Once the change reaches every location, the comparison group disappears and later reads fall back to a historical comparison.

Important limitations

  • Two stockout days at Harbor St excluded from the comparison calculation.
  • One comparison location ran a loyalty promotion in week 4; that week is weighted down.
  • Cannibalization is measured on category units, not individual customer switching.
  • Six weeks captures novelty decay but not a full seasonal cycle.
  • only 2 comparison locations (3+ preferred)
  • 6 weeks observed (8+ preferred for a full cycle)
  • 2 stockout day(s) in the test window
  • 2 noted confounder(s) mitigated but not eliminated

Why this design. A comparable group did not receive the change over the same dates. Strength depends on the two groups already moving together beforehand.

Every study answers the same five questions

  1. 1

    What decision was tested?

  2. 2

    What happened during the test?

  3. 3

    What does the comparison suggest would have happened anyway?

  4. 4

    How strong an answer can the design support?

  5. 5

    What should management do next?

What this sample evaluates

This sample evaluates a product launch using first-party location transactions, item costs and untreated comparison locations. It does not evaluate advertising, media activity, audiences or cross-platform behavior.