Methods
How we measure what a decision delivered.
Every estimate on ActionBasis comes with a range and a label that says how we know it. Every claimed result comes from a test agreed with you before it starts.
Reproducible measurement check
Can the readout distinguish lift, loss and no change?
An existing synthetic test provides the answer before measurement. Recalculate the readout from its inputs, including the cases it gets wrong. This checks one estimator; it is not a customer result or validation of a whole company model.
Known lift
20 of 20 intervals include the injected +10% change.
Known loss
18 of 20 intervals include the injected −15% change. Two miss it.
No change
39 of 40 runs detect no effect. One produces a false positive.
| Example | Known change | Estimate | 95% interval | Readout |
|---|---|---|---|---|
| Known liftpositive-1 | +10.00% | +10.24% | +9.24% to +11.25% | Change detected |
| Known lossnegative-501 | -15.00% | -14.63% | -15.35% to -13.90% | Change detected |
| No injected changenull-1001 | 0.00% | -0.05% | -0.91% to +0.81% | No effect detected |
| Interval misses truthnegative-503 | -15.00% | -14.17% | -14.99% to -13.35% | Does not recover known truth |
| False positivenull-1029 | 0.00% | +1.02% | +0.15% to +1.90% | Detects an effect that was not injected |
The published readouts were replayed from the exported inputs during the build.
Inspect the design, limits and replay instructions
Each run has 40 synthetic units in each arm, 13 weeks before and 13 after. Revenue is generated with independent store-week noise. The estimator compares each unit’s period totals with its own baseline, then compares the two arms. Its 95% normal-approximation interval uses unequal store-level variances.
The injected change is held apart from the inputs used by the estimator. These are seeded test cases, not held-out real-company data. This fixture does not test confounding, nonparallel trends, selection, spillovers, missing records or all forms of dependence. An interval can miss known truth, and a null case can appear significant; both failures remain in the download.
A detected effect does not authorize rollout. Costs, field design, comparable controls and live measurement are still required. No profit, payback or annual benefit is claimed. Fewer than five controls withholds the readout.
The download includes the period totals, native readouts, source hashes and an independent replay of the published estimate and interval from period totals. It does not regenerate the synthetic panel or replay the pre-trend checks. Unzip it and run node replay.mjs with Node.js 18 or later. It requires no installation, network request or company data.
Receipt SHA-256: a2d78eb12c02716d60622138588d81cfa5cd6d0acc19dc411fc682c68fe9fabb
Five rules
What we hold ourselves to.
Every estimate is a range
A decision is shown with the range of outcomes it is likely to produce. A wide range is a reason to test before you commit, and we say so.
Results come from matched controls
A change goes first to pilot stores, regions or customers. Control groups are matched to them on size, location, demand and recent trend. The effect is the difference between the two, reported with its interval.
The test is agreed before it starts
Before a pilot begins, you and we agree in writing what is measured, against which controls, for how long, and what result would count as a success.
Known answers first
We test methods on synthetic cases with known answers and retain failures. The checks below show interval misses and a false positive. Passing these checks alone does not qualify a method for your question.
We say when the data cannot answer
Some questions cannot be settled from history. We say which ones, and propose the test that would settle them, with its size and duration.
Worked example
The Saturday-hours pilot, step by step.
The numbers below are a fictional example for a specialty retailer with about 1,800 stores, separate from the downloadable measurement fixture.
- The question
- Does opening an hour longer on Saturdays add sales once the extra staff cost is paid?
- Agreed before the start
- 60 pilot stores and 60 control stores matched on sales, size, region and the last 12 weeks' trend. A 12-week period. Success means higher like-for-like sales net of the added staff cost.
- Check before the change
- In this fictional example, the groups followed similar trends over the previous 12 weeks. That is a diagnostic for the parallel-trends assumption, not proof that it holds after the change.
- Measured result
- Pilot stores sold 0.7% more than their controls, with an interval of +0.3% to +1.1%.
- Sales-equivalent cost adjustment
- Subtracting staff cost expressed as 0.2% of sales gives +0.1% to +0.9% on that same base. This is not an incremental profit interval: product margin, other added costs and displaced sales must also be accounted for.
- What changed next
- The measured range replaced the modelled one (+0.4% to +1.6%), and the next decisions were re-ranked.
Evidence labels
Every number carries its evidence.
Five labels separate what was recorded from what was calculated, estimated, simulated or set by a person. A simulated number is never shown as observed, and missing data shows as a gap.
- Observed
- Recorded in a source we can point to: a filing, a review, your sales data.
- Derived
- Calculated from observed values, with the calculation kept.
- Modelled
- Estimated by a fitted model and shown with its range.
- Simulated
- Produced by running a scenario forward. Kept apart from observed data.
- Assumed
- An input someone chose. Visible, editable and named in every result that uses it.
Measured value
What counts as value.
Value counts once a pilot or rollout has been measured against its controls, over the period agreed in advance, net of the costs the change added. Until then it stays in the value ledger as modelled value. Once measured, the ledger shows both figures side by side, and the ranking uses the measured one.
See the method on your own company.
Start with a briefing from public data. We send it to your work email.