Skip to main content

Controlled Pre-Experiment Data (CUPED) and how we use it

How PDQ sets parameters & prepares data for tests

In A/B testing, speed and accuracy are a balancing act - wait too long and you lose momentum; move too fast and you risk chasing false positives. CUPED (Controlled Pre-Experiment Data) tilts the balance in your favour: it reduces variance in your results by factoring in shopper behaviour before they're exposed to the test, so a real effect reaches significance sooner. PDQ adds a dynamic, per-shop optimization layer on top for maximum variance reduction without overfitting.

Lower variance makes the same lift easier to detect - so a real effect reaches significance sooner.

PDQ's "dynamic-outlier" mode

Unlike standard CUPED (a fixed covariate and static trimming rules), PDQ optimizes the CUPED configuration per shop:

  1. Pre-experiment window: 60 days of recent pre-test data.

  2. Outlier-removal options tested per shop: trim at the 99th percentile (TP99%), the 99.9th percentile (TP99.9%), or beyond 3 standard deviations (3SD).

  3. Rule selection: choose the rule yielding the highest correlation (θ) between the pre-experiment covariate and the outcome, within guardrails.

  4. Qualification: apply CUPED only if θ > 10%.

  5. Lock-in: store the selected rule and θ before the test starts, for reproducibility.

This adaptive approach maximizes variance reduction while minimizing the risk of overfitting or unstable adjustments.
​

The pre-experiment covariate

Initial checkout subtotal before discounts, captured the split-second checkout loads (Shopify's first payload with the cart value).

  • Why it's valid: recorded before any test exposure - a true pre-treatment measurement.

  • PDQ's advantage: because it's collected just before assignment, it's available for both first-time and returning shoppers - 100% of traffic - which many CUPED setups can't do. More data = more statistical power.

  • Why it's powerful: strongly correlated with ARPC for converters.

  • Challenges: zero-inflation (many sessions have a subtotal but zero ARPC due to abandonment) and a mixture distribution (converters vs non-converters), producing a nonlinear relationship.

The adjustment (in plain terms)

CUPED subtracts the part of each shopper's outcome that their pre-experiment cart value already predicted, leaving a lower-variance signal. The parameter θ quantifies how much variance is removed - you can request your shop's θ from your CSM; excluding outliers can sometimes double it, indicating substantially more noise control.
​

Business impact

  • Speed: up to ~15 days saved on sequential-bound tests.

  • Efficiency: cuts idea-to-decision cycles nearly in half, expanding experimentation bandwidth.

  • Rigor: sequential α-spending keeps Type-I error controlled even with variance reduction.

Good to know / caveats

  • Backend-computed: CUPED adjustments, dynamic outlier removal, and sequential bounds run in the stats-engine backend - raw exports alone can't reproduce official results.

  • Outlier definition differs by view: the Test Results dashboard trims on initial cart value; the legacy Deep A/B Test Analysis trims on final order subtotal, which can cause different numbers. Use Test Results for authoritative CUPED figures.

  • Deep A/B Test Analysis is exploratory slicing only - it does not provide statistical significance.

  • Significance on ARPC but not its components: ARPC = conversion × order value, so two small, individually non-significant shifts can combine (with their covariance) into a significant ARPC change - especially after variance reduction.

Future direction

  • CUPED V2: multi-covariate CUPED to capture even more variance.

CUPED isn't about changing your results - it's about revealing them faster and with greater clarity. For merchants running multiple tests per quarter, the time savings can be game-changing.
​

Next step: to enable CUPED for upcoming tests or review your shop's configuration, reach out to your PDQ CSM.

Did this answer your question?