In A/B testing, speed and accuracy are a balancing act - wait too long and you lose momentum; move too fast and you risk chasing false positives. CUPED (Controlled Pre-Experiment Data) tilts the balance in your favour: it reduces variance in your results by factoring in shopper behaviour before they're exposed to the test, so a real effect reaches significance sooner. PDQ adds a dynamic, per-shop optimization layer on top for maximum variance reduction without overfitting.
Lower variance makes the same lift easier to detect - so a real effect reaches significance sooner.
PDQ's "dynamic-outlier" mode
Unlike standard CUPED (a fixed covariate and static trimming rules), PDQ optimizes the CUPED configuration per shop:
Pre-experiment window: 60 days of recent pre-test data.
Outlier-removal options tested per shop: trim at the 99th percentile (TP99%), the 99.9th percentile (TP99.9%), or beyond 3 standard deviations (3SD).
Rule selection: choose the rule yielding the highest correlation (θ) between the pre-experiment covariate and the outcome, within guardrails.
Qualification: apply CUPED only if θ > 10%.
Lock-in: store the selected rule and θ before the test starts, for reproducibility.
This adaptive approach maximizes variance reduction while minimizing the risk of overfitting or unstable adjustments.
The pre-experiment covariate
Initial checkout subtotal before discounts, captured the split-second checkout loads (Shopify's first payload with the cart value).
Why it's valid: recorded before any test exposure - a true pre-treatment measurement.
PDQ's advantage: because it's collected just before assignment, it's available for both first-time and returning shoppers - 100% of traffic - which many CUPED setups can't do. More data = more statistical power.
Why it's powerful: strongly correlated with ARPC for converters.
Challenges: zero-inflation (many sessions have a subtotal but zero ARPC due to abandonment) and a mixture distribution (converters vs non-converters), producing a nonlinear relationship.
The adjustment (in plain terms)
CUPED subtracts the part of each shopper's outcome that their pre-experiment cart value already predicted, leaving a lower-variance signal. The parameter θ quantifies how much variance is removed - you can request your shop's θ from your CSM; excluding outliers can sometimes double it, indicating substantially more noise control.
Business impact
Speed: up to ~15 days saved on sequential-bound tests.
Efficiency: cuts idea-to-decision cycles nearly in half, expanding experimentation bandwidth.
Rigor: sequential α-spending keeps Type-I error controlled even with variance reduction.
Good to know / caveats
Backend-computed: CUPED adjustments, dynamic outlier removal, and sequential bounds run in the stats-engine backend - raw exports alone can't reproduce official results.
Outlier definition differs by view: the Test Results dashboard trims on initial cart value; the legacy Deep A/B Test Analysis trims on final order subtotal, which can cause different numbers. Use Test Results for authoritative CUPED figures.
Deep A/B Test Analysis is exploratory slicing only - it does not provide statistical significance.
Significance on ARPC but not its components: ARPC = conversion × order value, so two small, individually non-significant shifts can combine (with their covariance) into a significant ARPC change - especially after variance reduction.
Future direction
CUPED V2: multi-covariate CUPED to capture even more variance.
CUPED isn't about changing your results - it's about revealing them faster and with greater clarity. For merchants running multiple tests per quarter, the time savings can be game-changing.
Next step: to enable CUPED for upcoming tests or review your shop's configuration, reach out to your PDQ CSM.
