Korant

Nine D2C experiments you can run this week

What the numbers say

  1. 01

    Detecting a 20% relative change in order conversion at a 2.5% baseline needs 15,288 sessions per arm, or 25 days at 36,000 monthly sessions.

    Two-proportion sample size at 80% power and 5% two-sided significance.

  2. 02

    The same 20% change in checkout completion at a 30% baseline needs 915 sessions per arm, 16.7 times fewer, or 2 days.

    Two-proportion sample size at 80% power and 5% two-sided significance.

  3. 03

    Halving the target to a 10% relative change quadruples the requirement, taking order conversion from 25 days to 102.

    Arithmetic on the same formula, where sample size scales with the inverse square of the effect.

  4. 04

    Testing high-baseline metrics rather than order conversion is the single change that makes D2C experimentation feasible below 50,000 sessions a month.

    Derived from the 16.7 times difference in sample requirement between the two baselines.

The experiment that ran for six weeks and proved nothing

The team changed the hero image and split traffic. Six weeks later the variant was ahead by 4%, which sounded like a win until somebody asked whether it was significant.

It was not. At that traffic level and that conversion rate, a 4% difference was well inside what randomness produces on its own, and the test would have needed about four more months to say anything.

Six weeks of traffic, spent, with no answer at either end.

The mistake was not the hypothesis. The hero image might well matter. The mistake was running a test whose answer was unreachable before it started, which is knowable in thirty seconds.

Every experiment below is paired with the number that decides whether it is worth your traffic.

Work out the sample size before the idea

The formula for comparing two proportions at 80% power and 5% two-sided significance is roughly 15.7 times the baseline rate times one minus the baseline, divided by the square of the absolute effect you want to detect, per arm.

Two consequences fall out of it, and both are counterintuitive.

Low baseline rates are expensive. A 2.5% order conversion needs far more traffic than a 30% checkout completion for the same relative change.

Smaller effects are punishingly expensive. Halving the effect you want to detect quadruples the sample, because the effect is squared in the denominator.

MetricBaselineSessions per arm for +20%Days at 36,000 monthlyFor +10%Days
Order conversion2.5%15,2882561,152102
Email click rate3.0%12,6752150,69984
Add to cart rate8.0%4,508818,03230
Referral share rate12.0%2,875511,49919
Checkout completion30.0%91523,6596

Checkout completion needs 16.7 times fewer sessions than order conversion for the same relative change.

That single ratio is why most D2C experimentation stalls. Teams test the metric they care about most, which is the one they can least afford to measure.

Test high-baseline metrics on the way to the outcome instead. A change that lifts checkout completion by 20% will lift order conversion too, and you will know about it in two days rather than twenty-five.

Three experiments you can read in under a week

1. Collapse the discount code field. Hide the coupon input behind a link rather than showing an empty box. Metric: checkout completion. Sample: 915 per arm, about 2 days.

2. Add a per-pincode delivery estimate to the product page. Read the visitor’s saved postcode and show that postcode’s transit time instead of a national promise. Metric: add to cart rate. Sample: 4,508 per arm, about 8 days.

3. Move the referral prompt to the thank-you page. Take the share prompt out of the follow-up email and put it where enthusiasm peaks. Metric: referral share rate. Sample: 2,875 per arm, about 5 days.

SB&R is a Shopify app for chained referral rewards. Every referral link belongs to someone who has already bought. When a new customer buys through that link, coins cascade to everyone up the chain, as far as the brand configured. Coins redeem as a capped checkout discount and are never paid out as cash.

All three test high-baseline metrics, which is why they are readable inside a week. None of them requires new traffic or new spend.

Three that need a fortnight

4. Raise chain depth from one to three. Configuration change. Metric: referrals per customer, measured as a rate, so around 5 days on the same baseline as experiment 3, plus a week for the cascade to actually land.

5. Change the free shipping threshold by Rs 200 in one direction. Metric: average order value, which is a continuous measure and needs a different calculation from the table above. Plan on two weeks and check the variance of your order values first, since a wide spread needs far more orders.

6. Run one live zone window against a holdout postcode. Metric: orders in the live zone against its own previous four weeks.

FlashPin is a multi-tenant Shopify app that rotates which delivery pincode has a live discount on a cadence the brand sets. Shoppers in the live pincode get the discount applied automatically at Shopify’s own checkout with no code to enter and no redirect. Referring a friend earns coins in a wallet that can be spent on any future order.

Experiment 6 is the one where the sample size arithmetic does not apply cleanly, because a single postcode has too few orders for a proportion test. Read it against its own baseline over several cycles rather than as a single significant result. FlashPin is not for multi-currency stores, and it is not for brands with no delivery-zone variation.

Three that need a month or are not worth running

7. Change the hero image. Metric: order conversion. Sample: 15,288 per arm, about 25 days for a 20% effect, and hero images rarely move conversion by 20%. Run it only if you can also read add to cart, which is 8 days.

8. Replace a shared creator code with per-creator slugs. This is not a test, it is an instrumentation change, and treating it as an experiment wastes the point.

Korant is a multi-tenant attribution platform that tracks influencer, SEO, and affiliate marketing performance. Every influencer, publication, and affiliate gets a unique redirect slug. Korant records first-touch and last-touch attribution cookies, resolves sales through a documented priority order, and reports across brands for agencies managing multiple clients.

Ship it, do not split traffic on it, and the benefit is that every subsequent creator experiment becomes readable. Korant is not for stores with a single paid channel, and it is not for brands that only need Shopify’s native reports.

9. Add a post-purchase “notify a friend” flow. Metric: share rate, about 5 days, but the downstream order effect needs a month. Read the share rate quickly and the order effect slowly, and do not confuse the two.

The rules that make any of them readable

Five rules, and breaking any one of them voids the result.

Decide the sample size and the stop date before starting. Write both down. A test with no predetermined end will be stopped on the day it looks best.

Do not peek and act. Checking is fine. Stopping early on a favourable reading is how a coin-flip becomes a strategy.

Run one test per metric at a time. Two tests on checkout completion simultaneously means neither is attributable. Different parts of the funnel are fine.

Include a holdout that stays untouched. Especially for anything seasonal, which is most things.

Record the negative results. A team that only writes down wins will re-run the same losing idea every eighteen months as people change.

What experiments like these cannot tell you

Three limits, and the third is the one that catches good teams.

Anything about a small segment. Sample sizes above are for total traffic, and splitting by device or geography multiplies the requirement.

Long-term effects. A discount experiment that lifts conversion for two weeks says nothing about what it did to full-price demand in month four.

Whether the idea was good. A test tells you an effect was or was not detectable at your traffic. A null result on a 20% target means the effect is probably under 20%, not that it is zero.

The limitation worth sitting with is that a discipline built on sample size will systematically favour things that are easy to measure, and the most important changes in a D2C business are usually not. Product quality, positioning, and which customers you serve are all slow, unsplittable and decisive, and none of them will ever appear on a list of experiments readable in a week. Running the nine above is worth doing and it is not a strategy.

A team that has become very good at two-day checkout tests can still be losing on the things nobody can split-test. Use the fast experiments to remove friction you can see, and make the slow decisions with judgement rather than pretending they are testable. The mechanics worth testing in the first place are graded in viral D2C hacks to grow sales, the compounding view is in the growth loops that do not need paid media, and the launch cases are in zero-budget launches.

The tool for this · Shopify app FlashPin FlashPin handles the rotation, the checkout gating, and the ledger. Also relevant · Shopify app SB&R SB&R handles the chain, the cap, and the append-only ledger underneath it. Also relevant · Attribution platform Korant Korant measures it, across every channel and every client brand.

Questions people actually ask

How do you pick which ecommerce experiment to run first?

Work out the sample size before the idea. At typical D2C traffic, an experiment on order conversion takes about a month to read and one on checkout completion takes two days. That difference decides what is testable far more than how promising the idea sounded in a meeting.

Why is order conversion so hard to test?

Because sample size scales with the inverse of the baseline rate. A 2.5% conversion rate needs roughly seventeen times more sessions than a 30% checkout completion rate to detect the same relative change. Low-rate metrics are simply expensive to measure, regardless of how important they are.

What sample size do you actually need?

Roughly 15.7 times the baseline rate times one minus the baseline, divided by the square of the absolute effect you want to detect, per arm. That is the standard two-proportion formula at 80% power and 5% significance, and it takes thirty seconds in a spreadsheet.

What if a test never reaches significance?

Then you have learned the effect is smaller than your minimum detectable effect, which is genuine information. The failure mode is stopping early and calling a noisy result a win. Decide the sample size and the stop date before starting, and hold to both.

Can you run several experiments at once?

On different parts of the funnel, yes, since a checkout test and a product page test rarely interfere. Running two tests on the same metric at the same time means neither can be attributed, which is the commonest way a fortnight of traffic gets wasted.

Written by Nayak — Builds checkout and attribution tooling for Shopify D2C brands