Experimenting on an AI feature

Before you design anything, name the assumption your feature breaks. Ordinary A/B tests rest on four, and an AI feature usually snaps at least two of them.

The idea

A normal A/B test assumes four quiet things. The same input gives the same output. The thing being tested holds still while you measure it. One user's experience does not change the next user's. And the metric is something you can observe directly, like a click.

An AI feature breaks these one at a time, and each break has a different fix. Non-determinism means the unit of analysis is the user rather than the response, and the sample size has to come from variance you measured rather than variance you borrowed from the last test. A model that learns from live traffic will not hold still, so either you pin the version or you say plainly that the result describes a system as it learns. Outputs become tomorrow's training data and people adapt to the model, so a two week read and a two month read are answers to different questions. And quality is usually judged by a or by another model, which means the judge itself has to be checked against human labels before it is allowed to decide anything.

Sometimes you are not running an experiment at all. If nobody knows how accurate the thing is, the first study measures accuracy against human labels, and the rollout design comes after.

Plan one test

A rubric scored assistant, control at 62 points

Powering for a 3 point lift, two sided, 95 percent, 80 percent power. 30 users per arm per day.

Planner for one AI feature experiment Rubric scores for individual responses in the treatment arm over the test window, the control mean, and an enrolment bar against the sample size the design needs.

powered

during the test
mid test event

How it works

Start by naming the broken assumption, because the assumption picks the design. Then size the test from variance you have actually observed.

  1. Pick the unit of analysis, and make it the user. The same prompt returns a different answer on retry, so responses from one person are correlated with each other and are not independent samples. Average a person's responses into one number per person, then compare people. Counting responses inflates your sample by a factor you did not earn, and the test reports significance it has not got.
  2. Measure the two kinds of variance before you size anything. Run the feature on a few hundred logged inputs, several times each. The spread between people is one number. The spread between repeated runs on the same input is the other. Both go into the sample size, and the second one is the one teams forget.
  3. Compute the sample size from those numbers, never from the last test. The click metric on the last test had almost no within user noise. This one does, and borrowing its sample size is how a study ends up underpowered without anyone noticing.
  4. Decide what holds still. If the question is about the feature, pin the model version, the prompt, the retrieval index and the decoding settings for the whole window, and write the versions in the test plan. If it cannot be pinned, that is allowed, but then the write up says the result describes the system as it learns, and does not call it the feature's effect.
  5. Validate the judge before it decides anything. If quality comes from a rubric or from a model grading the output, get human labels on a stratified sample, report agreement, and check that the judge does not favour one arm's style. An unvalidated judge is a second untested model sitting in the measurement path.
  6. Choose the horizon on purpose. A two week read measures first contact. A two month read includes people learning to work the model, and any retraining on the traffic the feature itself produced. Both are legitimate; they are not the same question, and the write up has to say which one it answered.
  7. Freeze the analysis plan, then defend the window. Any change to the model, the provider, the prompt or the index during the run splits an arm into populations that cannot be pooled. Log every version change with a timestamp so you can see it happen.

The sizing arithmetic, with the planner's default settings worked through:

# two sided alpha 0.05, power 0.80
z_alpha = 1.9600      z_beta = 0.8416
K = 2 * (1.9600 + 0.8416)^2 = 2 * 7.8489 = 15.698

# variance of one user's mean score, k responses per user
var_user = sd_between^2 + sd_within^2 / k

  sd_between = 16      (person to person, rubric points)
  sd_within  = 18      (same user, run to run, the non-determinism)
  k          = 6       (responses per user in a 14 day window)

  var_user = 256 + 324/6 = 256 + 54 = 310

# users per arm, for a 3 point lift
n = K * var_user / delta^2
  = 15.698 * 310 / 9
  = 541 users per arm

# what the click metric on the last test would have needed
n = 15.698 * 256 / 9 = 447     <- borrowing this is the mistake

Two things fall out of that formula and both are worth saying out loud in an interview. Within user noise enters divided by k, so collecting more responses per person is a cheaper way to buy power than recruiting more people, up to the point where person to person spread dominates. And once sd_within is large, the sample size grows fast: doubling the spread from 18 to 36 takes the requirement from 541 to 824.

When the model is learning during the window, two more terms appear:

# cohort drift: users enrolled on day 1 and day 40 saw different models
R        = points the model gains across the window (capped at 12 here)
var_user = sd_between^2 + sd_within^2/k + R^2/12

# interference: one user's traffic trains the model the next user gets,
# so users are not independent. Inflate the requirement.
n = 1.2 * K * var_user / delta^2

The 1.2 is a planning judgement, not a derived constant. Say so when you use one. Notice also that the drift term is usually small next to person to person spread, which is the honest and slightly awkward point: a moving baseline costs you far more in interpretation than it costs you in sample size. The number still comes out; it just no longer means what people will assume it means.

When to use it

The whole discipline in one line: say which assumption your feature breaks, say what your design does about it, and say what your result therefore does and does not describe. A senior answer is three sentences long and names all three.

Watch out for

Worked example

The scenario, roughly as it gets asked: a moderation model reviews reported posts about ten times faster than the human queue. Nobody knows how accurate it is. Product wants to know whether to ship it. How would you run the experiment?

The first move is to say that this is not an experiment yet. The unknown is accuracy, not lift, so the first study is a measurement study. Sample the reported queue, stratified by category and by how borderline the case is, because a uniform sample will be dominated by the easy items and will flatter the model. Route those same items to the model and to human reviewers, with the reviewers blind to the model's call. Then report two error rates separately, never one accuracy number: how often the model removes something the humans kept, and how often it keeps something the humans removed. Add agreement between two human reviewers on the same items, because that is the ceiling. If two experienced reviewers agree only eighty percent of the time on borderline cases, a model that matches humans eighty percent of the time is at the ceiling, and no amount of model work moves it.

Only then does a rollout design make sense, and it is a shadow deployment before it is an A/B test. Run the model on live traffic, log its decision, act on nothing. That gives you volume, a per category error profile and an operating threshold chosen against a stated policy: at this threshold we remove this many lawful posts to catch this many violations. The trade is a decision for policy and legal, and your job is to hand them the curve, not to pick a point on it.

When you do finally run the live test, the outcome metric is not model accuracy. It is queue time, reviewer load, appeal rate and reinstatement rate on appeal, because a reinstated removal is your measured by the process that exists for that purpose. Randomise by reviewer shift or by queue rather than by post, since one reviewer handles many posts and their calibration drifts with what the model has been showing them, which is a spillover. Pin the model version for the window. And keep a small permanent holdout after launch, because the loop here is sharp: moderated content changes what gets posted, which changes what the model sees next month.

If the interviewer pushes on timelines, the honest answer is that the measurement study is days, the shadow run is weeks, and only the last stage needs the sample size arithmetic. Saying which stage answers which question is the part that reads as senior.

Check yourself

A summarisation feature is tested for two weeks. Each user gets about eight summaries, and the team analysed 41,000 summaries as 41,000 samples. The result is significant at p = 0.01. What is the first thing you say?

On day 9 of a 21 day test, the provider ships a new model version and it goes live in the treatment arm. What is the right call?