There is more of this guide.
Practice
10 questions- Experiment designYour offline evals say a new prompt for your LLM summarization feature beats the current one. The PM asks why you'd bother with an online A/B test at all. What's your answer?
- Experiment designYou're A/B testing two LLM-backed variants whose outputs differ run to run, even on identical inputs. What does that non-determinism do to your experiment design and your read of the results?
- Experiment designYour team proposes an LLM-as-judge quality score as the primary metric for an assistant A/B test. What would you validate before letting it decide the ship call?
- Experiment designYou're A/B testing a new prompt for an AI feature, and the model provider upgrades the underlying model halfway through the test. Why is that a problem for your results?
- Experiment designA more capable model wins your quality A/B but costs five times more per request. What else has to be true before you ship, and how do you put cost into the experiment readout?
- Experiment designYou're testing an AI-powered recommendation engine that learns from user behavior over time. How does the model's learning affect your experiment duration and when you lock the model?