Running A/B Experiments
Perspect's A/B framework is built for a tight iteration loop: frame a hypothesis, ship variants, measure against a pre-committed metric, and feed the outcome into the next iteration. The same tools are available to humans in the admin UI and to agents via MCP — most teams use a mix: humans frame the question, the Experimenter agent runs the mechanics and reports.
This guide walks through one full loop end-to-end with signups as the running example. The pattern generalizes to any conversion objective.
1. Frame the hypothesis
Every flag needs a hypothesis in a fixed form:
Changing X from A to B will move metric M by at least D, measurable within W.
For signups, that might be:
Changing the hero CTA copy from "Start free" to "Get started in 60 seconds" will lift the
user_signuprate on/by at least 10% over 14 days.
Resist the urge to run a test without this. A test without a pre-committed metric and effect size is not an experiment — it's a vibe.
2. Pick the goal event
An "objective" in Perspect is exactly one thing: the primary goal event on the flag. Goals are an array of { eventName, primary } entries; exactly one is marked primary. That event is the metric the experiment is judged on.
For "increase signups," the eventName is whatever your site fires on a completed signup — for example user_signup or signup.completed. Convention, not enum: the string must match what your site code actually emits.
Secondary goals (primary: false) are guardrails. List the things you'd notice if the variant quietly broke something: support_ticket_opened, checkout_abandoned, page_error. The results surface tracks them alongside the primary goal so a variant that wins on signups but doubles support load doesn't ship silently.
3. Confirm the event already fires
This is the precondition that sinks most tests before they start.
If your site doesn't emit user_signup yet, the experiment can't measure anything. Instrument first, then test. In practice:
- Find the success path in code (the post-signup redirect handler, or the
onSuccessof the form). - Add a
track("user_signup")call there. - Deploy.
- Verify exposures show up in observability before touching the flag.
The Experimenter agent enforces this as a hard precondition — it will refuse to create a flag whose primary goal event isn't already flowing. Humans should enforce it on themselves.
4. Build the variant
A variant is a code branch the site chooses at render time based on the flag evaluation. The SDK's ab_integration_guide tool returns the exact integration pattern for your codebase; the short version:
- Evaluate the flag for the current visitor.
- Read the variant key.
- Branch on it — different component, different copy, different layout.
Keep variants cheap. A variant that ships 200 KB of extra JS loses on performance before copy ever mattered. The Experimenter's "performance is a confound" principle is real — a slower variant can tank conversion for reasons unrelated to the change you meant to test.
5. Create and start the flag
Once the variant ships and the goal event is confirmed, create the flag. From the admin: Site → A/B → New flag. From an agent: ab_create_flag. Either way, the shape is:
{
"key": "hero-cta-copy-v1",
"description": "Test whether specific duration wording lifts signup conversion on the homepage hero.",
"variants": [
{ "key": "control", "weightBp": 5000 },
{ "key": "sixty-seconds", "weightBp": 5000 }
],
"defaultVariant": "control",
"goals": [
{ "eventName": "user_signup", "primary": true },
{ "eventName": "support_ticket_opened", "primary": false }
],
"trafficAllocationBp": 10000,
"attributionWindowDays": 7
}
A few things to notice:
- Weights are in basis points (10000 = 100%). 50/50 is
5000each. trafficAllocationBpis how much of total traffic enters the test at all. Start at 10000 (all) for anything pre-launch; ramp up gradually if the variant is risky.attributionWindowDaysbounds how long after exposure a conversion counts. Default 7. For slow-burn outcomes like subscription renewal, go longer; for same-session signups, 1–3 is fine.
Then start it: admin UI Start button, or ab_start_flag. Record the start time and the planned end time based on your declared window.
6. Monitor — without calling it early
The rollup cron aggregates exposures and goal events every 5 minutes, so results are near-realtime. But early results lie. A dramatic-looking gap on day one almost always shrinks as the sample grows.
What to actually watch in the first 24 hours:
- Sample Ratio Mismatch (SRM). If you declared 50/50 but see 55/45 exposures, something upstream is splitting traffic unevenly. That's a broken experiment, not a signal. Pause and diagnose.
- Traces for the losing variant. If one variant is underperforming, check its latency and error rate in observability before concluding the change is worse. You may be looking at a regression, not a test result.
- Events flowing at expected volume. A missing event pipeline produces zero conversions and looks like a flat test. Check exposures are being recorded and goal events are being recorded.
What to watch at the planned end:
- Did you hit the pre-declared sample size? If not, the test is underpowered. "Inconclusive, extend or drop" is the honest answer — not "no winner, ship control."
- Is the lift larger than your pre-declared minimum detectable effect? Statistical significance at an effect smaller than you said you cared about is not a business result.
Pull results via the admin UI or ab_get_results. The response includes per-variant exposures, per-goal conversions, conversion rate, and confidence. Read the number; don't let it read you.
7. Decide, then iterate
At the planned end, you have one of four outcomes:
| Outcome | Action |
|---|---|
| Variant wins by ≥ MDE, guardrails clean | Ship it. Remove the flag branch in code (keep the winner unconditionally), stop the flag. |
| Variant loses | Stop the flag, keep control, log what you learned. |
| Inconclusive, underpowered | Extend the window OR drop the test. Do NOT "ship the leader." |
| SRM or guardrail regression | Pause, fix the underlying issue, restart as a new version. |
Every concluded test — win, loss, or inconclusive — feeds the next hypothesis. The Experimenter agent is built to write both the result and the interpretation to memory so later experiments build on this one instead of rediscovering it.
8. How agents participate
If you run the Experimenter agent on your site, the loop above is something you can hand off in one message:
"Run an experiment to see if changing the hero CTA copy lifts signups. Primary metric is
user_signupon/. 14-day window, 50/50 split, 10% MDE."
The agent will:
- Check that
user_signupis already emitted, and tell you to instrument first if not. - Draft the hypothesis and ask for confirmation of the variant copy.
- Hand the variant implementation to the App Developer agent (it doesn't edit code itself).
- Create and start the flag once the variant is deployed.
- Monitor SRM, sample size, and guardrails.
- Report at the declared end with a concrete ship/drop/extend recommendation.
You can also drive the same MCP tools yourself: ab_create_flag, ab_start_flag, ab_get_results, ab_pause_flag, ab_resume_flag, ab_stop_flag. The ab_integration_guide tool returns the site-code wiring pattern.
What this framework intentionally does NOT do
- No auto-stop on "early significance." Peek-driven decisions are the single biggest source of false positives in A/B testing. Tests run for their pre-declared window.
- No built-in Bayesian posterior. Results are conversion rates with confidence intervals; if you need Bayesian reporting, layer it on the raw rollups.
- No multi-variate testing in v1. One flag tests one change. Interaction effects belong in a different framework.
- No separate "objective" abstraction. The primary goal event is the objective. A higher-level objectives registry is on the roadmap, but in practice tests map to events, and events are enough.
Where to go next
- A/B Testing SDK guide — the code-level API for evaluating flags and tracking events.
- Webhooks — subscribe to
ab.flag_started,ab.flag_stopped, and related lifecycle events if you need external systems to react. - Observability — the tracing and logs surfaces you'll lean on when a variant underperforms for reasons unrelated to the change.