How do I know if an A/B test winner from an AI tool is real?
It depends on what the AI did. Some tools simulate visitors instead of testing with real ones, and their winner is only as good as the simulation. If an AI tool wrote the variant and you ran it as a normal A/B test on real visitors, the winner is as real as any other test result. You can trust it as far as you trust the test's setup and statistics. Adaptive methods such as multi-armed bandits are also harder to read than a classic A/B test, since traffic shifts toward one variant while the test runs. Calling a winner takes more than hitting a significance number. You have to agree on what the test was meant to improve and what result is worth shipping. At Coframe, we work out what counts as a winner together with each customer, then help the team interpret the results.
What if the tool simulated users instead of testing real ones?
Some research systems skip live traffic. They use AI agents as stand-in visitors and report which variant the agents preferred. Two recent papers show the approach and how their authors check it.
- Shopify researchers describe SimGym, which runs browser agents modeled on real shopper data. They report "attaining 77% directional alignment with add-to-cart shifts observed across interface variants in real-buyer traffic," on UI theme changes at one e-commerce platform. That is agreement on the direction of a change in one metric.
- The authors of SimAB report 67% accuracy against 47 historical A/B tests. They report 83% on its high-confidence cases.
These are the authors' own reports on their own test sets, so read them as evidence that simulation can be useful and that the track record is still partial. A simulated winner is a prediction about how real people might behave. It has no visitors behind it, so the usual sample size and significance checks do not apply to it in the same way.
If you are shown a simulated winner, these checks help:
- Ask what the simulation was validated against and whether the comparison used real test outcomes on pages like yours.
- Ask how many past tests were used and how often the simulation picked the opposite winner.
- Treat the result as a way to screen ideas before spending traffic on them. Confirm anything you plan to ship with a live test.
How do you check a winner from a normal A/B test?
Start by finding out how the test ran. A variant written by AI and shown to real visitors in a split test gets judged by the same checks as any other test. The checks below come from general experimentation practice, and they do not depend on who wrote the variant.
- Was the sample size and stopping point set in advance? Evan Miller's classic essay on testing mistakes explains that if you keep checking and stop the moment you see a significant difference, the reported significance levels stop meaning what they say. His fix is simple: "Committing to a sample size completely mitigates the problem described here." Some tools use sequential methods built to allow early looks. If yours does, ask the vendor which method it uses and what it guarantees.
- Did traffic split the way it was supposed to? A sample ratio mismatch means the share of visitors in each group differs from the planned share. In their KDD 2019 paper on diagnosing it, Fabijan and co-authors warn that ignoring the SRM without knowing the root cause "may result in a bad product modification appearing to be good and getting shipped to users, or vice versa". We have also seen a split look lopsided when assignment was fine. Visitors who leave before a tracking script loads can go missing from one side's count in the customer's own analytics.
- Can you see the numbers behind the label? A result that says only "winner" is a claim until you can see the conversion counts, the sample size and how uncertain the estimate is. We have seen revenue per visitor rise by more than half in one arm on very few purchases. With so few conversions that is usually noise, and the overlap in the intervals showed it.
- Which metric won? Decide before launch which metric counts. A variant can raise trial starts while lowering direct purchases, so a win on the first metric is not a win on revenue.
- How many things were tested at once? Optimizely, in a 2015 post about its own stats engine, says "Testing too many goals and variations at once greatly increases errors due to false discovery". A winner picked from a long list of metrics deserves more caution than one picked on a single metric chosen before launch.
- Does the lift hold over time? Look at the daily lift as well as the total, and split the result by new and returning visitors and by device. A customer has described a variant that surges in the first days and then evens out, which fits what we know about early data. A big early gap on little data is a weak reading.
- Does everyone score the result the same way? Two tools can call the same data differently when they use different statistical methods. We have had a customer re-score a readout in their own tools and reach a different verdict from the same inputs. Agreeing on how a winner gets called before the test starts avoids that argument.
Sample size has a practical limit too. A low-traffic page can take a long time to reach a conclusive result, and faster variant production does not change that.
What changes when the tool uses a bandit?
A multi-armed bandit shifts traffic while the test runs. Adobe's documentation says bandit algorithms "use adaptive allocation," and as evidence builds "more traffic is directed toward better-performing treatments." A classic A/B test keeps the split fixed. Adobe says that fixed allocation "reduces susceptibility to biases such as the 'winner's curse.'"
The same document lists weaker statistical guarantees as a drawback of bandits: "Traditional hypothesis testing is harder to apply, and stopping rules are less clear." Adobe's guide suggests bandits for always-on campaigns and for maximizing conversions during the test. It points to an A/B test when you want clear, confident insights.
When a tool uses adaptive allocation, ask these questions:
- How does it decide when a variant has "won"?
- Does it report an estimate with an interval, and how was the interval computed?
- Can you see how traffic moved over time, so you can tell whether early results drove the allocation?
We can walk you through how Coframe works out winners with customers. Talk to us about your results.