5 scans per month · page heatmaps included.
AI powered hypothesis generation, ranked and falsifiable.
ABTestly is a code first A/B testing platform. Signals is its optional CRO audit add on: it reads your page and your GA4, then generates a ranked, evidence graded backlog of A/B test hypotheses. Every idea is graded on its evidence and put through a pass that argues against it.
Three evidence streams. One ranked backlog.
Seven stages with a shape. Three evidence streams converge into one multimodal payload, then Signals generates hypotheses, argues against its own work, and ranks what it found. Tap any stage to see inside it.
Read how each stage works in the docs →
The browser extension captures a page from your own logged in session · free during beta, within your monthly scan allowance.
Three things a best practice checklist will not do.
Anyone can print a list of ideas. The work is deciding what is worth your traffic, and being honest about how sure it is.
A structural finding reaches your backlog, but it is never Test ready on its own.
Signals grades on two signals, a structural finding on the page plus evidence that agrees with it: a behavioral pattern in your GA4, or customer voice you paste into your brand context. A structural finding alone leaves an idea labeled Needs evidence, however good it looks. A second signal puts it in contention, and it reaches Test ready only once its score also clears the threshold for its evidence tier; short of that it reads Needs evidence or Below threshold. A small mobile CTA is a flagged finding on its own, and a confirmed issue once your mobile conversion rate agrees. The evidence contract lists every threshold, and the idea board guide shows the badges in place.
Before a test slot is spent, the agent argues against itself.
A second pass names the strongest failure mode, the likeliest confounder, the alternative explanation, and the single test that would tell them apart. When it comes back with one, you get the argument against the recommendation, attached to the recommendation.
A standalone audit guesses once. Signals hears back.
Because Signals runs inside ABTestly, the audit and the test engine are one system. Once a linked experiment concludes, its outcome feeds back automatically as a prior for the next scan of that domain. The learning is in context, not model training.
Test ready is earned, never assumed.
The point of Signals is telling you what it does not know. These behaviors are built in, not bolted on.
The qualification gate
It tells you when your traffic cannot reach significance, and refuses to fake confidence about it.
Evidence tier, set per idea
Every hypothesis carries one of four evidence tiers: Verified, Full evidence, Structural plus qualitative, or Structural. The tier is derived from which evidence questions were credited on that idea, so two ideas from the same scan land in different tiers when your GA4 speaks to one and not the other.
It never invents a number
Thin evidence is labeled, not inflated. When a claim rests on page structure alone, Signals says so and names the cheapest action that would confirm it.
| Evidence tier | What earns it | Score it needs | Highest it can reach |
|---|---|---|---|
| Verified | Your GA4 and a first party heatmap both credited | 9 or more | 10, or 12 with customer voice |
| Full evidence | Your GA4 credited | 7 or more | 8, or 10 with customer voice |
| Structural + qualitative | Customer voice credited, no analytics | 6 or more | 8 |
| Structural | Page structure alone | No score reaches Test ready | 6 |
Clearing the score is necessary and not sufficient. An idea reads Test ready only when a second independent signal agrees with the structural finding and the score clears its tier. Miss either one and it reads Needs evidence, with the board naming the specific thing that would move it, down to "Connect GA4 so the scan can check this idea against funnel data."
What one idea looks like when it comes back.
Not a paragraph of advice. Every hypothesis in the backlog carries the same structured fields, so two ideas from different scans can be compared without rereading them.
To see these fields filled in, read an illustrated scan of a product page.
| On every idea | What it holds |
|---|---|
| Observation and proposed change | What it found on the page, and the specific change to make. One of five CRO pillars: Trust and Credibility, Offer Clarity, Friction, Urgency and Motivation, or Page Flow and Hierarchy |
| Mechanism | Why the change should work, tagged with up to four named frameworks rather than left as intuition |
| Research citation | Every idea carries one. None ship uncited |
| Expected lift and effort | Effort is Low, Medium or High, and it feeds the rank, so a cheap idea is not buried under an expensive one |
| Statistical note | What the numbers behind the idea do and do not support |
| Tracking spec | The events to record so the test can actually be measured |
| Interaction spec | How the variant should behave once it is built |
| Data source | Where the evidence came from, which is what sets the evidence tier |
A recommended run time in weeks arrives when your traffic supports computing one, capped at eight. Customer voice appears on an idea only where your brand context spoke to it. And when the falsification pass returns, four more fields come with it: the strongest failure mode, the likeliest confounder, the alternative explanation, and the single test that would tell them apart.
Every linked experiment that concludes sharpens the next scan.
Signals does not hand off a backlog and forget it. Ship a hypothesis in ABTestly, and when the experiment concludes its result returns as a prior for the next scan. Nothing to log by hand.
Signals runs on the Anthropic commercial API, whose terms do not use your data to train models. So this is a data feedback loop, not model fine tuning. Signals produces the backlog and the test plan; you stay in control of what ships.
What it needs from you, and what it cannot see.
None of this is a caveat buried in the docs. It changes what a scan can return, so it belongs next to the promise.
GA4, or a lower ceiling
Without analytics connected, no idea can reach Verified or Full evidence. The best available is Structural plus qualitative, and only when you paste customer voice into your brand context. Page structure on its own has no threshold to clear, so it never reaches Test ready however good the finding looks.
The GA4 it reads is the last 90 days
A scan pulls a 90 day window. If you redesigned the page inside that window, some of the behaviour being weighed belongs to the design you replaced. Signals cannot detect that for you, and it is worth knowing before you act on a funnel number.
The counter argument can be absent
Every idea goes through the falsification pass, and the argument arrives attached when the pass returns one. When the pass cannot complete, the idea still reaches you without it rather than being held back. So read a missing argument as missing, not as an idea that survived one.
Signals is free for every paid plan, while it is in Beta.
No extra charge and nothing to switch on. If your ABTestly plan is paid, Signals runs today at no cost, within a monthly scan allowance.
Each account gets a monthly scan allowance: 5 scans on Starter, 25 on Pro, 100 on Business and 100 on Enterprise. Every site in the account draws on it, and it resets on the 1st of each month (UTC). Unused scans do not carry over. A scan that fails does not count. When the allowance runs out, new scans are refused until it resets or you move to a plan with a larger one.
Signals will be priced by scans per month. One scan is one full audit run. These prices are not live yet, Signals stays free until Beta ends.
25 scans per month · page heatmaps included.
100 scans per month · page heatmaps included.
Experiment heatmaps are live and included with Pro and Business at no extra cost, covering the pages your tests run on. Signals reads that behaviour as evidence when it writes hypotheses, including on pages where the test has already finished. On Starter, where experiment heatmaps are not included, Signals brings heatmaps with it. The heatmap corroborates GA4 funnel data rather than replacing it, so it counts toward an idea's evidence tier only on a scan where GA4 also contributed.
See what Signals finds on your own site.
Point it at a page, connect GA4, and read the backlog it returns. Signals is free while it is in Beta, on any paid plan, within a monthly scan allowance.