Skip to main content
Beta. Signals is in Beta. This contract describes how grading works today, but wording can still change as we refine it. If something reads differently in your dashboard, trust the dashboard and email support@abtestly.com so we can fix the page. See the Signals overview.
Contract version 2.0, in force from the PXL v2 release · Last updated 2026-09-04 This page is the reference for how Signals decides how much weight a recommendation carries. It pulls together the grading rules that appear across Signals overview, Read a scan, Connect Google Analytics, and Build a test from an idea into one place, so you can see the whole model at once. Signals grades on two things: what evidence was credited on an idea, which becomes its evidence tier, and whether the idea’s PXL score clears the gate set for that tier. Everything below follows from those two ideas. Credited is not the same as connected: a source can reach the scan, appear as Used on its receipt, and be credited on no idea at all.

The evidence an idea can rest on

Every recommendation rests on the page structure, plus whichever of three further sources actually speaks to that idea. Structure is always there. The other three are there only when you connect or supply them, and each is scored separately below.

Structural evidence

The page structure: the DOM, meaning the headings, buttons, forms, copy, and layout the browser actually rendered. Structure is the first signal, and it is always present on every scan. It is enough to spot a buried call to action, a form asking for too much, or a value proposition that never appears above the fold.

Behavioral evidence

What visitors actually do, read from a connected Google Analytics 4 property, read only. Where they fall out of a funnel, how conversion differs by device, where revenue comes from by channel. Present only when GA4 contributes to the scan (Q8 below).

First party heatmap

Where visitors on your own pages click, including dead and rage clicks, which elements they reach and interact with, and how far down the page they scroll, from the ABTestly heatmap. It corroborates funnel data, it never replaces it: it is scored only alongside a credited GA4 answer, so a heatmap on a scan with no credited GA4 answer lifts nothing (Q7 below).

Customer qualitative voice

What your customers said in their own words, from what you supplied in Brand context: reviews, support themes, survey verbatims. This is a second signal in its own right, so an idea it is credited on can be marked Test ready with no GA4 at all (Q6 below).
When this page says a second independent signal agreed with the structure, it means GA4 funnel data or customer voice. The heatmap is not a second signal on its own: it strengthens GA4 rather than standing in for it.
The behavioral signal reads the last 90 days of the connected property. A page you capture today is compared against the visitors who came to it over that window, so a recent redesign, a seasonal peak, or a traffic source that has since changed can all sit inside the numbers an idea cites. Read the date range on the idea before treating a figure as current.
Structure tells you what the page is. Behavior tells you what is happening on it. Signals never invents a statistic when the data to support it is not there. When the number exists, you see it. When it does not, the idea says so plainly.

Who decides what

The AI evaluates the evidence. ABTestly computes the score. The model never hands back a score, a tier, a badge, a place on the board, or the next step an idea names. It answers eight yes or no questions about one idea, each with a short rationale, and ABTestly derives everything else from those answers and from what the scan actually had. An answer the evidence cannot support is overridden, and the rationale written for it is discarded rather than stored. If the model answers Q8 yes with a paragraph about funnel figures on a scan where GA4 never contributed, the point is not credited and the paragraph is not kept. The saved idea records the system’s reason in its place, so a stored idea can never argue against itself.

The four evidence tiers

The evidence tier says what was credited on one idea. It is not a measure of anyone’s confidence in that idea, and it is not a property of the scan: it is set per hypothesis, from the evidence the scan could credit on it. Every hypothesis carries exactly one tier, derived from the credited answers above, not chosen by you. From strongest to weakest: Because the tier belongs to the idea and not to the scan, two ideas on the same GA4 connected scan can land in different tiers: one the funnel data speaks to, one it does not. A source that contributed to the scan makes its question scorable for every idea on that scan; a connected source that returned nothing makes it scorable for none. It does not award the answer. The heatmap counts only alongside GA4. It corroborates funnel data, it never replaces it, so a heatmap on a scan with no credited GA4 answer lifts nothing. Used on the receipt is not the same as credited, and the two come apart by design. The receipt records what reached the run; the tier records what the scan could credit on one idea. A source can appear as Used on a scan’s receipt and be credited on no idea at all: GA4 can contribute and the scan still find it has nothing to say about a given idea, leaving that idea Structural, and a heatmap can be read and shown as Used while it is credited nowhere because there was no GA4 beside it.
Two tiers that appeared in earlier versions of this page, High confidence and Estimated, have been removed. Both described GA4 figures arriving by CSV export or by hand, and Signals has never had a way to take either in, so no scan could honestly earn them. Ideas scored before this contract may still show those names; see Ideas scored under the previous model.

How an idea becomes Test ready

An evidence tier says what was credited on the idea, not whether it is ready to run. Test ready takes two things, and both are required. First, a second independent signal has to agree with what the page structure showed. Second, the idea’s PXL score has to clear the threshold set by its tier:

The score is read against that scan’s maximum

The PXL score counts the points the idea earned. The number it is read against is the maximum that scan could reach, which is not a fixed scale: it is built from the evidence that was available on the run, so it is lower when a source is absent. So a card reads its score out of that scan’s own maximum. The same 7 points read as 7 of 12 on a scan with qualifying traffic and all three sources contributing, and as 7 of 8 on a scan that had qualifying traffic and customer voice but no GA4. Both are honest, because the denominator says what was genuinely reachable on that run.
Fourteen is the theoretical scale, not a reachable one. It counts Q5, user testing and interviews, which ABTestly does not run and therefore never scores. Under this contract an idea is read against its own scan’s maximum, not against 14. An idea scored under the previous model is read out of 14, and the badge says so; see Ideas scored under the previous model.

The board badges

The badge on the idea board reflects the outcome:
A Test ready badge means a second signal agreed and the score cleared its gate, not that the test will win. Nothing predicts a winner except running the experiment. The results page settles that.

Ideas scored under the previous model

Ideas scored before this contract took effect carry no version, and Signals will not vouch for their readiness. They read Unverified, captioned “scored under the previous model, readiness unverified”:
  • their readiness badge is suppressed, so a Test ready stored by the old model is never shown as Test ready
  • they are left out of the Test ready count in the scan header
  • their score is shown as it was stored, against the old 14 point scale, and it is not comparable with a score from this contract
  • they are never recomputed. The previous model wrote its reasoning as prose, with no structured answers underneath it, so there is nothing for the current arithmetic to run over. Scan the page again to get ideas scored under this contract.
An idea you added by hand is not affected: it reads as your own idea, never as one scored under the previous model.

Which experiments feed back

Signals learns from experiments, but not from every test you ship. The feedback rule is narrow and specific. An experiment feeds back into future scans only when both of these are true:
1

It was built from a Signals idea

The experiment came from an idea on the board through the Build this button. A test created outside the Signals flow does not feed the loop.
2

It ran to a conclusion

The experiment reached a verdict of won, lost, or inconclusive. An experiment still running, or one stopped before it concluded, does not feed back.
When both hold, that outcome feeds back into future scans of the same site as ground truth, so the next scan already knows what worked on your pages and what did not. This is a data feedback loop that carries past results forward as priors in the next scan, not model training. See Build a test from an idea for the handoff.
This is narrower than “every shipped test.” A test that was not built from a Signals idea, or that never reached a conclusion, does not feed the loop.

Fields a scan reads

A scan assembles evidence from several categories before it grades anything: The thirty four structural signals the DOM read captures are: page type and framework flags (React, Vue, Angular, single page app, hydration); h1 count and text, heading hierarchy issues, page title, meta description, and viewport meta; form field count, autocomplete presence, and mobile keyboard hints; CTA button count, primary CTA text, add to cart candidates, hidden add to cart count, and whether add to cart is gated by a variant; product schema, aggregate rating, visible review count, rating value, visible star rating, review count source, and trust badge count; image count and image alt coverage; hero copy, value propositions, navigation labels, and footer copy; HTTPS, bot wall detection and vendor, and a tech stack summary (platform, checkout providers, review widget, experiment tools, and script count). Each stored hypothesis is itself a structured record of twenty eight fields, and the split between them is the split described above. The model writes eighteen: title, pillar, observation, proposedChange, mechanism, mechanismTags, expectedLift, effort (Low, Medium, or High), researchCitation, dataSource, statisticalNote, qualitativeEvidence, brandRespects, opportunitySize (a point estimate with a low and high range and its assumptions), trackingSpec, recommendedRunTimeWeeks, interactionSpec, and pxlAnswers, its answers to the eight questions. ABTestly computes the other ten from those answers and from what the scan had: id (the rank), pxlScore, maxPossible, pxlBreakdown, confidenceTier, gateStatus, readinessReason, promotionPath, pxlVersion, and pxlDisplaySuppressed. The pxlAnswers that is stored is the normalised one, so an overridden answer carries the system’s reason in place of the model’s discarded rationale. Every generated idea is put through the adversarial pass, and one that comes back with an argument against it carries a falsificationCheck (failure mode, confounder, alternative hypothesis, discriminating test, and brand risk). The pass never blocks a scan: if it cannot complete, the ideas still reach your board without that block rather than the scan failing. An idea you added by hand never has one.

Model and data dates

  • Model provider. Signals runs on the Anthropic commercial API, whose terms do not use your data to train models.
  • Model version. Reasoning and hypothesis generation run on claude-sonnet-5. The screenshot reading pass runs on the vision model claude-haiku-4-5-20251001. Both are set in the worker configuration and can be overridden per environment.
  • GA4 data window. A scan reads the last 90 days of GA4 behavioral data, from 90 days ago through today, for every report in the pull.

Signals overview

The evidence, the closed loop, and the stages inside a scan.

Read a scan

The idea board, PXL ranking, campaigns, and triage.

Connect Google Analytics

One click, read only. Lets a scan read your funnel data and credit it on the ideas it speaks to.

Build a test from an idea

The handoff and the learning loop that feeds the next scan.
Last modified on September 4, 2026