> ## Documentation Index
> Fetch the complete documentation index at: https://docs.abtestly.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Signals evidence contract

> The precise contract for how Signals grades evidence: structural versus behavioral signals, the four evidence tiers, the PXL gate that decides the Test ready, Needs evidence, Below threshold, and Unverified board badges, which experiments feed back, and what a scan reads. Currently in Beta.

<Note>
  **Beta.** Signals is in Beta. This contract describes how grading works
  today, but wording can still change as we refine it. If something reads
  differently in your dashboard, trust the dashboard and email
  **[support@abtestly.com](mailto:support@abtestly.com)** so we can fix the page. See the [Signals
  overview](/signals/overview).
</Note>

**Contract version 2.0, in force from the PXL v2 release · Last updated
2026-09-04**

This page is the reference for how Signals decides how much weight a
recommendation carries. It pulls together the grading rules that appear
across [Signals overview](/signals/overview), [Read a
scan](/signals/read-a-scan), [Connect Google
Analytics](/signals/connect-google-analytics), and [Build a test from an
idea](/signals/build-a-test) into one place, so you can see the whole
model at once.

Signals grades on two things: what evidence was **credited** on an idea,
which becomes its **evidence tier**, and whether the idea's PXL score
clears the gate set for that tier. Everything below follows from those
two ideas. Credited is not the same as connected: a source can reach the
scan, appear as **Used** on its receipt, and be credited on no idea at
all.

## The evidence an idea can rest on

Every recommendation rests on the page structure, plus whichever of
three further sources actually speaks to that idea. Structure is always
there. The other three are there only when you connect or supply them,
and each is scored separately below.

<CardGroup cols={2}>
  <Card title="Structural evidence" icon="code">
    The page structure: the DOM, meaning the headings, buttons, forms,
    copy, and layout the browser actually rendered. Structure is the
    first signal, and it is **always present** on every scan. It is
    enough to spot a buried call to action, a form asking for too much,
    or a value proposition that never appears above the fold.
  </Card>

  <Card title="Behavioral evidence" icon="chart-line">
    What visitors actually do, read from a connected [Google Analytics
    4](/signals/connect-google-analytics) property, read only. Where
    they fall out of a funnel, how conversion differs by device, where
    revenue comes from by channel. Present **only when GA4 contributes to
    the scan** (Q8 below).
  </Card>

  <Card title="First party heatmap" icon="fire">
    Where visitors on your own pages click, including dead and rage
    clicks, which elements they reach and interact with, and how far
    down the page they scroll, from the ABTestly heatmap. It
    **corroborates funnel data, it never replaces it**: it is scored
    only alongside a credited GA4 answer, so a heatmap on a scan with
    no credited GA4 answer lifts nothing (Q7 below).
  </Card>

  <Card title="Customer qualitative voice" icon="comments">
    What your customers said in their own words, from what you supplied
    in **Brand context**: reviews, support themes, survey verbatims.
    This is a second signal in its own right, so an idea it is credited
    on can be marked Test ready **with no GA4 at all** (Q6 below).
  </Card>
</CardGroup>

When this page says a **second independent signal** agreed with the
structure, it means GA4 funnel data or customer voice. The heatmap is
not a second signal on its own: it strengthens GA4 rather than standing
in for it.

<Note>
  The behavioral signal reads the **last 90 days** of the connected
  property. A page you capture today is compared against the visitors who
  came to it over that window, so a recent redesign, a seasonal peak, or a
  traffic source that has since changed can all sit inside the numbers an
  idea cites. Read the date range on the idea before treating a figure as
  current.
</Note>

Structure tells you what the page is. Behavior tells you what is
happening on it. Signals never invents a statistic when the data to
support it is not there. When the number exists, you see it. When it
does not, the idea says so plainly.

## Who decides what

**The AI evaluates the evidence. ABTestly computes the score.** The
model never hands back a score, a tier, a badge, a place on the board,
or the next step an idea names. It answers **eight yes or no questions**
about one idea, each with a short rationale, and ABTestly derives
everything else from those answers and from what the scan actually had.

| #   | The question                                           | Worth | Answered by                                                                          |
| --- | ------------------------------------------------------ | ----- | ------------------------------------------------------------------------------------ |
| Q1  | Is the change above the fold, in the primary viewport? | 1     | the model                                                                            |
| Q2  | Is it noticeable within 5 seconds?                     | 1     | the model                                                                            |
| Q3  | Does it add or remove an element?                      | 1     | the model                                                                            |
| Q4  | Does it address a primary motivational variable?       | 1     | the model                                                                            |
| Q5  | Is it informed by user testing or interviews?          | 2     | nobody. ABTestly does not run user testing, so this point can never be scored.       |
| Q6  | Is it informed by customer qualitative voice?          | 2     | the model, and only when you supplied voice in Brand context                         |
| Q7  | Is it informed by the first party heatmap?             | 2     | the model, and only when the heatmap contributed to the scan **and** Q8 was credited |
| Q8  | Is it informed by GA4 funnel data?                     | 2     | the model, and only when GA4 contributed to the scan                                 |
| Q9  | Is the page above the traffic qualification threshold? | 1     | ABTestly, from the monthly sessions you gave at intake                               |
| Q10 | Is it tied to the primary programme goal?              | 1     | the model                                                                            |

An answer the evidence cannot support is **overridden**, and the
rationale written for it is discarded rather than stored. If the model
answers Q8 yes with a paragraph about funnel figures on a scan where
GA4 never contributed, the point is not credited and the paragraph is
not kept. The saved idea records the system's reason in its place, so a
stored idea can never argue against itself.

## The four evidence tiers

The **evidence tier** says what was credited on one idea. It is not a
measure of anyone's confidence in that idea, and it is not a property of
the scan: it is set **per hypothesis**, from the evidence the scan could
credit on it. Every hypothesis carries exactly one tier, derived from the
credited answers above, not chosen by you. From strongest to weakest:

| Tier                         | What was credited on this idea                                                                                      |
| ---------------------------- | ------------------------------------------------------------------------------------------------------------------- |
| **Verified**                 | GA4 funnel data was credited on the idea, and the first party heatmap was credited alongside it.                    |
| **Full evidence**            | GA4 funnel data was credited on the idea. The heatmap either was not available or was not credited on it.           |
| **Structural + qualitative** | Customer qualitative voice you supplied was credited on the idea. No funnel data credited behind it.                |
| **Structural**               | Page structure alone, the always present first signal. No GA4, no heatmap and no customer voice was credited on it. |

Because the tier belongs to the idea and not to the scan, **two ideas on
the same GA4 connected scan can land in different tiers**: one the
funnel data speaks to, one it does not. A source that contributed to the
scan makes its question scorable for every idea on that scan; a
connected source that returned nothing makes it scorable for none. It
does not award the answer.

The heatmap counts only alongside GA4. It corroborates funnel data, it
never replaces it, so a heatmap on a scan with no credited GA4 answer
lifts nothing.

**Used** on the receipt is not the same as credited, and the two come
apart by design. The receipt records what reached the run; the tier
records what the scan could credit on one idea. A source can appear as
**Used** on a scan's receipt and be credited on no idea at all: GA4
can contribute and the scan still find it has nothing to say about a
given idea, leaving that idea Structural, and a heatmap can be read
and shown as **Used** while it is credited nowhere because there was
no GA4 beside it.

<Note>
  Two tiers that appeared in earlier versions of this page, **High
  confidence** and **Estimated**, have been removed. Both described GA4
  figures arriving by CSV export or by hand, and Signals has never had a
  way to take either in, so no scan could honestly earn them. Ideas scored
  before this contract may still show those names; see [Ideas scored under
  the previous model](#ideas-scored-under-the-previous-model).
</Note>

## How an idea becomes Test ready

An evidence tier says what was credited on the idea, not whether it is
ready to run. Test ready takes **two** things, and both are required.
First, a second independent signal has to agree with what the page
structure showed. Second, the idea's **PXL score** has to clear the
threshold set by its tier:

| Tier                         | Test ready when the PXL score is at least                                            |
| ---------------------------- | ------------------------------------------------------------------------------------ |
| **Verified**                 | 9                                                                                    |
| **Full evidence**            | 7                                                                                    |
| **Structural + qualitative** | 6                                                                                    |
| **Structural**               | Never. A structural only idea is always surfaced for validation, never marked ready. |

## The score is read against that scan's maximum

The PXL score counts the points the idea earned. The number it is read
against is **the maximum that scan could reach**, which is not a fixed
scale: it is built from the evidence that was available on the run, so
it is lower when a source is absent.

| Available on the scan                                                                                                                                 | Adds   |
| ----------------------------------------------------------------------------------------------------------------------------------------------------- | ------ |
| The five questions every scan can answer: placement, noticeability, whether the change adds or removes an element, motivation, and the programme goal | 5      |
| Monthly sessions known and at least 1,000                                                                                                             | 1      |
| GA4 contributed                                                                                                                                       | 2      |
| The heatmap contributed **and** GA4 contributed                                                                                                       | 2      |
| Customer qualitative voice in Brand context                                                                                                           | 2      |
| **The most any scan can reach**                                                                                                                       | **12** |

So a card reads its score out of that scan's own maximum. The same 7
points read as 7 of 12 on a scan with qualifying traffic and all three
sources contributing, and as 7 of 8 on a scan that had qualifying
traffic and customer voice but no GA4. Both are honest, because the
denominator says what was genuinely reachable on that run.

<Note>
  Fourteen is the theoretical scale, not a reachable one. It counts Q5,
  user testing and interviews, which ABTestly does not run and therefore
  never scores. Under this contract an idea is read against its own
  scan's maximum, not against 14. An idea scored under the previous model
  is read out of 14, and the badge says so; see [Ideas scored under the
  previous model](#ideas-scored-under-the-previous-model).
</Note>

## The board badges

The badge on the [idea board](/signals/read-a-scan) reflects the outcome:

| Badge               | What it means                                                                                                                                                                                                                                                                                                         |
| ------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Test ready**      | A second signal agrees and the score cleared its tier's gate. A strong candidate to run, though only the experiment settles whether it wins.                                                                                                                                                                          |
| **Needs evidence**  | The idea is sound but has not cleared its gate, or was credited with no second signal and so has no gate to clear. It names either the cheapest real next step that would raise its evidence or the score it fell short by. [Connect GA4](/signals/connect-google-analytics) to give more ideas a chance to clear it. |
| **Below threshold** | The score fell short of the tier's gate, on a scan where another idea cleared its own gate. The gate is the fixed number per tier under [How an idea becomes Test ready](#how-an-idea-becomes-test-ready), not a fraction of the scan's maximum.                                                                      |
| **Unverified**      | Readiness could not be verified for this idea, so none is claimed. Either it was scored under the previous model, before ABTestly took over the arithmetic, or it was scored under this contract but its score failed an arithmetic check and was withheld.                                                           |

<Note>
  A Test ready badge means a second signal agreed and the score cleared
  its gate, not that the test will win. Nothing predicts a winner except
  running the experiment. The [results page](/results/reading-a-result)
  settles that.
</Note>

## Ideas scored under the previous model

Ideas scored before this contract took effect carry no version, and
Signals will not vouch for their readiness. They read **Unverified**,
captioned "scored under the previous model, readiness unverified":

* their readiness badge is suppressed, so a Test ready stored by the old
  model is never shown as Test ready
* they are left out of the Test ready count in the scan header
* their score is shown as it was stored, against the old 14 point scale,
  and it is not comparable with a score from this contract
* they are **never recomputed.** The previous model wrote its reasoning
  as prose, with no structured answers underneath it, so there is
  nothing for the current arithmetic to run over. Scan the page again to
  get ideas scored under this contract.

An idea you added by hand is not affected: it reads as your own idea,
never as one scored under the previous model.

## Which experiments feed back

Signals learns from experiments, but not from every test you ship. The
feedback rule is narrow and specific.

An experiment feeds back into future scans only when **both** of these
are true:

<Steps>
  <Step title="It was built from a Signals idea">
    The experiment came from an idea on the board through the [Build
    this](/signals/build-a-test) button. A test created outside the
    Signals flow does not feed the loop.
  </Step>

  <Step title="It ran to a conclusion">
    The experiment reached a verdict of **won, lost, or inconclusive**.
    An experiment still running, or one stopped before it concluded,
    does not feed back.
  </Step>
</Steps>

When both hold, that outcome feeds back into future scans of the **same
site** as ground truth, so the next scan already knows what worked on
your pages and what did not. This is a data feedback loop that carries
past results forward as priors in the next scan, not model training. See
[Build a test from an idea](/signals/build-a-test) for the handoff.

<Warning>
  This is narrower than "every shipped test." A test that was not built
  from a Signals idea, or that never reached a conclusion, does not feed
  the loop.
</Warning>

## Fields a scan reads

A scan assembles evidence from several categories before it grades
anything:

| Category                      | What is captured                                                                                                  | Notes                                                                            |
| ----------------------------- | ----------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------- |
| **Screenshots**               | Desktop, mobile, and a throttled connection capture of the page, each reconciled against what the structure read. | Three image blocks.                                                              |
| **Page structure (DOM)**      | Thirty four structural signals, listed in full below the table.                                                   | Always present.                                                                  |
| **Behavioral metrics (GA4)**  | Funnel drop off, conversion by device, revenue by channel, new against returning visitors, and busiest pages.     | Present only when a connected GA4 property returns data for the scan, read only. |
| **Business context**          | Your niche, average order value, the goal that matters, and your device split.                                    | Supplied at intake.                                                              |
| **Prior experiment outcomes** | Results from experiments built from earlier ideas on this same site that ran to a conclusion.                     | Feeds the learning loop above.                                                   |

The thirty four structural signals the DOM read captures are: page type
and framework flags (React, Vue, Angular, single page app, hydration);
h1 count and text, heading hierarchy issues, page title, meta
description, and viewport meta; form field count, autocomplete presence,
and mobile keyboard hints; CTA button count, primary CTA text, add to
cart candidates, hidden add to cart count, and whether add to cart is
gated by a variant; product schema, aggregate rating, visible review
count, rating value, visible star rating, review count source, and trust
badge count; image count and image alt coverage; hero copy, value
propositions, navigation labels, and footer copy; HTTPS, bot wall
detection and vendor, and a tech stack summary (platform, checkout
providers, review widget, experiment tools, and script count).

Each stored hypothesis is itself a structured record of twenty eight
fields, and the split between them is the split described above. The
model writes eighteen: `title`, `pillar`, `observation`,
`proposedChange`, `mechanism`, `mechanismTags`, `expectedLift`, `effort`
(Low, Medium, or High), `researchCitation`, `dataSource`,
`statisticalNote`, `qualitativeEvidence`, `brandRespects`,
`opportunitySize` (a point estimate with a low and high range and its
assumptions), `trackingSpec`, `recommendedRunTimeWeeks`,
`interactionSpec`, and `pxlAnswers`, its answers to the eight questions.
ABTestly computes the other ten from those answers and from what the
scan had: `id` (the rank), `pxlScore`, `maxPossible`, `pxlBreakdown`,
`confidenceTier`, `gateStatus`, `readinessReason`, `promotionPath`,
`pxlVersion`, and `pxlDisplaySuppressed`. The `pxlAnswers` that is
stored is the normalised one, so an overridden answer carries the
system's reason in place of the model's discarded rationale.

Every generated idea is put through the adversarial
pass, and one that comes back with an argument against it carries a
`falsificationCheck` (failure mode, confounder, alternative hypothesis,
discriminating test, and brand risk). The pass never blocks a scan: if
it cannot complete, the ideas still reach your board without that block
rather than the scan failing. An idea you added by hand never has one.

## Model and data dates

* **Model provider.** Signals runs on the **Anthropic commercial API**,
  whose terms do not use your data to train models.
* **Model version.** Reasoning and hypothesis generation run on
  `claude-sonnet-5`. The screenshot reading pass runs on the vision
  model `claude-haiku-4-5-20251001`. Both are set in the worker
  configuration and can be overridden per environment.
* **GA4 data window.** A scan reads the **last 90 days** of GA4
  behavioral data, from 90 days ago through today, for every report in
  the pull.

## Related pages

<CardGroup cols={2}>
  <Card title="Signals overview" icon="magnifying-glass" href="/signals/overview">
    The evidence, the closed loop, and the stages inside a scan.
  </Card>

  <Card title="Read a scan" icon="list-check" href="/signals/read-a-scan">
    The idea board, PXL ranking, campaigns, and triage.
  </Card>

  <Card title="Connect Google Analytics" icon="chart-line" href="/signals/connect-google-analytics">
    One click, read only. Lets a scan read your funnel data and credit
    it on the ideas it speaks to.
  </Card>

  <Card title="Build a test from an idea" icon="flask" href="/signals/build-a-test">
    The handoff and the learning loop that feeds the next scan.
  </Card>
</CardGroup>
