Methodology

How Undersignal works

Undersignal™ measures how a text is constructed to persuade. Not whether its claims are true, and not who wrote it. This page explains what the instrument measures, how it is calibrated, how we validate it, and where its limits are.

Versioned instrument · Every report records the version that produced it

Instrument version 2.61 · Page updated 2026-09-02

1. What the score measures

The Undersignal Risk Score (1 to 10) quantifies rhetorical construction intensity: the density and coordination of persuasion mechanics present in a text. A high score does not mean a text is false, malicious, or unlawful. It means the text is built with heavier persuasive machinery, and that machinery is identified, named, and quoted in the report so a reader can weigh it.

We score texts, not entities. The same organization can publish a 2 and an 8 in the same week. A score attaches to one analyzed text at one point in time, never to an author, outlet, or brand.

2. What the instrument is grounded in

The instrument draws on the Persuasion Knowledge Model (Friestad & Wright, 1994), the sponsorship-disclosure literature, work on narrative transportation and fictionality labeling, risk-as-feelings and fear-appeal research, identity-protective cognition, and the deceptive-design literature. Where these bodies of evidence disagree about what disclosure changes, the instrument follows the evidence per mechanism rather than applying one rule everywhere.

Each report discloses the construction: which mechanisms are present, where they appear, and how they coordinate. Where a report names a specific rhetorical device, it shows the passage that label rests on, and section 3 explains how far those names should be trusted. The reader's own judgment does the rest. We do not tell readers what to believe. We show them how the text is trying to make them believe it. Disclosure remains the product's operating principle, and the score credits disclosure exactly where the evidence supports it.

3. The measurement architecture

Every report scores twelve construction dimensions, grouped into five clusters. Two further readings are applied after the clusters: coercive pressure, drawn from the asks the text actually makes of the reader, and disguise, measured on its own short scale. Genre and disclosure are established from the text and modulate how construction counts rather than measuring it. The instrument keeps these separate because they behave differently, and a reader should know which is which.

Construction dimensions 12 dimensions · five clusters · drive the score

These measure how the text is built: how claims are sourced and framed, how emotion, fear, and identity are engaged, how the narrative is structured and controlled, and how authority is invoked. The clusters that dominate a text determine its construction pattern, the labeled fingerprint at the top of every report: Declarative, Neutral, Patterned, Elevated, or Structured. They also drive the composite score. The full index is below.

Coercive pressure Applied after the clusters · enters the score

The engine separately judges how hard the text pushes the reader toward immediate action, working from the asks the text actually makes: urgency, compliance demands, and coercive calls to act. It is not one of the twelve dimensions and is not part of any cluster. It is added to the score after the disclosure adjustment, so declaring persuasive intent never reduces it. A hard-sell close operates on the reader even when the intent behind it is fully declared. Disclosure discounts framing; it does not discount asks.

Disguise Applied after the clusters · own scale

How much the text works to appear as something other than what it is. Disguise runs on its own short scale rather than the 1 to 10 the dimensions use, is not part of any cluster, and is added after the disclosure adjustment alongside pressure. It also feeds the genre and disclosure judgment described in section 4.

Genre and disclosure signals Scored signals · modulate the score

These do not measure construction; they modulate how construction counts. The engine judges, per text, whether the text openly declares its persuasive purpose and whether it honors or betrays the genre it presents itself as. Section 4 explains how.

The 12 construction dimensions

Grouped by cluster. Each entry states what the dimension measures; the scoring rubric behind each one is proprietary (section 9).

Emotional ActivationCluster
Fear Amplification

How much the text leans on threat and danger to move the reader, relative to what its own evidence supports.

Emotional Loading

How much emotionally charged language carries the argument instead of the underlying facts.

Archetypal Framing

How far people and institutions are cast into story roles such as hero, villain, or victim to do the persuading.

Identity ExploitationCluster
Identity Pressure

The social cost the text attaches to disagreeing with it.

Tribal Signaling

How strongly the text frames its subject as us versus them and cues group loyalty.

Legitimacy ManipulationCluster
Authority Exploitation

How expert and institutional credibility is used to settle questions rather than inform them.

Manufactured Consensus

How much agreement is asserted or staged rather than demonstrated.

Money Context

How visibly financial interests connect to the narrative, and whether the beneficiaries are named.

Narrative ControlCluster
Suppression & Promotion

How competing explanations are sidelined or disqualified rather than engaged.

Centralized Narrative

How tightly every element of the text is organized to point at a single conclusion.

Narrative Structure

How much the content works as a story with characters and an arc rather than a set of claims.

Logical IntegrityCluster
Logical Fallacies

How much of the argument rests on reasoning that fails when stated plainly.

Device labels, and how far to trust them

Alongside the dimension scores, reports name specific rhetorical devices where they appear and quote the passage each name rests on. Device names come from a fixed, defined canon rather than being invented per report, and every proposed label is checked in code against the conditions that device actually requires before it is allowed to stand. A label that fails its check is dropped rather than reworded into something that passes.

This is the weakest layer of the instrument, and we would rather tell you than have you find out. Naming fine-grained rhetorical devices reliably is an unsolved problem across the field, and published results for automated classification of specific techniques sit well below what trained human coders achieve. Read a device label as an observation with its evidence attached, to be checked against the quoted passage. The dimension scores and the construction pattern are the measurements we stand behind. The device names are the working out, shown so you can audit them.

When a label is dropped, it means the specific charge was not supported by the evidence offered for it. It never means the text is free of that mechanism, and no part of a report presents a dropped label as a finding.

4. Genre and disclosure modulation

Disclosure is credited where the evidence says it operates. When a text openly declares its persuasive purpose, the score discounts borrowed credibility and concealed commercial interest: the legitimacy-manipulation cluster. It does not discount emotional, identity-based, narrative, or logical mechanics. The disclosure literature supports the first effect; for the others, the published evidence is that recognition does not reliably reduce the mechanism's force, and a fallacy is not repaired by announcing it. We set these coefficients from that literature rather than fitting them, because no ground truth exists for the composite score, and fitting free parameters against no criterion is how unfounded weights get made. The values are deliberately coarse and falsifiable per mechanism.

Disclosure is read from the page rather than assumed. Where a publisher prints a sponsorship or paid-content label, the instrument detects it and accounts for it. Where a text argues for the organization whose platform it is published on, and presents itself as promotion or as opinion, the venue itself is treated as disclosing that interest, because a reader can see whose site they are on. Neither signal stands in for the other, and both are recorded in the report.

The reverse case rests on different evidence. When persuasive machinery operates inside a form the reader trusts as neutral, the score weights that incongruence upward, across all mechanisms. This is not a disclosure effect but a covertness effect, grounded in the deceptive-design literature, where concealment is definitional to the harm. Genre judgment is made per text, from the text, never from the outlet's reputation.

5. Calibration

Calibration is normed against a screened reference corpus of real published texts spanning news, analysis, opinion, official statements, advocacy, promotional content, and oratory. It is distributional: the scale's endpoints are anchored to the observed distribution of construction intensity in the wild, so a score locates a text among real texts rather than against an arbitrary ideal.

Calibration constants are refit only at named, versioned milestones, never adjusted quietly in between. At the first fit, screening consisted of human-vetted quarantine review and duplicate removal; the pipeline's automated extraction-failure flag was added to the screen afterwards. We reconstructed that original fit population, measured what the omission had cost, and committed the reconstruction and its receipt to the repository rather than describing them. The upper anchor of the scale has been refit once since, at such a milestone. Every report states its instrument version so a score is read against the calibration that produced it.

The corpus is documented under a versioned methodology with per-tranche sampling frames, and grows by defined cohorts rather than convenience sampling. Benchmark comparisons in reports are suppressed wherever the reference cell is too small to support them.

6. Validation

Three layers, each ongoing:

Anchor validation

A frozen suite of real-world fixture texts with known construction profiles: openly persuasive content whose legitimacy-based construction must be credited for its disclosure, buried-disclosure advertising that must trigger the incongruence weighting, labeled sponsored content that must not, wire copy that must sit low. Every instrument revision must pass the full suite before deployment.

Stability measurement

Scoring runs several independent passes over the same text and publishes their median. This narrows run-to-run variation; it does not eliminate it, and we publish that behaviour rather than claiming determinism. Two sources of variation are known and named. The instrument's own judgment moves between passes, which is why the median exists. And the underlying model can return different responses to identical requests, which we have measured and cannot remove. Only a saved report returns unchanged, and it does so by retrieval rather than by recomputation (see section 7).

Independent validation program

An independent linguist reviews the instrument under contract: checking that the evidence a report cites actually supports the claim attached to it, testing the device definitions against real texts, and assessing the methodology itself. The review is adversarial by design and has changed the instrument repeatedly. Several of the limitations stated on this page exist because that review found them. A larger inter-rater reliability study with independent academic coders is designed but is not yet running, and we will not describe it as validation until it has been conducted.

7. Reproducibility

A report you have already run is stable. It keeps its score, its verdict, its evidence, and the version stamp of the instrument that produced it, and its own URL reopens that saved report unchanged. The saved report is the artifact of record. If a particular analysis matters to you, keep or export it.

Submitting the same source again is a new measurement rather than a replay of the old one. Whether a resubmission returns the report on file or scores fresh depends on whether the instrument version, the configuration, and your access to the earlier report all still match; when the instrument is revised, existing reports stay retrievable at their own URLs while a new submission is scored by the current version. Live pages also change between fetches. Scores are therefore generation-scoped, comparisons across instrument versions are not valid, and every report states its version for exactly this reason.

Every report carries a reproducibility record naming the instrument version, the model configuration, and a fingerprint of the source text as it was fetched. The configuration is fixed and versioned; the exact model is named in every report's reproducibility certificate rather than on this page. Every report also links back to this page, so the methodology a report was produced under is always one click from the report itself.

8. Limitations

Stated plainly, because a measurement instrument that hides its limits is not one:

ScopeWe measure construction, not truth. A factually accurate text can score high; a false one can score low. The score is orthogonal to fact-checking.
ScopeTexts, not settings. Delivery context (a captive audience, a trusted speaker, a platform's amplification) is outside the instrument. Where pressure appears in the text itself, it is scored; the room it was read in is not visible to us.
VarianceA fresh run is a fresh measurement, not a replay. Analysing identical text again can return a different number. Several passes and a published median narrow that variation without removing it, and the underlying model can return different responses to identical requests. The second decimal is not meaningful. The saved report is the artifact of record.
ScopeEnglish-language, document-scale. The instrument is built and calibrated for full documents in English. Fragments, transcripts of visual media, and other languages are outside current calibration.
VarianceVersion-scoped scores. Instrument revisions can shift scores. Cross-version comparisons are invalid by design; the version stamp on every report is the boundary.
CalibrationDisclosure credit is narrow by design. Declaring persuasive intent lowers legitimacy-based construction scores; it does not lower most others. On a large share of scored texts the disclosure adjustment has no effect at all, because the mechanisms it moderates are not among that text's dominant clusters.
EvidenceGenre does not order scores. At equal construction, our measurements did not show labelled opinion reliably scoring below straight news. That was measured on an earlier generation of the instrument and has not been re-measured on the current one. We record it as a measured result rather than a marketing convenience, and we will restate it when it is measured again.
EvidenceDevice names are the weakest layer. Named rhetorical devices are proposed, checked in code against the conditions each device requires, and dropped when they fail that check. They are shown as observations with the passage they rest on so you can judge them yourself. They are not findings, and a dropped label never means the text is free of that mechanism.
CalibrationThe bottom of the scale is a floor. Scores sit on a fixed scale with a defined lower bound, and a text built with very little persuasive machinery lands at or near it. A score at the floor means the instrument measured little construction. It does not mean the analysis failed, and it does not mean nothing was examined.
CalibrationSeveral modulation coefficients are provisional: set to satisfy fixed anchor cases, pending fit at the next versioned calibration milestone.

9. What we don't publish, and why

The dimension rubric's level anchors, weighting, calibration constants, and fixture contents are proprietary. So are the device canon as an enumerated list and the conditions each device must satisfy. This is a deliberate line, not evasiveness: publishing the scoring key would let a text be written to it, which would end the instrument's usefulness to the clients who rely on it. We publish what the instrument measures, how it is grouped and grounded, how it is calibrated and validated, and where it fails. That is the standard a diligence reader needs. The recipe stays held back.

Frequently asked questions

Will the same URL show the same report?

A report you have already run keeps its own URL, and that URL reopens it unchanged, with the same score, evidence and version stamp. Submitting the source again is a separate question. If nothing relevant has changed you will get the report on file; if the instrument has been revised since, or the earlier report is not one you have access to, the submission is scored fresh under the current version and can land differently. First-time submission of a dynamic page captures it at a moment in time, and pasting the article text directly gives you control over exactly what is scored. If a particular report matters, keep or export it.

Same pasted text, same score?

A fresh analysis of pasted text is a new measurement, scored under the current instrument version. Several independent passes are run and the published score is their median, which narrows the variation between runs, but no run is guaranteed to repeat exactly and the second decimal is not meaningful.

Does a high score mean the text is false?

No. Construction and accuracy are independent. A factually accurate text can score high because of how it is built; a false one can score low because it simply asserts claims without persuasive architecture.

Can a score describe an outlet or author?

No. Texts, not entities: one text, one score, one point in time. Naming the outlet or author is attribution for the text you analyzed, not a verdict on the entity that produced it.

Questions about the methodology, calibration approach, or research applications: hello@undersignal.ai

Run a free report →