Methodology

How Undersignal works

Undersignal measures how a text is constructed to persuade. Not whether its claims are true, and not who wrote it. This page explains what the instrument measures, how it is calibrated, how we validate it, and where its limits are.

Versioned instrument · Every report records the version that produced it

Instrument version 2.1 · Page updated 2026-07-29

1. What the score measures

The Undersignal Risk Score (1 to 10) quantifies rhetorical construction intensity: the density and coordination of persuasion mechanics present in a text. A high score does not mean a text is false, malicious, or unlawful. It means the text is built with heavier persuasive machinery, and that machinery is identified, named, and quoted in the report so a reader can weigh it.

We score texts, not entities. The same organization can publish a 2 and an 8 in the same week. A score attaches to one analyzed text at one point in time, never to an author, outlet, or brand.

2. What the instrument is grounded in

The instrument draws on the Persuasion Knowledge Model (Friestad & Wright, 1994), the sponsorship-disclosure literature, work on narrative transportation and fictionality labeling, risk-as-feelings and fear-appeal research, identity-protective cognition, and the deceptive-design literature. Where these bodies of evidence disagree about what disclosure changes, the instrument follows the evidence per mechanism rather than applying one rule everywhere.

Each report discloses the construction: the specific devices, where they appear, and how they coordinate. The reader's own judgment does the rest. We do not tell readers what to believe. We show them how the text is trying to make them believe it. Disclosure remains the product's operating principle, and the score credits disclosure exactly where the evidence supports it.

3. The measurement architecture

Every report shows 14 dimension scores: the 12 construction dimensions below, plus spread potential and disguise, which are reported on their own terms. Two further judgments, coercive pressure and genre, are scored signals rather than dimensions. The instrument separates these kinds of measurement because they behave differently, and a reader should know which is which.

Construction dimensions 12 dimensions · five clusters · drive the score

These measure how the text is built: how claims are sourced and framed, how emotion, fear, and identity are engaged, how the narrative is structured and controlled, and how authority is invoked. The clusters that dominate a text determine its construction pattern, the labeled fingerprint at the top of every report: Declarative, Neutral, Patterned, Elevated, or Structured. They also drive the composite score. The full index is below.

Coercive pressure Scored signal · enters the score

The engine separately judges how hard the text pushes the reader toward immediate action: urgency, compliance demands, and coercive calls to act. This judgment enters the score and is never discounted by disclosure. A hard-sell close operates on the reader even when the persuasive intent is fully declared, so pressure counts outside the genre modulation described in section 4. It is not a dimension; it is reported alongside them.

Spread potential Reported dimension · never scored

How far and fast a text is built to travel. It appears in every report's dimension scores as context for the reader, and it never enters the composite score.

Disguise Reported dimension · own axis

How much the text works to appear as something other than what it is. Disguise is scored on its own scale, separate from the 1 to 10 construction dimensions, and feeds the genre and disclosure judgment described in section 4.

Genre and disclosure signals Scored signals · modulate the score

These do not measure construction; they modulate how construction counts. The engine judges, per text, whether the text openly declares its persuasive purpose and whether it honors or betrays the genre it presents itself as. Section 4 explains how.

The 12 construction dimensions

Grouped by cluster. Each entry states what the dimension measures; the scoring rubric behind each one is proprietary (section 9).

Emotional ActivationCluster
Fear Amplification

How much the text leans on threat and danger to move the reader, relative to what its own evidence supports.

Emotional Loading

How much emotionally charged language carries the argument instead of the underlying facts.

Archetypal Framing

How far people and institutions are cast into story roles such as hero, villain, or victim to do the persuading.

Identity ExploitationCluster
Identity Pressure

The social cost the text attaches to disagreeing with it.

Tribal Signaling

How strongly the text frames its subject as us versus them and cues group loyalty.

Legitimacy ManipulationCluster
Authority Exploitation

How expert and institutional credibility is used to settle questions rather than inform them.

Manufactured Consensus

How much agreement is asserted or staged rather than demonstrated.

Money Context

How visibly financial interests connect to the narrative, and whether the beneficiaries are named.

Narrative ControlCluster
Suppression & Promotion

How competing explanations are sidelined or disqualified rather than engaged.

Centralized Narrative

How tightly every element of the text is organized to point at a single conclusion.

Narrative Structure

How much the content works as a story with characters and an arc rather than a set of claims.

Logical IntegrityCluster
Logical Fallacies

How much of the argument rests on reasoning that fails when stated plainly.

4. Genre and disclosure modulation

Disclosure is credited where the evidence says it operates. When a text openly declares its persuasive purpose, the score discounts borrowed credibility and concealed commercial interest: the legitimacy-manipulation cluster. It does not discount emotional, identity-based, narrative, or logical mechanics. The disclosure literature supports the first effect; for the others, the published evidence is that recognition does not reliably reduce the mechanism's force, and a fallacy is not repaired by announcing it. We set these coefficients from that literature rather than fitting them, because no ground truth exists for the composite score, and fitting free parameters against no criterion is how unfounded weights get made. The values are deliberately coarse and falsifiable per mechanism.

The reverse case rests on different evidence. When persuasive machinery operates inside a form the reader trusts as neutral, the score weights that incongruence upward, across all mechanisms. This is not a disclosure effect but a covertness effect, grounded in the deceptive-design literature, where concealment is definitional to the harm. Genre judgment is made per text, from the text, never from the outlet's reputation.

5. Calibration

Calibration is normed against a screened reference corpus of real published texts spanning news, analysis, opinion, official statements, advocacy, promotional content, and oratory. It is distributional: the scale's endpoints are anchored to the observed distribution of construction intensity in the wild, so a score locates a text among real texts rather than against an arbitrary ideal.

At the initial calibration, screening consisted of human-vetted quarantine review and duplicate removal; the pipeline's automated extraction-failure flag was added to the calibration screen after that fit. We reconstructed the original fit population and measured the effect of the omission: 1.3 percent of fit rows carried the flag, and removing them does not change either published constant. The reconstruction and its receipt are part of the repository, and future calibrations apply the full screen.

The corpus is documented under a versioned methodology with per-tranche sampling frames, and grows by defined cohorts rather than convenience sampling. Benchmark comparisons in reports are suppressed wherever the reference cell is too small to support them.

6. Validation

Three layers, each ongoing:

Anchor validation

A frozen suite of real-world fixture texts with known construction profiles: openly persuasive content whose legitimacy-based construction must be credited for its disclosure, buried-disclosure advertising that must trigger the incongruence weighting, labeled sponsored content that must not, wire copy that must sit low. Every instrument revision must pass the full suite before deployment.

Stability measurement

Scoring uses multiple independent passes with median aggregation, which narrows run-to-run variance; it does not eliminate it. Repeat analyses of the same text produce scores within a narrow band, and verdict and band classifications are substantially more stable than the decimal score. We publish this behavior rather than claiming determinism: only a saved report returns unchanged on retrieval (see section 7).

Independent validation program

A structured external validation effort, codebook-based inter-rater reliability on text classification and dimension-level human agreement studies conducted with independent academic raters, is in progress. Results will be published to this page as they are completed.

7. Reproducibility

A URL-based report is locked to its source: re-running the same URL surfaces the saved, versioned report. That report, its score, verdict, and evidence, stays retrievable unchanged at its URL, and every report records the instrument version that produced it.

A fresh analysis of pasted text is a new measurement, not a replay. Live pages change between fetches, the instrument itself is versioned and improves, and independent scoring passes vary within a band. Scores are therefore generation-scoped: comparisons across instrument versions are not valid, and reports state their version for exactly this reason.

8. Limitations

Stated plainly, because a measurement instrument that hides its limits is not one:

We measure construction, not truth. A factually accurate text can score high; a false one can score low. The score is orthogonal to fact-checking.
Texts, not settings. Delivery context (a captive audience, a trusted speaker, a platform's amplification) is outside the instrument. Where pressure appears in the text itself, it is scored; the room it was read in is not visible to us.
Run variance exists. Fresh analyses of identical text vary within a narrow band. Verdicts and bands are substantially more stable than the decimal score; second decimals are not meaningful.
English-language, document-scale. The instrument is built and calibrated for full documents in English. Fragments, transcripts of visual media, and other languages are outside current calibration.
Version-scoped scores. Instrument revisions can shift scores. Cross-version comparisons are invalid by design; the version stamp on every report is the boundary.
Disclosure credit is narrow by design. Declaring persuasive intent lowers legitimacy-based construction scores; it does not lower most others. On a large share of scored texts the disclosure adjustment has no effect at all, because the mechanisms it moderates are not among that text's dominant clusters.
Genre does not order scores. At equal construction, our measurements do not show labeled opinion reliably scoring below straight news. We record this as a measured result, not a marketing inconvenience.
Several modulation coefficients are provisional: set to satisfy fixed anchor cases, pending fit at the next versioned calibration milestone.

9. What we don't publish, and why

The dimension rubric's level anchors, weighting, calibration constants, and fixture contents are proprietary. This is a deliberate line, not evasiveness: publishing the scoring key would let any text be engineered to it, which would end the instrument's usefulness to the clients who rely on it. We publish what the instrument measures, how it is grouped and grounded, how it is calibrated and validated, and where it fails. That is the standard a diligence reader needs. The recipe stays held back.

Frequently asked questions

Will the same URL show the same report?

Yes. If you've submitted this URL before, you're seeing the saved result. Reports are saved and versioned, so resubmitting the same URL returns the report on file rather than a fresh run. First-time submission of a dynamic page captures it at a moment in time; pasting the article text directly gives you control over exactly what is scored.

Same pasted text, same score?

A fresh analysis of pasted text is a new measurement, scored under the current instrument version. Verdict and band classifications are substantially more stable across runs than the decimal score, but none of them is guaranteed to repeat exactly.

Does a high score mean the text is false?

No. Construction and accuracy are independent. A factually accurate text can score high because of how it is built; a false one can score low because it simply asserts claims without persuasive architecture.

Can a score describe an outlet or author?

No. Texts, not entities: one text, one score, one point in time. Naming the outlet or author is attribution for the text you analyzed, not a verdict on the entity that produced it.

Questions about the methodology, calibration approach, or research applications: hello@undersignal.ai

Run a free report →