AnnouncementTurnkeeper is working toward an open standard for sharing safety intelligence across platforms.

Turnkeeper Labs · Paper 002

Two Signals, One Source

When does apparent corroboration become double-counted evidence?

We are designing a synthetic-first study of what changes when several safety signals look independent but descend from the same observation, model, dataset, or reporting chain.

Synthetic first · WardBench live · No findings claimed

Evidence lineageDependence visible
SourceA
01Direct
02Derived
03Duplicate
Visible signals
3
Independent sources
1
Fig. 01 — three visible signals can still contain only one independent source
Current stageBenchmark available
Benchmark productWardBench WB01–WB12 (v1.5)
Reviewer studyGated
FindingsNone claimed

Research question

Count independent support, not interface objects.

How does making source dependence visible affect confidence and escalation decisions when multiple privacy-minimized signals appear to corroborate the same hypothesis?

A system can display three signals even when all three inherit the same mistake. Paper 002 isolates that gap: the difference between the number of observations a reviewer can see and the number of independent sources those observations actually represent.

Apparent agreement must not become stronger evidence unless the supporting sources are independent enough to justify the added weight.

Synthetic-first method

Matched evidence. One controlled difference.

The first stage uses generated cases only. A later reviewer study remains gated and has not recruited participants.

  1. Generate a known evidence history

    Create synthetic cases with known sources, derivation paths, contradictions, and ground-truth dependence relationships.

  2. Hold content constant and vary lineage

    Present matched evidence sets as independent, derived, or duplicated so only source dependence changes.

  3. Vary how provenance is shown

    Compare hidden lineage, a source list, and an explicit provenance graph while preserving identical underlying signals.

  4. Measure decisions without blending outcomes

    Keep confidence, escalation, dependence detection, reversal after correction, and review cost as separate measures.

Experimental factors

Four factors define the benchmark.

Signal count and source count are manipulated separately so multiplicity cannot stand in for independence.

Signal count

One, two, or four visible signals while the number of independent sources is controlled separately.

Dependence

Genuinely independent observations, transformations of one observation, or exact and near duplicates.

Lineage display

No source relationship, a flat source list, or an explicit derivation graph with confidence and timestamps.

Source quality

Matched high-reliability, mixed-reliability, and unknown-reliability sources without averaging them together.

Measurement ledger

Six outcomes, kept separate.

No composite safety score will hide which part of the decision changed.

  1. False corroboration rate

    How often multiple dependent signals are treated as support from multiple independent sources.

  2. Confidence inflation

    The confidence added by duplicate or derived signals beyond the matched single-source condition.

  3. Escalation flips

    Whether apparent corroboration moves a case across a preregistered review threshold.

  4. Lineage sensitivity

    How much a decision changes when the same evidence is shown with its source relationships made explicit.

  5. Correction recovery

    Whether removing or revoking a parent signal removes every conclusion that depended on it.

  6. Review cost

    Time, inspection steps, and uncertainty introduced by each provenance representation.

Preregistered direction

What the study is designed to falsify.

These are hypotheses to challenge before design freeze, not results.

  1. H1 · Multiplicity without independence inflates support

    When lineage is hidden, dependent signals will produce more threshold crossings than the matched single-source condition.

  2. H2 · Explicit lineage reduces false corroboration

    A provenance graph will reduce double-counting more than a flat source list while preserving genuinely independent support.

  3. H3 · Mixed source quality magnifies the error

    Dependence mistakes will be largest when a low-quality parent produces several polished downstream signals.

  4. H4 · Provenance-aware decisions recover more completely

    Revoking a parent signal will reverse more dependent conclusions when lineage survives aggregation.

Safety invariants

What evidence fusion must preserve.

Each invariant is a property the benchmark will try to break.

  1. Signals are not conclusions

    Multiple observations may guide investigation, but their count does not establish a finding.

  2. Lineage survives aggregation

    Every derived claim remains traceable to the source observations that support it.

  3. Dependence is not hidden confidence

    A transformed or duplicated signal cannot silently add the weight of a new independent source.

  4. Correction travels downstream

    A correction, expiry, or revocation reaches every conclusion that inherited the original signal.

  5. Contradictions remain visible

    Counterevidence is preserved beside supporting evidence instead of being flattened into one score.

  6. No automatic enforcement

    No benchmark output or reviewer response grants authority to act against a person or account.

Current evidence

A defensible plan, not a finished paper.

The research object exists. The benchmark and study results do not.

Defined now

  • Research question and claim boundary

  • Two-stage synthetic-first study structure

  • Factor and outcome ledger

  • Draft hypotheses and safety invariants

  • Public literature basis

  • Frozen WardBench synthetic suite (WB01–WB12, v1.5 assumption-driven) at /wardbench

Not started

  • Preregistered analysis and decision thresholds

  • Held-out confirmatory run

  • Human-participant recruitment or data collection

  • Production-system or real-world safety claims

Reviewer-study gate

No recruitment before these conditions are met.

Human-participant work is a separate phase, not an automatic continuation of the synthetic benchmark.

  • A frozen synthetic benchmark establishes that the manipulation works as intended.

  • A documented ethics, privacy, consent, and data-retention review approves the protocol.

  • Recruitment avoids live accusations, active cases, customer content, and identifiable minors.

  • The study can stop without affecting a participant's work, access, or employment.

Research boundary

What this study cannot establish.

  • That dependent signals cause real-world reviewers to make a specific decision

  • That a provenance graph is the best interface across every workflow

  • The prevalence of duplicate or derived safety signals in production systems

  • The correctness of any live allegation, identity link, or enforcement decision

  • Causal effects on platform safety, people, or protected groups

Literature basis

Established pieces. An unresolved system question.

The paper connects evidence-dependence research to privacy-minimized safety review; it does not claim the underlying problem is newly discovered.

arXiv:2311.08592

Structured adversarial evaluation (inspiration)

AART shows how to generate systematic adversarial examples for evaluation. WardBench applies that discipline to synthetic *cases*—dependence, trajectory, and counterevidence—not jailbreak prompts.

Open source →
CogSci 2024

Source dependence in human evidence judgments

Strittmatter, Pilditch, and Lagnado report that people can perceive source dependence yet often fail to discount it fully.

Open source →
eLife 2025

Correlated evidence in perceptual decisions

Tardiff, Kang, and Gold test how people combine evidence whose samples share a common source.

Open source →
OASIS STIX 2.1

Machine-readable relationship vocabulary

STIX 2.1 defines relationships such as derived-from and duplicate-of that can carry lineage without deciding how a reviewer should weight it.

Open source →
Technology Coalition

Shared signals require local investigation

The Lantern program describes exchanged signals as clues rather than proof and leaves investigation to the receiving company.

Open source →

Help make Paper 002 harder to fool.

Walk the frozen WardBench suite, then challenge the measures before any human-participant study.