Signal count
One, two, or four visible signals while the number of independent sources is controlled separately.
Turnkeeper Labs · Paper 002
When does apparent corroboration become double-counted evidence?
We are designing a synthetic-first study of what changes when several safety signals look independent but descend from the same observation, model, dataset, or reporting chain.
Synthetic first · WardBench live · No findings claimed
Research question
How does making source dependence visible affect confidence and escalation decisions when multiple privacy-minimized signals appear to corroborate the same hypothesis?
A system can display three signals even when all three inherit the same mistake. Paper 002 isolates that gap: the difference between the number of observations a reviewer can see and the number of independent sources those observations actually represent.
Apparent agreement must not become stronger evidence unless the supporting sources are independent enough to justify the added weight.
Synthetic-first method
The first stage uses generated cases only. A later reviewer study remains gated and has not recruited participants.
Create synthetic cases with known sources, derivation paths, contradictions, and ground-truth dependence relationships.
Present matched evidence sets as independent, derived, or duplicated so only source dependence changes.
Compare hidden lineage, a source list, and an explicit provenance graph while preserving identical underlying signals.
Keep confidence, escalation, dependence detection, reversal after correction, and review cost as separate measures.
Experimental factors
Signal count and source count are manipulated separately so multiplicity cannot stand in for independence.
One, two, or four visible signals while the number of independent sources is controlled separately.
Genuinely independent observations, transformations of one observation, or exact and near duplicates.
No source relationship, a flat source list, or an explicit derivation graph with confidence and timestamps.
Matched high-reliability, mixed-reliability, and unknown-reliability sources without averaging them together.
Measurement ledger
No composite safety score will hide which part of the decision changed.
How often multiple dependent signals are treated as support from multiple independent sources.
The confidence added by duplicate or derived signals beyond the matched single-source condition.
Whether apparent corroboration moves a case across a preregistered review threshold.
How much a decision changes when the same evidence is shown with its source relationships made explicit.
Whether removing or revoking a parent signal removes every conclusion that depended on it.
Time, inspection steps, and uncertainty introduced by each provenance representation.
Preregistered direction
These are hypotheses to challenge before design freeze, not results.
When lineage is hidden, dependent signals will produce more threshold crossings than the matched single-source condition.
A provenance graph will reduce double-counting more than a flat source list while preserving genuinely independent support.
Dependence mistakes will be largest when a low-quality parent produces several polished downstream signals.
Revoking a parent signal will reverse more dependent conclusions when lineage survives aggregation.
Safety invariants
Each invariant is a property the benchmark will try to break.
Multiple observations may guide investigation, but their count does not establish a finding.
Every derived claim remains traceable to the source observations that support it.
A transformed or duplicated signal cannot silently add the weight of a new independent source.
A correction, expiry, or revocation reaches every conclusion that inherited the original signal.
Counterevidence is preserved beside supporting evidence instead of being flattened into one score.
No benchmark output or reviewer response grants authority to act against a person or account.
Current evidence
The research object exists. The benchmark and study results do not.
Research question and claim boundary
Two-stage synthetic-first study structure
Factor and outcome ledger
Draft hypotheses and safety invariants
Public literature basis
Frozen WardBench synthetic suite (WB01–WB12, v1.5 assumption-driven) at /wardbench
Preregistered analysis and decision thresholds
Held-out confirmatory run
Human-participant recruitment or data collection
Production-system or real-world safety claims
Reviewer-study gate
Human-participant work is a separate phase, not an automatic continuation of the synthetic benchmark.
A frozen synthetic benchmark establishes that the manipulation works as intended.
A documented ethics, privacy, consent, and data-retention review approves the protocol.
Recruitment avoids live accusations, active cases, customer content, and identifiable minors.
The study can stop without affecting a participant's work, access, or employment.
Research boundary
That dependent signals cause real-world reviewers to make a specific decision
That a provenance graph is the best interface across every workflow
The prevalence of duplicate or derived safety signals in production systems
The correctness of any live allegation, identity link, or enforcement decision
Causal effects on platform safety, people, or protected groups
Literature basis
The paper connects evidence-dependence research to privacy-minimized safety review; it does not claim the underlying problem is newly discovered.
AART shows how to generate systematic adversarial examples for evaluation. WardBench applies that discipline to synthetic *cases*—dependence, trajectory, and counterevidence—not jailbreak prompts.
Open source →Strittmatter, Pilditch, and Lagnado report that people can perceive source dependence yet often fail to discount it fully.
Open source →Tardiff, Kang, and Gold test how people combine evidence whose samples share a common source.
Open source →STIX 2.1 defines relationships such as derived-from and duplicate-of that can carry lineage without deciding how a reviewer should weight it.
Open source →The Lantern program describes exchanged signals as clues rather than proof and leaves investigation to the receiving company.
Open source →Walk the frozen WardBench suite, then challenge the measures before any human-participant study.