AnnouncementTurnkeeper is working toward an open standard for sharing safety intelligence across platforms.

WardBench · Turnkeeper Labs · Synthetic safety benchmark

Three warnings do not always mean three sources.

WardBench tests whether safety systems can trace warning signs back to their origins, preserve conflicting evidence, and avoid treating repeated information as independent confirmation.

12 synthetic casesassumption drivenNo live user data
Method and limits

WardBench uses controlled, synthetic test cases to evaluate how accurately Turnkeeper reconstructs safety signals. It does not use live customer data or authorize enforcement decisions.

WardBench v1.5 result

Naive counting fails 9 of 12 casesv1.5 · 12 cases

Lineage-aware reconstruction correctly handles all 12.

Naive signal count
75%
Lineage-aware reconstruction
0%
Ground-truth matches
12/12
Assumption-tagged
4/12
Frozen synthetic suite

Why it matters

Signal volume can create false confidence.

A safety reviewer needs to know whether several warnings confirm one another—or simply repeat the same underlying information.

01

Repeated evidence

One source can travel through several systems and return looking like consensus.

02

Missing context

Conflicting evidence, changed timelines, and revoked signals can alter the case.

03

Human decisions

Reviewers need an explainable case, not a larger pile of disconnected alerts.

What WardBench tests

A focused test for safer case reconstruction.

WardBench examines whether a system can preserve where evidence came from, how it changed, and what might point toward a different conclusion.

Available in this benchmarkv1.5 · gated
  • Twelve synthetic cases (WB01–WB12), including four assumption-tagged v1.5 additions
  • A direct comparison between naive counting and lineage-aware reconstruction
In scope
What comes after validationGated / roadmap
  • Practitioner-validated case set (next after sharp conversations)
  • Human-participant reviewer study (Labs research gate)
  • Model Lab persistence, live traffic, or prevalence claims
Not claimed

Help us challenge the benchmark.

We are inviting safety practitioners to test whether these synthetic cases reflect the difficult situations reviewers encounter in practice.