AnnouncementTurnkeeper is working toward an open standard for sharing safety intelligence across platforms.

Turnkeeper Labs · Paper 001

One Bad Link

What changes downstream when one identity connection is wrong?

We are building a synthetic benchmark and metrics for measuring how one incorrect identity link transfers risk to unrelated entities, changes review decisions, and persists—or fails to disappear—after the link is removed.

Research in progress · Synthetic evidence · No held-out findings

Counterfactual graph statesOne controlled edit
Clean
+ One false link
Corrected
The graph is deliberately simple here. The benchmark varies topology, size, link confidence, position, policy, and recovery behavior.
Current stagePhase 1 revise — freeze blocked
Held-out studyNot run
DataSynthetic only
FindingsNone claimed

Research question

Measure the consequence, not only the match.

How does the graph position of a single false identity link affect downstream contamination under predefined identity-to-risk propagation policies in synthetic identity graphs?

Conventional matching metrics tell us whether a link is correct. This study asks what that error changes after evidence begins moving through a system — and whether correction actually removes every dependent conclusion.

An uncertain identity association may guide investigation, but it must not silently become evidence that the linked person engaged in harmful behavior.

Synthetic method

One edit. Paired counterfactuals. Inspectable policies.

  1. Generate a clean graph

    Create a seeded synthetic graph with known entities, observations, and provenance.

  2. Inject one false link

    Join an elevated-risk component to an unrelated benign component with one controlled edit.

  3. Compare downstream outputs

    Evaluate transparent propagation policies on identical clean and corrupted graph pairs.

  4. Remove the link and test recovery

    Check whether scores, explanations, cases, and review decisions return to the clean state.

Measurement ledger

Six outcomes, kept separate.

The study will not compress all downstream effects into one universal harm score.

  1. False-link blast radius

    How many entities, cases, scores, explanations, and decisions change.

  2. Innocent-entity contamination

    How often a ground-truth benign entity is changed by the false link.

  3. Intervention flips

    Whether an unrelated entity crosses a preregistered human-review threshold.

  4. Transferred risk

    How much score moves to an unrelated entity across the injected link.

  5. Attribution concentration

    How much of an affected score depends on a path crossing the false link.

  6. Recovery completeness

    How much affected state returns to the clean counterfactual after correction.

Safety invariants

What the benchmark is built to pressure-test.

Each invariant is a property the study tries to break, not a property we claim as established.

  1. No silent transfer

    An uncertain association must not become an unmarked risk signal on a linked entity.

  2. Separate confidence

    Link confidence and risk conclusion stay distinct values, never collapsed into one score.

  3. Bounded influence

    A single link has a declared, limited ceiling on how far its effect can travel.

  4. Counterfactual traceability

    Any affected output can be traced to the paths that depend on the link in question.

  5. Complete reversibility

    Removing a link removes every conclusion that depended on it, not only its score contribution.

  6. No automatic enforcement

    Nothing in the study proposes autonomous action against any person or account.

Current evidence

Built enough to test the design — not enough to claim a result.

Development runs test the benchmark. They are not held-out findings.

Implemented locally

  • Deterministic synthetic graph generator

  • Clean, corrupted, and corrected graph triplets

  • Eight transparent propagation policies

  • Six downstream-contamination metrics

  • Full-recompute and link-record-only recovery comparison

  • Scenario cards and pre-registration draft v0.1

  • Contract tests and development-only analysis artifacts

Not yet tested

  • Held-out confirmatory hypotheses

  • Explicit peripheral, hub, and bridge injection positions

  • Matched true-link positive-control / correct-link utility

  • Real-world prevalence or base rates

  • Production-system performance

  • Demographic fairness or subgroup effects

Before design freeze

Phase 1 verdict: revise.

Methodology review returned REVISE. The preregistration stays open and the held-out study remains closed until these are resolved.

  1. Injection position factor

    Peripheral, hub, and bridge positions must be explicit experimental conditions, not incidental variation.

  2. Correct-link utility control

    A matched true-link condition is required so contamination is read against benefit, not in isolation.

  3. Comparative estimands

    Policy comparisons need estimands defined before any run, with the contrast stated in advance.

  4. Literature verification

    The closest prior work must be located and cited before the preregistration is frozen.

Research boundary

What this study cannot establish.

  • The prevalence or magnitude of downstream harm in production systems

  • Which propagation policy is best across all systems and datasets

  • Causal effects on real-world people, moderators, or platform outcomes

  • Performance outside the declared synthetic graph families

  • Fairness across protected classes or sensitive attributes

  • A production prevention mechanism or real-world harm reduction

Progress log

The work, including the gates.

  1. Research direction selected

    Narrowed Paper 001 to one controlled false identity merge and its downstream effects.

    Aug 11, 2026 · Complete

  2. Development skeleton implemented

    Added the generator, graph triplets, policies, metrics, recovery modes, tests, and runner.

    Aug 11, 2026 · Complete

  3. Phase 1 methodology review

    Identified missing injection-position and correct-link utility conditions before freeze.

    Aug 11, 2026 · Revise

  4. Design freeze and literature investigation

    Resolve the open factors, freeze the preregistration, then verify the closest prior work.

    Next · Planned

Open research

Help us make the study harder to fool.

We welcome methodological challenge, domain criticism, and independent reproduction before we look at any held-out result.

Methods review

Challenge the design, the propagation policies, and the preregistered hypotheses before anything is run.

Domain review

Tell us where the synthetic graph families diverge from the evidence you work with in practice.

Independent reproduction

Run the generator and metrics yourself and report where our results fail to reproduce.

Contribute to the One Bad Link study.

Review the method, challenge the assumptions, or reproduce the benchmark independently — before any confirmatory run.