Methods review
Challenge the design, the propagation policies, and the preregistered hypotheses before anything is run.
Turnkeeper Labs · Paper 001
What changes downstream when one identity connection is wrong?
We are building a synthetic benchmark and metrics for measuring how one incorrect identity link transfers risk to unrelated entities, changes review decisions, and persists—or fails to disappear—after the link is removed.
Research in progress · Synthetic evidence · No held-out findings
Research question
How does the graph position of a single false identity link affect downstream contamination under predefined identity-to-risk propagation policies in synthetic identity graphs?
Conventional matching metrics tell us whether a link is correct. This study asks what that error changes after evidence begins moving through a system — and whether correction actually removes every dependent conclusion.
An uncertain identity association may guide investigation, but it must not silently become evidence that the linked person engaged in harmful behavior.
Synthetic method
Create a seeded synthetic graph with known entities, observations, and provenance.
Join an elevated-risk component to an unrelated benign component with one controlled edit.
Evaluate transparent propagation policies on identical clean and corrupted graph pairs.
Check whether scores, explanations, cases, and review decisions return to the clean state.
Measurement ledger
The study will not compress all downstream effects into one universal harm score.
How many entities, cases, scores, explanations, and decisions change.
How often a ground-truth benign entity is changed by the false link.
Whether an unrelated entity crosses a preregistered human-review threshold.
How much score moves to an unrelated entity across the injected link.
How much of an affected score depends on a path crossing the false link.
How much affected state returns to the clean counterfactual after correction.
Safety invariants
Each invariant is a property the study tries to break, not a property we claim as established.
An uncertain association must not become an unmarked risk signal on a linked entity.
Link confidence and risk conclusion stay distinct values, never collapsed into one score.
A single link has a declared, limited ceiling on how far its effect can travel.
Any affected output can be traced to the paths that depend on the link in question.
Removing a link removes every conclusion that depended on it, not only its score contribution.
Nothing in the study proposes autonomous action against any person or account.
Current evidence
Development runs test the benchmark. They are not held-out findings.
Deterministic synthetic graph generator
Clean, corrupted, and corrected graph triplets
Eight transparent propagation policies
Six downstream-contamination metrics
Full-recompute and link-record-only recovery comparison
Scenario cards and pre-registration draft v0.1
Contract tests and development-only analysis artifacts
Held-out confirmatory hypotheses
Explicit peripheral, hub, and bridge injection positions
Matched true-link positive-control / correct-link utility
Real-world prevalence or base rates
Production-system performance
Demographic fairness or subgroup effects
Before design freeze
Methodology review returned REVISE. The preregistration stays open and the held-out study remains closed until these are resolved.
Peripheral, hub, and bridge positions must be explicit experimental conditions, not incidental variation.
A matched true-link condition is required so contamination is read against benefit, not in isolation.
Policy comparisons need estimands defined before any run, with the contrast stated in advance.
The closest prior work must be located and cited before the preregistration is frozen.
Research boundary
The prevalence or magnitude of downstream harm in production systems
Which propagation policy is best across all systems and datasets
Causal effects on real-world people, moderators, or platform outcomes
Performance outside the declared synthetic graph families
Fairness across protected classes or sensitive attributes
A production prevention mechanism or real-world harm reduction
Progress log
Narrowed Paper 001 to one controlled false identity merge and its downstream effects.
Aug 11, 2026 · Complete
Added the generator, graph triplets, policies, metrics, recovery modes, tests, and runner.
Aug 11, 2026 · Complete
Identified missing injection-position and correct-link utility conditions before freeze.
Aug 11, 2026 · Revise
Resolve the open factors, freeze the preregistration, then verify the closest prior work.
Next · Planned
Open research
We welcome methodological challenge, domain criticism, and independent reproduction before we look at any held-out result.
Challenge the design, the propagation policies, and the preregistered hypotheses before anything is run.
Tell us where the synthetic graph families diverge from the evidence you work with in practice.
Run the generator and metrics yourself and report where our results fail to reproduce.
Review the method, challenge the assumptions, or reproduce the benchmark independently — before any confirmatory run.