v1 · evidence hygiene

Citation Dedupe v1

How TWOG collapses duplicate citation refs, preserves merged provenance, and prevents repeated chunks from masquerading as independent evidence.

activeapplies to · briefs and candidate referencesclaims · provenance repair
What this governs

Citation Dedupe v1 keeps the evidence count honest. Multiple chunks can support a claim, but they should not look like multiple independent papers when they came from the same source.

Evidence volume is not evidence diversity.

01 · Collect refs

Source briefs and candidate snapshots may contain citation labels, source URLs, identifiers, research object IDs, titles, and chunk IDs.

02 · Build keys

The dedupe layer compares DOI, PMID, PMCID, NCT ID, research object ID, normalized title, and chunk provenance.

03 · Choose primary

Duplicate groups retain one primary public citation while preserving merged citation IDs and duplicate matches.

04 · Preserve support

The system keeps the supported claim, evidence kind, source sections, and chunk IDs so source density is not lost.

05 · Flag gaps

Missing identifiers, unresolved refs, stale validation refs, and weak provenance become visible repair tasks.

Why dedupe is needed

A single full-text article can produce many chunks and many local citation labels. Without dedupe, one paper can appear to be an entire evidence base.

What gets merged

Duplicate citation labels are merged by durable identifiers first, then research object IDs, then normalized titles when identifiers are missing.

What does not get erased

Chunk-level provenance, duplicate labels, source sections, and supported claims remain visible so reviewers can see both evidence diversity and evidence density.

How unresolved refs should read

Unresolved citation labels should be treated as public defects. They can remain visible temporarily, but they should trigger citation repair before a record is promoted.

What a reader can verify
Identifier keys
DOI, PMID, PMCID, NCT ID, research object ID, normalized title, and source URL.
Duplicate group
Primary citation ID, duplicate citation IDs, duplicate count, and matched keys.
Merged provenance
Research object IDs, chunk IDs, source sections, and source brief IDs.
Supported claim
The exact claim the citation is being used to support or challenge.
Repair status
Whether the public reference is resolved, partially resolved, or still unresolved.
How to interpret this method
  • A high citation count can collapse into one source after dedupe.
  • Duplicate chunks can strengthen source coverage but not independent support.
  • Missing identifiers should reduce public confidence until repaired.
  • Dedupe should preserve provenance, not hide it.
What this method does not certify

Citation Dedupe v1 improves evidence hygiene. It does not evaluate whether the source itself is high quality or sufficient for validation.