v1 · evidence hygiene
Citation Dedupe v1
How TWOG collapses duplicate citation refs, preserves merged provenance, and prevents repeated chunks from masquerading as independent evidence.
Citation Dedupe v1 keeps the evidence count honest. Multiple chunks can support a claim, but they should not look like multiple independent papers when they came from the same source.
Evidence volume is not evidence diversity.
A single full-text article can produce many chunks and many local citation labels. Without dedupe, one paper can appear to be an entire evidence base.
Duplicate citation labels are merged by durable identifiers first, then research object IDs, then normalized titles when identifiers are missing.
Chunk-level provenance, duplicate labels, source sections, and supported claims remain visible so reviewers can see both evidence diversity and evidence density.
Unresolved citation labels should be treated as public defects. They can remain visible temporarily, but they should trigger citation repair before a record is promoted.
- Identifier keys
- DOI, PMID, PMCID, NCT ID, research object ID, normalized title, and source URL.
- Duplicate group
- Primary citation ID, duplicate citation IDs, duplicate count, and matched keys.
- Merged provenance
- Research object IDs, chunk IDs, source sections, and source brief IDs.
- Supported claim
- The exact claim the citation is being used to support or challenge.
- Repair status
- Whether the public reference is resolved, partially resolved, or still unresolved.
- A high citation count can collapse into one source after dedupe.
- Duplicate chunks can strengthen source coverage but not independent support.
- Missing identifiers should reduce public confidence until repaired.
- Dedupe should preserve provenance, not hide it.
Citation Dedupe v1 improves evidence hygiene. It does not evaluate whether the source itself is high quality or sufficient for validation.