When a DOI is present
An exact DOI agreement is the strongest identity signal. Other fields are still compared so contradictory metadata remains visible instead of being hidden by the identifier.
VERIFICATION METHODOLOGY
CiteWise separates text parsing, scholarly identity resolution, and field-level verification. The result is a reviewable assessment, not a claim that a confidence number is ground truth.
Last reviewed: September 9, 2026
This page describes the current production approach. Matching thresholds still require calibration against a larger labeled corpus.
PIPELINE
The original reference remains immutable. Normalization protects identifiers while removing noise that would harm retrieval.
Rules extract DOI and bibliographic fields. Parsing confidence measures extraction quality, not whether the cited work is real.
CiteWise issues bounded queries to relevant scholarly providers using the strongest available identifiers and metadata.
Candidates are grouped and ranked as complete works. Fields from competing works are not silently combined.
For the selected identity, title, authors, year, venue, pages, volume, issue, and identifiers receive separate evidence assessments.
The response exposes the decision, confidence, contradictions, candidates, diagnostics, and deterministic formatted outputs.
SOURCES
Coverage and metadata differ by discipline and provider. CiteWise records where each value came from, keeps provider failures distinct from empty results, and does not assume that the most complete record is automatically correct.
MATCHING
An exact DOI agreement is the strongest identity signal. Other fields are still compared so contradictory metadata remains visible instead of being hidden by the identifier.
Title similarity, author overlap, publication year, venue, volume, issue, and pages contribute according to availability and reliability. Missing fields are not treated as contradictions.
Contradictions reduce confidence. Near-tied candidates produce an ambiguous decision rather than an arbitrary winner.
Timeouts, rate limits, and upstream errors are diagnostic states. They do not become a false “not found” result.
DECISIONS
| Decision | Interpretation |
|---|---|
| Verified | The leading work has strong, consistent evidence and no material unresolved contradiction. |
| Probably verified | The leading work is convincing, but the evidence is not strong enough for the highest decision. |
| Ambiguous | Two or more plausible works remain too close to choose safely. |
| Conflicting metadata | A likely identity exists, but important fields disagree across the input or evidence. |
| Not found | Available providers returned no sufficiently plausible work. This does not prove the work does not exist. |
| Insufficient evidence | The input or available provider data is too limited for a defensible decision. |
| Possible hallucination | The reference contains enough detail to search, but no coherent scholarly identity can be supported. |
AI BOUNDARY
Deterministic extraction and provider retrieval run first. When those signals are incomplete, an optional model can propose bounded missing-field or search hints. Suggested values remain labeled as suggestions and require provider evidence before they can support verification. The model is not treated as a bibliographic authority.
LIMITATIONS
Start with 10 provider-backed checks each month.