A Methodological Framework for Real-World Performance Studies of Clinical Variant Classification Platforms at Early Organizational Stages

Vladimir Mitev

Under consideration at a peer-reviewed journal. The manuscript has not been peer reviewed.

Published
1 July 2026
Record type
Preprint
Version
Version 1.0

Abstract

Abstract

Early-stage variant classification platforms need a performance method that is stronger than an internal comparison but feasible before a multi-adjudicator regulatory study. A single peer laboratory is difficult to treat as ground truth when qualified laboratories can disagree. Formal adjudication, however, requires people and quality infrastructure that a small developer may not yet have.

The paper proposes a three-layer performance model, an evidence-strength-ordered approach to ground truth, a six-category disposition taxonomy for disagreements and a phased route toward later regulatory evidence. Each observation is assigned to analytical, classification or clinical performance so that the source of a discrepancy remains visible.

The method is demonstrated on 97 scored cases across three classifier versions. The demonstration shows how the framework operates. It does not validate the platform or establish superiority over another system.

Evidence boundary

Evidence boundary

This is a methodological proposal with a single-platform demonstration. It uses one peer comparator, develops the taxonomy on the same cohorts used to illustrate it and carries conflicts of interest that require independent replication. The preprint has not been peer reviewed.

01

Four parts of the framework

The performance model separates analytical, classification and clinical questions. Ground truth is assembled from more than one source and weighted by evidence strength. Disagreements are assigned to six methodological dispositions with different corrective actions. A phased pathway preserves early evidence for later, more formal evaluation.

02

What the demonstration contributes

The 97 scored cases show how a team can record discrepancies, attribute them to the correct performance layer and avoid treating every platform-comparator difference as a classifier defect.

The paper claims a method. Validation would require independent cohorts, broader comparators and adjudication that is separate from the team that developed the framework.