Overview
As commercial genetic testing becomes increasingly accessible, the sheer volume of data generated by multi-gene panels and exome sequencing has exploded. While modern laboratory techniques excel at identifying molecular deviations from the reference genome, determining the actual clinical consequence of those deviations remains one of the most profound challenges in modern medicine. A sequence report may identify a dozen genetic mutations, but identifying a mutation is not a diagnosis; establishing pathogenicity is the diagnosis. This is where the public archive known as ClinVar becomes an indispensable tool for the clinical geneticist. Maintained by the National Center for Biotechnology Information (NCBI) at the National Institutes of Health (NIH), ClinVar is a globally centralized database that aggregates submissions from commercial testing laboratories, academic research institutions, and specialized clinical clinics regarding the relationship between human genetic variations and observable health statuses. It tracks whether a specific variant (e.g., a particular missense mutation in the MYBPC3 gene) has been historically classified as Pathogenic, Benign, or, as is often the case, a Variant of Uncertain Significance (VUS). Without access to this aggregated historical context, interpreting a modern genetic report is essentially impossible. However, manually cross-referencing a lengthy PDF genetic report against the ClinVar database is a tedious, error-prone, and time-consuming process. A clinician must manually type complex alphanumeric variant nomenclatures into a web portal, risking transcription errors that could pull up the clinical history for an entirely different genetic mutation. Sanjeevni automates this critical verification step entirely, integrating the vast wealth of ClinVar data directly into the physician's unified workspace.
The Variant Classification Challenge
As commercial genetic testing becomes increasingly accessible, the sheer volume of data generated by multi-gene panels and exome sequencing has exploded. While modern laboratory techniques excel at identifying molecular deviations from the reference genome, determining the actual clinical consequence of those deviations remains one of the most profound challenges in modern medicine. A sequence report may identify a dozen genetic mutations, but identifying a mutation is not a diagnosis; establishing pathogenicity is the diagnosis. This is where the public archive known as ClinVar becomes an indispensable tool for the clinical geneticist.
Automated Text-Based Extraction and Cross-Referencing
Sanjeevni’s integration with ClinVar is built upon our advanced Natural Language Processing (NLP) architecture. When a physician or patient uploads a standard commercial genetic testing PDF report, Sanjeevni’s engine scans the document specifically searching for standardized genetic nomenclature and gene symbols. It identifies the targeted gene, the specific molecular alteration, and the classification provided by the issuing laboratory at the time the report was printed.
Monitoring VUS Reclassifications
The dynamic nature of genetic classification creates a significant logistical nightmare for busy genetics clinics. When a patient is discharged with an unresolved VUS, the clinic bears a moral, and increasingly legal, responsibility to inform the patient if that variant is ever reclassified to Pathogenic in the future. However, actively monitoring thousands of historic patient files for potential ClinVar updates is an administrative impossibility for human staff.