Translate this page into:
Bioinformatic tools in the clinical diagnosis of genetic disorders in children: Focus on next-generation sequencing
-
Received: ,
Accepted: ,
How to cite this article: Sait H, Adarsha N, Pandey M. Bioinformatic tools in the clinical diagnosis of genetic disorders in children: Focus on next-generation sequencing. J Pediatr Endocrinol Diabetes. doi: 10.25259/JPED_35_2026
Abstract
Advances in next-generation sequencing (NGS) have transformed the diagnostic approach to pediatric endocrine disorders, facilitating the detection of monogenic causes and contributing to the understanding of complex endocrine traits. Although specialized laboratories largely perform genomic data analysis, clinicians require a working understanding of variant interpretation to critically evaluate reports and integrate genetic findings into patient care. Bioinformatic tools play a crucial role in supporting this process by enabling systematic variant filtering, prioritization, and clinical interpretation. While bioinformatic tools are utilized across all stages of NGS analysis, including primary and secondary processing, this review focuses specifically on the clinically relevant aspects of tertiary analysis, particularly variant interpretation and prioritization. Emphasis is placed on resources most applicable to clinicians in routine practice. A representative clinical scenario is used to illustrate the practical application of these tools in translating genomic data into clinically meaningful insights. In conclusion, familiarity with these tools will help clinicians translate genomic data into meaningful clinical applications.
Keywords
Bioinformatics
Next-generation sequencing
Pediatric endocrine disorders
Phenotype-driven analysis
Variant interpretation
INTRODUCTION
Genetic disorders contribute significantly to pediatric endocrine disorders, with many conditions arising from monogenic defects affecting hormone synthesis, binding proteins, transcription factors, ion channels, membrane receptors, and intracellular signal transduction pathways[1] as well as from complex polygenic variation. To date, approximately 450 endocrine disorders have been identified in which genetic factors contribute substantially to disease pathogenesis.[2]
Conventionally, the diagnosis of endocrine disorders has relied on clinical, biochemical, and imaging-based evaluation. However, with advances in molecular genetics and the increasing availability of next-generation sequencing (NGS) —including gene panels, whole-exome sequencing (WES), and whole-genome sequencing (WGS) —genetic analysis has become integral to the diagnostic workup, especially in cases with inconclusive findings or associated extra-endocrine manifestations.
While genetic test reports typically include variant classifications based on established guidelines, interpretations may vary across laboratories due to differences in classification approaches, available evidence, database usage, and interpreter-dependent factors. A study has reported discordance in variant pathogenicity classification in up to 46% of cases, with 37% of these discrepancies having potential clinical impact.[3] Consequently, it is important for clinicians to have a basic understanding of variant interpretation, to critically evaluate reports, correlate findings with the patient’s phenotype, and, where appropriate, contribute to variant reclassification over time.
This review provides an overview of selected bioinformatic resources and tools useful to clinicians for interpreting genetic variants and understanding genomic reports for monogenic pediatric endocrine disorders, particularly those identified through NGS-based tests. Rather than providing an exhaustive technical discussion, this review focuses on practical tools that can assist clinicians in the routine interpretation of genetic test reports.
CASE SCENARIO
A 5-month-old male infant, first child of a nonconsanguineous couple with an unremarkable antenatal and perinatal history, presented with poor feeding, lethargy, and decreased activity following a minor illness. He had recurrent focal seizures with secondary generalization since 3 months of age (3–4 episodes/week), partially controlled with levetiracetam (20 mg/kg/day), along with global developmental delay. Social smile and maternal recognition were achieved at 4 months; he had visual fixation and tracking but no neck control or vocalization. Hearing was normal. There was no relevant family history.
On examination, he was drowsy, dehydrated, and tachypneic with acidotic breathing. His weight was 5.5 kg (–2.7 standard deviation [SD]) and length 62 cm (–1.9 SD). No dysmorphism was noted. Neurological examination showed generalized hypotonia, poor head control, and exaggerated deep tendon reflexes.
Investigations revealed hyperglycemia (random plasma glucose 580 mg/dL), hyponatremia (sodium 125 mmol/L), normal potassium levels (4.5 mmol/L), and high anion gap metabolic acidosis (pH 7.12, bicarbonate 9 mmol/L, partial pressure of carbon dioxide 20 mmHg, and anion gap 25) with elevated serum ketones, consistent with moderate diabetic ketoacidosis. Glycated hemoglobin was 12.2%, and fasting C-peptide was <0.1 ng/mL. Magnetic resonance imaging of the brain revealed mild cerebral atrophy, and electroencephalography showed hypsarrhythmia.
The child was managed with intravenous fluids and insulin infusion as per standard diabetic ketoacidosis protocol, followed by transition to subcutaneous basal —bolus insulin. WES identified a heterozygous missense variant in the KCNJ11 gene (NM_000525.4: c.602G>A; p. Arg201His), consistent with autosomal dominant developmental delay, epilepsy, and neonatal diabetes (DEND) syndrome (OMIM #618856). The evaluation of this variant and its implications for management are discussed subsequently.
DEFINITION OF BIOINFORMATICS
In the context of clinical genomics, bioinformatic tools refer to computational resources, software programs, and curated biological databases that enable the analysis, annotation, prioritization, and interpretation of genetic variants generated through high-throughput sequencing technologies such as NGS. These tools provide information on gene function, population variation, disease associations, and predicted pathogenicity, thereby assisting clinicians in the identification of genetic variants for diagnosis, prognosis, and personalized therapeutic strategies.
BIOINFORMATICS ANALYSIS OF THE NGS DATA
NGS workflows across targeted panels, WES, and WGS share core steps—DNA extraction, library preparation, sequencing, and analysis—but differ mainly in enrichment strategies and genomic coverage, balancing depth and coverage for specific clinical needs. Library preparation involves fragmentation, adapter ligation, and (where applicable) target enrichment, followed by high-throughput sequencing on platforms such as Illumina, generating data whose interpretation is central to downstream bioinformatics analysis.[4]
From a clinical perspective, appropriate sequencing depth and coverage are critical to ensure reliable variant detection. WES typically requires a mean depth of ~100×, with ≥95 —98% of target regions covered at ≥20×, to compensate for capture-related variability. In contrast, WGS provides more uniform genomic coverage and can achieve reliable variant detection at a lower mean depth of ~30×, with ≥95% of the genome covered at ≥10×.
The raw data generated from sequencing must then undergo structured bioinformatics processing to enable accurate variant identification and clinical interpretation. NGS bioinformatics involves three key stages:[5]
Primary: conversion of raw sequencing signals into FASTA with quality (FASTQ) reads with quality control
Secondary: alignment to the reference genome and variant calling to generate variant call format (VCF) files
Tertiary: variant annotation, filtering, and prioritization.
While the initial steps ensure data accuracy and variant detection, the most clinically meaningful stage is tertiary analysis, which focuses on variant annotation, filtering, and prioritization—integrating genomic data with phenotype and inheritance patterns to identify clinically relevant variants.
This review will primarily concentrate on the clinical variant prioritization and interpretation. Readers can refer to Flow Chart 1 and Table 1 for an overview of bioinformatics tools used across primary, secondary, and tertiary NGS data analysis.

| Phase | Step | Description | Common tools/Software | Input | Output | Technical notes |
|---|---|---|---|---|---|---|
| Pre- analytical | Sample collection and DNA extraction | Isolation of high-quality DNA | Qubit, NanoDrop | Biological sample | Purified DNA | DNA quality critical (A260/280~1.8) |
| Library preparation | Fragmentation, adapter ligation | Illumina kits, Ion kits | DNA | Sequencing library | WES includes a capture step | |
| Target enrichment (WES only) | Capture of exonic regions | Agilent SureSelect, IDT xGen | Library | Enriched library | Capture bias possible | |
| Primary analysis | Base calling | Converts signals to bases | Illumina DRAGEN, bcl2fastq | Raw signals | Sequence reads | Platform-dependent |
| Demultiplexing | Separates pooled samples | bcl2fastq, DRAGEN | Raw reads | Sample FASTQ | Barcode accuracy important | |
| Quality scoring | Assigns Phred scores | DRAGEN, Ion torrent suite (platform software) | Reads | FASTQ | Q30 >85% desirable | |
| Secondary analysis | Quality control | Assesses read quality | FastQC, MultiQC | FASTQ | QC reports | Detects contamination |
| Trimming/filtering | Removes adapters, low-quality bases | Cutadapt, Trimmomatic | FASTQ | Clean FASTQ | Improves alignment | |
| Alignment | Maps reads to a reference genome | BWA-MEM, Bowtie2 | FASTQ | SAM/BAM | Use GRCh38 | |
| Sorting & indexing | Organizes BAM file | SAMtools | BAM | Sorted BAM | Required for downstream analysis | |
| Duplicate marking | Removes PCR duplicates | Picard MarkDuplicates | BAM | Dedup BAM | Reduces bias | |
| BQSR | Corrects base quality errors | GATK BaseRecalibrator | BAM | Recalibrated BAM | Improves variant accuracy | |
| Variant calling | Detects SNPs/Indels | GATK HaplotypeCaller, DeepVariant | BAM | VCF/gVCF | Core diagnostic step | |
| Joint genotyping | Multi-sample variant calling | GATK GenotypeGVCFs | gVCF | Cohort VCF | Improves sensitivity | |
| Variant filtering | Removes false positives | GATK VariantFiltration | VCF | Filtered VCF | Based on QC metrics | |
| CNV detection | Identifies copy number changes | CNVkit, ExomeDepth (WES), CNVnator (WGS) | BAM | CNV calls | Better in WGS | |
| SV detection | Detects structural variants | Manta, LUMPY | BAM | SV calls | Strong in WGS | |
| Tertiary analysis | Variant annotation | Adds biological context | ANNOVAR, VEP | VCF | Annotated VCF | Includes gene, frequency |
| Variant prioritization | Filters clinically relevant variants | Exomiser, PhenIX | Annotated VCF | Candidate variants | Uses phenotype (HPO) | |
| Clinical classification | Assigns pathogenicity | ACMG guidelines | Variants | Classified variants | Standardized interpretation | |
| Clinical reporting | Generates diagnostic report | Custom pipelines, LIMS | Final variants | Clinical report | CAP/CLIA compliance |
BAM: Binary alignment map, BQSR: Base quality score recalibration, CAP/CLIA: College of American Pathologists/Clinical Laboratory Improvement Amendments, CNV: Copy number variants, DNA: Deoxyribonucleic acid, FASTQ: FASTA with Quality, HPO: Human Phenotype Ontology, Indels: Insertion and deletions, PCR: Polymerase chain reaction, QC: Quality control, Indels: Insertion and deletions, SNP: Single nucleotide polymorphism, SV: Structural variants, SAM: sequence alignment map, VCF: variant call format, WES: Whole exome sequencing, WGS: Whole genome sequencing
Bioinformatic tools used in clinical variant prioritization and interpretation can be broadly categorized into several groups based on their primary function. Numerous tools, databases, and genome browsers are available within each of these categories; however, this review will focus on selected resources that are freely accessible, widely used, and most relevant for clinical practice. The broad categories include:
Population variant frequency databases
Diseases and gene disease association databases
Computational prediction tools
Variant interpretation and classification platforms
Phenotype-driven and integrative tools.
Together, these tools help clinicians move from raw sequencing data to clinically meaningful interpretation by integrating genomic findings with phenotypic and functional evidence. A summary of all these resources and tools is provided in Table 2.
| Databases | Data type | Features | Limitations | URL |
|---|---|---|---|---|
| 1. Population databases | ||||
| Genome Aggregation Database (gnomAD v4.1.0) | Population allele frequency (WES/WGS) | -Large-scale population allele frequency resource - Broad ancestral diversity -Gene-level constraint metrics (quantifies a gene’s intolerance to variant) -High-quality variant annotation and filtering - Includes nuclear, mitochondrial, and structural variants |
- Uneven population representation (underrepresentation of Asian and African populations) | https://gnomad.broadinstitute.org/ |
| Genome India project | Population allele frequency (WGS) |
- National initiative (~10,000 genomes) - Captures ethnic, linguistic, and tribal diversity - Creation of Indian reference genome dataset |
- Dataset still evolving; not fully publicly accessible - Limited immediate clinical applicability - Insufficient detailed genotype phenotype correlation |
https://genomeindia.in/ |
| IndiGenomes | Population allele frequency (WGS) |
-India-specific genomic database: 1000 genomes (CSIR initiative) - Represents diverse Indian populations -Identifies population-specific variants - Improves variant interpretation in Indian patients |
- Limited sample size compared to global databases - Incomplete representation of all Indian subpopulations |
https://clingen.igib.res.in/indigen/ |
| 2. Disease and gene disease association databases | ||||
| Online Mendelian Inheritance in Man (OMIM) | Gene—disease association | -Detailed gene—phenotype relationships and inheritance pattern - Structured summaries of each disorder and gene function |
- Selective curation; not a comprehensive variant repository | https://www.omim.org/ |
| ClinVar | Variant—disease association | - Aggregates global submissions - Provides clinical significance (ACMG/AMP classification) with review status (star system) - Freely accessible and regularly updated |
- Conflicting interpretations between submitters - Not all variants curated/validated |
https://www.ncbi.nlm.nih.gov/clinvar/ |
| Human Gene Mutation Database (HGMD) |
Published disease-causing variants | - Comprehensive collection of published gene mutations -Focus on disease-causing variants - Includes literature references for each variant - Widely used in clinical and research settings - Available in both a publicly free version and subscription-based professional version |
- Free version may include limited or outdated evidence - Not strictly ACMG/AMP classified |
https://www.hgmd.cf.ac.uk/ |
| Leiden Open Variation Database (LOVD) | Gene-specific variant database | -Locus-specific, expert-curated databases -Detailed gene-wise variant information - Includes phenotype and clinical data (when available) -Open-access |
- Not available for all disease causing genes - Data quality varies by curator |
https://www.lovd.nl/ |
| DECIPHER | Genotype —phenotype data | - Contains anonymized phenotype-linked genotype data - Includes wide range of variants including germline and mosaic variants |
- Incomplete clinical information in some entries | https://deciphergenomics.org/ |
| EndoGene | Variant—endocrine disease association | - Focused on endocrine genetics - Both deep (panel) and broader (WES) variant coverage - Helps identify recurrent or conserved pathogenic variants - Variant prioritization in suspected endocrine disorders |
- Predominantly Russian cohort - Limited generalizability; risk of population-specific bias |
https://zenodo.org/records/14554719 |
| 3. Computational prediction tools | ||||
| REVEL (Rare Exome Variant Ensemble Learner) |
Missense variants | -Ensemble-based predictor (combining multiple tools) - Higher accuracy than individual predictors |
- Limited to missense variants | https://sites.google.com/site/revelgenomics/ |
| BayesDel | Coding & non-coding variants | -Deleteriousness meta-score using Bayesian framework | - Limited clinical validation | https://fenglab.chpc.utah.edu/BayesDel/BayesDel.html |
| AlphaMissense | Missense variants | - Deep learning model (based on AlphaFold concepts) - Predicts pathogenicity of missense variants genome-wide |
- Limited clinical validation | https://alphamissense.hegelab.org/ |
| CADD (Combined annotation dependent depletion) |
SNVs, Indels | - Integrates different features into one deleteriousness score | - Limited specificity for variant types | https://cadd.gs.washington.edu/ |
| SpliceAI | Splice-site variants | -Deep learning-based splicing prediction - Detects cryptic splice sites and exon skipping - High sensitivity for intronic variants |
- May overpredict deep intronic effects | https://spliceailookup.broadinstitute.org/ |
| 4. Phenotype driven and integrative tools | ||||
| 4a) Phenotype driven diagnostic tools | ||||
| Face2Gene | Phenotype (image based analysis) |
- AI-based dysmorphology analysis (DeepGestalt) - Suggests possible genetic syndromes -Supports phenotype-driven diagnosis - Available in free and subscription-based versions |
- Population bias (less accurate in underrepresented groups) - Dependent on image quality and phenotyping |
https://www.face2gene.com/ |
| Phenomizer | Phenotype-driven diagnosis | - Matches HPO terms to possible diseases - Ranks differential diagnoses - Useful in clinical evaluation |
- Limited accessibility | https://compbio.charite.de/phenomizer/ |
| Genetic Disease Diagnosis Platform (GDDP) | AI-based phenotype- genotype integration | - Integrates clinical and genomic data - AI-driven prioritization - Supports rare disease diagnosis |
- Limited validation | https://gddp.research.cchmc.org/ |
| 4b) Variant interpretation platform | ||||
| Franklin* | Variant annotation and interpretation | - Automated ACMG/AMP classification - Integrates multiple databases (ClinVar, gnomAD, etc.) -User-friendly clinical interface |
- Limited access | https://franklin.genoox.com/ |
| VarSome | Variant annotation and interpretation | - Aggregates multiple databases/prediction tools - ACMG/AMP classification with evidence breakdown |
- Conflicting source data - Limited access |
https://varsome.com/ |
| InterVar | ACMG-based variant classification | -Semi-automated ACMG/AMP criteria assignment | - Limited database integration - Requires manual expert review |
http://wintervar.wglab.org/ |
| 4c) Genotype—phenotype integrative prioritization tools | ||||
| Exomiser | Genotype- phenotype integration | - Combines HPO terms with variant data - Prioritizes candidate variants/genes -Uses cross-species phenotype data - Useful in rare disease diagnosis |
- Dependent on phenotype quality - May miss novel gene associations |
https://www.sanger.ac.uk/tool/exomiser/ |
POPULATION FREQUENCY DATABASES
Population frequency databases are among the most important resources used in the initial filtering strategy during genetic variant assessment. Most population databases derive their data from large cohorts of apparently healthy or unaffected individuals. These databases contain information on the frequency of genetic variants across different populations and help determine whether a variant is rare or common. This distinction is crucial because rare variants are more likely to be associated with Mendelian disorders, whereas common variants present in the general population are usually benign. Thus, these resources help clinicians distinguish potentially disease-causing variants from the millions of largely benign variants present in every human genome. However, it is important to exercise caution when using population frequency data. Although pathogenic variants associated with Mendelian disorders are typically rare, most rare variants are not pathogenic. Therefore, rarity alone is insufficient to establish pathogenicity and must be interpreted alongside other lines of evidence.
The most frequently used databases are discussed below.
Genome aggregation database
It is currently one of the most frequently accessed global population reference datasets for variant interpretation (gnomAD, https://gnomad.broadinstitute.org).[6] In addition to single-nucleotide variants (SNVs) and small insertions or deletions, gnomAD also provides data on structural variants and mitochondrial DNA variants. In this database, variants are often considered rare if their minor allele frequency (MAF) is <1% (0.01), whereas common variants are typically defined as those with a MAF >5%. A limitation of the database is the relatively low representation of certain populations, including South Asians (approximately 3%), which may lead to underestimation of true allele frequencies. As a result, variants that are relatively common in these populations may appear artificially rare, potentially complicating accurate clinical interpretation.
Indian population databases
To improve the accuracy of variant interpretation in individuals of Indian ancestry, population-specific genetic databases have been developed to capture the unique and diverse genetic architecture of Indian populations. These are described below.
Genome India project
This is a large national initiative supported by the Department of Biotechnology, Government of India. The project generated a comprehensive catalog of genetic variations from more than 10,000 individuals representing the diverse ethnic and geographic groups across India. Although the dataset is currently available primarily as controlled-access research data and does not yet function as a fully open variant browser, it is expected to become one of the largest reference datasets for the Indian population. It may eventually serve a role similar to gnomAD for variant interpretation in Indian patients.[7]
IndiGenomes
This freely accessible database , (https://clingen.igib.res. in/indigen/) based on WGS of over 1000 individuals from diverse Indian populations, provides allele-frequency information that can assist in the interpretation of genetic variants in individuals of Indian ancestry.[8]
DISEASES AND GENE DISEASE ASSOCIATION DATABASES
Online Mendelian Inheritance in Man (OMIM)
OMIM (https://www.omim.org) is a frequently updated, comprehensive catalog of human genes and genetic disorders. Its primary focus is on the relationship between genetic variation and clinical phenotype. The database contains structured summaries describing the clinical features, inheritance patterns, and molecular basis of Mendelian disorders, along with selected disease-causing variants associated with specific genes. It often serves as the first point of reference for clinicians seeking a concise overview of gene-disease associations.
ClinVar
ClinVar is a freely accessible public database that aggregates information on human genetic variants and their relationship to disease. It is maintained by the National Centre for Biotechnology Information. ClinVar (www.ncbi.nlm.nih.gov/clinvar/) contains submissions from clinical testing laboratories, research groups, and expert panels worldwide, and currently includes data on millions of genetic variants. ClinVar primarily provides information on the clinical significance of genetic variants, with classifications such as pathogenic, likely pathogenic, variant of uncertain significance (VUS), likely benign, or benign. For clinicians, ClinVar is particularly useful in determining whether a detected variant has been previously reported and how it has been interpreted by other laboratories. This can provide valuable supporting evidence during the process of variant classification.[9] However, certain limitations should be considered when using ClinVar. The database relies on submitted interpretations and does not independently validate all entries, which can occasionally result in discrepancies or conflicting classifications for the same variant.
Leiden Open (source) Variation Database
This is a widely used platform for creating locus-specific gene databases (LOVD, https://www.lovd.nl/3.0/home) that compile variants identified in individual genes. This platform allows the collection, expert curation, and display of genomic variants along with associated clinical information, and supports both gene-centered and patient-centered views for data exploration.[10] In contrast to databases that list only variant summaries, LOVD frequently includes case-level clinical data, such as the patient or family in whom the variant was identified, associated phenotypic features, details of the genetic testing performed, and the variants detected. This combination of genotype and phenotype information can help clinicians assess the clinical relevance of a variant.
Human Gene Mutation Database (HGMD)
HGMD, (https://www.hgmd.cf.ac.uk) is one of the most comprehensive, expert-curated repositories of published germline variants associated with human inherited diseases. It compiles peer-reviewed evidence for over 570,000 reported mutations, supporting rapid and informed variant classification. HGMD is available in both a public (free) version and a professional subscription version; however, the public version is updated less frequently and lacks certain advanced features available in the full version.[11]
DECIPHER
This database (https://www.deciphergenomics.org) contains information on anonymized, phenotype-linked genomic data from patients with rare diseases. It integrates genomic variants with detailed clinical features encoded using Human Phenotype Ontology (HPO) terms, facilitating genotype — phenotype correlation and variant interpretation. The platform contains information on a wide range of variant types (sequence variants, short tandem repeats, copy-number variants, and large structural variants), including germline and mosaic variation in the nuclear and mitochondrial genomes.[12]
EndoGene
This publicly accessible database (https://zenodo.org/records/14554719) focuses on genetic variants associated with endocrine disorders. It comprises data from 5926 patients of Russian descent with over 450 endocrine and related conditions. It includes information on 2711 clinically relevant genetic variants identified through both targeted gene panels (covering 220 –382 genes) and WES.[13] Clinically, EndoGene is valuable for validating rare and recurrent pathogenic variants in conserved endocrine genes that may be shared across populations. However, caution is warranted when extrapolating these findings to other ethnic groups, as population-specific allele frequencies and genetic backgrounds may increase the risk of misclassification or false-positive interpretations.
In addition, literature databases such as PubMed (https://pubmed.ncbi.nlm.nih.gov) provide critical supporting evidence by enabling the identification of previously reported variants and associated phenotypes in research studies.
COMPUTATIONAL PREDICTION TOOLS
In silico prediction tools are computational algorithms used to estimate the potential pathogenicity of genetic variants. Most prediction algorithms integrate multiple sources of information, including evolutionary conservation, protein structure, biophysical properties, and genomic context, to assess the likelihood that a variant may disrupt gene or protein function. The two main categories of such tools are those that predict whether a missense change damages the protein’s function or structure, and those that predict whether it affects splicing.
At present, the most widely used prediction tools are ensemble or meta-predictors, based on machine-learning approaches. These tools combine outputs from multiple individual prediction algorithms along with additional genomic features—such as conservation scores and allele frequency—to generate a single aggregate pathogenicity score. Examples of commonly used tools include REVEL, a meta-predictor frequently used to assess missense variants, which assigns higher scores (typically >0.5 or >0.75) to variants with a greater likelihood of pathogenicity. Other ensemble predictors include BayesDel. More recently, deep learning-based tools such as AlphaMissense have been developed to evaluate missense variants by integrating protein structural context with evolutionary information. For variants that may affect RNA splicing, specialized tools such as SpliceAI are widely used and are considered among the most effective methods for predicting splice-altering variants.[14] Combined annotation-dependent depletion (CADD) is another widely used in silico tool that combines multiple biological evidences to predict the deleteriousness of SNVs and small insertion/deletions (indels). Higher CADD scores (e.g., Phred-scaled score 20) indicate a higher probability that a variant is deleterious or pathogenic.[15]
Despite their sophistication, these tools should be interpreted with caution. In silico predictions provide supporting evidence but should not be used in isolation to determine variant pathogenicity or as stand-alone diagnostic criteria.
FUNCTIONAL EVIDENCE IN VARIANT INTERPRETATION
Functional evidence provides direct insight into the biological impact of genetic variants and represents a higher level of evidence compared to in silico predictions. Experimental studies—including enzyme activity assays, electrophysiological analyses, and cell-based functional assays—can demonstrate the effect of a variant on protein function. However, unlike population and disease databases, functional data are not consolidated within a single comprehensive resource. While selected databases, such as ClinVar, may include summaries of functional studies submitted by laboratories, and specialized resources such as MaveDB (https://www.mavedb.org) provide high-throughput functional data for specific genes, the availability of such data remains limited and variable across variants. Consequently, much of the functional evidence used in clinical interpretation is derived from published studies, making literature databases such as PubMed essential for identifying experimentally validated variant effects.
VARIANT INTERPRETATION AND CLASSIFICATION PLATFORMS
After gathering supporting evidence from population databases, disease databases, in silico prediction tools, and functional evidence, the next critical step is the systematic interpretation and classification of genetic variants. To ensure consistency and clinical reliability, standardized frameworks have been developed to guide this process. The most widely adopted guidelines are those proposed by the American College of Medical Genetics and Genomics (ACMG) and the Association for Molecular Pathology (AMP), which provide a structured approach for the interpretation of sequence variants.[16] These guidelines integrate multiple lines of evidence—including population frequency, computational predictions, functional studies, segregation data, and phenotype correlation—to classify variants into five categories: Pathogenic, likely pathogenic, VUS, likely benign, and benign. A similar standard has also been developed by ACMG and Clinical Genome Resource (ClinGen) for the interpretation and reporting of constitutional copy number variants.[17]
To further refine variant interpretation and strengthen gene-disease validity assessment, ClinGen was established as an international initiative to systematically curate clinically relevant genes and variants (https://www.clinicalgenome.org/start). ClinGen expert panels evaluate the strength of evidence supporting gene –disease relationships and provide gene and disease-specific recommendations for variant interpretation. These curated resources help standardize variant classification across laboratories, reduce interpretative discrepancies, and improve diagnostic accuracy.[18] As a result, the use of ClinGen-curated data contributes to higher rates of definitive molecular diagnoses and a reduction in the number of variants classified as VUS.
These guidelines provide a standardized framework for variant classification, and their effective application in clinical practice requires integration with a detailed patient phenotype.
The following section outlines bioinformatic tools that facilitate this genotype —phenotype correlation and support clinically meaningful interpretation.
PHENOTYPE-DRIVEN AND INTEGRATIVE TOOLS
A range of tools has been developed to standardize phenotypic descriptions, enable phenotype-driven diagnosis, and integrate clinical features with genomic data for effective variant prioritization. These approaches not only improve diagnostic yield but also enhance the clinical relevance of NGS findings.
Phenotype-driven diagnostic tools (clinical-first approach)
Phenotype-driven tools play a crucial role in the initial stages of genetic diagnosis, particularly in children with complex or non-specific clinical presentations. These tools rely on standardized phenotypic descriptors using (HPO, https://hpo.jax.org), which provides a structured vocabulary to convert clinical features (e.g., “short stature,” “micropenis,” and “hypoglycemia”) into computable terms. On entering HPO terms, tools such as Face2Gene (https://www.face2gene.com), Phenomiser (http://compbio.charite.de/phenomizer), and Genetic Disease Diagnosis Platforms (GDDP) (https://gddp.research.cchmc.org) analyze phenotypic similarity and generate ranked lists of possible diagnoses or candidate genes. This helps clinicians formulate a focused differential diagnosis and guides targeted genetic testing or variant evaluation. In addition, Face2Gene allows the incorporation of facial photographs of patients with dysmorphism, alongside clinical features. By combining image-based pattern recognition with HPO terms, it refines the differential diagnosis and improves diagnostic accuracy.
The use of these tools is particularly valuable when the phenotype is evolving, syndromic features are subtle, or clinical findings are overlapping. They enhance clinical reasoning, reduce diagnostic uncertainty, and facilitate more efficient and targeted analysis of genomic data. However, the accuracy of these tools depends on the quality and completeness of phenotypic input, and incorrect or incomplete HPO annotation may lead to suboptimal prioritization.
Open-access variant interpretation platforms
Open-access variant interpretation platforms simplify what is otherwise a time-consuming, stepwise process of manual curation. Instead of individually analyzing population frequencies, disease databases, literature evidence, and in silico prediction tools, platforms such as VarSome (https://varsome.com), InterVar (https://wintervar.wglab. org), and Franklin by Genoox (https://franklin.genoox. com/clinical-db/home) integrate these multiple data sources and automatically apply ACMG criteria to provide variant classification along with supporting evidence. This significantly reduces the manual workload and allows clinicians to rapidly interpret variants in a structured and standardized manner.
However, an important caveat is that these platforms should be used as decision-support tools rather than definitive arbiters. Automated classifications may vary between platforms, can be influenced by database limitations, and may not fully capture clinical context or emerging evidence. Therefore, final interpretation must always involve clinician review, correlation with phenotype, and, where required, expert judgment.
Genotype —phenotype integrated prioritization tools
For clinicians interested in analyzing sequencing data independently, genotype–phenotype integrated tools provide a powerful interface to combine patient-specific genomic data with clinical features. Platforms such as Exomiser (https://www.sanger.ac.uk/tool/exomiser/) and Franklin by Genoox allow users to upload variant files (typically VCF) along with HPO terms to perform comprehensive analysis. These tools integrate multiple layers of evidence—including variant frequency, predicted functional impact, gene –disease associations, inheritance models, and phenotype similarity— to rank candidate variants in a clinically meaningful manner.
These platforms are particularly useful in exome/genome datasets with large numbers of variants, enabling efficient narrowing down to a small set of clinically relevant candidates. By incorporating phenotype-driven filtering into genomic data, they significantly enhance diagnostic yield and reduce the burden of manual interpretation.
APPROACH TO THE CASE SCENARIO
The initial clinical and biochemical evaluation raised suspicion for a syndromic form of diabetes. To standardize phenotypic characterization and facilitate computational analysis, key clinical features were mapped to HPO terms, including hyperglycemia, diabetic ketoacidosis, infantile onset, developmental delay, hypotonia, and hypsarrhythmia. Phenotype-driven diagnostic tools (such as Phenomizer/Face2Gene/GDDP) prioritized differential diagnoses, with developmental DEND syndrome emerging as the most likely candidate, followed by transient neonatal diabetes mellitus and other syndromic insulin resistance disorders.
Subsequently, genetic testing using WES identified a heterozygous missense variant in the KCNJ11 gene (c.602G>A; p. Arg201His). This gene encodes the Kir6.2 subunit of the ATP-sensitive potassium (K+ATP) channel, which plays a critical role in pancreatic β-cell insulin secretion and neuronal excitability.
A structured variant interpretation was performed in accordance with ACMG —AMP guidelines. The variant was absent in population databases such as gnomAD and IndiGenomes (PM2), supporting rarity. It had been previously reported in disease databases, including ClinVar and LOVD (PP5), suggesting prior association with disease. Multiple in silico prediction tools (REVEL, CADD, and AlphaMissense) indicated a deleterious effect (PP3). In addition, the variant is located in a functionally critical domain of the protein (PM1), and other pathogenic missense variants affecting the same codon have been reported (PM5). The gene is known to have a missense mutation mechanism (PP2). Based on this cumulative evidence (PM1, PM2, PM5, PP2, PP3, and PP5), the variant was classified as likely pathogenic.
Correlation of the genetic findings with the clinical phenotype confirmed the diagnosis of DEND syndrome due to a KCNJ11 mutation. This diagnosis had immediate therapeutic implications. The patient was transitioned from insulin therapy to oral sulfonylurea (glibenclamide), which directly targets the K+ATP channel defect, resulting in improved glycemic control and allowing discontinuation of insulin.
Genetic counseling was provided to the family, explaining the autosomal dominant inheritance pattern, with most cases being de novo, and a recurrence risk of <1% due to the possibility of germline mosaicism [Flow Chart 2].

CONCLUSION
Bioinformatic tools have become integral to the interpretation of genomic data in pediatric endocrine disorders, supporting clinicians in interpreting complex sequencing data for clinical application. Effective use of bioinformatic resources—including databases, computational tools, and integrative platforms—requires a practical understanding of their application and limitations. Familiarity with these tools empowers clinicians to critically evaluate genetic reports, correlate findings with clinical phenotypes, and make informed decisions, ultimately improving diagnostic accuracy and patient outcomes.
GLOSSARY
Gene panels: A technique that involves analyzing multiple genes simultaneously to identify genetic variants associated with specific diseases or conditions
WES: A technique that targets and sequences only the protein-coding regions (1 —2% of the genome) of an individual’s DNA, known as the exome
WGS: A technique that determines the entire DNA sequence, including protein-coding regions and non-coding regions of the genome
Variant: Alteration in the DNA sequence compared to a reference, replacing the term “mutation”
FASTQ: Standard text-based format for storing raw NGS data, containing both nucleotide sequences and their corresponding base quality scores
VCF: Standard text file format to store information about DNA sequence variations found in a sample compared to a reference genome
Variant annotation: Process of attaching biological, functional, and clinical information to raw genetic variants identified in a VCF file
Sequencing depth: Number of times a particular nucleotide is read during the sequencing process
Sequencing coverage: Pertains to the proportion of the genome or targeted region that has been sequenced at least once
MAF: The frequency at which the second most common allele (the “minor” allele) occurs in a given population for a specific single-nucleotide polymorphism
Sequence variant: A permanent change in the nucleotide sequence in the DNA sequence compared to a reference sequence. These changes include single-nucleotide substitutions, small deletions, insertions, duplications, and splicing changes
Copy number variant: A structural genomic alteration where a segment of DNA, typically larger than 1000 base pairs, is present in a variable number of copies compared to a reference genome. These include large deletions or duplications.
Ethical approval:
Institutional Review Board approval is not required.
Declaration of patient consent:
Patient’s consent is not required as there are no patients in this study.
Conflicts of interest:
There are no conflicts of interest.
Use of artificial intelligence (AI)-assisted technology for manuscript preparation:
The authors confirm that there was no use of artificial intelligence (AI)-assisted technology for assisting in the writing or editing of the manuscript, and no images were manipulated using AI.
Financial support and sponsorship: Nil.
References
- Williams Textbook of Endocrinology. Acta Endocrinol (Buchar). 2019;15:416.
- [CrossRef] [Google Scholar]
- National Center for Advancing Translational Sciences. Endocrine Diseases. Genetic and Rare Diseases Information Center (GARD) - an NCATS Program. Available from: https://rarediseases.info.nih.gov/diseases/diseases-by-category/8/endocrine-diseases [Last accessed on 2026 Apr 23]
- [Google Scholar]
- Variant classification concordance using the ACMG-AMP variant interpretation guidelines across nine genomic implementation research studies. Am J Hum Genet. 2020;107:932-41.
- [CrossRef] [PubMed] [Google Scholar]
- Library construction for next-generation sequencing: Overviews and challenges. Biotechniques. 2014;56:66, 68
- [CrossRef] [PubMed] [Google Scholar]
- Bioinformatics and computational tools for next-generation sequencing analysis in clinical genetics. J Clin Med. 2020;9:132.
- [CrossRef] [PubMed] [Google Scholar]
- Variant interpretation using population databases: Lessons from gnomAD. Hum Mutat. 2022;43:1012-30.
- [CrossRef] [PubMed] [Google Scholar]
- Mapping genetic diversity with the GenomeIndia project. Nat Genet. 2025;57:767-73.
- [CrossRef] [Google Scholar]
- IndiGenomes: A comprehensive resource of genetic variants from over 1000 Indian genomes. Nucleic Acids Res. 2021;49:D1225-32.
- [CrossRef] [PubMed] [Google Scholar]
- ClinVar: Improving access to variant interpretations and supporting evidence. Nucleic Acids Res. 2018;46:D1062-7.
- [CrossRef] [PubMed] [Google Scholar]
- The LOVD3 platform: Efficient genome-wide sharing of genetic variants. Eur J Hum Genet. 2021;29:1796-803.
- [CrossRef] [PubMed] [Google Scholar]
- The human gene mutation database (HGMD®): Optimizing its use in a clinical diagnostic or research setting. Hum Genet. 2020;139:1197-207.
- [CrossRef] [PubMed] [Google Scholar]
- DECIPHER: Supporting the interpretation and sharing of rare disease phenotype-linked variant data to advance diagnosis and research. Hum Mutat. 2022;43:682-97.
- [CrossRef] [PubMed] [Google Scholar]
- EndoGene database: Reported genetic variants for 5,926 Russian patients diagnosed with endocrine disorders. Front Endocrinol (Lausanne). 2025;16:1472754.
- [CrossRef] [PubMed] [Google Scholar]
- Insights on variant analysis in silico tools for pathogenicity prediction. Front Genet. 2022;13:1010327.
- [CrossRef] [PubMed] [Google Scholar]
- CADD: Predicting the deleteriousness of variants throughout the human genome. Nucleic Acids Res. 2019;47:D886-94.
- [CrossRef] [PubMed] [Google Scholar]
- Standards and guidelines for the interpretation of sequence variants: A joint consensus recommendation of the American college of medical genetics and genomics and the association for molecular pathology. Genet Med. 2015;17:405-24.
- [CrossRef] [PubMed] [Google Scholar]
- Technical standards for the interpretation and reporting of constitutional copy-number variants: A joint consensus recommendation of the American College of Medical Genetics and Genomics (ACMG) and the Clinical Genome Resource (ClinGen) Genet Med. 2020;22:245-57.
- [CrossRef] [PubMed] [Google Scholar]
- Electronic address: Splon@bcm.edu; ClinGen Consortium. The clinical genome resource (ClinGen): Advancing genomic knowledge through global curation. Genet Med. 2025;27:101228.
- [Google Scholar]


