CIBERER

HPO2TISSUE

Tissue prioritization for RNA-seq from HPO phenotypes · differential diagnosis · network analysis

HPO 2026-02-16GTEx v11Tabula Sapiens v2Orphanet 2025-12OMIM 2026-02-16GenCC 2026-06gnomAD v4.1.1STRING v12ReactomeMGIGLOWgenes v1MONDO 2025-12DGIdb v5DrugCentral 2023Open Targets 26.03PhenoDP 2025

status column reflects how each extracted ID landed in your live hp.obo: direct = predictable and still primary; alt_remapped = renamed since the model was trained, automatically substituted (an extra hpo_id_pt column then appears with the original ID PhenoTagger emitted, for audit); out_of_vocab = PhenoTagger emitted an ID absent from your ontology (rare; usually filtered).


PhenoTagger v1.2's BERT head was trained on HPO ~Jan 2024 and can only emit IDs from a fixed lable.vocab . Below is how that label space maps to your current hp.obo.


                  

IC (Information Content) = -log2(n_genes(h) / N_genes_total) — computed empirically from genes_to_phenotype.txt annotations. Very general HPOs like HP:0000118 (Phenotypic abnormality) have IC ≈ 0; very specific HPOs like HP:0001212 (Prominent fingertip pads, ~50 annotated genes) have IC ≈ 7. Range in this database: 0 (a term annotated to every gene) to ~12.3 (a single-gene term). The blue bar visualizes the value relative to the maximum IC observed in the database.


Unmapped HPOs:

                  

UpSet plot of the gene sets of each input HPO (requires ≥2 HPOs with mapped genes). Visual exploration of all 2^N − 1 overlap patterns; the consensus / pairwise sets used downstream are picked in the next tab.

Coverage histogram

Distribution of multi-phenotype support per gene — blue bars contribute to the current primary gene set.


Genes shared across input HPOs

Genes present in ≥ N of the input HPO gene sets, with the HPO terms sharing them. Defaults to all selected HPOs (full intersection); lower it to relax the overlap.

Shared genes — copyable list

Individual HPO gene sets

One row per input HPO term. Yellow-highlights the gene in every set that contains it; the 'Contains?' column flags hits at the row level.

Score = value of the scoring method picked above. Rows: gene sets (consensus by coverage + singletons + ALL ∩ + pairs). Tissues ordered by anatomical group (Blood/Fibroblasts/Lymphocytes on the left in red ), separated by dashed black lines.

Raw log2(TPM+1) of the primary gene set. The clustered heatmap groups genes by hierarchical clustering (profile similarity); the boxplot shows per-tissue distributions, one point per gene. Clinically accessible tissues highlighted in red .

Rank-based gene-set scoring (singscore) of the primary set against every tissue, with a permutation-calibrated p-value (parametric tail of the random gene-set null).

Singscore (rank-based) of each input HPO's genes separately. Useful when HPOs point to different systems.


Top 3 tissues per HPO

Top 3 tissues per HPO

Clinically accessible tissues (wet-lab biopsy options) for the primary gene set. Per gene: log2(TPM+1), GTEx rank percentile, tissue specificity, gene coverage and the contributing inputs. Always uses GTEx regardless of the active backend.





Each point = one gene. Diagonal line = identical expression. Genes well above the line = preferential in blood; well below = in fibroblasts.

Random Walk with Restart over 13 evidence networks (PPI, pathways, co-expression, function, phenotype, mouse models, regulation, complex, co-essentiality, drug, co-localization, genomic, co-citation), one per knowledge category. Per-gene scores are weighted by disease-aware recall × specificity (cross-validated on the seed set). Candidates outside the seed set with high score are plausible novel disease-gene predictions.
Method: de la Fuente et al., Int. J. Mol. Sci. 2023, 24:1661 [doi]GLOWgenes repo . Networks (CC BY-NC-SA): GLOWgenesNets v1 (figshare 21408393).


Networks are pre-loaded at app startup. Set CV = 0 to skip cross-validation (uniform weights, faster but less informative).


PanelApp panels most loaded with our top-N predictions

For each Genomics England PanelApp panel, the bar shows how many of our live top- Top-N candidates are also in that panel's top-N pre-computed by the GLOWgenes authors. Top-of-list panels suggest which clinical entities your phenotype is most aligned with — independent of which seeds you started from.


PanelApp panel — pre-computed ranking

Pick any panel above (or below) to inspect the published GLOWgenes ranking for that disease as-is, plus how those genes rank in our live computation.

Rank vs RWR score (coloured by PanelApp context)

X = rank (1 = best, on the left), Y = RWR score. Colour: black seed gene from your input; red in the top-N of the panel picked above; blue known to at least one panel; grey novel candidate. Hover for top supporting network.



Predicted novel candidates (live ranking)

Genes ranked by combined RWR score across the 13 evidence networks. Your seed genes are excluded from this list — every row is a network-proximal candidate that is not in your input set. panelapp_seed_count = number of PanelApp panels in which this gene is a canonical (rank-0) seed → high values mean a known pleiotropic disease gene; zero means a fully novel network-derived prediction.


Download full ranking

Method-level diagnostics, not clinical outputs. Per-network weights show which knowledge category drove the ranking (networks with recall = 0 contributed nothing). Top genes × network heatmap shows whether each top candidate is supported by many networks (robust) or just one (potentially spurious).

Per-network weights (disease-aware)

Top genes × network heatmap (raw RWR scores)

STRING physical PPI v12 (experimentally observed, no text-mining) on the primary set + top interactors. Nodes in blue (orange label) = seeds; grey = interactors. STRING score ∈ [0, 1000] = interaction confidence. Blue edges = seed-seed; grey = seed-neighbor; thickness ∝ score.


Type: seed-seed = edge between two input genes (internal to the module); seed-neighbor = edge between an input gene and an external interactor (potential novel candidate).

For each gene in the set, its top 30 partners by expression-profile correlation across the ~70 GTEx tissues. Column in_hpo_universe indicates whether the partner already has HPO annotation (those that do NOT are potential novel candidates).

Same logic as GTEx, but the correlation is computed across the ~74 Tabula Sapiens tissues. Different sample structure surfaces partners that GTEx may miss (retina, cornea, cochlea cell types, etc.).

Pathways enriched in the primary set (hypergeometric test vs Reactome human universe, BH FDR). Plot shows only pathways with p_hyper < 0.05; X axis = fold enrichment (matches/expected); Y axis = -log10(p_hyper); color = p_adj_BH; size = number of matches in the set.


For each gene in the set, the pathways it participates in.

Drug repurposing layer over the primary gene set. Drug-target evidence integrates DGIdb v5, DrugCentral and Open Targets (rare-disease subset). Drug IDs canonicalised to ChEMBL when possible (InChIKey → name fallback).

Per-gene list of known drug interactions across the three sources.

Top drugs by # genes hit

Hypergeometric test: drugs hitting >= 2 of the primary-set genes more than expected by chance, BH-adjusted across all candidate drugs.


Personalised PageRank on the bipartite drug-target graph (seeds = primary set). Top drugs by network proximity to the phenotype, regardless of direct-hit count. Bar colour: approved / investigational / unknown .


MP terms from the mouse KO of each gene. Mouse→human mapping by uppercasing the symbol (direct orthologs).

Two-stage ranking. Stage 1 (all diseases, ~1 s): IC-Jaccard + hypergeometric p-value. Stage 2 (top-N from Stage 1, ~30 s): Phenomizer-style Resnik MICA semantic similarity with permutation p-value (requires hp.obo ). Use p_adj_BH_phen as primary FDRp_adj_BH is just the Stage-1 screen.

Higher = finer p-value resolution. 0 = similarity only.

Significant candidates plot

Each point is a candidate disease. X axis = phenotypic overlap with the input HPOs (IC-weighted Jaccard; higher = better match of the symptom set). When Phenomizer is on, Y axis = phenotypic similarity score (Phenomizer; higher = closer match in the HPO ontology) and color encodes the empirical p-value (darker = more confident). Without Phenomizer, the Y axis is the hypergeometric significance directly. Size = fraction of the input phenotype information covered by the disease. Hover over a point to see disease ID, name and metrics.


Full table

Phenomizer only refines the Top-N candidates set above. Non-refined rows have NA in p_phen / p_adj_BH_phen — toggle the checkbox to include them in the table.

Columns: disease_id | disease_name | inheritance (MoI: AD/AR/X-linked, extracted from phenotype.hpoa) | n_match (input HPOs found in the disease) | n_disease (total HPOs annotated to the disease) | ic_jaccard = ΣIC(matches) / ΣIC(union) | recall_ic = ΣIC(matches) / ΣIC(input) | p_hyper (stage 1, hypergeometric) | p_adj_BH (FDR on p_hyper, stage 1) | sim_qd/sim_dq/sim_sym (Phenomizer Resnik asymmetric/symmetric) | p_phen (empirical p-value by random query permutation) | p_adj_BH_phen (FDR on p_phen, stage 2 — primary ) | match_hpos (list of matched HPO IDs).


Cross-references (MONDO)

Click a row in the table above to see the MONDO unified identifier and the equivalent IDs in OMIM, Orphanet, MeSH, DOID, NCIT, MedGen, UMLS. Useful for cross-resource navigation when sharing a candidate diagnosis with a collaborator who uses a different ontology.

Combined phenotype-only gene prioritiser. Ranks candidate genes by blending the PhenoDP disease ranking (IC + Phi, collapsed to the best disease per gene) with an empirical patient channel — how frequently the query phenotypes occur in each gene's real phenopacket-store patients (IC-weighted). The 0.9 / 0.1 blend was tuned by 5-fold cross-validation on phenopacket-store (held out by publication), lifting shortlist recall without diluting rank-1.


Evidence scatter

Each point is a gene; colour = blended ranking score. Genes high on both axes are supported by curated disease annotation and by real patients.


Full ranking

Score = 0.9 · PhenoDP + 0.1 · patient (per-query normalised, 0–1). patient = IC-weighted empirical frequency of the query phenotypes in the gene's phenopacket-store patients (0 when the gene has no such data).

PhenoDP-style disease ranker. Ranks diseases from the input HPO terms by Combined = IC-based + Phi-based similarity (Jiang-Conrath IC-weighted best-match + phi coefficient over ancestor-expanded term sets), after a fast IC-Jaccard prefilter. Native R reimplementation of PhenoDP's analytical components (Yang et al., Genome Medicine 2025); the GCN/semantic component is omitted as it needs the upstream pretrained embeddings.


Ranked diseases

Each point is a disease; colour = combined score. Diseases high on both axes are supported by IC and Phi evidence alike.


Wide table of every gene in the primary set. Coverage, HPO specificity (IDF), top tissues, blood/fibroblast log2(TPM+1), gnomAD constraint, and the full per-tissue expression vector.

One row per gene: diseases in which it appears + association_type + presence in GTEx.

Phenotypes annotated to the selected disease(s) in phenotype.hpoa (HPO consortium). Reverse lookup: what HPO terms each disease presents with. ic is the HPO term's information content (rarer = more diagnostic); n_genes is how many genes are annotated to that HPO across the whole corpus (lower = more specific phenotype). source tags how the HPO was linked: hpoa = direct from phenotype.hpoa (most authoritative), g2p = genes_to_phenotype tagged with this exact disease, gene = HPOs of the disease's genes (weaker, but catches diseases not curated in hpoa — the via_gene column shows which gene linked the HPO).

UpSet plot of the gene sets of each selected disease (requires ≥2 diseases with mapped genes). Visual exploration of all 2^N − 1 overlap patterns; the consensus / pairwise sets used downstream are picked in the next tab.

Genes shared across selected diseases

Genes present in ≥ N of the selected-disease gene sets, with the diseases sharing them. Defaults to all selected diseases (full intersection); lower it to relax the overlap.

Shared genes — copyable list

Individual disease gene sets

One row per selected disease. Yellow-highlights the gene in every set that contains it; the 'Contains?' column flags hits at the row level.

Gene set × tissue heatmap. Tissues by anatomical group; clinically accessible on the left in red .

Raw log2(TPM+1) of the primary gene set. The clustered heatmap groups genes by hierarchical clustering (profile similarity); the boxplot shows per-tissue distributions, one point per gene. Clinically accessible tissues highlighted in red .

Rank-based gene-set scoring (singscore) of the primary set against every tissue, with a permutation-calibrated p-value (parametric tail of the random gene-set null).

Clinically accessible tissues (wet-lab biopsy options) for the primary gene set. Per gene: log2(TPM+1), GTEx rank percentile, tissue specificity, gene coverage and the contributing inputs. Always uses GTEx regardless of the active backend.

Random Walk with Restart over 13 evidence networks (PPI, pathways, co-expression, function, phenotype, mouse models, regulation, complex, co-essentiality, drug, co-localization, genomic, co-citation), one per knowledge category. Per-gene scores are weighted by disease-aware recall × specificity (cross-validated on the seed set). Candidates outside the seed set with high score are plausible novel disease-gene predictions.
Method: de la Fuente et al., Int. J. Mol. Sci. 2023, 24:1661 [doi]GLOWgenes repo . Networks (CC BY-NC-SA): GLOWgenesNets v1 (figshare 21408393).



PanelApp panels most loaded with our top-N predictions

How many of our live top- Top-N candidates appear in each Genomics England PanelApp panel (sorted descending). Top panels indicate which clinical entities your disease set is most aligned with.


PanelApp panel — pre-computed ranking
Rank vs RWR score (coloured by PanelApp context)

X = rank (1 = best, on the left), Y = RWR score. Black = seed; red = in panel above; blue = known to any panel; grey = novel.



Predicted novel candidates (live ranking)

Genes ranked by combined RWR score across the 13 evidence networks. Seed genes are excluded — every row is a network-proximal candidate that is not in your input set. panelapp_seed_count = number of PanelApp panels in which this gene is a canonical (rank-0) seed → high = known pleiotropic gene; zero = fully novel prediction.


Download full ranking
Per-network weights

Top genes × network heatmap (raw RWR scores)

STRING physical PPI v12 (experimentally observed, no text-mining) on the primary set + top interactors. Nodes in blue (orange label) = seeds; grey = interactors. STRING score ∈ [0, 1000] = interaction confidence. Blue edges = seed-seed; grey = seed-neighbor; thickness ∝ score.


For each gene in the set, its top 30 partners by expression-profile correlation across the ~70 GTEx tissues. Column in_hpo_universe indicates whether the partner already has HPO annotation (those that do NOT are potential novel candidates).

Same logic as GTEx, but the correlation is computed across the ~74 Tabula Sapiens tissues. Different sample structure surfaces partners that GTEx may miss (retina, cornea, cochlea cell types, etc.).

Pathways enriched in the primary set (hypergeometric test vs Reactome human universe, BH FDR). Plot shows only pathways with p_hyper < 0.05; X axis = fold enrichment (matches/expected); Y axis = -log10(p_hyper); color = p_adj_BH; size = number of matches in the set.


For each gene in the set, the pathways it participates in.

Drug repurposing layer over the primary gene set. Drug-target evidence integrates DGIdb v5, DrugCentral and Open Targets (rare-disease subset). The 'From Open Targets' sub-tab is anchored on the selected disease ID (resolved via ORPHA/OMIM/MONDO ↔ EFO crosswalk).

Per-gene list of known drug interactions across the three sources.

Top drugs by # genes hit

Hypergeometric test: drugs hitting >= 2 of the primary-set genes more than expected by chance, BH-adjusted across all candidate drugs.


Personalised PageRank on the bipartite drug-target graph (seeds = primary set). Top drugs by network proximity to the phenotype, regardless of direct-hit count. Bar colour: approved / investigational / unknown .


Open Targets clinical_indication evidence anchored on the selected disease ID. The barchart counts distinct drugs by maximum clinical stage reached.


MP terms from the mouse KO of each gene. Mouse→human mapping by uppercasing the symbol (direct orthologs).

Wide table of every gene in the primary set. Coverage, HPO specificity (IDF), top tissues, blood/fibroblast log2(TPM+1), gnomAD constraint, and the full per-tissue expression vector.

Input summary: which gene symbols matched the HUGO universe (HPO + disease + GTEx). n_hpos = number of HPO terms annotating this gene; n_diseases = number of OMIM/ORPHA/DECIPHER entries listing this gene; LOEUF / pLI from gnomAD constraint.


Unmapped symbols (typo or alias):

                  

Reverse lookup: which HPO terms are loaded with the input gene list. Higher counts at the top → those phenotypes are over-represented in your variant set.

Top HPO terms (sorted by # input genes annotated)

Reverse lookup: which OMIM/ORPHA/DECIPHER entities are loaded with the input gene list. Higher counts at the top → those disease entities are over-represented in your variant set.

Top diseases (HPO + Orphanet integrated)

Raw log2(TPM+1) of the primary gene set. The clustered heatmap groups genes by hierarchical clustering (profile similarity); the boxplot shows per-tissue distributions, one point per gene. Clinically accessible tissues highlighted in red .

Rank-based gene-set scoring (singscore) of the primary set against every tissue, with a permutation-calibrated p-value (parametric tail of the random gene-set null).

Clinically accessible tissues (wet-lab biopsy options) for the primary gene set. Per gene: log2(TPM+1), GTEx rank percentile, tissue specificity, gene coverage and the contributing inputs. Always uses GTEx regardless of the active backend.

Random Walk with Restart over 13 evidence networks (PPI, pathways, co-expression, function, phenotype, mouse models, regulation, complex, co-essentiality, drug, co-localization, genomic, co-citation), one per knowledge category. Per-gene scores are weighted by disease-aware recall × specificity (cross-validated on the seed set). Candidates outside the seed set with high score are plausible novel disease-gene predictions.
Method: de la Fuente et al., Int. J. Mol. Sci. 2023, 24:1661 [doi]GLOWgenes repo . Networks (CC BY-NC-SA): GLOWgenesNets v1 (figshare 21408393).



PanelApp panels most loaded with our top-N predictions

PanelApp panel — pre-computed ranking
Rank vs RWR score (coloured by PanelApp context)

X = rank (1 = best, on the left), Y = RWR score. Black = seed; red = in panel above; blue = known to any panel; grey = novel.



Predicted novel candidates (live ranking)

Genes ranked by combined RWR score across the 13 evidence networks. Your input list is excluded — every row is a network-proximal candidate.


Download full ranking
Per-network weights

Top genes × network heatmap (raw RWR scores)

STRING physical PPI v12 (experimentally observed, no text-mining) on the primary set + top interactors. Nodes in blue (orange label) = seeds; grey = interactors. STRING score ∈ [0, 1000] = interaction confidence. Blue edges = seed-seed; grey = seed-neighbor; thickness ∝ score.


For each gene in the set, its top 30 partners by expression-profile correlation across the ~70 GTEx tissues. Column in_hpo_universe indicates whether the partner already has HPO annotation (those that do NOT are potential novel candidates).

Same logic as GTEx, but the correlation is computed across the ~74 Tabula Sapiens tissues. Different sample structure surfaces partners that GTEx may miss (retina, cornea, cochlea cell types, etc.).

Pathways enriched in the primary set (hypergeometric test vs Reactome human universe, BH FDR). Plot shows only pathways with p_hyper < 0.05; X axis = fold enrichment (matches/expected); Y axis = -log10(p_hyper); color = p_adj_BH; size = number of matches in the set.


For each gene in the set, the pathways it participates in.

Drug repurposing layer over the primary gene set. Drug-target evidence integrates DGIdb v5, DrugCentral and Open Targets (rare-disease subset). Drug IDs canonicalised to ChEMBL when possible (InChIKey → name fallback).

Per-gene list of known drug interactions across the three sources.

Top drugs by # genes hit

Hypergeometric test: drugs hitting >= 2 of the primary-set genes more than expected by chance, BH-adjusted across all candidate drugs.


Personalised PageRank on the bipartite drug-target graph (seeds = primary set). Top drugs by network proximity to the phenotype, regardless of direct-hit count. Bar colour: approved / investigational / unknown .


MP terms from the mouse KO of each gene. Mouse→human mapping by uppercasing the symbol (direct orthologs).

Wide table of every gene in the primary set. Coverage, HPO specificity (IDF), top tissues, blood/fibroblast log2(TPM+1), gnomAD constraint, and the full per-tissue expression vector.