Neoantigens

Identifying novel immunological targets created by tumor-specific splicing events.

Beyond sequence identification, SPLISOFORMS adds a unique layer of Structural PTM-Sensitivity Analysis to neoantigen discovery.

The Discovery Engine

Junction Extraction

Overlapping 8–11 mer peptides are extracted from every novel splice junction (novel exon, alternative splice site, intron retention, frameshift). A sliding window ensures all possible proteasomal cleavage frames spanning the junction are covered.

MHCflurry 2.2 Presentation Model

Binding affinities (IC50, nM) and two complementary scores are predicted via MHCflurry 2.2:

  • Presentation score (0–1) — joint probability that a peptide is processed and displayed on MHC-I. Primary metric for ranking neoantigen candidates; ≥ 0.5 is considered likely presented.
  • Processing score (0–1) — proteasomal cleavage probability scored with 10-aa flanking context on each side of the junction. High values at a splice junction suggest the novel sequence actively enhances cleavage.

HLA-C alleles fall back to affinity-only mode due to sparser MHCflurry training data. All alleles are passed as a single-genotype call so the strongest cross-allele prediction is returned per peptide.

Binding Categories

Peptides are classified as strong binders (<50 nM IC50), weak binders (50–500 nM), or non-binders (>500 nM). Only strong and weak binders, plus any peptide with a presentation score ≥ 0.5, are stored and displayed.

Flanking Sequences

Ten amino acids of N-terminal and C-terminal flanking context are supplied to the MHCflurry processing model, allowing it to score proteasomal cleavage efficiency at the splice junction with full sequence context.

HLA Allele Panel

Predictions are run across a curated panel of 35 HLA class-I alleles selected to maximise global population coverage across HLA-A, HLA-B, and HLA-C loci. The panel follows the IEDB reference set recommendations and adds alleles with high frequency in African and Asian populations to reduce Eurocentric bias.

HLA-A13 alleles
AlleleApprox. population frequency
HLA-A*01:01~16% European
HLA-A*02:01~28% European, ~10% Asian
HLA-A*03:01~14% European
HLA-A*11:01~22% Asian, ~5% European
HLA-A*23:01~9% African
HLA-A*24:02~18% Asian, ~5% European
HLA-A*26:01~5% European
HLA-A*29:02~6% African, ~4% European
HLA-A*30:01~8% African
HLA-A*31:01~7% Asian, ~5% European
HLA-A*32:01~6% European
HLA-A*33:01~8% Asian
HLA-A*68:01~7% African, ~3% European
HLA-B16 alleles
AlleleApprox. population frequency
HLA-B*07:02~13% European
HLA-B*08:01~10% European
HLA-B*13:02~7% Asian
HLA-B*15:01~9% European
HLA-B*15:02~8% Asian (SJS-associated)
HLA-B*15:03~10% sub-Saharan African
HLA-B*18:01~8% European
HLA-B*27:05~8% European (AS-associated)
HLA-B*35:01~8% European / Latino
HLA-B*38:01~4% European
HLA-B*40:01~9% Asian
HLA-B*44:02~10% European
HLA-B*44:03~7% African
HLA-B*51:01~7% Asian / Mediterranean
HLA-B*53:01~10% sub-Saharan African
HLA-B*58:01~8% Asian (allopurinol-HSR)
HLA-C6 alleles
AlleleApprox. population frequency
HLA-C*03:04~10% European / Asian
HLA-C*04:01~12% African / European
HLA-C*06:02~9% European
HLA-C*07:01~18% European
HLA-C*07:02~14% European
HLA-C*12:03~5% European

MHCflurry 2.0 HLA-C coverage is sparser than HLA-A/B. These alleles fall back to affinity-only mode when the presentation model lacks training data.

HLA-A — 13 allelesHLA-B — 16 allelesHLA-C — 6 allelesTotal — 35 alleles

Priority Ranking (P1–P4)

Candidates are ranked on immunological evidence, not on splice/PTM origin. The headline priority tier is a transparent conjunction of two orthogonal axes: multi-tool MHC presentation consensus and BigMHC immunogenicity. Splice origin, PTM disruption, expression and fold quality are reported as decorating flags — so we can state, for example, "of P1 epitopes, X % are splice-enabled" rather than gating the tier on it.

P1 · High-confidence

Strong multi-tool presentation (≥ 2 of 3 tools call the peptide a strong binder) AND predicted immunogenic (BigMHC-IM ≥ 0.5). The candidates most likely to be presented and elicit a T-cell response.

P2 · Strong presentation

Strong multi-tool presentation, but not predicted immunogenic. Robustly presented; immunogenicity support is absent.

P3 · Supported

Moderate presentation (a single strong call, or ≥ 2 weak calls) AND predicted immunogenic.

P4 · Candidate

Presentation evidence from a single tool or without immunogenicity support. Retained as a candidate but the weakest immunological evidence.

Presentation consensus (the backbone)

Three predictors are run per peptide × allele pair — MHCflurry 2.x, NetMHCpan-4.2 (EL) and BigMHC EL — and vote at two stringencies (strong / weak) using field-standard eluted-ligand %rank cutoffs. The count of votes sets the presentation tier.

Strong

≥ 2 of 3 tools call the peptide a strong binder: MHCflurry %rank ≤ 0.5, NetMHCpan-4.2 EL %rank ≤ 0.5, or BigMHC EL ≥ 0.5.

Moderate

Exactly one strong call, or ≥ 2 tools presenting at the weak stringency (%rank ≤ 2.0 / BigMHC EL ≥ 0.25).

Weak

A single tool presents the peptide at the weak stringency.

None

No tool presents the peptide — excluded from the candidate set.

Decorating flags

Orthogonal annotations layered on top of any tier. They describe why a candidate is interesting (splice origin, PTM disruption) or add supporting evidence (expression, fold quality) — but they never change the priority tier.

Immunogenic

BigMHC-IM immunogenicity score ≥ 0.5 — predicted T-cell recognition. This is the immunogenicity axis of the P-tier and is also surfaced as a standalone flag.

Splice-enabled

The epitope exists only because of the splice event (a canonical inhibitory PTM site is lost at/near the junction). The resource's unique angle — reported, not used to rank.

Expressed

The isoform is detected in the underlying expression data (≥ 1 sample with full-length support). Currently derived from the ccRCC long-read set.

PTM-disrupting

The novel sequence ablates a canonical PTM site inside the epitope (ptm_canonical_residue_lost = TRUE).

High-fidelity fold

The epitope window is confidently modelled (avg pLDDT > 70, min pLDDT > 50) — the structural context is trustworthy.

Interpretation. All tiers are predicted, not experimentally validated. Presentation (the P-tier backbone) is well benchmarked; immunogenicity relies on a single model (BigMHC-IM) and should be read as supporting evidence. Because neoepitopes are junction-derived they are novel relative to the canonical protein, but a peptide may still occur elsewhere in the normal proteome — a proteome-wide uniqueness filter is a planned addition.

Epitope Atlas

The Epitope Atlas (accessible via the Results page toggle) provides a proteome-wide view of all predicted neoantigens in a single searchable, filterable table — complementing the per-isoform neoantigen panel.

Deduplication

Each peptide × isoform pair is deduplicated using DISTINCT ON, keeping only the best-affinity allele hit. This prevents the same junction peptide from inflating counts across all 35 alleles.

Sortable Columns

Sort by priority tier, consensus %rank, immunogenicity, IC50, pLDDT, junction type, gene, or isoform. Default order is priority ascending (P1 first). Column headers show sort direction and carry inline tooltips explaining each metric.

Filters

Filter simultaneously by gene symbol, priority tier (P1–P4), decorating flag (immunogenic, splice-enabled, expressed, PTM-disrupting, high-fidelity fold), binding category, HLA allele (dynamically loaded), junction type, and minimum pLDDT.

Summary Statistics

Header chips show real-time counts per priority tier (P1–P3), immunogenic and expressed peptides, plus unique peptides and unique genes matching the active filter set.