Neoantigens
Identifying novel immunological targets created by tumor-specific splicing events.
Beyond sequence identification, SPLISOFORMS adds a unique layer of Structural PTM-Sensitivity Analysis to neoantigen discovery.
The Discovery Engine
Junction Extraction
Overlapping 8–11 mer peptides are extracted from every novel splice junction (novel exon, alternative splice site, intron retention, frameshift). A sliding window ensures all possible proteasomal cleavage frames spanning the junction are covered.
MHCflurry 2.2 Presentation Model
Binding affinities (IC50, nM) and two complementary scores are predicted via MHCflurry 2.2:
- Presentation score (0–1) — joint probability that a peptide is processed and displayed on MHC-I. Primary metric for ranking neoantigen candidates; ≥ 0.5 is considered likely presented.
- Processing score (0–1) — proteasomal cleavage probability scored with 10-aa flanking context on each side of the junction. High values at a splice junction suggest the novel sequence actively enhances cleavage.
HLA-C alleles fall back to affinity-only mode due to sparser MHCflurry training data. All alleles are passed as a single-genotype call so the strongest cross-allele prediction is returned per peptide.
Binding Categories
Peptides are classified as strong binders (<50 nM IC50), weak binders (50–500 nM), or non-binders (>500 nM). Only strong and weak binders, plus any peptide with a presentation score ≥ 0.5, are stored and displayed.
Flanking Sequences
Ten amino acids of N-terminal and C-terminal flanking context are supplied to the MHCflurry processing model, allowing it to score proteasomal cleavage efficiency at the splice junction with full sequence context.
HLA Allele Panel
Predictions are run across a curated panel of 35 HLA class-I alleles selected to maximise global population coverage across HLA-A, HLA-B, and HLA-C loci. The panel follows the IEDB reference set recommendations and adds alleles with high frequency in African and Asian populations to reduce Eurocentric bias.
| Allele | Approx. population frequency |
|---|---|
| HLA-A*01:01 | ~16% European |
| HLA-A*02:01 | ~28% European, ~10% Asian |
| HLA-A*03:01 | ~14% European |
| HLA-A*11:01 | ~22% Asian, ~5% European |
| HLA-A*23:01 | ~9% African |
| HLA-A*24:02 | ~18% Asian, ~5% European |
| HLA-A*26:01 | ~5% European |
| HLA-A*29:02 | ~6% African, ~4% European |
| HLA-A*30:01 | ~8% African |
| HLA-A*31:01 | ~7% Asian, ~5% European |
| HLA-A*32:01 | ~6% European |
| HLA-A*33:01 | ~8% Asian |
| HLA-A*68:01 | ~7% African, ~3% European |
| Allele | Approx. population frequency |
|---|---|
| HLA-B*07:02 | ~13% European |
| HLA-B*08:01 | ~10% European |
| HLA-B*13:02 | ~7% Asian |
| HLA-B*15:01 | ~9% European |
| HLA-B*15:02 | ~8% Asian (SJS-associated) |
| HLA-B*15:03 | ~10% sub-Saharan African |
| HLA-B*18:01 | ~8% European |
| HLA-B*27:05 | ~8% European (AS-associated) |
| HLA-B*35:01 | ~8% European / Latino |
| HLA-B*38:01 | ~4% European |
| HLA-B*40:01 | ~9% Asian |
| HLA-B*44:02 | ~10% European |
| HLA-B*44:03 | ~7% African |
| HLA-B*51:01 | ~7% Asian / Mediterranean |
| HLA-B*53:01 | ~10% sub-Saharan African |
| HLA-B*58:01 | ~8% Asian (allopurinol-HSR) |
| Allele | Approx. population frequency |
|---|---|
| HLA-C*03:04 | ~10% European / Asian |
| HLA-C*04:01 | ~12% African / European |
| HLA-C*06:02 | ~9% European |
| HLA-C*07:01 | ~18% European |
| HLA-C*07:02 | ~14% European |
| HLA-C*12:03 | ~5% European |
MHCflurry 2.0 HLA-C coverage is sparser than HLA-A/B. These alleles fall back to affinity-only mode when the presentation model lacks training data.
Priority Ranking (P1–P4)
Candidates are ranked on immunological evidence, not on splice/PTM origin. The headline priority tier is a transparent conjunction of two orthogonal axes: multi-tool MHC presentation consensus and BigMHC immunogenicity. Splice origin, PTM disruption, expression and fold quality are reported as decorating flags — so we can state, for example, "of P1 epitopes, X % are splice-enabled" rather than gating the tier on it.
Strong multi-tool presentation (≥ 2 of 3 tools call the peptide a strong binder) AND predicted immunogenic (BigMHC-IM ≥ 0.5). The candidates most likely to be presented and elicit a T-cell response.
Strong multi-tool presentation, but not predicted immunogenic. Robustly presented; immunogenicity support is absent.
Moderate presentation (a single strong call, or ≥ 2 weak calls) AND predicted immunogenic.
Presentation evidence from a single tool or without immunogenicity support. Retained as a candidate but the weakest immunological evidence.
Presentation consensus (the backbone)
Three predictors are run per peptide × allele pair — MHCflurry 2.x, NetMHCpan-4.2 (EL) and BigMHC EL — and vote at two stringencies (strong / weak) using field-standard eluted-ligand %rank cutoffs. The count of votes sets the presentation tier.
≥ 2 of 3 tools call the peptide a strong binder: MHCflurry %rank ≤ 0.5, NetMHCpan-4.2 EL %rank ≤ 0.5, or BigMHC EL ≥ 0.5.
Exactly one strong call, or ≥ 2 tools presenting at the weak stringency (%rank ≤ 2.0 / BigMHC EL ≥ 0.25).
A single tool presents the peptide at the weak stringency.
No tool presents the peptide — excluded from the candidate set.
Decorating flags
Orthogonal annotations layered on top of any tier. They describe why a candidate is interesting (splice origin, PTM disruption) or add supporting evidence (expression, fold quality) — but they never change the priority tier.
BigMHC-IM immunogenicity score ≥ 0.5 — predicted T-cell recognition. This is the immunogenicity axis of the P-tier and is also surfaced as a standalone flag.
The epitope exists only because of the splice event (a canonical inhibitory PTM site is lost at/near the junction). The resource's unique angle — reported, not used to rank.
The isoform is detected in the underlying expression data (≥ 1 sample with full-length support). Currently derived from the ccRCC long-read set.
The novel sequence ablates a canonical PTM site inside the epitope (ptm_canonical_residue_lost = TRUE).
The epitope window is confidently modelled (avg pLDDT > 70, min pLDDT > 50) — the structural context is trustworthy.
Interpretation. All tiers are predicted, not experimentally validated. Presentation (the P-tier backbone) is well benchmarked; immunogenicity relies on a single model (BigMHC-IM) and should be read as supporting evidence. Because neoepitopes are junction-derived they are novel relative to the canonical protein, but a peptide may still occur elsewhere in the normal proteome — a proteome-wide uniqueness filter is a planned addition.
Epitope Atlas
The Epitope Atlas (accessible via the Results page toggle) provides a proteome-wide view of all predicted neoantigens in a single searchable, filterable table — complementing the per-isoform neoantigen panel.
Deduplication
Each peptide × isoform pair is deduplicated using DISTINCT ON, keeping only the best-affinity allele hit. This prevents the same junction peptide from inflating counts across all 35 alleles.
Sortable Columns
Sort by priority tier, consensus %rank, immunogenicity, IC50, pLDDT, junction type, gene, or isoform. Default order is priority ascending (P1 first). Column headers show sort direction and carry inline tooltips explaining each metric.
Filters
Filter simultaneously by gene symbol, priority tier (P1–P4), decorating flag (immunogenic, splice-enabled, expressed, PTM-disrupting, high-fidelity fold), binding category, HLA allele (dynamically loaded), junction type, and minimum pLDDT.
Summary Statistics
Header chips show real-time counts per priority tier (P1–P3), immunogenic and expressed peptides, plus unique peptides and unique genes matching the active filter set.