SPLISOFORMS Documentation
The SPLISOFORMS splicing proteomics knowledgebase provides a unified view of how alternative splicing events reshape the protein landscape. Use these guides to understand the biological logic and technical algorithms powering our analytical discovery engine.
Datasets
The foundation of SPLISOFORMS is the GRCh38 reference proteome, supplemented by additional long-read transcriptomes from multiple experimental contexts. All models are processed through the same structural prediction and analysis pipeline.
Reference Proteome
GRCh38 · GENCODE v47
Canonical isoforms anchored to the GENCODE v47 / Ensembl 113 reference on GRCh38, with MANE Select transcripts flagged as the per-gene ground truth. This serves as the primary structural baseline and is used for all comparative structural impact calculations and canonical domain annotation.
Long-Read Datasets
Experimental datasets containing both annotated transcripts and novel isoforms discovered through long-read sequencing.
ccRCC Cohort
Primary Dataset
PacBio MAS-seq long-read sequencing from ten clear cell renal cell carcinoma patients. Covers primary tumour, matched metastasis, and adjacent normal kidney tissue. Isoform models derived from SQANTI3-classified assemblies. View Publication
Breast Cancer
Head et al. 2026
Nanopore long-read ESPRESSO assemblies from healthy breast tissue, primary breast tumour, and cultured fibroblasts. Includes both annotated GENCODE transcripts and novel isoforms not present in reference databases. View Publication
System Overview
SPLISOFORMS is a splicing proteomics knowledgebase that transforms genomic transcriptomic data into three-dimensional conformational, immunological, and regulatory insights. By integrating structural predictions from custom AlphaFold 3 models across multiple tissue contexts, the platform identifies high-impact perturbations that affect protein stability, functional domains, neoantigen potential, and post-translational regulation.
Every protein-coding transcript and novel isoform is modelled with the same AlphaFold 3 workflow for cross-comparable structures. Isoforms are reconciled to canonical Ensembl identifiers (with MANE Select prioritisation), ensuring a unified GENCODE v47 (Ensembl 113)reference architecture.