SPLISOFORMS Documentation

The SPLISOFORMS splicing proteomics knowledgebase provides a unified view of how alternative splicing events reshape the protein landscape. Use these guides to understand the biological logic and technical algorithms powering our analytical discovery engine.

Datasets

The foundation of SPLISOFORMS is the GRCh38 reference proteome, supplemented by additional long-read transcriptomes from multiple experimental contexts. All models are processed through the same structural prediction and analysis pipeline.

Reference Proteome

GRCh38 · GENCODE v47

Canonical isoforms anchored to the GENCODE v47 / Ensembl 113 reference on GRCh38, with MANE Select transcripts flagged as the per-gene ground truth. This serves as the primary structural baseline and is used for all comparative structural impact calculations and canonical domain annotation.

GENCODE v47Ensembl 113MANE SelectGRCh38

Long-Read Datasets

Experimental datasets containing both annotated transcripts and novel isoforms discovered through long-read sequencing.

ccRCC Cohort

Primary Dataset

PacBio MAS-seq long-read sequencing from ten clear cell renal cell carcinoma patients. Covers primary tumour, matched metastasis, and adjacent normal kidney tissue. Isoform models derived from SQANTI3-classified assemblies. View Publication

Primary TumourMetastasisNormal KidneyPacBio MAS-seq

Breast Cancer

Head et al. 2026

Nanopore long-read ESPRESSO assemblies from healthy breast tissue, primary breast tumour, and cultured fibroblasts. Includes both annotated GENCODE transcripts and novel isoforms not present in reference databases. View Publication

Healthy BreastTumourFibroblastsNanopore · ESPRESSO

System Overview

SPLISOFORMS is a splicing proteomics knowledgebase that transforms genomic transcriptomic data into three-dimensional conformational, immunological, and regulatory insights. By integrating structural predictions from custom AlphaFold 3 models across multiple tissue contexts, the platform identifies high-impact perturbations that affect protein stability, functional domains, neoantigen potential, and post-translational regulation.

Every protein-coding transcript and novel isoform is modelled with the same AlphaFold 3 workflow for cross-comparable structures. Isoforms are reconciled to canonical Ensembl identifiers (with MANE Select prioritisation), ensuring a unified GENCODE v47 (Ensembl 113)reference architecture.

84k+Isoforms Indexed
0Neoantigen Predictions
277k+Pfam & TED Domains
84k+AF3 Models Computed