# MobiDB Public Dataset — Release 2026_07

## Overview

This directory contains the MobiDB database release datasets, available as both standard [MobiDB JSON Lines](https://mobidb.org) and AI-optimized [Apache Parquet](https://parquet.apache.org/) files.

MobiDB is a comprehensive database of protein intrinsic disorder and related features. It integrates curated, derived, homology-based, and predicted annotations from multiple sources with explicit provenance tracking.

**Homepage:** [https://mobidb.org](https://mobidb.org)
**Release:** 2026_07
**License:** CC-BY-4.0

---

## Dataset Tiers & Provenance

MobiDB categorizes proteins and annotations into hierarchical quality tiers based on evidence:

- **Gold Tier (`level = 2`)**: High-confidence, manually curated annotations from dedicated databases (DisProt, IDEAL, PhasepDB, ELM, FuzDB) plus high-confidence homology transfer from curated proteins.
- **Silver Tier (`level >= 1`)**: A superset combining all **Gold** data with **experimentally derived structural data** from PDB structures via SIFTS (missing residues, mobile/flexible regions, inter-chain contacts, and disorder-to-order binding modes) and high-confidence **AlphaFold** disorder/pLDDT/LIP predictions.
- **All / Predictions (`level = 0`)**: Genome-wide computational predictions across all of UniProt (Swiss-Prot, TrEMBL, splice variants) generated via MobiDB-lite.

---

## Release Files

| File | Format | Size | Description |
|------|--------|------|-------------|
| `releases/2026_07/mobidb_silver.mjson.gz` | MobiDB JSON Lines (gzip-compressed) | 3.3 GB | Complete MobiDB JSON records for the Silver tier (level >= 1, combining curated, homology, PDB-derived, and AlphaFold annotations). Full nested document structure. |
| `releases/2026_07/mobidblite_sprot.mjson.gz` | MobiDB JSON Lines (gzip-compressed) | 270.0 MB | MobiDB-lite consensus disorder predictions for UniProt Swiss-Prot (reviewed canonical sequences). |
| `releases/2026_07/mobidblite_sprot_varsplic.mjson.gz` | MobiDB JSON Lines (gzip-compressed) | 31.5 MB | MobiDB-lite consensus disorder predictions for UniProt Swiss-Prot alternative splice isoforms. |
| `releases/2026_07/mobidblite_trembl.mjson.gz` | MobiDB JSON Lines (gzip-compressed) | 106.5 GB | MobiDB-lite consensus disorder predictions for UniProt TrEMBL (unreviewed sequences). |
| `releases/2026_07/proteins_silver_2026_07_annotations.parquet` | Apache Parquet | 834.0 MB | One row per annotation region with provenance metadata (evidence level, source, feature type). |
| `releases/2026_07/proteins_silver_2026_07_residues.parquet` | Apache Parquet | 2.5 GB | One row per residue position with binary annotation indicators and score columns for each feature. |

---

## Parquet Dataset Schemas

### proteins_silver_2026_07_annotations.parquet

One row per annotation region. Each annotation has explicit provenance columns (`evidence`, `feature`, `source`) so that curated, derived, homology-based, and predicted annotations are clearly distinguished.

**Size:** 834.0 MB (151,965,897 rows)

#### Schema

| Column | Type | Description |
|--------|------|-------------|
| `acc` | `string` | UniProt accession identifier for the protein. |
| `length` | `int32` | Length of the protein sequence in amino acids. |
| `feature` | `string` | Annotation feature type. Values include: disorder (intrinsically disordered r... |
| `evidence` | `dictionary<values=string, indices=int32, ordered=0>` | Provenance class of the annotation. 'curated' = experimentally supported/manu... |
| `source` | `string` | Annotation source tool or database. Examples: disprot, ideal, uniprot, elm, p... |
| `source_id` | `string` | External identifier from the source database (e.g., DisProt entry ID). Empty ... |
| `start` | `int32` | Start position of the annotated region (1-indexed, inclusive). |
| `end` | `int32` | End position of the annotated region (1-indexed, inclusive). |
| `region_id` | `string` | Region-specific identifier (e.g., Pfam accession PF04947). Empty if not appli... |
| `region_name` | `string` | Human-readable region name (e.g., 'Poxvirus VLTF3, late transcription factor'... |
| `content_count` | `int32` | Total number of residues annotated by this feature across all its regions for... |
| `content_fraction` | `float` | Fraction of the protein sequence covered by this feature (content_count / len... |
| `mobidb_release` | `dictionary<values=string, indices=int32, ordered=0>` | MobiDB release version identifier (e.g., '2026_07'). |


### proteins_silver_2026_07_residues.parquet

One row per residue position. Columns are dynamically generated based on the annotation features present in the dataset.

**Size:** 2.5 GB (888,855,030 rows)

#### Schema

| Column | Type | Description |
|--------|------|-------------|
| `acc` | `string` | UniProt accession identifier for the protein. |
| `position` | `int32` | Residue position within the protein sequence (1-indexed). |
| `aa` | `dictionary<values=string, indices=int32, ordered=0>` | Single-letter amino acid code at this position. |
| `curated-binding_mode_disorder_to_disorder-fuzdb` | `uint8` | Annotation feature column |
| `curated-binding_mode_disorder_to_disorder-priority` | `uint8` | Annotation feature column |
| `curated-disorder-disprot` | `uint8` | Annotation feature column |
| `curated-disorder-ideal` | `uint8` | Annotation feature column |
| `curated-disorder-uniprot` | `uint8` | Annotation feature column |
| `curated-lip-dibs` | `uint8` | Annotation feature column |
| `curated-lip-disprot` | `uint8` | Annotation feature column |
| `curated-lip-elm` | `uint8` | Annotation feature column |
| `curated-lip-ideal` | `uint8` | Annotation feature column |
| `curated-lip-lmpid` | `uint8` | Annotation feature column |
| `curated-lip-mfib` | `uint8` | Annotation feature column |
| `curated-lip-momap` | `uint8` | Annotation feature column |
| `curated-phase_separation-phasepro` | `uint8` | Annotation feature column |
| `derived-binding_mode_context_dependent-mobi` | `uint8` | Annotation feature column |
| `derived-binding_mode_context_dependent-priority` | `uint8` | Annotation feature column |
| `derived-binding_mode_disorder_to_disorder-mobi` | `uint8` | Annotation feature column |
| `derived-binding_mode_disorder_to_disorder-priority` | `uint8` | Annotation feature column |
| `derived-binding_mode_disorder_to_order-mobi` | `uint8` | Annotation feature column |
| `derived-binding_mode_disorder_to_order-priority` | `uint8` | Annotation feature column |
| `derived-fixed-th_90` | `uint8` | Annotation feature column |
| `derived-lip-flipper` | `uint8` | Annotation feature column |
| `derived-lip-merge` | `uint8` | Annotation feature column |
| `derived-lip-priority` | `uint8` | Annotation feature column |
| `derived-lip-th_90` | `uint8` | Annotation feature column |
| `derived-missing_residues-mobi` | `uint8` | Annotation feature column |
| `derived-missing_residues-priority` | `uint8` | Annotation feature column |
| `derived-missing_residues-th_90` | `uint8` | Annotation feature column |
| `derived-missing_residues_context_dependent-th_90` | `uint8` | Annotation feature column |
| `derived-mobile-mobi` | `uint8` | Annotation feature column |
| `derived-mobile-th_90` | `uint8` | Annotation feature column |
| `derived-mobile_context_dependent-th_90` | `uint8` | Annotation feature column |
| `derived-observed-mobi` | `uint8` | Annotation feature column |
| `derived-observed-priority` | `uint8` | Annotation feature column |
| `derived-observed-th_90` | `uint8` | Annotation feature column |
| `homology-binding_mode_disorder_to_disorder-fuzdb` | `uint8` | Annotation feature column |
| `homology-binding_mode_disorder_to_disorder-priority` | `uint8` | Annotation feature column |
| `homology-disorder-disprot` | `uint8` | Annotation feature column |
| `homology-disorder-ideal` | `uint8` | Annotation feature column |
| `homology-disorder-uniprot` | `uint8` | Annotation feature column |
| `homology-domain-gene3d` | `uint8` | Annotation feature column |
| `homology-domain-merge` | `uint8` | Annotation feature column |
| `homology-domain-pfam` | `uint8` | Annotation feature column |
| `homology-lip-dibs` | `uint8` | Annotation feature column |
| `homology-lip-disprot` | `uint8` | Annotation feature column |
| `homology-lip-elm` | `uint8` | Annotation feature column |
| `homology-lip-ideal` | `uint8` | Annotation feature column |
| `homology-lip-lmpid` | `uint8` | Annotation feature column |
| `homology-lip-mfib` | `uint8` | Annotation feature column |
| `homology-lip-momap` | `uint8` | Annotation feature column |
| `homology-msa_conservation-psiblast_score` | `float` | Annotation feature column |
| `homology-msa_entropy-psiblast_score` | `float` | Annotation feature column |
| `homology-msa_information_content-psiblast_score` | `float` | Annotation feature column |
| `homology-msa_occupancy-psiblast_score` | `float` | Annotation feature column |
| `homology-phase_separation-phasepro` | `uint8` | Annotation feature column |
| `prediction-coiled_coil-uniprot` | `uint8` | Annotation feature column |
| `prediction-compact-mobidb_lite_sub` | `uint8` | Annotation feature column |
| `prediction-cysteine_rich-mobidb_lite_sub` | `uint8` | Annotation feature column |
| `prediction-disorder-dis465` | `uint8` | Annotation feature column |
| `prediction-disorder-disHL` | `uint8` | Annotation feature column |
| `prediction-disorder-espD` | `uint8` | Annotation feature column |
| `prediction-disorder-espN` | `uint8` | Annotation feature column |
| `prediction-disorder-espX` | `uint8` | Annotation feature column |
| `prediction-disorder-glo` | `uint8` | Annotation feature column |
| `prediction-disorder-iupl` | `uint8` | Annotation feature column |
| `prediction-disorder-iups` | `uint8` | Annotation feature column |
| `prediction-disorder-mobidb_lite_region` | `uint8` | Annotation feature column |
| `prediction-disorder-mobidb_lite_score` | `float` | Annotation feature column |
| `prediction-disorder-priority` | `uint8` | Annotation feature column |
| `prediction-disorder-th_50` | `uint8` | Annotation feature column |
| `prediction-extended-mobidb_lite_sub` | `uint8` | Annotation feature column |
| `prediction-glycine_rich-mobidb_lite_sub` | `uint8` | Annotation feature column |
| `prediction-low_complexity-mobidb_lite_sub` | `uint8` | Annotation feature column |
| `prediction-low_complexity-seg` | `uint8` | Annotation feature column |
| `prediction-negative_polyelectrolyte-mobidb_lite_sub` | `uint8` | Annotation feature column |
| `prediction-plddt-alphafold_region` | `uint8` | Annotation feature column |
| `prediction-plddt-alphafold_score` | `float` | Annotation feature column |
| `prediction-polar-mobidb_lite_sub` | `uint8` | Annotation feature column |
| `prediction-polyampholyte-mobidb_lite_sub` | `uint8` | Annotation feature column |
| `prediction-positive_polyelectrolyte-mobidb_lite_sub` | `uint8` | Annotation feature column |
| `prediction-proline_rich-mobidb_lite_sub` | `uint8` | Annotation feature column |
| `prediction-signal_peptide-uniprot` | `uint8` | Annotation feature column |
| `prediction-transmembrane-uniprot` | `uint8` | Annotation feature column |


---

## Evidence Levels

MobiDB distinguishes four provenance classes for annotations:

| Evidence | Description |
|----------|-------------|
| `curated` | Experimentally supported/manual annotations from databases like DisProt, IDEAL, UniProt |
| `derived` | Annotations inferred from experimental structural data (e.g., PDB missing residues) |
| `homology` | Transferred annotations via PSI-BLAST sequence similarity |
| `prediction` | Computational predictions (e.g., MobiDB-lite, ESpritz, IUPred, disHL) |

> **Important:** A model should not silently mix evidence levels. A curated annotation, a homology transfer, and a MobiDB-lite prediction carry very different confidence levels.

---

## Coordinate Convention

All positions use **1-based inclusive** coordinates. Both `start` and `end` positions are included in the annotated region.

---

## Usage Examples

### Python + DuckDB (remote query, no download required)

```python
import duckdb

# Query all curated disorder annotations
duckdb.sql("""
    SELECT acc, feature, source, start, end, content_fraction
    FROM 'https://data.mobidb.org/releases/2026_07/*_annotations.parquet'
    WHERE feature = 'disorder'
      AND evidence = 'curated'
    LIMIT 10
""").show()
```

### Python + Pandas

```python
import pandas as pd

df = pd.read_parquet("https://data.mobidb.org/releases/2026_07/*_annotations.parquet")

# Get all disorder annotations for human TP53
p53 = df[(df["acc"] == "P04637") & (df["feature"] == "disorder")]
print(p53[["evidence", "source", "start", "end", "content_fraction"]])
```

### Python + Polars

```python
import polars as pl

df = pl.read_parquet("https://data.mobidb.org/releases/2026_07/*_annotations.parquet")

# Filter curated annotations only
curated = df.filter(pl.col("evidence") == "curated")
print(curated.head())
```

---

## Citation

If you use MobiDB data in your research, please cite:

> Piovesan D, et al. MobiDB: 10 years of intrinsically disordered proteins.
> *Nucleic Acids Res.* 2023;51(D1):D438-D444.
> DOI: [10.1093/nar/gkac1065](https://doi.org/10.1093/nar/gkac1065)

---

## Related Resources

- **MobiDB website:** [https://mobidb.org](https://mobidb.org)
- **MobiDB API:** [https://mobidb.org/help/apidoc](https://mobidb.org/help/apidoc)
- **Bioschemas markup:** Available on mobidb.org protein pages
- **IDP Knowledge Graph:** [https://idpkg.mobidb.org](https://idpkg.mobidb.org)
