M02 · In Build
SENTRA
Data ingestion and synthetic corpus pipeline
Investment
$1,950,000
Share of programme
13.6%
Timeline
Q2 2025 — Q4 2026
Delivery team
14 engineers / 9 data specialists
74% complete

In plain English
The refinery that turns messy, mixed-up source material into clean training data — and keeps a receipt showing where every document came from.
What the software does
- Collects text, tables, code and transcripts from licensed sources and files them in one consistent format.
- Throws away duplicates and low-quality material before anyone pays to train on it.
- Records a licence and a lineage trail for every document, so an auditor can ask 'where did this come from?' and get an answer.
- Writes additional practice material where real data is scarce or legally restricted.
How it works, step by step
- 01Gather with permissionEvery source is contracted or openly licensed, and the permission record travels with the document.
- 02Clean and gradeNear-identical copies are removed and everything left is scored for quality, so noise never reaches the training run.
- 03Serve the right mixA sampler decides what the next training run actually sees, weighted towards the subjects where the model is currently weakest.
A simple analogy
It is a water treatment plant for information: the same river goes in, but what comes out is safe to drink and you can prove where it was drawn.
Why it matters
Regulated buyers will not deploy a model whose training data cannot be accounted for. The pipeline, not the weights, is the durable asset.
How SENTRA works
Inside the module
Messy source material goes in; clean, licence-attested training data comes out.
Input
Licensed corpora
Contracted and openly licensed text, code and tables
Customer archives
Partner-supplied documents, with permission recorded
Transcripts
Audio and video converted to searchable text
M02 pipeline · select a stage
1/4Everything is collected into one consistent format
Mixed formats — PDFs, spreadsheets, code repositories, transcripts — are normalised so downstream steps only ever handle one shape of data.
Output
Training shards
Curated, weighted data ready for NEXUS Core
Synthetic data
Generated practice material where real data is restricted
Lineage record
Auditable evidence of where every document came from
Stage by stage, in detail
01
Licensed acquisition
Sources are contracted or permissively licensed, and every document enters with a machine-readable licence record.
02
Normalisation
Multi-format extraction converts documents, tables, code and transcripts into a single structured representation.
03
Dedup and scoring
Near-duplicate detection and quality scoring strip noise before a single training token is spent on it.
04
Governed synthesis
Synthetic instruction and reasoning data is generated under policy constraints, then verified before admission.
05
Curriculum sampling
A sampler tied to evaluation regressions decides what the next training run actually sees.
4.2PB
Corpus ingested
Licensed and permissive data normalised to date
31%
Noise removed
Effective corpus noise cut by near-duplicate detection
100%
Attested documents
Every document carries a signed licence attestation
12
Languages queued
Multilingual expansion in progress
Questions answered
SENTRA FAQ — how the AI works, in plain terms
Common investor questions about what this module does, how it does it, and why it is funded as part of the programme.
Scope
SENTRA turns raw, messy, multi-source data into training-grade corpora with a full provenance chain. Every document carries a licence record, a quality score and a lineage hash that survives all the way into model checkpoints.
The synthetic arm generates task-specific instruction and reasoning data under policy constraints, expanding coverage in domains where licensed human data is scarce or legally restricted.
This module is what makes the platform defensible: the pipeline, not just the weights, is the compounding asset.
Contracted deliverables
- Petabyte-scale ingestion and dedup cluster
- Provenance ledger with per-document licence attestation
- Synthetic generation and verification pipeline
- Data quality scoring and curriculum sampler
Achieved to date
- 4.2PB of licensed and permissive data ingested and normalised
- Provenance ledger live with per-document attestation
- Near-duplicate detection cutting effective corpus noise by 31%
Currently in production
- Synthetic reasoning corpus for scarce industrial domains
- Automated curriculum sampler tied to eval regressions
- Multilingual expansion across 12 additional languages
Where it is used
Defensible training data
The pipeline, not just the weights, is the compounding asset competitors cannot copy quickly.
Audit-ready lineage
Regulators and enterprise buyers can trace any model behaviour back to source documents.
Scarce-domain coverage
Synthetic generation fills industrial domains where licensed human data barely exists.
Platform dependencies
- Feeds all NEXUS training runs
- Supplies annotation corpora to ORACLE
- Licence evidence surfaced through AEGIS
Key risks and mitigations
Licensing regime tightens
Provenance ledger allows surgical removal of any source without retraining from scratch.
Synthetic data quality drift
Every synthetic batch is verified against held-out human-labelled reference sets.
