All modules

M02 · In Build

SENTRA

Data ingestion and synthetic corpus pipeline

Investment

$1,950,000

Share of programme

13.6%

Timeline

Q2 2025 — Q4 2026

Delivery team

14 engineers / 9 data specialists

74% complete

Data refinery pipeline visualisation for the SENTRA module
M02 · SENTRAData ingestion and synthetic corpus pipeline

In plain English

The refinery that turns messy, mixed-up source material into clean training data — and keeps a receipt showing where every document came from.

What the software does

  • Collects text, tables, code and transcripts from licensed sources and files them in one consistent format.
  • Throws away duplicates and low-quality material before anyone pays to train on it.
  • Records a licence and a lineage trail for every document, so an auditor can ask 'where did this come from?' and get an answer.
  • Writes additional practice material where real data is scarce or legally restricted.

How it works, step by step

  1. 01Gather with permissionEvery source is contracted or openly licensed, and the permission record travels with the document.
  2. 02Clean and gradeNear-identical copies are removed and everything left is scored for quality, so noise never reaches the training run.
  3. 03Serve the right mixA sampler decides what the next training run actually sees, weighted towards the subjects where the model is currently weakest.

A simple analogy

It is a water treatment plant for information: the same river goes in, but what comes out is safe to drink and you can prove where it was drawn.

Why it matters

Regulated buyers will not deploy a model whose training data cannot be accounted for. The pipeline, not the weights, is the durable asset.

How SENTRA works

Inside the module

Messy source material goes in; clean, licence-attested training data comes out.

Input

  • Licensed corpora

    Contracted and openly licensed text, code and tables

  • Customer archives

    Partner-supplied documents, with permission recorded

  • Transcripts

    Audio and video converted to searchable text

M02 pipeline · select a stage

1/4

Everything is collected into one consistent format

Mixed formats — PDFs, spreadsheets, code repositories, transcripts — are normalised so downstream steps only ever handle one shape of data.

Format normalisationLicence capture at sourceContent addressing

Output

  • Training shards

    Curated, weighted data ready for NEXUS Core

  • Synthetic data

    Generated practice material where real data is restricted

  • Lineage record

    Auditable evidence of where every document came from

Stage by stage, in detail

01

Licensed acquisition

Sources are contracted or permissively licensed, and every document enters with a machine-readable licence record.

02

Normalisation

Multi-format extraction converts documents, tables, code and transcripts into a single structured representation.

03

Dedup and scoring

Near-duplicate detection and quality scoring strip noise before a single training token is spent on it.

04

Governed synthesis

Synthetic instruction and reasoning data is generated under policy constraints, then verified before admission.

05

Curriculum sampling

A sampler tied to evaluation regressions decides what the next training run actually sees.

4.2PB

Corpus ingested

Licensed and permissive data normalised to date

31%

Noise removed

Effective corpus noise cut by near-duplicate detection

100%

Attested documents

Every document carries a signed licence attestation

12

Languages queued

Multilingual expansion in progress

Petabyte object storageSpark / Ray processingMinHash near-dup detectionSigned provenance ledgerPolicy-constrained synthesisCurriculum sampler

Questions answered

SENTRA FAQ — how the AI works, in plain terms

Common investor questions about what this module does, how it does it, and why it is funded as part of the programme.

Scope

SENTRA turns raw, messy, multi-source data into training-grade corpora with a full provenance chain. Every document carries a licence record, a quality score and a lineage hash that survives all the way into model checkpoints.

The synthetic arm generates task-specific instruction and reasoning data under policy constraints, expanding coverage in domains where licensed human data is scarce or legally restricted.

This module is what makes the platform defensible: the pipeline, not just the weights, is the compounding asset.

Contracted deliverables

  • Petabyte-scale ingestion and dedup cluster
  • Provenance ledger with per-document licence attestation
  • Synthetic generation and verification pipeline
  • Data quality scoring and curriculum sampler

Achieved to date

  • 4.2PB of licensed and permissive data ingested and normalised
  • Provenance ledger live with per-document attestation
  • Near-duplicate detection cutting effective corpus noise by 31%

Currently in production

  • Synthetic reasoning corpus for scarce industrial domains
  • Automated curriculum sampler tied to eval regressions
  • Multilingual expansion across 12 additional languages

Where it is used

Defensible training data

The pipeline, not just the weights, is the compounding asset competitors cannot copy quickly.

Audit-ready lineage

Regulators and enterprise buyers can trace any model behaviour back to source documents.

Scarce-domain coverage

Synthetic generation fills industrial domains where licensed human data barely exists.

Platform dependencies

  • Feeds all NEXUS training runs
  • Supplies annotation corpora to ORACLE
  • Licence evidence surfaced through AEGIS

Key risks and mitigations

Licensing regime tightens

Provenance ledger allows surgical removal of any source without retraining from scratch.

Synthetic data quality drift

Every synthetic batch is verified against held-out human-labelled reference sets.