M06 · In Design
VERTEX
Edge deployment and compute fabric
Investment
$1,800,000
Share of programme
12.6%
Timeline
Q2 2026 — Q1 2028
Delivery team
12 engineers / 3 hardware specialists
29% complete

In plain English
The plumbing that makes the platform actually run — scheduling scarce hardware efficiently and shrinking models so they fit on modest machines.
What the software does
- Packs training and serving jobs onto expensive hardware so very little of it sits idle.
- Compresses large models into smaller versions that still perform, for sites with limited hardware.
- Runs the platform in a customer's own datacentre, at a remote site, or on a factory floor with poor connectivity.
- Moves work between locations automatically when hardware fails or costs spike.
How it works, step by step
- 01ScheduleJobs are bin-packed across available accelerators, mixing reserved and opportunistic capacity to cut cost per unit of work.
- 02ShrinkDistillation and quantisation produce lighter models that keep the accuracy the task actually needs.
- 03Deploy anywhereOne deployment description runs the same stack in the cloud, on-premise, or at a disconnected edge site.
A simple analogy
It is the logistics network behind the product: the same goods, delivered wherever the customer is, at the lowest sensible cost.
Why it matters
Compute is the largest recurring cost in the programme. Every efficiency here improves gross margin permanently.
How VERTEX works
Inside the module
Model artefacts and demand go in; fast, affordable inference anywhere comes out.
Input
Model artefacts
Trained checkpoints from NEXUS Core
Workloads
Training runs and live inference demand
Targets
Cloud, on-premise, air-gapped and edge hardware
M06 pipeline · select a stage
1/4Models are optimised for the hardware they will run on
Distillation and quantisation shrink the model so it runs on constrained devices without a meaningful loss of quality.
Output
Optimised builds
One model, running on any supported hardware
Stable latency
Predictable response times under production load
Lower unit cost
Less compute spent per answer produced
Stage by stage, in detail
01
Distil
3B–13B student models are distilled from the production model for constrained hardware.
02
Quantise
A quantisation toolchain targets 8GB-class devices without collapsing task quality.
03
Schedule
A hardware-aware fabric bin-packs workloads across heterogeneous accelerators, spot and reserved capacity.
04
Deploy anywhere
Cloud, on-premise and air-gapped sites run the same API with signed model bundles.
05
Operate
Fleet telemetry and remote health management keep distributed sites observable and updatable.
88%
Quality retained
3B distillation prototype versus pilot model task quality
8GB
Device target
Class of hardware the quantisation toolchain targets
2
Launch partners
Joined the scheduler design review
Questions answered
VERTEX FAQ — how the AI works, in plain terms
Common investor questions about what this module does, how it does it, and why it is funded as part of the programme.
Scope
Many of the highest-value buyers cannot send data to a shared cloud. VERTEX makes that a configuration choice rather than a rewrite, packaging distilled 3B–13B models with a hardware-aware scheduler.
The compute fabric bin-packs workloads across heterogeneous accelerators, spot capacity and reserved clusters, and is the single largest lever on gross margin as inference volume scales.
Air-gapped deployment ships with signed model bundles and an offline update path for defence, utilities and healthcare buyers.
Contracted deliverables
- Distilled model family for constrained hardware
- Heterogeneous scheduling and bin-packing fabric
- Air-gapped and on-premise deployment bundles
- Fleet telemetry and remote health management
Achieved to date
- 3B distillation prototype at 88% of pilot-model task quality
- Scheduler design review completed with two launch partners
Currently in production
- Cross-accelerator bin-packing scheduler
- Signed offline update channel for air-gapped sites
- Quantisation toolchain for 8GB-class devices
Where it is used
Air-gapped deployment
Defence, utilities and healthcare buyers who cannot send data to a shared cloud.
Field and edge inference
Low-latency inference on constrained devices at industrial sites.
Gross-margin control
The single largest lever on unit economics as inference volume scales.
Platform dependencies
- Distils NEXUS production weights
- Runs AEGIS policy locally at each site
- Billing and telemetry surfaced through ATLAS
Key risks and mitigations
Hardware partner slippage
Validation spread across multiple accelerator vendors rather than a single supplier.
Quality loss at aggressive quantisation
Per-target quality floors block any bundle that drops below the contracted threshold.
