All modules

M06 · In Design

VERTEX

Edge deployment and compute fabric

Investment

$1,800,000

Share of programme

12.6%

Timeline

Q2 2026 — Q1 2028

Delivery team

12 engineers / 3 hardware specialists

29% complete

Distributed edge compute fabric visualisation for the VERTEX module
M06 · VERTEXEdge deployment and compute fabric

In plain English

The plumbing that makes the platform actually run — scheduling scarce hardware efficiently and shrinking models so they fit on modest machines.

What the software does

  • Packs training and serving jobs onto expensive hardware so very little of it sits idle.
  • Compresses large models into smaller versions that still perform, for sites with limited hardware.
  • Runs the platform in a customer's own datacentre, at a remote site, or on a factory floor with poor connectivity.
  • Moves work between locations automatically when hardware fails or costs spike.

How it works, step by step

  1. 01ScheduleJobs are bin-packed across available accelerators, mixing reserved and opportunistic capacity to cut cost per unit of work.
  2. 02ShrinkDistillation and quantisation produce lighter models that keep the accuracy the task actually needs.
  3. 03Deploy anywhereOne deployment description runs the same stack in the cloud, on-premise, or at a disconnected edge site.

A simple analogy

It is the logistics network behind the product: the same goods, delivered wherever the customer is, at the lowest sensible cost.

Why it matters

Compute is the largest recurring cost in the programme. Every efficiency here improves gross margin permanently.

How VERTEX works

Inside the module

Model artefacts and demand go in; fast, affordable inference anywhere comes out.

Input

  • Model artefacts

    Trained checkpoints from NEXUS Core

  • Workloads

    Training runs and live inference demand

  • Targets

    Cloud, on-premise, air-gapped and edge hardware

M06 pipeline · select a stage

1/4

Models are optimised for the hardware they will run on

Distillation and quantisation shrink the model so it runs on constrained devices without a meaningful loss of quality.

DistillationINT8 / FP8 quantisationKernel compilation

Output

  • Optimised builds

    One model, running on any supported hardware

  • Stable latency

    Predictable response times under production load

  • Lower unit cost

    Less compute spent per answer produced

Stage by stage, in detail

01

Distil

3B–13B student models are distilled from the production model for constrained hardware.

02

Quantise

A quantisation toolchain targets 8GB-class devices without collapsing task quality.

03

Schedule

A hardware-aware fabric bin-packs workloads across heterogeneous accelerators, spot and reserved capacity.

04

Deploy anywhere

Cloud, on-premise and air-gapped sites run the same API with signed model bundles.

05

Operate

Fleet telemetry and remote health management keep distributed sites observable and updatable.

88%

Quality retained

3B distillation prototype versus pilot model task quality

8GB

Device target

Class of hardware the quantisation toolchain targets

2

Launch partners

Joined the scheduler design review

Knowledge distillationINT4 / INT8 quantisationHeterogeneous bin-packing schedulerSigned offline update channelFleet telemetry

Questions answered

VERTEX FAQ — how the AI works, in plain terms

Common investor questions about what this module does, how it does it, and why it is funded as part of the programme.

Scope

Many of the highest-value buyers cannot send data to a shared cloud. VERTEX makes that a configuration choice rather than a rewrite, packaging distilled 3B–13B models with a hardware-aware scheduler.

The compute fabric bin-packs workloads across heterogeneous accelerators, spot capacity and reserved clusters, and is the single largest lever on gross margin as inference volume scales.

Air-gapped deployment ships with signed model bundles and an offline update path for defence, utilities and healthcare buyers.

Contracted deliverables

  • Distilled model family for constrained hardware
  • Heterogeneous scheduling and bin-packing fabric
  • Air-gapped and on-premise deployment bundles
  • Fleet telemetry and remote health management

Achieved to date

  • 3B distillation prototype at 88% of pilot-model task quality
  • Scheduler design review completed with two launch partners

Currently in production

  • Cross-accelerator bin-packing scheduler
  • Signed offline update channel for air-gapped sites
  • Quantisation toolchain for 8GB-class devices

Where it is used

Air-gapped deployment

Defence, utilities and healthcare buyers who cannot send data to a shared cloud.

Field and edge inference

Low-latency inference on constrained devices at industrial sites.

Gross-margin control

The single largest lever on unit economics as inference volume scales.

Platform dependencies

  • Distils NEXUS production weights
  • Runs AEGIS policy locally at each site
  • Billing and telemetry surfaced through ATLAS

Key risks and mitigations

Hardware partner slippage

Validation spread across multiple accelerator vendors rather than a single supplier.

Quality loss at aggressive quantisation

Per-target quality floors block any bundle that drops below the contracted threshold.