All modules

M03 · In Build

ORACLE

Multimodal reasoning and vision stack

Investment

$2,250,000

Share of programme

15.7%

Timeline

Q4 2025 — Q3 2027

Delivery team

17 engineers / 5 researchers

41% complete

Multimodal perception and reasoning visualisation for the ORACLE module
M03 · ORACLEMultimodal reasoning and vision stack

In plain English

Gives the platform eyes and ears. It reads scanned paperwork, photographs, diagrams, charts and audio, and hands the meaning to the reasoning model.

What the software does

  • Reads scanned forms, invoices and handwritten notes and pulls out the fields that matter.
  • Inspects photographs and video for defects, damage, wear or safety breaches.
  • Interprets charts, engineering drawings and screenshots rather than just describing them.
  • Transcribes spoken audio and links it to whatever was being discussed on screen.

How it works, step by step

  1. 01Take in the raw signalAn image, page, video frame or audio clip is captured exactly as it exists in the real workflow — no manual re-keying.
  2. 02Translate it into meaningThe content is converted into the same internal representation the language model already understands.
  3. 03Reason across formats togetherA photo, the maintenance manual and the sensor log can be weighed in a single answer instead of three separate reviews.

A simple analogy

It is the difference between an assistant who can only read your emails and one who can also look at the photo you attached.

Why it matters

Most real enterprise work is not clean text. Perception is what lets the platform touch the documents and imagery businesses actually hold.

How ORACLE works

Inside the module

Pictures, scans, audio and video go in; the same reasoning applies to all of them.

Input

  • Images & video

    Photographs, inspection footage, camera streams

  • Documents

    Scanned forms, plans, contracts and handwriting

  • Signals

    Audio recordings, telemetry and sensor feeds

M03 pipeline · select a stage

1/4

Each input type is read by a specialist encoder

A photograph, a scanned page and an audio clip are each processed by a component built for that medium, producing a machine-readable description of the content.

Vision encodersDocument layout parsingSpeech recognition

Output

  • Structured records

    Fields and tables extracted from unstructured sources

  • Unified embeddings

    Searchable representations shared with the core model

  • Confidence flags

    Explicit uncertainty instead of a silent guess

Stage by stage, in detail

01

Encode

Images, scanned documents, video frames, charts and sensor series pass through modality-specific encoders.

02

Project

A shared projection layer maps every modality into the NEXUS token space for single-context reasoning.

03

Ground

Outputs are tied back to the exact source region — pixel span, table cell or time window.

04

Structure

Contracts, invoices, drawings and filings are emitted as structured records with citations.

05

Explain

Visual inspection workflows return a justification, not just a classification label.

94%

Span accuracy

Region-level citation prototype on internal document sets

5

Modalities

Image, document, video, chart and sensor series

41%

Module complete

Encoders integrated with the pilot model

Vision transformer encodersShared cross-modal projectorOCR + layout parsingTemporal video adapterTime-series fusion layerRegion-level citation format

Questions answered

ORACLE FAQ — how the AI works, in plain terms

Common investor questions about what this module does, how it does it, and why it is funded as part of the programme.

Scope

ORACLE adds perception. A shared projection layer maps images, scanned documents, video frames, charts and time-series sensor streams into the NEXUS token space, letting one model reason across modalities in a single context.

Document intelligence is the first commercial surface: contracts, invoices, engineering drawings and compliance filings parsed to structured records with citation back to the source pixel region.

The same stack drives visual inspection workflows where the model must explain its judgement, not just classify.

Contracted deliverables

  • Unified multimodal projector and encoder set
  • Document intelligence service with region-level citation
  • Video and time-series reasoning adapters
  • Grounded visual explanation output format

Achieved to date

  • Image and document encoders integrated with the 13B pilot
  • Region-level citation prototype at 94% span accuracy

Currently in production

  • Video temporal reasoning adapter
  • Sensor and time-series fusion layer
  • Chart and schematic understanding benchmark suite

Where it is used

Contract and filing intelligence

Structured extraction with a citation back to the source region, suitable for legal review.

Engineering drawing review

Schematic understanding that flags deviations and explains the reasoning behind each flag.

Explainable visual inspection

Quality and safety inspection where an auditable justification is a procurement requirement.

Platform dependencies

  • Projects into the NEXUS token space
  • Annotation corpora from SENTRA
  • Surfaced to customers through ATLAS

Key risks and mitigations

Annotation cost overruns

Synthetic augmentation from SENTRA reduces the human-labelled volume required.

Modality integration regressions

Per-modality benchmark suites gate every merge into the shared projector.