MOTHER EXO
A frontier world model — not text tokens, but Eyes, Ears, a Mouth, a Brain, memory and understanding, all at once.
MOTHER EXO is our answer to a frontier model for the physical world. Where a language model predicts the next word, EXO perceives, understands and acts across every modality simultaneously — seeing, hearing, speaking, reasoning, remembering, and reading the emotion in a scene. Built on a frozen CORE-7B mind and a frozen MOTHER DeepVision eye, with dozens of task heads trained on real on-node data, EXO is one model that does everything for us at once: a true world model, observe-and-advise, human-in-the-loop, non-kinetic by design.
Eyes, Ears, a Mouth, a Brain — at once
One model with every sense, running simultaneously — not a pipeline of tools, but a single world model.
MOTHER DeepVision ViT — detects, tracks and segments objects, faces and relationships in real time across camera, drone and satellite feeds.
Audio and speech understanding — transcribes, diarises and interprets what's said and the sounds of a scene.
Natural speech and expression — EXO responds in voice and communicates state, not just text.
A frozen CORE-7B mind reasons over everything it senses, plans, and grounds decisions in retrieved knowledge.
Persistent memory and self-learning heads let EXO accumulate context over time and improve from experience.
Scene-graph and relations heads build a structured understanding of who is doing what to whom, and why.
Expression and affect heads read emotion-based patterns — the mood and intent behind behaviour, not just the pixels.
Flight, humanoid and driving heads turn understanding into advised action — always observe-and-advise, human-in-the-loop.
Capabilities
One model, every modality — at once
EXO fuses vision, audio, language and action in a single world model. It doesn't switch between tools; it perceives and reasons simultaneously.
A living world model
EXO maintains an internal model of the scene — objects, people, relationships, motion and intent — and predicts how it will evolve.
Self-learning
Head-only training on real on-node corpora lets EXO keep learning new capabilities without retraining the frozen mind and eye.
Emotion-aware
Expression and affect heads surface emotion-based patterns, so EXO understands the human meaning of a scene.
Grounded reasoning
The CORE-7B backbone reasons over what EXO senses and grounds decisions in retrieved knowledge, with citations and audit.
Observe-and-advise
Across flight, driving and humanoid domains EXO advises a human operator — it never selects targets or takes kinetic action.
Architecture
- 1
Frozen CORE-7B ⊕ frozen DeepVision ViT
EXO shares MOTHER CORE as its reasoning mind and MOTHER DeepVision as its visual cortex — both frozen — so the expensive backbones are trained once and reused.
- 2
Head-only training
Dozens of lightweight task heads (detect, track, worldmodel, face, relations, memory, speech, vision, audio, expression, flight, humanoid, driving) are trained on real on-node data, so new senses are added cheaply and safely.
- 3
Real-data evaluation
Every head is evaluated on real video and images — detect · track · relations · face · assess · worldmodel · expression — plus reasoning and refusal, with measured scores on the /gb10/exo dashboard.
- 4
Sovereign GB10 node
EXO trains and infers on a sovereign GB10 Blackwell node; telemetry streams live to the operations dashboard.
Specifications
| Composition | Frozen CORE-7B ⊕ frozen DeepVision ViT ⊕ 14 trained heads |
| Bodies | humanoid · drone (Strike 0) · vehicle · defence |
| Capability heads | detect · track · worldmodel · face · relations · memory · speech · vision · audio · expression · flight · humanoid · driving |
| Detection data | COCO + LVIS + aircraft — 682k |
| Face data | MS-Celeb / CelebA / LFW — 5.4M crops |
| Teleop data | Open-X + GR00T — 150k real steps |
| Audio / reasoning | LibriSpeech + RAVDESS/CREMA-D 6.8k+ · sovereign QA 861k |
| Serving | load_exo.py · exo_node :8000 · humanoid_act / flight_act / drive |
| Safety | Observe-and-advise · HITL · Strike = 0 |
The 14 trained weights — function · dataset
| Weight | Kind | Function | Dataset | Trained |
|---|---|---|---|---|
| reasoning | text | Situational-awareness QA · threat classification · summaries · relationship & decision support | mother_core_v31 · golden_datasets · military/space converted · reasoning CoT/emotion | ✓ |
| detect | vision | 1,233-class presence + per-class boxes, parallel box decoding | COCO + LVIS + RarePlanes + military — 682k | ✓ |
| track | vision | Appearance embeddings — keep object identity across frames (contrastive) | perception frames · stage_F temporal · cosmos_reason1 | ✓ |
| worldmodel | world | Predict next-step latent / plausible futures (contrastive predictive) | on-node video + latent rollouts | ✓ |
| face | vision | Face representation + consent-only identity | MS-Celeb / CelebA / LFW — 5.4M crops (opt-in gallery) | ✓ |
| relations | scene-graph | Spatial relations — on / next_to / under classification | scene-graph annotations | ✓ |
| memory | memory | MOTHERrag Semantic-Pyramid long-term memory — write & recall | MOTHERrag corpus — 800k | ✓ |
| speech | speech | 20-language ID + Text-to-Voice; greets known faces by name | LibriSpeech + neural-codec | ✓ |
| vision | vision | Self-supervised visual foundation shared by the vision heads | SigLIP-SO400M · SimCLR over all on-node imagery | ✓ |
| audio | audio | ASR — radio signals + human speech → text (feeds reasoning) | LibriSpeech + RAVDESS / CREMA-D — 6.8k+ | ✓ |
| expression | vision | Reads affect — happy/sad/angry/worried/calm (NOT identity) | RAVDESS / CREMA-D / affect crops | ✓ |
| flight | autonomy | ISR velocity policy in the 6-DOF sim — observe-and-advise, Strike 0 | sim rollouts (6-DOF ISR) | ✓ |
| humanoid | autonomy | Embodiment motor policy behaviour-cloned from REAL teleop; HITL | Open-X + GR00T — 150k real teleop steps | ✓ |
| driving | autonomy | Perception → planning → control policy (first-pass); observe-and-advise | driving corpus (first-pass) | ✓ |
Model evaluation — industry benchmarks
| Benchmark | Metric | Result | Status |
|---|---|---|---|
| Detection — COCO | mAP50 | — | ○ to run |
| Tracking — MOT17 | MOTA / IDF1 | — | ○ to run |
| Face — LFW / IJB-C | accuracy | — | ○ to run |
| ASR — LibriSpeech | WER | — | ○ to run |
| Manipulation | task success | — | ○ to run |
| Locomotion | fall-rate | — | ○ to run |
| World-model | rollout FID | — | ○ to run |
Benchmarks scheduled on this build — measured results are recorded here as each run completes.
What it's for
Common operating picture
EXO fuses live camera, drone and sensor feeds into a single understood scene — objects, people, relationships and intent — for decision support.
Autonomous ISR (advise)
Supervises multi-UAV observation, mapping flight paths and surfacing what matters, while a human authorises every action.
Embodied assistance
Drives humanoid and mobile platforms in observe-and-advise mode, reading the environment and proposing safe actions.
World understanding
A general perception-and-reasoning engine any MOTHER product can call to understand what is happening in the physical world.
Safety & governance
- Strike = 0 across every domain — flight, driving and humanoid. EXO never selects targets, controls fire, or completes a kill-chain.
- Observe-and-advise only: EXO informs a human decision; a human acts.
- Human-in-the-loop is mandatory for any outward or physical action.
Frequently asked
What is MOTHER EXO?
MOTHER EXO is a frontier world model that perceives, understands and advises across vision, audio, speech, language and action simultaneously. It is built on a frozen CORE-7B reasoning mind and a frozen MOTHER DeepVision visual cortex, with dozens of task heads trained on real on-node data.
How is a world model different from a language model?
A language model predicts the next text token. A world model like EXO builds and maintains an internal model of the physical scene — objects, people, relationships, motion, emotion and intent — and reasons and advises across every modality at once.
Is MOTHER EXO a weapon?
No. EXO is observe-and-advise and non-kinetic by design, with Strike = 0 across flight, driving and humanoid domains. It supports human decisions and never selects targets or takes kinetic action.
What can MOTHER EXO do?
See, hear, speak, reason, remember, understand relationships, read emotion-based patterns, and advise action across flight, driving and humanoid platforms — all from one model, at once.
How does EXO keep learning?
EXO uses head-only training on real on-node corpora, so new capabilities are added by training lightweight heads without retraining the frozen mind and eye.
The MOTHER model family
A sovereign British reasoning model — the mind behind MOTHER.
The eyes of MOTHER — real-time detection, tracking and segmentation across every feed.
The sovereign general-purpose assistant — everyday chat and document making.
Describe it — MOTHER Code builds the website, app or game.
Build on MOTHER EXO
Sovereign, on-node and observe-and-advise by design. Talk to us about access and deployment.