MOTHER DeepVision
The eyes of MOTHER — real-time detection, tracking and segmentation across every feed.
MOTHER DeepVision is our sovereign computer-vision pipeline and the visual cortex behind MOTHER EXO. It detects, tracks, segments and relates objects, faces and expressions across camera, drone and satellite imagery in real time — YOLOv8, RF-DETR and Segment Anything, trained on real data with TorchGeo for geospatial feeds. DeepVision is what lets MOTHER see.
Capabilities
Object detection
High-throughput detection with YOLOv8 and RF-DETR across camera, drone and satellite feeds.
Multi-object tracking
Persistent tracks across frames and cameras, with re-identification and motion estimation.
Segmentation
Pixel-accurate masks with the Segment Anything Model for precise scene understanding.
Faces & expression
Face detection and expression/affect reading that feed EXO's emotion understanding.
Relations & scene graphs
Builds structured relationships between detected entities — who is near, holding or interacting with what.
Geospatial vision
TorchGeo-powered analysis of satellite and aerial imagery for mapping and change detection.
Architecture
- 1
Detector ensemble
YOLOv8 for speed, RF-DETR for accuracy and SAM for segmentation — combined into one perception pipeline.
- 2
Frozen ViT backbone
A MOTHER DeepVision Vision Transformer provides the shared visual features that MOTHER EXO reasons over — trained once, reused everywhere.
- 3
Training pipeline
Python 3.10+ / PyTorch training with real datasets and TorchGeo for geospatial feeds; heads evaluated on real frames and clips.
Specifications
| Backbone | SigLIP-SO400M ViT · 896px · frozen (feeds EXO) |
| Measured | kNN 1.0 on retrieval eval |
| Detectors | YOLOv8 · RF-DETR · Segment Anything (SAM) |
| Tasks | detect · track · segment · face · expression · relations |
| Geospatial | TorchGeo (satellite / aerial) |
| Arms | Vision (live) · T2V (placeholder) · Audio/Dialog (EXO heads) |
| Serving | Python 3.10+ · PyTorch · sovereign on-node |
Models & arms — function · dataset
| Weight | Kind | Function | Dataset | Trained |
|---|---|---|---|---|
| SigLIP-SO400M ViT | backbone | Self-supervised visual foundation (896px) — the shared features EXO reasons over | all on-node imagery · kNN 1.0 | ✓ |
| YOLOv8 | detector | Fast object detection | COCO / LVIS | ✓ |
| RF-DETR | detector | High-accuracy detection | COCO / remote-sensing | ✓ |
| SAM | segmenter | Pixel-accurate masks | Segment Anything | ✓ |
| face · expression · relations | heads | Consent face-ID · affect · scene-graph — on the backbone | consent gallery · affect · scene-graph | ✓ |
| T2V | arm | Text-to-video (latent diffusion) | tokenizer TBD | placeholder |
| Audio / Dialog | arm | ASR + text-to-voice — lives in the EXO speech/audio heads | LibriSpeech + codec | in EXO |
Model evaluation — industry benchmarks
| Benchmark | Metric | Result | Status |
|---|---|---|---|
| ImageNet kNN | top-1 | 1.0 | ◉ measured |
| COCO detection | mAP50 | — | ○ to run |
| ADE20k segmentation | mIoU | — | ○ to run |
| LVIS | mAP | — | ○ to run |
| MOT17 tracking | MOTA | — | ○ to run |
Benchmarks scheduled on this build — measured results are recorded here as each run completes.
What it's for
Surveillance & ISR
Real-time detection and tracking across large camera estates and drone feeds for situational awareness.
Geospatial intelligence
Change detection and object analysis on satellite and aerial imagery with TorchGeo.
Perception for EXO
The visual cortex that supplies MOTHER EXO with the features it needs to build a world model.
Safety & governance
- Observation and decision-support only — DeepVision informs, humans decide.
- Runs on sovereign infrastructure with no data egress.
- Evaluated on real imagery; performance is measured, never asserted.
Frequently asked
What is MOTHER DeepVision?
MOTHER DeepVision is Media Stream AI's sovereign computer-vision pipeline — real-time detection, tracking, segmentation, face and expression analysis and scene relations using YOLOv8, RF-DETR and Segment Anything, and the visual cortex behind MOTHER EXO.
What models does DeepVision use?
An ensemble of YOLOv8 (speed), RF-DETR (accuracy) and the Segment Anything Model (segmentation), plus a frozen DeepVision Vision Transformer backbone that feeds MOTHER EXO.
Can DeepVision analyse satellite imagery?
Yes — DeepVision uses TorchGeo for geospatial analysis of satellite and aerial imagery, including mapping and change detection.
How does DeepVision relate to MOTHER EXO?
DeepVision is EXO's eyes: its frozen ViT backbone provides the visual features EXO reasons over to build its world model.
The MOTHER model family
A frontier world model — not text tokens, but Eyes, Ears, a Mouth, a Brain, memory and understanding, all at once.
A sovereign British reasoning model — the mind behind MOTHER.
The sovereign general-purpose assistant — everyday chat and document making.
Describe it — MOTHER Code builds the website, app or game.
Build on MOTHER DeepVision
Sovereign, on-node and observe-and-advise by design. Talk to us about access and deployment.