Computer vision platform
Real Vision
Cameras that understand what they see. Real-time detection, multi-camera identity tracking, licence-plate recognition and vehicle damage baselining — narrated in language, running entirely on the client's own hardware.
In production · on-premises and cloud GPU · 2026
~347
detection classes
4
model tiers, one GPU
7
safety event kinds
0
footage leaves site
01 The problem
- CCTV records everything and tells you nothing. Finding an event means a person scrubbing footage.
- Damage disputes come down to one person's word against another's, with no evidence trail.
- Off-the-shelf detection models recognise a fixed list of objects, which never matches what a specific site actually cares about.
02 The approach
- Open-vocabulary detection rather than a fixed class list, so the system can be pointed at what a site actually cares about without retraining.
- A tiered model stack on one GPU: tracking every frame, plate OCR and re-identification only when triggered, a fast vision-language model for live narration, and a heavy one on demand for deep analysis.
- Everything runs on the client's own machine. No footage leaves the premises.
03 Architecture
Tiered inference — one GPU, four tiers
Tier 0 every frame YOLOWorld + ByteTrack ~30–50 ms Tier 1 triggered plate OCR · re-ID · Florence-2 on detection Tier 2 gated live VLM narration on scene change Tier 3 on demand heavy VLM deep-look · summary operator request
04 What it does
Open-vocabulary detection
YOLOWorld in hybrid mode unions a general class list with a domain-specific one — roughly 347 classes, deduplicated and order-stable, with domain classes added first so they can never be dropped by the cap.
Multi-camera tracking & re-ID
ByteTrack lanes per camera with cross-view identity fusion, plus deep OSNet appearance embeddings that fall back to histogram matching when the GPU backend is unavailable — so it degrades instead of breaking.
Plate recognition
ONNX-based ANPR with region-aware cleanup that conservatively corrects O/0 and I/1 confusions against known regional plate layouts, and tags each read with its region.
Vehicle condition baselining
Entry damage baseline per vehicle, canonicalised diffing on exit, and evidence photos captured automatically — so a dispute has a record.
Safety events
A seven-kind alarm pipeline with arbitration between the detector and the vision-language model, plus owner notification over Telegram, webhook or email with the evidence photo attached.
Learns without forgetting
A self-correcting label memory that survives restarts, and a continuous training loop behind a quality gate so corrections improve the model instead of drifting it.
Built with
- Python 3.12
- Flask
- YOLOWorld
- ByteTrack
- OpenCV
- Ollama VLM
- Florence-2
- OSNet re-ID
- fast-alpr
- SQLite WAL
- CUDA
- Docker
Deployed for automotive service businesses in the Gulf. Client details are confidential and are not published here.