Architecture
Complete System Architecture
Five clearly separated layers. Requests and control signals travel downward; telemetry, heartbeats and generated tokens travel upward. Click any component for its purpose, responsibilities, technologies, inputs, outputs and failure behaviour.
Interactive system architecture
Layer 1 — Applications
Layer 2 — API
authentication · request validation · streaming
Layer 3 — Control Plane (Python Cluster Controller)
Layer 4 — Compute Nodes
Layer 5 — Distributed Runtime
Reference
Layer Reference
The same architecture as a text diagram, useful for docs and issues.
LAYER 1 APPLICATIONS
Project Scout
Other AI Applications
Developer Clients
↓
LAYER 2 API
OpenAI-Compatible API
Authentication
Request Validation
Streaming
↓
LAYER 3 CONTROL PLANE
Python Cluster Controller
├── Node Registry
├── Scheduler
├── Health Monitor
├── Job Manager
├── Model Manager
├── Admin API
└── Event Stream ↓
LAYER 4 COMPUTE NODES
Node Agent 01
Node Agent 02
Node Agent 03
Node Agent N
↓
LAYER 5 DISTRIBUTED RUNTIME
Distributed Inference Engine
│
▼
llama.cpp
│
▼
GGUF
FLOW
↓ API requests / control signals
↑ telemetry / heartbeats / tokensThe controller is currently deployed on a cloud platform and acts purely as the control plane. It does not perform inference — the distributed runtime on the compute nodes does.
Build Your Own AI Compute Cluster
Connect your machines. Deploy the Node Agent. Start the controller. Build a distributed AI environment around the hardware you already have.