Architecture

Complete System Architecture

Five clearly separated layers. Requests and control signals travel downward; telemetry, heartbeats and generated tokens travel upward. Click any component for its purpose, responsibilities, technologies, inputs, outputs and failure behaviour.

Interactive system architecture

Layer 1 — Applications

↓ requests / control↑ telemetry / tokens

Layer 2 — API

authentication · request validation · streaming

↓ requests / control↑ telemetry / tokens

Layer 3 — Control Plane (Python Cluster Controller)

↓ requests / control↑ telemetry / tokens

Layer 4 — Compute Nodes

↓ requests / control↑ telemetry / tokens

Layer 5 — Distributed Runtime

llama.cpp
GGUF model weights
Control flowHealth / telemetryInference dataManagement

Reference

Layer Reference

The same architecture as a text diagram, useful for docs and issues.

LAYER 1  APPLICATIONS
  Project Scout
  Other AI Applications
  Developer Clients
        ↓
LAYER 2  API
  OpenAI-Compatible API
  Authentication
  Request Validation
  Streaming
        ↓
LAYER 3  CONTROL PLANE
  Python Cluster Controller
  ├── Node Registry
  ├── Scheduler
  ├── Health Monitor
  ├── Job Manager
  ├── Model Manager
  ├── Admin API
  └── Event Stream
        ↓
LAYER 4  COMPUTE NODES
  Node Agent 01
  Node Agent 02
  Node Agent 03
  Node Agent N
        ↓
LAYER 5  DISTRIBUTED RUNTIME
  Distributed Inference Engine
          │
          ▼
      llama.cpp
          │
          ▼
        GGUF

FLOW
  ↓ API requests / control signals
  ↑ telemetry / heartbeats / tokens

The controller is currently deployed on a cloud platform and acts purely as the control plane. It does not perform inference — the distributed runtime on the compute nodes does.

Build Your Own AI Compute Cluster

Connect your machines. Deploy the Node Agent. Start the controller. Build a distributed AI environment around the hardware you already have.