Introduction
System overview
The Distributed AI Compute Platform coordinates multiple computers as an inference cluster. A central controller owns identity, health, scheduling and job state; Node Agents report hardware and execute assigned work; the distributed runtime performs inference through llama.cpp.
Control plane
API validation, node registry, scheduling, health evaluation, job state and management operations.
Compute plane
Node Agents, local model files, llama.cpp processes and private-network inference traffic.
The controller coordinates execution; it does not perform inference. The dashboard communicates with the controller, never directly with worker nodes.
System design
Architecture and network boundary
Applications / OpenAI SDK
│ HTTPS REST + SSE
▼
OpenAI-Compatible API
│
▼
Cluster Controller
├─ Node registry
├─ Scheduler
├─ Health monitor
├─ Job manager
└─ Admin API
│ gRPC control channel
▼
Node Agents on private LAN
│
▼
llama.cpp primary ↔ RPC workers → GGUF modelHTTP API
Port 8000 by default
Node control
gRPC port 50051
Runtime peer
llama.cpp RPC port 50052
Runtime peer addresses are intended to stay inside private CIDR ranges. Exact firewall, TLS and deployment settings remain operator responsibilities.
Runtime
Three-phase execution lifecycle
Setup and validation
The controller selects eligible nodes and sends an execution plan. Agents validate model identity, local readiness and participating RPC workers before computation starts.
Distributed execution
The primary llama.cpp process evaluates the model with participating RPC workers over the private LAN according to the selected split.
Token streaming
Runtime output flows through the Node Agent and controller, then reaches the client as an OpenAI-compatible Server-Sent Events stream.
Pre-execution node loss can trigger re-scoring and a revised plan. In-flight recovery depends on runtime and workload state; uninterrupted generation is not guaranteed.
Operations
Node enrollment
Enrollment is controlled by a token generated and distributed by the cluster administrator.
.\node-agent.exe --foreground \ --controller-url "YOUR_CONTROLLER_URL" \ --enrollment-token "YOUR_ENROLLMENT_TOKEN" \ --node-name "YOUR_NODE_NAME"
- 1.The agent discovers local CPU, memory and supported accelerator details.
- 2.It sends an enrollment request to the controller over the configured control channel.
- 3.The controller validates the administrator-issued enrollment token and creates the node identity.
- 4.The node begins heartbeats, reports resources and becomes schedulable only after it is ready.
Development
Controller configuration
| Variable | Purpose | Example placeholder |
|---|---|---|
| DATABASE_URL | Controller state database connection | YOUR_DATABASE_URL |
| AI_API_KEY | Bearer key for client inference requests | YOUR_API_KEY |
| NODE_ENROLLMENT_SECRET | Administrator-controlled enrollment secret | YOUR_ENROLLMENT_SECRET |
| GRPC_PORT | Node control channel | 50051 |
| PORT | HTTP API listener | 8000 |
Keep credentials outside source control. The examples intentionally use placeholders rather than deployable secrets.
Operations
Health and monitoring
Agents periodically report health and resource telemetry. The controller translates missing or delayed heartbeats into node state changes and excludes unavailable capacity from new scheduling decisions.
Heartbeat current ──────────────► READY / BUSY
Heartbeat delayed ─────────────► DEGRADED
Heartbeat unavailable ─────────► OFFLINE
│
└─ excluded from new schedulingTiming thresholds are deployment configuration, not a universal guarantee. Consult the controller configuration used by your cluster.
Reference
API surface
| Method | Endpoint | Purpose |
|---|---|---|
| POST | /v1/chat/completions | Create a chat completion, optionally streamed with SSE. |
| GET | /v1/models | List models registered as available to the API. |
| GET | /health | Read controller service health. |
| GET | /api/admin/nodes | Inspect enrolled nodes and their current state. |
| POST | /api/admin/nodes/{node_id}/drain | Stop assigning new work to a node. |
| GET | /api/admin/jobs | Inspect recent and active jobs. |
Reference
Node states and event catalog
Node state machine
| State | Meaning | Scheduler eligible |
|---|---|---|
| UNINITIALIZED | Discovering local hardware and runtime dependencies. | No |
| REGISTERING | Completing the controller handshake. | No |
| READY | Healthy and available for assignment. | Yes |
| BUSY | Executing an inference job. | No |
| DRAINING | Completing current work without accepting new jobs. | No |
| DEGRADED | Heartbeat delayed or a non-critical fault was reported. | Fallback only |
| OFFLINE | No active controller connection or heartbeat. | No |
| ERROR | An unrecoverable runtime or hardware error occurred. | No |
Important events
| Event | Category | Meaning |
|---|---|---|
| NODE_ENROLLED | Lifecycle | Node authenticated and entered the registry. |
| HEARTBEAT_TIMEOUT | Health | Heartbeat exceeded the allowed window; node leaves normal scheduling. |
| JOB_DISPATCHED | Execution | An execution plan was assigned to participating nodes. |
| RPC_WORKER_STARTED | Runtime | A secondary llama.cpp RPC worker became ready on the LAN. |
| TOKEN_STREAM_CHUNK | Inference | A generated token chunk was forwarded toward the client. |
| JOB_COMPLETED | Execution | The inference job finished and its final state was recorded. |
| NODE_DRAINED | Management | A node stopped accepting new work for maintenance. |
Support
Troubleshooting
| Symptom | Likely cause | Check |
|---|---|---|
| gRPC connection refused | Controller is unreachable or the control port is blocked. | Verify controller availability, address, TLS settings and firewall rules. |
| RPC worker disconnected | The LAN connection dropped or a worker process stopped. | Check the worker address, LAN path and llama.cpp RPC process. |
| CUDA out of memory | The execution split exceeds available GPU memory. | Use a smaller model or quantization, or revise the execution plan. |
| Public IP security violation | A runtime peer address is outside the permitted private network range. | Keep compute-plane traffic on an approved private LAN. |
When diagnosing a node, capture its agent version, controller version, node state, recent events and whether failure occurs before or during execution.