Documentation

Technical Reference

Understand the system boundary first, then follow a request from enrollment through distributed execution, monitoring and recovery.

Introduction

System overview

The Distributed AI Compute Platform coordinates multiple computers as an inference cluster. A central controller owns identity, health, scheduling and job state; Node Agents report hardware and execute assigned work; the distributed runtime performs inference through llama.cpp.

Control plane

API validation, node registry, scheduling, health evaluation, job state and management operations.

Compute plane

Node Agents, local model files, llama.cpp processes and private-network inference traffic.

The controller coordinates execution; it does not perform inference. The dashboard communicates with the controller, never directly with worker nodes.

System design

Architecture and network boundary

Applications / OpenAI SDK
          │ HTTPS REST + SSE
          ▼
OpenAI-Compatible API
          │
          ▼
Cluster Controller
  ├─ Node registry
  ├─ Scheduler
  ├─ Health monitor
  ├─ Job manager
  └─ Admin API
          │ gRPC control channel
          ▼
Node Agents on private LAN
          │
          ▼
llama.cpp primary ↔ RPC workers → GGUF model

HTTP API

Port 8000 by default

Node control

gRPC port 50051

Runtime peer

llama.cpp RPC port 50052

Runtime peer addresses are intended to stay inside private CIDR ranges. Exact firewall, TLS and deployment settings remain operator responsibilities.

Runtime

Three-phase execution lifecycle

01

Setup and validation

The controller selects eligible nodes and sends an execution plan. Agents validate model identity, local readiness and participating RPC workers before computation starts.

02

Distributed execution

The primary llama.cpp process evaluates the model with participating RPC workers over the private LAN according to the selected split.

03

Token streaming

Runtime output flows through the Node Agent and controller, then reaches the client as an OpenAI-compatible Server-Sent Events stream.

Pre-execution node loss can trigger re-scoring and a revised plan. In-flight recovery depends on runtime and workload state; uninterrupted generation is not guaranteed.

Operations

Node enrollment

Enrollment is controlled by a token generated and distributed by the cluster administrator.

powershellplaceholder values
.\node-agent.exe --foreground \
  --controller-url "YOUR_CONTROLLER_URL" \
  --enrollment-token "YOUR_ENROLLMENT_TOKEN" \
  --node-name "YOUR_NODE_NAME"
  1. 1.The agent discovers local CPU, memory and supported accelerator details.
  2. 2.It sends an enrollment request to the controller over the configured control channel.
  3. 3.The controller validates the administrator-issued enrollment token and creates the node identity.
  4. 4.The node begins heartbeats, reports resources and becomes schedulable only after it is ready.

Development

Controller configuration

VariablePurposeExample placeholder
DATABASE_URLController state database connectionYOUR_DATABASE_URL
AI_API_KEYBearer key for client inference requestsYOUR_API_KEY
NODE_ENROLLMENT_SECRETAdministrator-controlled enrollment secretYOUR_ENROLLMENT_SECRET
GRPC_PORTNode control channel50051
PORTHTTP API listener8000

Keep credentials outside source control. The examples intentionally use placeholders rather than deployable secrets.

Operations

Health and monitoring

Agents periodically report health and resource telemetry. The controller translates missing or delayed heartbeats into node state changes and excludes unavailable capacity from new scheduling decisions.

Heartbeat current ──────────────► READY / BUSY
Heartbeat delayed ─────────────► DEGRADED
Heartbeat unavailable ─────────► OFFLINE
                                  │
                                  └─ excluded from new scheduling

Timing thresholds are deployment configuration, not a universal guarantee. Consult the controller configuration used by your cluster.

Reference

API surface

MethodEndpointPurpose
POST/v1/chat/completionsCreate a chat completion, optionally streamed with SSE.
GET/v1/modelsList models registered as available to the API.
GET/healthRead controller service health.
GET/api/admin/nodesInspect enrolled nodes and their current state.
POST/api/admin/nodes/{node_id}/drainStop assigning new work to a node.
GET/api/admin/jobsInspect recent and active jobs.
IMPLEMENTEDCore inference API
PLANNEDSome administrative controls may vary by build.

Reference

Node states and event catalog

Node state machine

StateMeaningScheduler eligible
UNINITIALIZEDDiscovering local hardware and runtime dependencies.No
REGISTERINGCompleting the controller handshake.No
READYHealthy and available for assignment.Yes
BUSYExecuting an inference job.No
DRAININGCompleting current work without accepting new jobs.No
DEGRADEDHeartbeat delayed or a non-critical fault was reported.Fallback only
OFFLINENo active controller connection or heartbeat.No
ERRORAn unrecoverable runtime or hardware error occurred.No

Important events

EventCategoryMeaning
NODE_ENROLLEDLifecycleNode authenticated and entered the registry.
HEARTBEAT_TIMEOUTHealthHeartbeat exceeded the allowed window; node leaves normal scheduling.
JOB_DISPATCHEDExecutionAn execution plan was assigned to participating nodes.
RPC_WORKER_STARTEDRuntimeA secondary llama.cpp RPC worker became ready on the LAN.
TOKEN_STREAM_CHUNKInferenceA generated token chunk was forwarded toward the client.
JOB_COMPLETEDExecutionThe inference job finished and its final state was recorded.
NODE_DRAINEDManagementA node stopped accepting new work for maintenance.

Support

Troubleshooting

SymptomLikely causeCheck
gRPC connection refusedController is unreachable or the control port is blocked.Verify controller availability, address, TLS settings and firewall rules.
RPC worker disconnectedThe LAN connection dropped or a worker process stopped.Check the worker address, LAN path and llama.cpp RPC process.
CUDA out of memoryThe execution split exceeds available GPU memory.Use a smaller model or quantization, or revise the execution plan.
Public IP security violationA runtime peer address is outside the permitted private network range.Keep compute-plane traffic on an approved private LAN.

When diagnosing a node, capture its agent version, controller version, node state, recent events and whether failure occurs before or during execution.

Build Your Own AI Compute Cluster

Connect your machines. Deploy the Node Agent. Start the controller. Build a distributed AI environment around the hardware you already have.