v0.1.0 (experimental) · self-hosted infrastructure
One AI System.
Multiple Computers.
Distributed Intelligence.
Build a distributed AI compute cluster from the machines you already own. Connect heterogeneous computers, coordinate workloads through a centralized controller, and execute AI inference using a distributed runtime.
Self-hosted • Heterogeneous Hardware • Fault-Aware • OpenAI-Compatible API
Conceptual cluster topology
ONLINEAI Application
Project Scout / clients
AI API Layer
OpenAI-compatible
Cluster Control Plane
registry · scheduler · health
NODE 01
RTX 3050
ONLINENODE 02
RTX 3060
BUSYNODE 03
RTX 3050
READYDistributed Inference
llama.cpp · GGUF
Overview
What Is the Distributed AI Compute Platform?
The platform transforms multiple independent computers into a coordinated AI inference cluster. Instead of treating each computer as an isolated machine, it provides the coordination layer needed to run AI workloads across them.
TRADITIONAL SETUP PC 1 → AI PC 2 → AI PC 3 → AI Each machine works independently.
DISTRIBUTED SETUP
Distributed AI Cluster
│
┌──────────────┼──────────────┐
▼ ▼ ▼
PC 1 PC 2 PC 3
└──────────────┼──────────────┘
▼
Shared AI WorkloadThe platform provides the coordination layer required to turn these machines into a unified compute environment.
Motivation
Why This Platform Exists
Running AI workloads on your own hardware runs into a set of recurring problems.
Expensive AI Hardware
Modern AI workloads can require significant GPU memory and compute.
Idle Hardware
Personal computers often have compute resources that remain unused.
GPU Memory Constraints
A single consumer GPU may not have enough memory for certain workloads or configurations.
Heterogeneous Machines
Different computers have different CPUs, GPUs, RAM and VRAM.
Single-Machine Failure
Depending on one machine can make an AI service unavailable when that machine fails.
Infrastructure Complexity
Connecting machines manually gives you no health monitoring, node registration, workload coordination, telemetry, recovery or centralized control.
This platform was designed to provide a software layer for coordinating those machines into a managed distributed inference environment.
Components
The Platform Has Four Core Layers
Each layer has a distinct responsibility. The controller coordinates; it does not perform inference.
Node Agent
Lightweight Go service on every compute machine: identity, enrollment, hardware reporting, heartbeats, job execution.
Component detailCluster Controller
Python control plane: node registry, scheduler, health monitor, job manager, model manager, admin API.
Component detailDistributed Runtime
The execution layer that actually runs inference across participating machines via llama.cpp and GGUF models.
Component detailAI API Layer
OpenAI-compatible API so applications consume the cluster without talking to individual nodes.
Component detailCluster Dashboard
Centralized visibility into cluster health, nodes, hardware, jobs, models, events and drain/resume operations. The dashboard talks only to the controller’s admin API — never directly to worker nodes.
Node Agent
Deploy a New Compute Node
Install the Node Agent on another computer and add it to the cluster.
.\install.ps1 ` -ControllerURL "YOUR_CONTROLLER_URL" ` -EnrollmentToken "YOUR_ENROLLMENT_TOKEN"
Platform buttons download the latest uploaded build for that platform. The enrollment token should be generated securely by the cluster administrator and never committed or published.
Benefits
Why Use a Distributed AI Cluster?
Reuse Existing Hardware
Use available computers instead of requiring a single high-end machine.
Flexible Capacity
Additional nodes can contribute additional compute resources.
Centralized Control
Manage the cluster through one control plane.
Hardware Awareness
Track CPU, GPU, RAM and VRAM characteristics of participating nodes.
Fault Awareness
Detect unhealthy or disconnected nodes.
Developer-Friendly API
Applications consume the AI service through an API rather than managing machines.
Benefits depend on workload characteristics, network performance, hardware configuration and runtime behaviour. No performance numbers are claimed on this site.
Interactive
Cluster Simulator
Add nodes, remove nodes, simulate a failure and dispatch a test job to see how cluster state and events change.
Cluster
healthy 3/3 · vram 20 GB (simulated)node-01
4 GBnode-02
12 GBnode-03
4 GBThis simulator is a frontend visualisation only. It is not connected to any live cluster or production infrastructure.
event stream
FAQ
Frequently Asked Questions
No. The architecture is designed to support heterogeneous compute nodes, subject to runtime and workload compatibility.
Build Your Own AI Compute Cluster
Connect your machines. Deploy the Node Agent. Start the controller. Build a distributed AI environment around the hardware you already have.