v0.1.0 (experimental) · self-hosted infrastructure

One AI System.
Multiple Computers.
Distributed Intelligence.

Build a distributed AI compute cluster from the machines you already own. Connect heterogeneous computers, coordinate workloads through a centralized controller, and execute AI inference using a distributed runtime.

Self-hosted • Heterogeneous Hardware • Fault-Aware • OpenAI-Compatible API

ONLINEREADYBUSYOFFLINE

Conceptual cluster topology

ONLINE

AI Application

Project Scout / clients

request

AI API Layer

OpenAI-compatible

control

Cluster Control Plane

registry · scheduler · health

dispatch

NODE 01

RTX 3050

ONLINE

NODE 02

RTX 3060

BUSY

NODE 03

RTX 3050

READY
tokens

Distributed Inference

llama.cpp · GGUF

Overview

What Is the Distributed AI Compute Platform?

The platform transforms multiple independent computers into a coordinated AI inference cluster. Instead of treating each computer as an isolated machine, it provides the coordination layer needed to run AI workloads across them.

TRADITIONAL SETUP

PC 1 → AI
PC 2 → AI
PC 3 → AI

Each machine works independently.
DISTRIBUTED SETUP

              Distributed AI Cluster
                       │
        ┌──────────────┼──────────────┐
        ▼              ▼              ▼
      PC 1            PC 2           PC 3
        └──────────────┼──────────────┘
                       ▼
              Shared AI Workload
Node discovery
Hardware detection
Health monitoring
Resource telemetry
Workload scheduling
Distributed inference
Failure detection
Recovery and redistribution
API access
Centralized monitoring

The platform provides the coordination layer required to turn these machines into a unified compute environment.

Motivation

Why This Platform Exists

Running AI workloads on your own hardware runs into a set of recurring problems.

Expensive AI Hardware

Modern AI workloads can require significant GPU memory and compute.

Idle Hardware

Personal computers often have compute resources that remain unused.

GPU Memory Constraints

A single consumer GPU may not have enough memory for certain workloads or configurations.

Heterogeneous Machines

Different computers have different CPUs, GPUs, RAM and VRAM.

Single-Machine Failure

Depending on one machine can make an AI service unavailable when that machine fails.

Infrastructure Complexity

Connecting machines manually gives you no health monitoring, node registration, workload coordination, telemetry, recovery or centralized control.

This platform was designed to provide a software layer for coordinating those machines into a managed distributed inference environment.

Components

The Platform Has Four Core Layers

Each layer has a distinct responsibility. The controller coordinates; it does not perform inference.

Node Agent

Lightweight Go service on every compute machine: identity, enrollment, hardware reporting, heartbeats, job execution.

Component detail

Cluster Controller

Python control plane: node registry, scheduler, health monitor, job manager, model manager, admin API.

Component detail

Distributed Runtime

The execution layer that actually runs inference across participating machines via llama.cpp and GGUF models.

Component detail

AI API Layer

OpenAI-compatible API so applications consume the cluster without talking to individual nodes.

Component detail

Cluster Dashboard

Centralized visibility into cluster health, nodes, hardware, jobs, models, events and drain/resume operations. The dashboard talks only to the controller’s admin API — never directly to worker nodes.

Node Agent

Deploy a New Compute Node

Install the Node Agent on another computer and add it to the cluster.

Windows

node-agent-windows.exe

Download for Windows

Linux

node-agent-linux

Download for Linux

Source

Build from source

View Source Code
powershellinstallation example
.\install.ps1 `
  -ControllerURL "YOUR_CONTROLLER_URL" `
  -EnrollmentToken "YOUR_ENROLLMENT_TOKEN"

Platform buttons download the latest uploaded build for that platform. The enrollment token should be generated securely by the cluster administrator and never committed or published.

Benefits

Why Use a Distributed AI Cluster?

Reuse Existing Hardware

Use available computers instead of requiring a single high-end machine.

Flexible Capacity

Additional nodes can contribute additional compute resources.

Centralized Control

Manage the cluster through one control plane.

Hardware Awareness

Track CPU, GPU, RAM and VRAM characteristics of participating nodes.

Fault Awareness

Detect unhealthy or disconnected nodes.

Developer-Friendly API

Applications consume the AI service through an API rather than managing machines.

Benefits depend on workload characteristics, network performance, hardware configuration and runtime behaviour. No performance numbers are claimed on this site.

Interactive

Cluster Simulator

Add nodes, remove nodes, simulate a failure and dispatch a test job to see how cluster state and events change.

Cluster

healthy 3/3 · vram 20 GB (simulated)

node-01

4 GB
ONLINE

node-02

12 GB
ONLINE

node-03

4 GB
ONLINE

This simulator is a frontend visualisation only. It is not connected to any live cluster or production infrastructure.

event stream

22:41:57 SIM_READY frontend visualisation only

FAQ

Frequently Asked Questions

No. The architecture is designed to support heterogeneous compute nodes, subject to runtime and workload compatibility.

Build Your Own AI Compute Cluster

Connect your machines. Deploy the Node Agent. Start the controller. Build a distributed AI environment around the hardware you already have.