Features

Capabilities, Observability and Roadmap

What the platform does today, what it lets you see while it runs, and what is planned but not yet built.

Benefits

Why Use a Distributed AI Cluster?

Reuse Existing Hardware

Use available computers instead of requiring a single high-end machine.

Flexible Capacity

Additional nodes can contribute additional compute resources.

Centralized Control

Manage the cluster through one control plane.

Hardware Awareness

Track CPU, GPU, RAM and VRAM characteristics of participating nodes.

Fault Awareness

Detect unhealthy or disconnected nodes.

Developer-Friendly API

Applications can consume the AI service through an API rather than managing individual machines.

Benefits depend on workload characteristics, network performance, hardware configuration and runtime behaviour.

Use Cases

What Can It Be Used For?

Personal AI Infrastructure

Run AI workloads across your own computers.

AI Development

Create a private inference environment for experimentation.

LLM Applications

Provide an API backend for AI applications.

Distributed AI Research

Experiment with distributed inference and heterogeneous compute.

Local Compute Pools

Coordinate idle machines within a trusted network.

Developer Labs

Build and test distributed AI infrastructure concepts.

Security

Security & Trust Model

Node authentication
Enrollment tokens
Secure communication
API authentication
Controller-managed access
No direct dashboard-to-node communication
Node identity
Controlled enrollment
Graceful node draining
Dashboard
   │
   │ Authenticated API
   ▼
Controller
   │
   │ Authenticated/Secure Node Protocol
   ▼
Node Agents

All tokens, keys and URLs shown across this website are placeholders. Real secrets must never be committed or published.

Observability

Know What Your Cluster Is Doing

Node health

ONLINEOFFLINEDRAININGBUSYREADY

Hardware

CPU / RAM / GPU / VRAM

Jobs

QUEUED / RUNNING / COMPLETED / FAILED / CANCELLED

Events

Node joined · Node disconnected · Job started · Job completed · Node drained · Recovery triggered

event stream

22:41:42 NODE_JOINED node-01
22:41:45 NODE_JOINED node-02
22:41:48 MODEL_READY node-01
22:41:51 JOB_CREATED job-92f3
22:41:54 JOB_DISPATCHED job-92f3
22:41:57 TOKENS_STREAM job-92f3

Reference

Node States

StateMeaning
ONLINENode connected and healthy
READYNode available for work
BUSYNode currently executing work
DRAININGNode finishing work and accepting no new work
OFFLINENode disconnected
UNHEALTHYNode failed health checks

Stack

Technology Stack

Node Layer

GoNode Agent

Controller

PythonFastAPIPostgreSQLRESTgRPC / secure node communication

Inference

Distributed inference runtimellama.cppGGUF models

Frontend

ReactTypeScriptTailwind CSS

Deployment

Cloud-hosted controllerWindows / Linux compute nodes

Benchmarks

Benchmark Your Own Cluster

No benchmark numbers are published here. These metric cards are placeholders that can consume real measurements from your own cluster.

Time to First Token

-- ms

Total Latency

-- ms

Tokens / Second

-- tok/s

Active Nodes

--

Total VRAM

-- GB

throughput chart · awaiting data

Benchmark results depend on hardware, network, model, quantization, workload and cluster size.

Roadmap

Roadmap

IMPLEMENTED

Completed

  • ✓ Node Agent
  • ✓ Hardware detection
  • ✓ Node registration
  • ✓ Heartbeats
  • ✓ Telemetry
  • ✓ Cluster Controller
  • ✓ Scheduling
  • ✓ Distributed inference integration
  • ✓ API layer
  • ✓ Dashboard integration
  • ✓ Fault / recovery testing
PLANNED

Future (planned, not completed)

  • ○ More operating systems
  • ○ Improved scheduling
  • ○ More inference runtimes
  • ○ Advanced observability
  • ○ Automated model distribution
  • ○ Better network optimization
  • ○ Production-grade authentication
  • ○ Expanded benchmarking
  • ○ Additional model formats

Build Your Own AI Compute Cluster

Connect your machines. Deploy the Node Agent. Start the controller. Build a distributed AI environment around the hardware you already have.