LAN clustering
Serve a model no single machine can hold by splitting its layers across several machines on your local network, as one provider on Tenzro Network 1.
A homelab with a gaming PC, a laptop and a Mac mini may not have one machine big enough for a large model, but together they might. LAN clustering splits a model's layers across machines on the same local network and serves it as a single provider. To the rest of the network the cluster looks like any other endpoint; inside, the head machine runs the model as a pipeline across its members.
How it works
The model's transformer layers are divided into contiguous ranges, one range per member. A token flows from the head through each member in turn: each runs its layers and passes only the boundary activation, one small vector per token, to the next. Weights never cross the wire during serving, and quantisation is chosen once for the whole model.
This is layer-wise pipeline parallelism, chosen because it tolerates ordinary Ethernet and Wi-Fi. Schemes that split every layer across machines need data-centre interconnects and do not hold up on a home network.
Because each member is simply a device executor for its range, members can mix backends: one machine on CUDA, another on Metal, another on Vulkan or CPU, all in one pipeline.
When a cluster forms
A cluster costs decode speed, since every token crosses the network between stages, so the node forms one only when it has to:
| Decision | Meaning |
|---|---|
RunLocal | The model fits on the biggest single member; it serves there alone. |
ClusterRequired | No single member can hold it; a cluster forms. |
ClusterForced | It would fit on one machine, but you asked for a cluster. |
# Automatic: single machine if it fits, cluster if it does not
tenzro model serve qwen3.5-27b
# Split even though it fits one machine (trades speed for memory headroom)
tenzro model serve qwen3.5-27b --cluster
# Never cluster
tenzro model serve qwen3.5-27b --force-singleFinding members
Clustering needs no manual configuration. An AI node willing to join clusters advertises its backend, capability key and runtime build on its provider announcement (see Hardware discovery). When you serve a model, the head reads the model's shape from its file header, gathers candidates from those announcements and from peers on the local segment, and plans the placement. Nodes that do not advertise are never auto-clustered.
# Who is on your network and could join a cluster
tenzro cluster members
# How a downloaded model would be placed, before you commit hardware
tenzro cluster preview qwen3.5-27bThe preview shows the discovered members, the proposed split and any member that was rejected, with the reason.
Who can join
Each candidate passes two gates.
Hardware gate. A member must run exactly the same inference runtime build as the rest of the cluster, since the pipeline protocol between stages has no version negotiation, and it must have memory for at least one layer. Rejections name the reason: build mismatch, insufficient memory or not reachable.
Network gate. Every member must be directly reachable for per-token traffic: on the same local segment, or publicly dialable or hole-punched. Members reachable only through a relay or behind a symmetric NAT are excluded. The planner then orders stages as a nearest-neighbour chain over measured latency and bandwidth to keep transfer cost down.
Layer assignment
Layers are assigned in proportion to each member's memory: more memory earns more layers, every admitted member gets at least one, and remainders go to the largest members first. The plan is a pure function of the model shape and the member list, so every member computes the identical layout without a coordination round.
Dry-run a placement with tenzro_clusterPlan, which reads no node state and is open:
{
"jsonrpc": "2.0", "id": 1, "method": "tenzro_clusterPlan",
"params": {
"model": { "layers": 64, "hidden_dim": 5120, "total_vram_gb": 32.0 },
"members": [ "<member>", "<member>" ],
"user_forced": false
}
}The result reports the fit decision, whether a cluster forms, the activation bytes per token and the layer range for each member. From the CLI, tenzro cluster plan --layers <n> --hidden-dim <n> --total-vram-gb <gb> --members members.json does the same.
From plan to pipeline
Serving and planning are one decision. When the plan forms a cluster, the head opens a session to each member over the authenticated peer-to-peer overlay, loads the model across them in pipeline order and exposes the result as one logical provider. Members never expose the runtime's pipeline socket on the network, and a member accepts pipeline sessions only from the head of a cluster it has been planned into.
If a member drops, the head tears down the half-open sessions, plans again over the members still reachable and reloads. Re-planning is deterministic, so the replacement layout is the same whichever head computes it.
tenzro_serveModel accepts force_single to keep serving on one machine, and an explicit cluster_members list to replace discovery with a fixed set of machines. Serving is an operator (admin) action.
Related
- Distributed MoE: split a mixture-of-experts model by expert across the wider network.
- Model serving
- Networking