PETALS

Distributed Post-Quantum LLM Inference

No GPU? No problem. Petals shards massive model layers across the Polygone network. Your machine handles a slice. The network handles the rest.

Architecture

DISTRIBUTED
PIPELINE

0–8
Relay Alice
Embedding
9–16
Relay Bob
Attention
17–24
Relay Carol
FFN
25–32
You
Decode

Each relay transfers encrypted hidden states (ML-KEM-1024) to the next. No relay sees the full model output. Computation is blind.

Usage

RUN A RELAY OR
START A SESSION

# Serve layers 0-8 on your machine
$ polygone-petals serve --layers 0-8 --listen 0.0.0.0:4003
✓ Relay participating on /ip4/0.0.0.0/tcp/4003
# Ask the distributed brain a question
$ polygone-petals chat --prompt "Are we anonymous?" --relays peer1:4003,peer2:4003
→ "You are invisible. The wave doesn't ask for permission."