/HYPERTRAIN · CHALLENGE ID hypertrain · PREVIEW
Hypertrain. One model, many miners.
Subnet 100 miners train one shared model together with decentralized training, syncing once per round. Not live yet: everything here is a labelled preview.
RUN #4 · 400M DECODER · SCHEDULED
Preview - simulated data
Run #4 has not started.
It is scheduled for 2026-10-20: 96 rounds, nothing trained yet. Loss, clusters and miners appear once it starts.
EXPAND ARCHITECTURECOLLAPSE ARCHITECTUREPreview - simulated data
OPTIMIZERS
- Inner: AdamW, run locally by every miner for 500 steps
- Outer: Nesterov · lr 0.7 · momentum 0.9, once per round
- Sync: int8 pseudo-gradients
HOW DECENTRALIZED TRAINING WORKS · FOUR STATIONS
Train a lot. Talk rarely.
Decentralized training keeps the network quiet: miners only talk once per outer round, not after every step.
01 · INNER STEPSTrain locallyEach miner runs many optimizer steps on its own GPUs, with no network traffic in between.
02 · PSEUDO-GRADIENTMeasure the driftWhen the round ends, a miner's change is how far its weights moved from the shared model.
03 · OUTER SYNCSync once a roundMiners exchange pseudo-gradients once per outer round. That is a fraction of the bandwidth DDP needs.
04 · OUTER STEPNesterov updateAn outer Nesterov optimizer applies the combined update to the one shared model. Then the next round starts.
MOTIONcyan coin hops station to station · the sync tile blinks twice while the miners exchange pseudo-gradients · the outer-step tile flashes cream when the shared model updates
INSERT COIN
Not live yet. Bounty is.
Hypertrain is a preview built on simulated data, so there is nothing to mine here yet. Bounty is open today.
