Mining for Bounty is open. Hypertrain is in preview.

HOW TO MINE
/HYPERTRAIN · CHALLENGE ID hypertrain · PREVIEW

Hypertrain. One model, many miners.

Subnet 100 miners train one shared model together with decentralized training, syncing once per round. Not live yet: everything here is a labelled preview.

RUN AT A GLANCE
Preview - simulated data
CURRENT LOSS3.86Global loss after round 30, the last sync. Lower is better.
OUTER ROUND31/ 48One round is many local steps, then one sync. This round is still running.
ACTIVE CLUSTERS5/ 6Clusters training or syncing now. 15 of 18 nodes are up.
TOTAL THROUGHPUT20.1Ktok/sTokens per second, summed over the active clusters.
BANDWIDTH SAVED400xvs DDPLess data crosses the network than DDP would send for the same rounds.
TOKENS TRAINED1.4B/ 3.0B48% of the run's token target.
RUN #3Stagepreview, not liveModel150M decoderStatusrunningOuterNesterov · lr 0.7 · μ 0.9SAMPLED 10-08
150M DECODER
Preview - simulated data
  1. Input token ids x1 to xt.
  2. Token embedding, 50,000 by 768. Positions use RoPE, applied to queries and keys inside attention.
  3. Decoder block repeated 12 times, pre-norm: RMSNorm, causal multi-head self-attention, residual add, RMSNorm, SwiGLU feed-forward, residual add.
  4. Attention: input h_t of 768 values, projections W_Q, W_K, W_V, RoPE on Q and K, 12 heads of 64 dimensions with a causal mask, concat, output projection W_O gives u_t.
  5. Feed-forward: 768 to 3,072 through gate and up projections multiplied element-wise after SiLU, then down to 768.
  6. Final RMSNorm.
  7. LM head from 768 to 50,000 scores, tied with the token embedding.
  8. Softmax gives next-token probabilities.
next-token probabilitiessoftmaxLM head 768 → 50ktied with embedding ERMSNormDecoder block × 12 (pre-norm)residual streamMLP · SwiGLU FFN768 → 3,072 → 768RMSNormCausal multi-headself-attentionRMSNormToken embedding · 50k × 768positions: RoPE (in attention)x1 x2 …xt input token ids
PARAMS
150M
LAYERS
12
D_MODEL
768
HEADS
12 × 64
CONTEXT
2,048
VOCAB
50,000
EXPAND ARCHITECTUREPreview - simulated data
Zoom: SwiGLU feed-forwardyt  ∈ ℝ⁷⁶⁸Wdown  · 3,072 → 768SiLU(Wgate h)768 → 3,072Wup h768 → 3,072ht ′ ∈ ℝ⁷⁶⁸Zoom: causal multi-head self-attentionut  ∈ ℝ⁷⁶⁸WO  · 768 × 768concat 12 × 64 = 76812 heads × dhead  = 64mask Mhead i (1 … 12)softmax(QKᵀ/√64 + M) VQKVRoPERoPEWQ WK WV ht  ∈ ℝ⁷⁶⁸
Decentralized sync: one outer roundOuter: Nesterovlr 0.7 · μ 0.9Δθ = θ − θk int8 pseudo-gradientInner: AdamW × 500local steps on every clustershared weights θnew θ
OPTIMIZERS
  • Inner: AdamW, run locally by every miner for 500 steps
  • Outer: Nesterov · lr 0.7 · momentum 0.9, once per round
  • Sync: int8 pseudo-gradients
TRAINING CURVE · GLOBAL LOSS

Is the shared model getting better?

Loss is how wrong the model is on held-out text. Every dot is one outer round; lower is better.
GLOBAL LOSS · ONE DOT PER OUTER ROUND
Preview - simulated data
CURRENT LOSS3.861
SINCE ROUND 1-65%
LOWEST3.72 · r26
HIGHEST AFTER R110.10 · r2
STARTED AT10.904
SYNCED ROUNDS30 of 48
TRAINING NOW · ESTIMATESimulated. Not measured from real machines.
TOKENS / S (EST.)
TOKENS TRAINED (EST.)
Preview - simulated data

Line chart of global loss at the end of each outer round. Loss falls from 10.90 at round 1 to 3.86 at round 30, with small ups and downs on the way; the lowest is 3.72 at round 26. Use the left and right arrow keys to read each round.

24681012135791113151719212325272931OUTER ROUNDLOSSlow 3.72 · r26spike 4.05 · r28
Hover the curve or focus it and press the arrow keys to read a round. Round 31 is still running, so it has no point yet (the dashed line).
THE GROUPS · CLUSTERS

Who is training, and how busy are they?

A cluster is a group of machines that trains together inside one location. Each one does its own local steps between syncs.
CLUSTERS · MACHINES THAT TRAIN TOGETHER
Preview - simulated data
Each square is one GPU. Each block is one node.TRAININGSYNCINGOFFLINE
Emberc-01TRAINING
US EastH100 80GB
NODES · GPUS
4 · 32
THROUGHPUT
6,965 tok/s
INNER STEPS THIS ROUND127 / 500
LAST SYNC
round 3010-01 20:00 UTC
TOKEN SHARE40.7%
587.8M tokens
Tundrac-02TRAINING
EU WestA100 80GB
NODES · GPUS
4 · 32
THROUGHPUT
6,798 tok/s
INNER STEPS THIS ROUND258 / 500
LAST SYNC
round 3010-01 20:00 UTC
TOKEN SHARE27.5%
397.0M tokens
Deltac-03SYNCING
AP SouthRTX 4090
NODES · GPUS
3 · 12
THROUGHPUT
1,444 tok/s
INNER STEPS THIS ROUND500 / 500
LAST SYNC
round 3010-01 20:00 UTC
TOKEN SHARE6.5%
94.6M tokens
Prismc-04TRAINING
US WestL40S
NODES · GPUS
3 · 12
THROUGHPUT
1,805 tok/s
INNER STEPS THIS ROUND458 / 500
LAST SYNC
round 3010-01 20:00 UTC
TOKEN SHARE8.3%
119.9M tokens
Quartzc-05TRAINING
EU CentralA100 80GB
NODES · GPUS
2 · 16
THROUGHPUT
3,176 tok/s
INNER STEPS THIS ROUND176 / 500
LAST SYNC
round 3010-01 20:00 UTC
TOKEN SHARE15.4%
221.9M tokens
Zephyrc-06OFFLINE
AP NortheastRTX 4090
NODES · GPUS
2 · 4
THROUGHPUT
—
INNER STEPS THIS ROUND0 / 500
LAST SYNC
round 2210-01 14:40 UTC
TOKEN SHARE1.6%
23.6M tokens
THE ROSTER · PARTICIPANTS AND ROUNDS

Every miner, every round.

Tokens are the work a miner contributed. Share is its part of all tokens trained so far.
PARTICIPANTS · EVERY MINER IN THE RUN
Preview - simulated data
#MINERCLUSTERGPUSSTATUSROUNDSTOKENSTOK/SSHARE
1miner-015d5M…xnwWEmber8 x H100 80GBTRAINING30172.6M2,45911.9%
2miner-045u6D…euMsEmber8 x H100 80GBOFFLINE29152.0M—10.5%
3miner-025jFf…1t13Ember8 x H100 80GBTRAINING27139.6M2,2109.7%
4miner-035Trb…jWNnEmber8 x H100 80GBTRAINING23123.6M2,2968.6%
5miner-055nAd…T6pfTundra8 x A100 80GBTRAINING30117.6M1,6758.1%
6miner-155WiT…Sjc9Quartz8 x A100 80GBTRAINING30114.8M1,6357.9%
7miner-165Vde…bSxWQuartz8 x A100 80GBTRAINING28107.1M1,6357.4%
8miner-065Xuj…XVdXTundra8 x A100 80GBTRAINING2699.4M1,6336.9%
9miner-085MaF…YwYkTundra8 x A100 80GBTRAINING2494.1M1,6756.5%
10miner-075otV…EuZdTundra8 x A100 80GBTRAINING2286.0M1,6706.0%
11miner-1258cd…wCaHPrism4 x L40STRAINING3043.0M6123.0%
12miner-145njH…PXF7Prism4 x L40STRAINING2739.4M6242.7%
13miner-135imq…9KMbPrism4 x L40STRAINING2837.5M5732.6%
14miner-105AcB…nACADelta4 x RTX 4090SYNCING2933.3M4902.3%
15miner-0956kh…VGMiDelta4 x RTX 4090SYNCING3033.2M4732.3%
16miner-115s4v…HDJcDelta4 x RTX 4090SYNCING2528.1M4811.9%
17miner-175SfF…xcKsZephyr2 x RTX 4090OFFLINE2212.9M—0.9%
18miner-185d1T…iwokZephyr2 x RTX 4090OFFLINE1810.7M—0.7%
RECENT ROUNDS · NEWEST FIRST
Preview - simulated data
ROUNDSTARTEDNODESCOMPUTESYNCSENTDDP WOULD SENDSAVEDLOSS
3110-01 20:00 UTCRUNNING1638m 56s1m 04s————
3010-01 19:20 UTC1639m 11s0m 49s12.2 GB4.8 TB395x3.861
2910-01 18:40 UTC1638m 46s1m 14s12.2 GB4.8 TB395x4.133
2810-01 18:00 UTC1639m 14s0m 46s11.4 GB4.8 TB420x4.046
2710-01 17:20 UTC1638m 52s1m 08s11.7 GB4.8 TB409x3.743
2610-01 16:40 UTC1639m 15s0m 45s11.8 GB4.8 TB406x3.716
2510-01 16:00 UTC1639m 04s0m 56s11.5 GB4.8 TB416x3.773
2410-01 15:20 UTC1639m 00s1m 00s12.2 GB4.8 TB393x3.804
2310-01 14:40 UTC1639m 06s0m 54s12.5 GB4.8 TB385x3.978
2210-01 14:00 UTC1838m 50s1m 10s14.1 GB5.4 TB382x4.078
HOW DECENTRALIZED TRAINING WORKS · FOUR STATIONS

Train a lot. Talk rarely.

Decentralized training keeps the network quiet: miners only talk once per outer round, not after every step.
01 · INNER STEPSTrain locallyEach miner runs many optimizer steps on its own GPUs, with no network traffic in between.
02 · PSEUDO-GRADIENTMeasure the driftWhen the round ends, a miner's change is how far its weights moved from the shared model.
03 · OUTER SYNCSync once a roundMiners exchange pseudo-gradients once per outer round. That is a fraction of the bandwidth DDP needs.
04 · OUTER STEPNesterov updateAn outer Nesterov optimizer applies the combined update to the one shared model. Then the next round starts.
MOTIONcyan coin hops station to station · the sync tile blinks twice while the miners exchange pseudo-gradients · the outer-step tile flashes cream when the shared model updates
INSERT COIN

Not live yet. Bounty is.

Hypertrain is a preview built on simulated data, so there is nothing to mine here yet. Bounty is open today.

OPEN /BOUNTY