Mining for Bounty is open. Hypertrain is in preview.

HOW TO MINE
/HYPERTRAIN · CHALLENGE ID hypertrain · PREVIEW

Hypertrain. One model, many miners.

Subnet 100 miners train one shared model together with decentralized training, syncing once per round. Not live yet: everything here is a labelled preview.

RUN AT A GLANCE
Preview - simulated data
FINAL LOSS4.50Final global loss, after round 16. Lower is better.
OUTER ROUND16/ 16All 16 rounds synced. The run is finished.
CLUSTERS3finishedClusters that took part in this run. 6 nodes, all finished.
AVG THROUGHPUT9.0Ktok/sTokens per second over the run, summed over the clusters that took part.
BANDWIDTH SAVED402xvs DDPLess data crossed the network than DDP would have sent for the same rounds.
TOKENS TRAINED (FINAL)292.7MFinal total. The run reached its token target.
RUN #1Stagepreview, not liveModel20M decoderStatuscompletedOuterNesterov · lr 0.7 · μ 0.9SAMPLED 10-08
20M DECODER
Preview - simulated data
  1. Input token ids x1 to xt.
  2. Token embedding, 50,000 by 256. Positions use RoPE, applied to queries and keys inside attention.
  3. Decoder block repeated 6 times, pre-norm: RMSNorm, causal multi-head self-attention, residual add, RMSNorm, SwiGLU feed-forward, residual add.
  4. Attention: input h_t of 256 values, projections W_Q, W_K, W_V, RoPE on Q and K, 4 heads of 64 dimensions with a causal mask, concat, output projection W_O gives u_t.
  5. Feed-forward: 256 to 1,024 through gate and up projections multiplied element-wise after SiLU, then down to 256.
  6. Final RMSNorm.
  7. LM head from 256 to 50,000 scores, tied with the token embedding.
  8. Softmax gives next-token probabilities.
next-token probabilitiessoftmaxLM head 256 → 50ktied with embedding ERMSNormDecoder block × 6 (pre-norm)residual streamMLP · SwiGLU FFN256 → 1,024 → 256RMSNormCausal multi-headself-attentionRMSNormToken embedding · 50k × 256positions: RoPE (in attention)x1 x2 …xt input token ids
PARAMS
20M
LAYERS
6
D_MODEL
256
HEADS
4 × 64
CONTEXT
2,048
VOCAB
50,000
EXPAND ARCHITECTUREPreview - simulated data
Zoom: SwiGLU feed-forwardyt  ∈ ℝ⁷⁶⁸Wdown  · 1,024 → 256SiLU(Wgate h)256 → 1,024Wup h256 → 1,024ht ′ ∈ ℝ⁷⁶⁸Zoom: causal multi-head self-attentionut  ∈ ℝ⁷⁶⁸WO  · 256 × 256concat 4 × 64 = 2564 heads × dhead  = 64mask Mhead i (1 … 4)softmax(QKᵀ/√64 + M) VQKVRoPERoPEWQ WK WV ht  ∈ ℝ⁷⁶⁸
Decentralized sync: one outer roundOuter: Nesterovlr 0.7 · μ 0.9Δθ = θ − θk int8 pseudo-gradientInner: AdamW × 500local steps on every clustershared weights θnew θ
OPTIMIZERS
  • Inner: AdamW, run locally by every miner for 500 steps
  • Outer: Nesterov · lr 0.7 · momentum 0.9, once per round
  • Sync: int8 pseudo-gradients
TRAINING CURVE · GLOBAL LOSS

Is the shared model getting better?

Loss is how wrong the model is on held-out text. Every dot is one outer round; lower is better.
GLOBAL LOSS · ONE DOT PER OUTER ROUND
Preview - simulated data
FINAL LOSS4.504
SINCE ROUND 1-60%
LOWEST4.50 · r16
HIGHEST AFTER R19.34 · r2
STARTED AT11.228
SYNCED ROUNDS16 of 16
RUN FINISHED · FINAL FIGURESSimulated. Static: nothing is training.
AVG TOKENS / S
TOKENS TRAINED (FINAL)
Preview - simulated data

Line chart of global loss at the end of each outer round. Loss falls from 11.23 at round 1 to 4.50 at round 16, with small ups and downs on the way; the lowest is 4.50 at round 16. Use the left and right arrow keys to read each round.

468101212345678910111213141516OUTER ROUNDLOSSlow 4.50 · r16spike 4.82 · r14
Hover the curve or focus it and press the arrow keys to read a round.
THE GROUPS · CLUSTERS

Who trained this run?

A cluster is a group of machines that trains together inside one location. Each one does its own local steps between syncs.
CLUSTERS · MACHINES THAT TRAIN TOGETHER
Preview - simulated data
Each square is one GPU. Each block is one node.FINISHED
Emberc-01FINISHED
US EastH100 80GB
NODES · GPUS
2 · 16
AVG THROUGHPUT
4,918 tok/s
INNER STEPS IN THE LAST ROUND500 / 500
LAST SYNC
round 1609-08 10:40 UTC
TOKEN SHARE53.5%
156.6M tokens
Tundrac-02FINISHED
EU WestA100 80GB
NODES · GPUS
2 · 16
AVG THROUGHPUT
3,156 tok/s
INNER STEPS IN THE LAST ROUND500 / 500
LAST SYNC
round 1609-08 10:40 UTC
TOKEN SHARE35.4%
103.7M tokens
Deltac-03FINISHED
AP SouthRTX 4090
NODES · GPUS
2 · 8
AVG THROUGHPUT
923 tok/s
INNER STEPS IN THE LAST ROUND500 / 500
LAST SYNC
round 1609-08 10:40 UTC
TOKEN SHARE11.1%
32.4M tokens
THE ROSTER · PARTICIPANTS AND ROUNDS

Every miner, every round.

Tokens are the work a miner contributed. Share is its part of all tokens trained so far.
PARTICIPANTS · EVERY MINER IN THE RUN
Preview - simulated data
#MINERCLUSTERGPUSSTATUSROUNDSTOKENSTOK/SSHARE
1miner-015iTY…EjHqEmber8 x H100 80GBFINISHED1696.0M2,56532.8%
2miner-025i3z…hoh5Ember8 x H100 80GBFINISHED1160.6M2,35320.7%
3miner-035b7r…rGG8Tundra8 x A100 80GBFINISHED1660.3M1,61020.6%
4miner-045rg2…h9qgTundra8 x A100 80GBFINISHED1243.4M1,54614.8%
5miner-055tRa…xJ4wDelta4 x RTX 4090FINISHED1617.6M4716.0%
6miner-065pKN…NpFzDelta4 x RTX 4090FINISHED1414.8M4525.1%
RECENT ROUNDS · NEWEST FIRST
Preview - simulated data
ROUNDSTARTEDNODESCOMPUTESYNCSENTDDP WOULD SENDSAVEDLOSS
1609-08 10:00 UTC639m 14s0m 46s619 MB240 GB388x4.504
1509-08 09:20 UTC639m 01s0m 59s609 MB240 GB394x4.810
1409-08 08:40 UTC639m 14s0m 46s594 MB240 GB404x4.820
1309-08 08:00 UTC638m 54s1m 06s592 MB240 GB406x4.745
1209-08 07:20 UTC639m 14s0m 46s576 MB240 GB416x4.889
1109-08 06:40 UTC638m 58s1m 02s582 MB240 GB413x5.241
1009-08 06:00 UTC639m 15s0m 45s594 MB240 GB404x5.423
909-08 05:20 UTC638m 48s1m 12s591 MB240 GB406x5.508
809-08 04:40 UTC639m 14s0m 46s624 MB240 GB384x6.210
709-08 04:00 UTC638m 45s1m 15s622 MB240 GB386x6.191
HOW DECENTRALIZED TRAINING WORKS · FOUR STATIONS

Train a lot. Talk rarely.

Decentralized training keeps the network quiet: miners only talk once per outer round, not after every step.
01 · INNER STEPSTrain locallyEach miner runs many optimizer steps on its own GPUs, with no network traffic in between.
02 · PSEUDO-GRADIENTMeasure the driftWhen the round ends, a miner's change is how far its weights moved from the shared model.
03 · OUTER SYNCSync once a roundMiners exchange pseudo-gradients once per outer round. That is a fraction of the bandwidth DDP needs.
04 · OUTER STEPNesterov updateAn outer Nesterov optimizer applies the combined update to the one shared model. Then the next round starts.
MOTIONcyan coin hops station to station · the sync tile blinks twice while the miners exchange pseudo-gradients · the outer-step tile flashes cream when the shared model updates
INSERT COIN

Not live yet. Bounty is.

Hypertrain is a preview built on simulated data, so there is nothing to mine here yet. Bounty is open today.

OPEN /BOUNTY