Mining for Bounty is open. Hypertrain is in preview.

HOW TO MINE
/HYPERTRAIN · CHALLENGE ID hypertrain · PREVIEW

Hypertrain. One model, many miners.

Subnet 100 miners train one shared model together with decentralized training, syncing once per round. Not live yet: everything here is a labelled preview.

RUN AT A GLANCE
Preview - simulated data
FINAL LOSS4.05Final global loss, after round 24. Lower is better.
OUTER ROUND24/ 24All 24 rounds synced. The run is finished.
CLUSTERS4finishedClusters that took part in this run. 10 nodes, all finished.
AVG THROUGHPUT16.9Ktok/sTokens per second over the run, summed over the clusters that took part.
BANDWIDTH SAVED401xvs DDPLess data crossed the network than DDP would have sent for the same rounds.
TOKENS TRAINED (FINAL)811.9MFinal total. The run reached its token target.
RUN #2Stagepreview, not liveModel50M decoderStatuscompletedOuterNesterov · lr 0.7 · μ 0.9SAMPLED 10-08
50M DECODER
Preview - simulated data
  1. Input token ids x1 to xt.
  2. Token embedding, 50,000 by 512. Positions use RoPE, applied to queries and keys inside attention.
  3. Decoder block repeated 8 times, pre-norm: RMSNorm, causal multi-head self-attention, residual add, RMSNorm, SwiGLU feed-forward, residual add.
  4. Attention: input h_t of 512 values, projections W_Q, W_K, W_V, RoPE on Q and K, 8 heads of 64 dimensions with a causal mask, concat, output projection W_O gives u_t.
  5. Feed-forward: 512 to 1,536 through gate and up projections multiplied element-wise after SiLU, then down to 512.
  6. Final RMSNorm.
  7. LM head from 512 to 50,000 scores, tied with the token embedding.
  8. Softmax gives next-token probabilities.
next-token probabilitiessoftmaxLM head 512 → 50ktied with embedding ERMSNormDecoder block × 8 (pre-norm)residual streamMLP · SwiGLU FFN512 → 1,536 → 512RMSNormCausal multi-headself-attentionRMSNormToken embedding · 50k × 512positions: RoPE (in attention)x1 x2 …xt input token ids
PARAMS
50M
LAYERS
8
D_MODEL
512
HEADS
8 × 64
CONTEXT
2,048
VOCAB
50,000
EXPAND ARCHITECTUREPreview - simulated data
Zoom: SwiGLU feed-forwardyt  ∈ ℝ⁷⁶⁸Wdown  · 1,536 → 512SiLU(Wgate h)512 → 1,536Wup h512 → 1,536ht ′ ∈ ℝ⁷⁶⁸Zoom: causal multi-head self-attentionut  ∈ ℝ⁷⁶⁸WO  · 512 × 512concat 8 × 64 = 5128 heads × dhead  = 64mask Mhead i (1 … 8)softmax(QKᵀ/√64 + M) VQKVRoPERoPEWQ WK WV ht  ∈ ℝ⁷⁶⁸
Decentralized sync: one outer roundOuter: Nesterovlr 0.7 · μ 0.9Δθ = θ − θk int8 pseudo-gradientInner: AdamW × 500local steps on every clustershared weights θnew θ
OPTIMIZERS
  • Inner: AdamW, run locally by every miner for 500 steps
  • Outer: Nesterov · lr 0.7 · momentum 0.9, once per round
  • Sync: int8 pseudo-gradients
TRAINING CURVE · GLOBAL LOSS

Is the shared model getting better?

Loss is how wrong the model is on held-out text. Every dot is one outer round; lower is better.
GLOBAL LOSS · ONE DOT PER OUTER ROUND
Preview - simulated data
FINAL LOSS4.055
SINCE ROUND 1-63%
LOWEST4.05 · r24
HIGHEST AFTER R110.04 · r2
STARTED AT11.083
SYNCED ROUNDS24 of 24
RUN FINISHED · FINAL FIGURESSimulated. Static: nothing is training.
AVG TOKENS / S
TOKENS TRAINED (FINAL)
Preview - simulated data

Line chart of global loss at the end of each outer round. Loss falls from 11.08 at round 1 to 4.05 at round 24, with small ups and downs on the way; the lowest is 4.05 at round 24. Use the left and right arrow keys to read each round.

4681012135791113151719212324OUTER ROUNDLOSSlow 4.05 · r24spike 6.19 · r10
Hover the curve or focus it and press the arrow keys to read a round.
THE GROUPS · CLUSTERS

Who trained this run?

A cluster is a group of machines that trains together inside one location. Each one does its own local steps between syncs.
CLUSTERS · MACHINES THAT TRAIN TOGETHER
Preview - simulated data
Each square is one GPU. Each block is one node.FINISHED
Emberc-01FINISHED
US EastH100 80GB
NODES · GPUS
3 · 24
AVG THROUGHPUT
7,543 tok/s
INNER STEPS IN THE LAST ROUND500 / 500
LAST SYNC
round 2409-18 16:00 UTC
TOKEN SHARE43.4%
352.0M tokens
Tundrac-02FINISHED
EU WestA100 80GB
NODES · GPUS
3 · 24
AVG THROUGHPUT
4,869 tok/s
INNER STEPS IN THE LAST ROUND500 / 500
LAST SYNC
round 2409-18 16:00 UTC
TOKEN SHARE27.1%
220.0M tokens
Prismc-03FINISHED
US WestL40S
NODES · GPUS
2 · 8
AVG THROUGHPUT
1,223 tok/s
INNER STEPS IN THE LAST ROUND500 / 500
LAST SYNC
round 2409-18 16:00 UTC
TOKEN SHARE7.7%
62.6M tokens
Quartzc-04FINISHED
EU CentralA100 80GB
NODES · GPUS
2 · 16
AVG THROUGHPUT
3,223 tok/s
INNER STEPS IN THE LAST ROUND500 / 500
LAST SYNC
round 2409-18 16:00 UTC
TOKEN SHARE21.8%
177.2M tokens
THE ROSTER · PARTICIPANTS AND ROUNDS

Every miner, every round.

Tokens are the work a miner contributed. Share is its part of all tokens trained so far.
PARTICIPANTS · EVERY MINER IN THE RUN
Preview - simulated data
#MINERCLUSTERGPUSSTATUSROUNDSTOKENSTOK/SSHARE
1miner-0151xG…picnEmber8 x H100 80GBFINISHED24137.4M2,44716.9%
2miner-025HQ8…h3zzEmber8 x H100 80GBFINISHED20118.9M2,54114.6%
3miner-035U7N…QmWyEmber8 x H100 80GBFINISHED1695.7M2,55511.8%
4miner-045zyp…9iDgTundra8 x A100 80GBFINISHED2491.4M1,62711.3%
5miner-095aGa…W2dfQuartz8 x A100 80GBFINISHED2490.5M1,61111.1%
6miner-105FwP…Rc4hQuartz8 x A100 80GBFINISHED2386.8M1,61210.7%
7miner-055oYB…ypsUTundra8 x A100 80GBFINISHED1865.5M1,5548.1%
8miner-065CKn…A4MXTundra8 x A100 80GBFINISHED1663.2M1,6887.8%
9miner-075HMz…v7UmPrism4 x L40SFINISHED2432.4M5774.0%
10miner-0853uV…4nTmPrism4 x L40SFINISHED2030.2M6463.7%
RECENT ROUNDS · NEWEST FIRST
Preview - simulated data
ROUNDSTARTEDNODESCOMPUTESYNCSENTDDP WOULD SENDSAVEDLOSS
2409-18 15:20 UTC1038m 57s1m 03s2.5 GB1.0 TB393x4.055
2309-18 14:40 UTC1038m 58s1m 02s2.6 GB1.0 TB390x4.139
2209-18 14:00 UTC1039m 07s0m 53s2.4 GB1.0 TB412x4.093
2109-18 13:20 UTC1039m 03s0m 57s2.4 GB1.0 TB419x4.270
2009-18 12:40 UTC1039m 14s0m 46s2.5 GB1.0 TB397x4.499
1909-18 12:00 UTC1039m 01s0m 59s2.5 GB1.0 TB399x4.202
1809-18 11:20 UTC1038m 52s1m 08s2.5 GB1.0 TB400x4.268
1709-18 10:40 UTC1038m 53s1m 07s2.6 GB1.0 TB390x4.686
1609-18 10:00 UTC1038m 52s1m 08s2.5 GB1.0 TB395x4.718
1509-18 09:20 UTC1039m 09s0m 51s2.5 GB1.0 TB400x4.943
HOW DECENTRALIZED TRAINING WORKS · FOUR STATIONS

Train a lot. Talk rarely.

Decentralized training keeps the network quiet: miners only talk once per outer round, not after every step.
01 · INNER STEPSTrain locallyEach miner runs many optimizer steps on its own GPUs, with no network traffic in between.
02 · PSEUDO-GRADIENTMeasure the driftWhen the round ends, a miner's change is how far its weights moved from the shared model.
03 · OUTER SYNCSync once a roundMiners exchange pseudo-gradients once per outer round. That is a fraction of the bandwidth DDP needs.
04 · OUTER STEPNesterov updateAn outer Nesterov optimizer applies the combined update to the one shared model. Then the next round starts.
MOTIONcyan coin hops station to station · the sync tile blinks twice while the miners exchange pseudo-gradients · the outer-step tile flashes cream when the shared model updates
INSERT COIN

Not live yet. Bounty is.

Hypertrain is a preview built on simulated data, so there is nothing to mine here yet. Bounty is open today.

OPEN /BOUNTY