One hot line stops a shire
Ivan put one global atomic counter in shire 0's scratchpad, had all 32 shires hammer it, and
found that shire 0 got 6% of its fair share and finished only after the other 31
. The obvious reading is
that the shire cache short-changes whoever owns the line. It does not. The atomic is shared out to within
half a percent (every shire between 0.998 and 1.004 of an even split), including with the shire that hosts it.
What that shire loses is something else entirely: with the clock steady at 600 MHz, its own memory path stops
— not slows, stops — for as long as the hammering lasts. The switch is sharp: 21 requesters in one other shire
leave it at 96% of its rate, 22 stop it, fewer than the 32 minions a shire has. It reproduces on all three cards we
could run it on (aifoundry2, aifoundry3 and aifoundry1 card 1), in three passes on each.
Checked on three cards (26 September 2026): this page's claims were
re-measured under a pre-registered plan, in three passes on each of aifoundry2, aifoundry3 (both firmware 1.3.1;
aifoundry3 held at 600 MHz) and aifoundry1 card 1 (firmware 1.2.0), and the sweeps' tables and charts now show those
passes. Several held with sharper numbers: the edge is exactly where section 3's inequality puts it, 22
requesters, and the pacing knee just above its 9,704 cycles. The one that did not: 31 requesters, one per shire,
saturate the line's bank but did not stop the one minion of its shire that was reading (section 7, corrected). The
reading-only rate in section 6 was qualified, and the bus error of section 8 was not repeated. The cards agree
throughout, except one aifoundry2 launch in section 1. Of 17 claims tested here, this page counts 14 held, 1 did not hold (corrected), 1 qualified and 1 not repeated; the hub’s scoreboard, 17 “a test behind it failed”. Record: docs/reports/data/2026-09-25-claims-v3.
Terms used on this page
Minions are the chip's small in-order RISC-V cores, 32 to a shire, and each shire's 4 MB of SRAM,
the shire cache, holds a 512 KB L2 cache, a 1 MB slice of the chip-wide L3 and a 2.5 MB scratchpad that any shire
can address (SCP in the errata). Every DRAM line is cached in the L3 slice of one home shire, chosen by
physical-address bits 10:6, and a scratchpad word's home is the shire whose scratchpad holds it. A hot line
is one line that many minions hammer at once and the host shire is the shire whose cache holds it; requests
from a shire's own minions are neighbourhood requests, those arriving from other shires over the mesh are L3-slave
requests, and amoaddg.w is a global atomic add, performed at the line's home. The PRM is Esperanto's
Programmer's Reference Manual; more in the
hub's glossary.
1. The atomic itself is fair
Every minion in the launch hammers one 4-byte word with amoaddg.w. All of them meet at a
chip-wide barrier first, so the window is the same window for everyone, and each one reports how many atomics
it completed inside it. A shire's share is its count divided by what an even split would give.
Why shares spread out with one minion per shire
Moving the line's home to shire 7, 15 or 31 changes nothing, so shire 0 is not special (one aifoundry2 pass with the line in shire 7 spread to 0.994–1.046; its other passes and the other cards stayed within 0.999–1.003). Shares spread out in one case only, the opposite of starvation. With a single minion per shire the bank still retires one atomic every 10 cycles, but each shire has only one request in the queue, so the time a request spends crossing the mesh decides how often each shire gets its turn. Shares fall step by step with distance from the line, and the host shire comes first because its request crosses no mesh links at all.
The table, shire by shire
In the a2 / a3 / a1c1 columns of this page's tables (aifoundry2, aifoundry3, aifoundry1 card 1), each card's value is the mean of its three passes, and a single number means the three cards agree to the digits shown.
2. The thing that actually starves
So we changed what the host shire does. Instead of joining the atomic, its 32 minions stream their own memory — 64-byte-strided loads over a private slice, either of their shire's scratchpad or of DRAM — while the other 31 shires hammer a line that lives in that same shire cache. The baseline is the identical loop with nobody else launched.
Inside the host shire, in three steps. Each step sets the N and P of section 3's sliders, and moving those sliders sets this drawing.
The table
These are not slow-downs, they are stops. The host shire's count is the same number whether the window is 5 milliseconds or 100:
The window-length chart and table
The table
About seventeen loads per minion get through, five warm-up loads before the window opens and twelve inside it, and then nothing: none through the 120 million remote atomics of a 2-second window, in all 21 such runs of the 23 September passes on the two cards. The 32 minions of that shire, a full thirty-second of the chip, are stalled in a load instruction that does not return. Nothing reports an error; the kernel finishes when the hammering stops.
A caveat: one session where the clock moved
The stop holds at a steady 600 MHz. In one aifoundry2 session (the first rerun attempt of 23 September, discarded for the energy figures) the ten 2-second runs whose clock moved between 600 and 800 MHz let the host's loads through at 1–7% of their rate, while the two at 600 MHz stopped at 384; one session does not establish that the clock is the cause.
3. The threshold is the bank going saturated
It takes one other shire. Two shires run, the host and one other, each with the same number of minions n. The host's n minions read their own scratchpad, and the other shire's n hammer a scratchpad word in the host shire. The host's rate is compared with n times the per-minion rate of its 32-minion loop alone, which is why a single host minion reads 111%.
There is no gradual degradation. From 2 to 20 remote requesters the host shire stays within about a percent of untouched (98.8–99.97% of its rate alone with the same number of minions, on all three cards), because each remote atomic takes about 216 cycles to make its round trip and twenty of them in flight still ask for less than the bank can retire. Twenty-one keep the bank at 10.3 cycles per atomic and leave the host at 95.6–95.7%; by twenty-two the measured cost reaches 10.0 cycles per atomic — the bank's floor, the queue full — and the host shire's own requests stop getting resources. One inequality says when:
N × 10 cycles < P + tround trip ⟹ the bank is not saturated
4. Esperanto knew: two errata describe it
Two entries in Esperanto's ET-SoC-1 errata describe it (numbered 4.1 and 4.2 in the document's body, 3.1 and 3.2 in its table of contents), and both are marked Postponed:
The two errata, quoted
That bears on both questions worth sending back. Would setting l3_yield restore the
host? Not certainly. Erratum 4.1 says it does not when the host and the mesh want the same address. The
case measured here is different addresses, which is what the yield was built for. But erratum 4.2 says the
yield has no sub-bank granularity, so a host request to the sub-bank the hot line saturates can still be
skipped indefinitely, and each host minion's stream reaches that sub-bank sooner or later. It was not tested.
Does it hold when the line is DRAM-backed rather than in scratchpad? Yes, and slightly worse: 192–240 of
the host's loads get through instead of 384 (section 2), at every window from 5 to 100 ms and on all three cards. The erratum's title says same SCP address
, but
the behaviour is not specific to the scratchpad: both kinds of address tested, a scratchpad word and a DRAM
line homed in that shire, do it.
5. The workaround, priced
Erratum 4.1's workaround is to slow the other shires down. It works, and near the knee it is nearly free — the bank was already saturated, so the first thing pacing takes away is queueing, not throughput:
6. What contention costs: 32× the time and 17× the energy
Spreading the same work over 32 lines, one per shire, changes everything: the same 1,024 minions running the same instruction do 32 times the work.
The table
What a sub-bank slot works out to
Ten cycles per atomic is the shire cache retiring one request at a time. If a sub-bank accepts a request
every two cycles, as a reading of the shire-cache spec in docs/research/counters-and-dram.md has it
(not measured here), a global atomic occupies about five sub-bank slots. The aggregate of about 60 million
atomics a second matches the figure Ivan quoted, which settles that it was an aggregate and not a per-shire
number.
What it costs in energy
Contention is not just slow, it is expensive per unit of work. Board power over idle on aifoundry2, in one session, measured at 10 Hz across three back-to-back launches of each case, with the die at 72–73 °C; the last column adds the 23 September passes on aifoundry2 and aifoundry3 (the three-card check did not repeat the power runs):
The table
The rates in this table are timed by the power runs themselves, so they differ slightly from the placement table's above (1,919 against 1,920 million a second). The last column pools this session with the energy manual's 23 September passes: .
A thousand minions stalled on a saturated line cost the card , which is the clearest possible statement that waiting is cheap and the loss is throughput, not power. Per operation the contended line costs times the energy of the same atomic spread over 32 lines.
7. What to do instead
- One global atomic per shire, never per minion, and never spin on the line. At 32 participants a barrier's arrivals serialise in 320 cycles, against a chip-wide barrier of about 5,000 cycles ( with one minion per shire): six percent. But 31 shires hammering one line already saturate its bank. Wait on a per-shire flag or a credit instead. The relay's first barrier, with every shire's leader polling one counter, hung in the one development run we made, on a card not recorded (Hand it to the next shire). At 1,024 participants the arrivals alone take 10,240 cycles, twice the whole barrier, with the hosting shire stopped throughout.
- Keep hot shared lines out of shires that compute. The cost is not paid by the code touching the line. It is paid by whatever else lives in that shire, and at a steady 600 MHz it is total. Make the owner a shire that does no compute: one of 32 is 3.1% of the chip. Not the master shire, whose scratchpad the firmware uses.
- If you must, pace it:
- Spread instead of sharing when you can. One line per shire is the same instruction at 32 times the rate and a seventeenth of the energy per operation.
- Measure the victim from user mode. Syscall 9,
SYSCALL_PMC_SC_SAMPLE, reads the cache-bank counters of any shire, so a kernel can watch the host shire's banks without touching firmware (docs/research/counters-and-dram.md).
8. Method, and what is not established
Method, and what is not established, in full
- The probe is
NB_HOTLINEinworkloads/nocbench, new for this. Every participant warms up, meets at a chip-wide barrier, then loops until a cycle deadline, counting completed operations. Counting completions inside a common window is what makes "share" mean anything; a fixed-count race would confound share with finishing order. - The home shire of a DRAM line is
PA[10:6], so the probe allocates a 2 KB-aligned region and uses the line at offset s×64. Scratchpad lines use the PRM's format 0, shire ID in bits [29:23]. The region is verified aligned at launch, and the 32-fold speed-up of one line per shire is the check that the address really does select the bank. - A global atomic cannot take the local path at all. Addressed through the self ID
0x7F,amoaddgraises a kernel bus error (seen during development, on a card not recorded; the three-card check did not repeat it). That is why section 1 comes out fair: from the host shire the atomic leaves through the same L3-slave port as everyone else's, so the arbitration that ranks the two classes never applies to it. The starvation in section 2 needs an ordinary load or store, which is what a computing shire issues. - Every sweep ran on aifoundry2 and aifoundry3 on 22 September, one launch per configuration, and again on
25–26 September in the three-card check: three passes on each of aifoundry2, aifoundry3 and aifoundry1 card 1,
each pass the 42 configurations plus 35 the check added (the edge at 19 and 21–23 requesters, the host's own
N-minion rates alone, pauses of 9,000–10,500 cycles, 5, 40 and 100 ms windows for all four homes, one requester per
shire, and warm-ups of 0 and 10), in a shuffled order, at 600 MHz. The tables and charts show those passes; the
energy table in section 6 is aifoundry2's first session (aifoundry2 and aifoundry3 are in the 23 September reruns of its last column). A result that
names no card held on all three, in every pass or on the pass means; where one rests on one card, one session or
fewer runs, the text says so. In the 75 configurations that have a host shire, its count is the same in all nine
launches (three passes on three cards) in 22 and within 2% in all but eight: the host reading scratchpad under a
DRAM-homed hot line (192–240 loads) and the paced runs at 9,800–12,000 cycles, at the knee (3–10% apart).
Raw data in
docs/reports/data/2026-09-25-claims-v3/raw/(<card>/lat/p*/hl/, one JSON line per launch;tools/claims-v3/lat/README.mdsays what each file holds), and the first sweeps indocs/reports/data/2026-09-22-hotline-aifoundry2/anddocs/reports/data/2026-09-22-hotline-aifoundry3/. - Why
l3_yieldwas left alone. It is a shire-cache configuration register on a shared lab card, and the errata suggest it would not fully help. It is also out of reach from user mode:l3_yield_priorityis a 5-bit field ofsc_reqq_ctl, an M-mode register that no syscall writes. And the 21 queue entries reserved for L3-slave requests look like deadlock avoidance, so changing the priority risks back-pressure on the mesh. - Not established: why Ivan's number was 6% rather than either of ours. Our reading is that his shire 0 did something local as well as the atomic — a poll, a flag read, the "am I last?" check a barrier makes — and that the local part is what was starved; but we did not run his code. The card's thermal telemetry gives the host only a chip-wide average of the shire sensors (plus the extremes since reset), in whole degrees, so the spatial question — whether a hot line makes its shire measurably hotter — cannot be answered with this instrument. The per-shire sensors exist on the die; Spatial temperature says what it would take to read them.
Reproduce this
To reproduce, with DATA a run directory on aifoundry2 and DATA3 one
on aifoundry3:
workloads/nocbench/run_hotline.sh DATA 6000000 # 42 configurations, each well under a second
workloads/nocbench/run_hotline.sh DATA3 6000000 # the same sweep, on aifoundry3
tools/ettelem/run_hotline_power.sh DATA # board power and rails, three launches per case
python3 tools/ettelem/analyze_hotline_power.py DATA --out DATA/power.json
# the three-card passes: unit hl of tools/claims-v3/lat/block.sh (tools/claims-v3/lat/README.md), then
V3=docs/reports/data/2026-09-25-claims-v3/raw
python3 workloads/nocbench/analyze_hotline.py --v3 $V3 --power DATA/power.json \
--context DATA/context.json --barrier docs/reports/data/2026-09-18-nocbench-aifoundry2/barrier-chip1.jsonl \
--stop-runs docs/reports/data/2026-09-23-reruns-aifoundry2-warm/hotline-pass*/runs.jsonl \
docs/reports/data/2026-09-23-reruns-aifoundry3/hotline-pass*/runs.jsonl \
--out DATA/hotline.json
python3 tools/ettelem/analyze_reruns.py docs/reports/data/2026-09-23-reruns-aifoundry2-warm \
docs/reports/data/2026-09-23-reruns-aifoundry3 --out reruns.json # the passes behind the ranges in section 6
Version history and provenance
The errata text, quoted by hand from the errata document, comes from context.json.
The window table is measured in the three-card passes (analyze_hotline.py --v3; each card's row of a
configuration is the mean of its passes), and the chart's 2-second windows are the 23 September reruns' starved runs
(--stop-runs). The barrier with one minion per shire is the three-card check's (results/lat.json, LAT-N4, three passes on each card; the 18 September nocbench run on aifoundry2 gave 4,995 cycles); the round trip and the bank's 10
cycles are computed from the sweep. Versions (one line per date; every wording is in
the source's history):
22 September, first published; 24 September, the threshold a range; 25 September, the bank rule necessary but not
sufficient, the chip barrier 4,995 cycles, not 4,997, and the stop stated for a steady 600 MHz clock; 26 September, three
passes on three cards (the edge at 22 requesters, the knee just above 9,704 cycles), and the claim about one requester
per shire corrected; 27 September, charts; 28 September, repeats cut and the note's counts given by both rules. Record:
docs/reports/data/2026-09-25-claims-v3/.
Related reports
- Hand it to the next shire — the companion: share data between shires through a slab of scratchpad instead of one hot line.
- The ET-SoC-1 energy manual, section 6 — these energies re-measured on aifoundry2 and aifoundry3, with ranges.
- On-chip communication — the chip-wide barrier on three cards (8.3 µs from global atomics and credits, 2.3 µs on the hardware tree) and the reduction trees that can replace one global counter.
- Anatomy of a memory access — how physical-address bits 10:6 pick a line's home shire, and what an L3 hit costs per mesh hop.
- Sparse compute — a divergence probe whose 2,048 harts take their work from one global atomic counter.
- Spatial temperature — the per-shire temperature sensors that could tell whether a hot line heats its shire.