One hot line stops a shire

22 September 2026, re-measured 25–26 September · three cards: aifoundry2, aifoundry3 and aifoundry1 card 1 · workloads/nocbench --test hotline, a new many-to-one probe · following up a result reported in Discord by Ivan, another ET-SoC-1 programmer (first analysed, with his 6% taken at face value, in L2 mainline starvation) · part of the ET-SoC-1 measurement reports

Ivan put one global atomic counter in shire 0's scratchpad, had all 32 shires hammer it, and found that shire 0 got 6% of its fair share and finished only after the other 31. The obvious reading is that the shire cache short-changes whoever owns the line. It does not. The atomic is shared out to within half a percent (every shire between 0.998 and 1.004 of an even split), including with the shire that hosts it. What that shire loses is something else entirely: with the clock steady at 600 MHz, its own memory path stops — not slows, stops — for as long as the hammering lasts. The switch is sharp: 21 requesters in one other shire leave it at 96% of its rate, 22 stop it, fewer than the 32 minions a shire has. It reproduces on all three cards we could run it on (aifoundry2, aifoundry3 and aifoundry1 card 1), in three passes on each.

Checked on three cards (26 September 2026): this page's claims were re-measured under a pre-registered plan, in three passes on each of aifoundry2, aifoundry3 (both firmware 1.3.1; aifoundry3 held at 600 MHz) and aifoundry1 card 1 (firmware 1.2.0), and the sweeps' tables and charts now show those passes. Several held with sharper numbers: the edge is exactly where section 3's inequality puts it, 22 requesters, and the pacing knee just above its 9,704 cycles. The one that did not: 31 requesters, one per shire, saturate the line's bank but did not stop the one minion of its shire that was reading (section 7, corrected). The reading-only rate in section 6 was qualified, and the bus error of section 8 was not repeated. The cards agree throughout, except one aifoundry2 launch in section 1. Of 17 claims tested here, this page counts 14 held, 1 did not hold (corrected), 1 qualified and 1 not repeated; the hub’s scoreboard, 17 “a test behind it failed”. Record: docs/reports/data/2026-09-25-claims-v3.

Terms used on this page

Minions are the chip's small in-order RISC-V cores, 32 to a shire, and each shire's 4 MB of SRAM, the shire cache, holds a 512 KB L2 cache, a 1 MB slice of the chip-wide L3 and a 2.5 MB scratchpad that any shire can address (SCP in the errata). Every DRAM line is cached in the L3 slice of one home shire, chosen by physical-address bits 10:6, and a scratchpad word's home is the shire whose scratchpad holds it. A hot line is one line that many minions hammer at once and the host shire is the shire whose cache holds it; requests from a shire's own minions are neighbourhood requests, those arriving from other shires over the mesh are L3-slave requests, and amoaddg.w is a global atomic add, performed at the line's home. The PRM is Esperanto's Programmer's Reference Manual; more in the hub's glossary.

Share of the atomic the host shire gets
against a fair share of 1.000; spread across all 32 shires below
Its own memory operations, contended
of what the same 32 minions do with nobody hammering
Remote requesters needed to flip it
Energy per atomic, contended vs spread
same instructions, 32× less work done

1. The atomic itself is fair

Every minion in the launch hammers one 4-byte word with amoaddg.w. All of them meet at a chip-wide barrier first, so the window is the same window for everyone, and each one reports how many atomics it completed inside it. A shire's share is its count divided by what an even split would give.

Share of one contended atomic, shire by shire, against mesh distance from the line (a DRAM line homed in shire 0)

Why shares spread out with one minion per shire

Moving the line's home to shire 7, 15 or 31 changes nothing, so shire 0 is not special (one aifoundry2 pass with the line in shire 7 spread to 0.994–1.046; its other passes and the other cards stayed within 0.999–1.003). Shares spread out in one case only, the opposite of starvation. With a single minion per shire the bank still retires one atomic every 10 cycles, but each shire has only one request in the queue, so the time a request spends crossing the mesh decides how often each shire gets its turn. Shares fall step by step with distance from the line, and the host shire comes first because its request crosses no mesh links at all.

The table, shire by shire

In the a2 / a3 / a1c1 columns of this page's tables (aifoundry2, aifoundry3, aifoundry1 card 1), each card's value is the mean of its three passes, and a single number means the three cards agree to the digits shown.

2. The thing that actually starves

So we changed what the host shire does. Instead of joining the atomic, its 32 minions stream their own memory — 64-byte-strided loads over a private slice, either of their shire's scratchpad or of DRAM — while the other 31 shires hammer a line that lives in that same shire cache. The baseline is the identical loop with nobody else launched.

Inside the host shire, in three steps. Each step sets the N and P of section 3's sliders, and moving those sliders sets this drawing.

Inside the host shire: its own loads and the remote atomics meet at one bank

The table

These are not slow-downs, they are stops. The host shire's count is the same number whether the window is 5 milliseconds or 100:

The window-length chart and table
Stops, not slows: operations completed in the window, host against the remote atomics hammering it

The table

About seventeen loads per minion get through, five warm-up loads before the window opens and twelve inside it, and then nothing: none through the 120 million remote atomics of a 2-second window, in all 21 such runs of the 23 September passes on the two cards. The 32 minions of that shire, a full thirty-second of the chip, are stalled in a load instruction that does not return. Nothing reports an error; the kernel finishes when the hammering stops.

A caveat: one session where the clock moved

The stop holds at a steady 600 MHz. In one aifoundry2 session (the first rerun attempt of 23 September, discarded for the energy figures) the ten 2-second runs whose clock moved between 600 and 800 MHz let the host's loads through at 1–7% of their rate, while the two at 600 MHz stopped at 384; one session does not establish that the clock is the cause.

3. The threshold is the bank going saturated

It takes one other shire. Two shires run, the host and one other, each with the same number of minions n. The host's n minions read their own scratchpad, and the other shire's n hammer a scratchpad word in the host shire. The host's rate is compared with n times the per-minion rate of its 32-minion loop alone, which is why a single host minion reads 111%.

There is no gradual degradation. From 2 to 20 remote requesters the host shire stays within about a percent of untouched (98.8–99.97% of its rate alone with the same number of minions, on all three cards), because each remote atomic takes about 216 cycles to make its round trip and twenty of them in flight still ask for less than the bank can retire. Twenty-one keep the bank at 10.3 cycles per atomic and leave the host at 95.6–95.7%; by twenty-two the measured cost reaches 10.0 cycles per atomic — the bank's floor, the queue full — and the host shire's own requests stop getting resources. One inequality says when:

N × 10 cycles < P + tround trip ⟹ the bank is not saturated

Will this hot line stop its shire? The host's own throughput against the load the remote minions put on its bank

4. Esperanto knew: two errata describe it

Two entries in Esperanto's ET-SoC-1 errata describe it (numbered 4.1 and 4.2 in the document's body, 3.1 and 3.2 in its table of contents), and both are marked Postponed:

The two errata, quoted

That bears on both questions worth sending back. Would setting l3_yield restore the host? Not certainly. Erratum 4.1 says it does not when the host and the mesh want the same address. The case measured here is different addresses, which is what the yield was built for. But erratum 4.2 says the yield has no sub-bank granularity, so a host request to the sub-bank the hot line saturates can still be skipped indefinitely, and each host minion's stream reaches that sub-bank sooner or later. It was not tested. Does it hold when the line is DRAM-backed rather than in scratchpad? Yes, and slightly worse: 192–240 of the host's loads get through instead of 384 (section 2), at every window from 5 to 100 ms and on all three cards. The erratum's title says same SCP address, but the behaviour is not specific to the scratchpad: both kinds of address tested, a scratchpad word and a DRAM line homed in that shire, do it.

5. The workaround, priced

Erratum 4.1's workaround is to slow the other shires down. It works, and near the knee it is nearly free — the bank was already saturated, so the first thing pacing takes away is queueing, not throughput:

Pacing the 992 remote minions: what each side gets

6. What contention costs: 32× the time and 17× the energy

Spreading the same work over 32 lines, one per shire, changes everything: the same 1,024 minions running the same instruction do 32 times the work.

The table
What a sub-bank slot works out to

Ten cycles per atomic is the shire cache retiring one request at a time. If a sub-bank accepts a request every two cycles, as a reading of the shire-cache spec in docs/research/counters-and-dram.md has it (not measured here), a global atomic occupies about five sub-bank slots. The aggregate of about 60 million atomics a second matches the figure Ivan quoted, which settles that it was an aggregate and not a per-shire number.

What it costs in energy

Contention is not just slow, it is expensive per unit of work. Board power over idle on aifoundry2, in one session, measured at 10 Hz across three back-to-back launches of each case, with the die at 72–73 °C; the last column adds the 23 September passes on aifoundry2 and aifoundry3 (the three-card check did not repeat the power runs):

The table

The rates in this table are timed by the power runs themselves, so they differ slightly from the placement table's above (1,919 against 1,920 million a second). The last column pools this session with the energy manual's 23 September passes: .

The last column drawn: energy per operation over every pass, one contended line against the same atomic spread over 32

A thousand minions stalled on a saturated line cost the card , which is the clearest possible statement that waiting is cheap and the loss is throughput, not power. Per operation the contended line costs times the energy of the same atomic spread over 32 lines.

7. What to do instead

8. Method, and what is not established

Method, and what is not established, in full
Reproduce this

To reproduce, with DATA a run directory on aifoundry2 and DATA3 one on aifoundry3:

workloads/nocbench/run_hotline.sh DATA 6000000           # 42 configurations, each well under a second
workloads/nocbench/run_hotline.sh DATA3 6000000          # the same sweep, on aifoundry3
tools/ettelem/run_hotline_power.sh DATA                  # board power and rails, three launches per case
python3 tools/ettelem/analyze_hotline_power.py DATA --out DATA/power.json
# the three-card passes: unit hl of tools/claims-v3/lat/block.sh (tools/claims-v3/lat/README.md), then
V3=docs/reports/data/2026-09-25-claims-v3/raw
python3 workloads/nocbench/analyze_hotline.py --v3 $V3 --power DATA/power.json \
    --context DATA/context.json --barrier docs/reports/data/2026-09-18-nocbench-aifoundry2/barrier-chip1.jsonl \
    --stop-runs docs/reports/data/2026-09-23-reruns-aifoundry2-warm/hotline-pass*/runs.jsonl \
                docs/reports/data/2026-09-23-reruns-aifoundry3/hotline-pass*/runs.jsonl \
    --out DATA/hotline.json
python3 tools/ettelem/analyze_reruns.py docs/reports/data/2026-09-23-reruns-aifoundry2-warm \
    docs/reports/data/2026-09-23-reruns-aifoundry3 --out reruns.json   # the passes behind the ranges in section 6
Version history and provenance

The errata text, quoted by hand from the errata document, comes from context.json. The window table is measured in the three-card passes (analyze_hotline.py --v3; each card's row of a configuration is the mean of its passes), and the chart's 2-second windows are the 23 September reruns' starved runs (--stop-runs). The barrier with one minion per shire is the three-card check's (results/lat.json, LAT-N4, three passes on each card; the 18 September nocbench run on aifoundry2 gave 4,995 cycles); the round trip and the bank's 10 cycles are computed from the sweep. Versions (one line per date; every wording is in the source's history): 22 September, first published; 24 September, the threshold a range; 25 September, the bank rule necessary but not sufficient, the chip barrier 4,995 cycles, not 4,997, and the stop stated for a steady 600 MHz clock; 26 September, three passes on three cards (the edge at 22 requesters, the knee just above 9,704 cycles), and the claim about one requester per shire corrected; 27 September, charts; 28 September, repeats cut and the note's counts given by both rules. Record: docs/reports/data/2026-09-25-claims-v3/.