Spatial temperature on the ET-SoC-1: 35 sensors on the die, one average at the host

Technical brief · 22 September 2026, checked on three cards 26 September · firmware source et-platform 353f20e, open RTL core-et b38a1a3 · a measured result holds on aifoundry2 and aifoundry3 unless the text names a card; the three-card check (aifoundry2, aifoundry3 and aifoundry1-c1) is stated where it applies · part of the ET-SoC-1 measurement reports
Can the ET-SoC-1 measure how heat moves across its 1,088-core die? In hardware, yes; with the stock firmware, no. The die carries 35 temperature sensors, one in every tile of its 6×6 grid except the PCIe shire. The host receives a 34-shire average plus two peak-hold extremes (and the I/O shire’s own reading), all in whole degrees: the firmware’s integer conversion truncates every sensor’s reading to 1 °C before it is averaged. The high leaks one number: each time it rose, on each of the three cards checked, the hottest shire read °C above the average (in whole degrees), but the host never learns which shire that was. This brief covers where the sensors are, where the firmware discards the detail, and what it would take to get a per-shire heat map.

Checked on three cards (26 September 2026): this brief’s claims were re-measured on aifoundry2, aifoundry3 and aifoundry1-c1 under a pre-registered plan. Of 8 claims tested here, this page counts 2 held, 5 corrected and 1 not confirmed; the hub’s scoreboard, 6 “a test behind it failed” and 2 “fewer than three repeats”. What held: the peak-hold reset and the missing per-shire temperature line. What was corrected: after a reset the high rose to 2 or 3 °C above the average (3 in 61–83% of the rises, card by card, against the 80% or more of 3 or 4 predicted), and the low sat 1 or 2 °C under the minimum, not always 1. The idle voltage map was not confirmed (too few complete captures) and reads differently on each card. Record: docs/reports/data/2026-09-25-claims-v3.

Temperature sensors on the die
35
34 minion shires + 1 I/O shire
Resolution in hardware
12-bit (0.061 °C)
Moortec PVT IP; the firmware keeps whole °C
Layout
6×6 grid
1 sensor per tile, except the PCIe shire (~3.7 mm grid)
What the host sees today
Collapsed
Current mean + since-reset extremes, whole °C; the high gives the hottest shire’s excess at a peak, not its place
Terms used on this page

A shire is a tile of 32 minion cores (small RISC-V cores with vector and tensor units) sharing 4 MB of SRAM (L2, an L3 slice and scratchpad); the 34 minion shires, the I/O and PCIe shires and 8 memory shires sit on the chip’s mesh network. The service processor (SP) is the on-chip management core; its second-stage firmware (BL2) reads the sensors and answers the host. The PMIC is the board’s power controller; PVT = process, voltage and temperature monitors. More in the hub’s glossary.

1. Sensor layout and architecture

The ET-SoC-1 has an on-die process, voltage and temperature (PVT) monitoring subsystem built from Moortec IP. The Programmer’s Reference Manual (PRM §1.6) counts 36 temperature sensors; the open RTL (core-et/rtl/inc/pvt_defines.vh, citing RTLMIN-3619) drops the PCIe shire’s, and the firmware enables 35 channels. So there are 35 temperature sensors on the die:

The 6×6 shire mesh, as the latency measurements place it

The 32 compute shires are placed on the 6×6 grid by TensorSend latency (marty1885’s map, checked on aifoundry2 in On-chip communication). Each cell in this grid is one tile, about 3.7 mm on a side: the tile pitch measured on Esperanto’s die plot and scaled to the 570 mm² die (Heat per millimetre §2).

Reading this diagram: orientation and the inferred cells

The mesh as latency places it (x across, y down). Sn is compute shire n and TSn its sensor channel. In this frame the eight memory shires sit one step beyond the top row (0–3, at x = 1–4) and the bottom row (4–7); the PRM calls those sides west and east, so this drawing is the die turned a quarter. The grey cells hold the master, spare, PCIe and I/O shires; which is which is not measured (the heat-per-mm die study infers I/O and PCIe at (0,4) and (0,5)). TS32 (master), TS33 (spare) and TS34 (I/O) are read like the others; the PCIe shire has no sensor.

The sensors' hardware registers, in full

The sensors’ hardware registers

Five PVT controllers read the 35 sensors, sampling continuously at 12 bits (their addresses are under Implementation notes). Every sensor has its own hardware registers:

2. The firmware reduction bottleneck

The firmware source: what it can read, and what it sends

The SP’s PVT driver can read each shire’s sensor on its own. In ServiceProcessorBL2/driver/pvt_controller.c:

int pvt_get_min_shire_ts_sample(PVTC_MINSHIRE_e min_id,
                                TS_Sample *ts_sample);
// exists, no caller:
int pvt_get_and_print(uint8_t print_ts, uint8_t print_vm,
                      PVT_PRINT_e print_select,  /* e.g. PVT_PRINT_MINSHIRE_ALL */
                      uint16_t *data, uint32_t *num_bytes);

But when it answers the host’s management interface (dev_mngt_service / libDM.so), the SP’s thermal service (thermal_pwr_mgmt.c, line 732) reduces the data:

// ServiceProcessorBL2/services/thermal_pwr_mgmt.c (abridged)
int get_module_current_temperature(struct current_temperature_t *temperature) {
    pvt_get_minion_avg_temperature(&pmic_temperature);
    pvt_get_minion_avg_low_high_temperature(&pvt_temperature);
    pvt_get_ioshire_ts_sample(&pvt_temperature);

    // The host's packet receives only:
    // integer mean of the 34 shires' whole-degree readings
    temperature->minshire_avg  = pvt_temperature.current;
    // highest value any shire's peak-hold register captured since the last reset
    temperature->minshire_high = pvt_temperature.high;
    // lowest value any shire's peak-hold register captured since the last reset
    temperature->minshire_low  = pvt_temperature.low;
    // the I/O shire's own sensor (also sent with its high and low)
    temperature->ioshire_current = ...;
    // the same minion-shire average, not a PMIC reading
    temperature->pmic_sys = pmic_temperature;
}
The sensor-to-host pipeline that code implements

Tab to (or hover) one of the five packet fields on the right to light what feeds it. The average and peak-hold boxes track whichever session (one of aifoundry2's on 20 September, or a card's reset window of the three-card check) and moment the time slider below ("What the host's temperature fields did") is set to; until you drag it, the raw-code slider follows along too, showing a 12-bit code that would truncate to the current mean. Drag it to try any other sample.

The field named pmic_sys is not a board sensor either: the firmware fills it with the same 34-shire average (), and the SP does not forward the PMIC’s own regulator temperatures. The service processor’s stats packet has a system-temperature slot meant for the PMIC, but the firmware writes 0 into it (a TODO in thermal_pwr_mgmt.c), and it read 0 in all samples of 20 September and in every committed sample from aifoundry2 and aifoundry3.

So the host never sees the current spread between shires: the low and high are peak-hold values. The firmware resets them together with the service processor’s own statistics, so both cover the same window (so the firmware source reads; on the cards, ). In aifoundry2’s telemetry of 20 September the low read °C in all samples, 1 °C under the service processor’s minimum of the average ( °C), while the average ran from °C. The high only ratchets, but it leaks one number. Every rise followed a new record of the average and settled exactly °C above the service processor’s maximum of the average (written below as maximum/high): when the log began, then during it (each after the record), and in the uncontrolled Horace session (the came s after its record). So at those peaks the hottest shire’s sensor read about °C above the 34-shire mean, in the firmware’s whole degrees. That was this evening’s figure, not a constant: . The high is a single hardware sample, so sensor noise can add to it. Which shire it was, the host cannot tell.

The host's temperature fields over time, charted: aifoundry2's sessions and a reset window on each card
What the host’s temperature fields did: aifoundry2 on 20 September, and one reset window on each of three cards

The four whole-degree fields the host received, one step line each, and the band between the service processor’s maximum of the average and the high, labelled with its width. The short ticks at the top mark each rise of the high. The load step was sampled at 10 Hz; the two Horace logs were committed thinned to 2 Hz, so their step times are good to 0.5 s. The three-card check’s windows (25–26 September, 10 Hz; for each card the first window of its first pass) start at a reset of the service processor’s statistics and of the low and high, and a 7 s random-data burst begins 3 s in; each is drawn from its first sample with the statistics set. Move the time slider, or hover or tap the chart: away from a rise, the high only says what the hottest shire read at the last peak, nothing about now.

Why voltage reaches the host shire by shire, and temperature does not
Why voltage reaches the host shire by shire, and temperature does not

The Power and temperature report (20 Sep, §3 The per-shire voltage map) reads per-shire voltages this way: with the SP’s log level at DEBUG, the firmware prints one line per shire per pass (MS %2d Voltage [mV]: VDD_MNN: %d [%d, %d] …). In one idle capture on aifoundry2 (20 September) the minion rail read 517–521 mV across the 34 shires, in whole millivolts; the monitors’ low/high captures span 513–522 mV. That line is printed because it sits inside the GET_MINION_VM macro, which the per-pass power update calls. For temperature the firmware already has the matching DEBUG line (MS %2d Temp [C]: %d [%d, %d], in pvt_print_min_shire_temperature_sampled_values()), but it is reached only through pvt_print_all(), which nothing calls, so the host is left with the chip-wide aggregate.

3. What is needed to measure spatial thermal flow

To watch heat move across the grid (say, by running multiply-add kernels on the shires in one corner and watching the heat spread to their neighbours), the following pieces are needed:

What each of the five pieces needs, table by table
ComponentStatusWhat is needed
1. Sensor hardware Available now 35 Moortec temperature sensors already run continuously on the silicon, at 12-bit resolution in hardware.
2. SP firmware telemetry patch Needs code Change ServiceProcessorBL2 to send the individual readings:
  1. Option A (trace buffer): add a 35-entry array, int16_t shire_temps[35], to the sp_stats periodic trace buffer.
  2. Option B (new command): add a management command, for example DM_CMD_GET_SPATIAL_TEMPERATURE, returning All_MinShire_samples.
  3. Option C (no protocol change): call the existing pvt_print_temperature_sampled_values(PVTC_MINION_SHIRE) from the SP’s per-pass loop, one line, and parse the trace like the voltage map (whole degrees, since the existing print uses converted values; printing the raw code needs one more changed line). It also writes a CRITICAL “MinShire Average Temp” line on every call.
Whichever option is used, export the raw 12-bit code (or millidegrees computed before the /1000), not sample.current, which is already truncated to 1 °C.
3. Firmware signing and flashing Untested BL1 and BL2 verify a certificate and signature against a public-key hash in OTP (crypto_rot.c / VaultIP) unless an OTP chicken bit is set, and neither the signing tool nor a test key is in the open tree. Whether these lab cards accept a rebuilt image is unknown until someone tries, and a bad image can brick the boot path, so reflashing is a lab-admin decision (the firmware caveat).
4. Host telemetry extraction Ready to build Extend tools/ettelem (which samples libDM at 10 Hz, about 45 Hz at most, and dumps the SP trace buffer on demand with ettelem sptrace) to unpack the 35 temperature channels into a 6×6 matrix.
5. 2D heat-diffusion reconstruction Not yet fitted The fitted thermal model is lumped: six RC stages, 1.47 °C/W in total, fitted to one day’s runs on one card (aifoundry2) in its chassis, driven by the chip-wide average. A 2D model (∂T/∂t = α∇²T + P(x, y)/C − (T − Tamb)/τ) needs lateral conductances and per-cell capacitances that only per-shire readings could fit.
Implementation notes: the PVT controllers' register map

Implementation notes

The five PVT controllers (PVTC 0–4) are mapped in the service processor’s address space at 0x00_5400_0000 + n·0x1_0000, and all five sample continuously at 12 bits (RUN_1 mode):

ControllerBase addressSensorsChannels active
PVTC 00x00_5400_0000TS 0–78 / 8
PVTC 10x00_5401_0000TS 8–158 / 8
PVTC 20x00_5402_0000TS 16–238 / 8
PVTC 30x00_5403_0000TS 24–318 / 8
PVTC 40x00_5404_0000TS 32–343 / 8 (mask 0xF8)

4. Current best workaround: temporal reconstruction

Until a modified SP firmware can be flashed (whether these cards require a signed image is untested), the measurements so far rely on temporal reconstruction: reading more out of the timing of the chip-wide average (see the Horace experiment, §8, A model from flips to temperature):

  1. Sub-degree de-quantization: the moments when the whole-degree average steps are sharp at 10 Hz, and fitting the heating power that reproduces them recovers the temperature below a degree (the Horace experiment, §2).
  2. Flip-to-heat modelling from the RTL: counting the toggles the operands cause in the open RTL of the multiply-add datapath predicts the board power, and from it the die temperature, with models fitted on aifoundry2 (the Horace experiment, §8).

Both sharpen the chip-wide average; neither says where on the die the heat is.

Sources

Sources, in full
Version history

← All ET-SoC-1 measurement reports