Spatial temperature on the ET-SoC-1: 35 sensors on the die, one average at the host
Checked on three cards (26 September 2026): this brief’s claims were re-measured on aifoundry2, aifoundry3 and aifoundry1-c1 under a pre-registered plan. Of 8 claims tested here, this page counts 2 held, 5 corrected and 1 not confirmed; the hub’s scoreboard, 6 “a test behind it failed” and 2 “fewer than three repeats”. What held: the peak-hold reset and the missing per-shire temperature line. What was corrected: after a reset the high rose to 2 or 3 °C above the average (3 in 61–83% of the rises, card by card, against the 80% or more of 3 or 4 predicted), and the low sat 1 or 2 °C under the minimum, not always 1. The idle voltage map was not confirmed (too few complete captures) and reads differently on each card. Record: docs/reports/data/2026-09-25-claims-v3.
Terms used on this page
A shire is a tile of 32 minion cores (small RISC-V cores with vector and tensor units) sharing 4 MB of SRAM (L2, an L3 slice and scratchpad); the 34 minion shires, the I/O and PCIe shires and 8 memory shires sit on the chip’s mesh network. The service processor (SP) is the on-chip management core; its second-stage firmware (BL2) reads the sensors and answers the host. The PMIC is the board’s power controller; PVT = process, voltage and temperature monitors. More in the hub’s glossary.
1. Sensor layout and architecture
The ET-SoC-1 has an on-die process, voltage and temperature (PVT) monitoring subsystem built from Moortec IP. The Programmer’s Reference Manual (PRM §1.6) counts 36 temperature sensors; the open RTL (core-et/rtl/inc/pvt_defines.vh, citing RTLMIN-3619) drops the PCIe shire’s, and the firmware enables 35 channels. So there are 35 temperature sensors on the die:
- 34 minion shires (MinShire 0–33): the 32 compute shires (1,024 minions) plus the master shire and the spare shire. Each has its own temperature sensor next to its process detector and voltage monitor.
- 1 I/O shire: one temperature sensor, alongside the 5 PVT controllers.
The 6×6 shire mesh, as the latency measurements place it
The 32 compute shires are placed on the 6×6 grid by TensorSend latency (marty1885’s map, checked on aifoundry2 in On-chip communication). Each cell in this grid is one tile, about 3.7 mm on a side: the tile pitch measured on Esperanto’s die plot and scaled to the 570 mm² die (Heat per millimetre §2).
Reading this diagram: orientation and the inferred cells
The mesh as latency places it (x across, y down). Sn is compute shire n and TSn its sensor channel. In this frame the eight memory shires sit one step beyond the top row (0–3, at x = 1–4) and the bottom row (4–7); the PRM calls those sides west and east, so this drawing is the die turned a quarter. The grey cells hold the master, spare, PCIe and I/O shires; which is which is not measured (the heat-per-mm die study infers I/O and PCIe at (0,4) and (0,5)). TS32 (master), TS33 (spare) and TS34 (I/O) are read like the others; the PCIe shire has no sensor.
The sensors' hardware registers, in full
The sensors’ hardware registers
Five PVT controllers read the 35 sensors, sampling continuously at 12 bits (their addresses are under Implementation notes). Every sensor has its own hardware registers:
TS_SDIF_DATA: the current conversion result.TS_SMPL_HILO: hardware peak capture. It holds the highest and lowest readings since it was last reset (throughTS_HILO_RESET), with no software polling.- The sensors run in
RUN_1mode (uncalibrated). A raw 12-bit sample (0 to 4095) converts to Celsius in steps of 0.061 °C:T(°C) = (57400 + (249400 * sample) / 4096 - 124700) / 1000
The firmware evaluates this in integer arithmetic (pvt_ts_conversion), so every per-shire value it holds is already whole degrees; the 0.061 °C step exists only in the raw 12-bit code inTS_SDIF_DATA.
2. The firmware reduction bottleneck
The firmware source: what it can read, and what it sends
The SP’s PVT driver can read each shire’s sensor on its own. In ServiceProcessorBL2/driver/pvt_controller.c:
int pvt_get_min_shire_ts_sample(PVTC_MINSHIRE_e min_id,
TS_Sample *ts_sample);
// exists, no caller:
int pvt_get_and_print(uint8_t print_ts, uint8_t print_vm,
PVT_PRINT_e print_select, /* e.g. PVT_PRINT_MINSHIRE_ALL */
uint16_t *data, uint32_t *num_bytes);
But when it answers the host’s management interface (dev_mngt_service / libDM.so), the SP’s thermal service (thermal_pwr_mgmt.c, line 732) reduces the data:
// ServiceProcessorBL2/services/thermal_pwr_mgmt.c (abridged)
int get_module_current_temperature(struct current_temperature_t *temperature) {
pvt_get_minion_avg_temperature(&pmic_temperature);
pvt_get_minion_avg_low_high_temperature(&pvt_temperature);
pvt_get_ioshire_ts_sample(&pvt_temperature);
// The host's packet receives only:
// integer mean of the 34 shires' whole-degree readings
temperature->minshire_avg = pvt_temperature.current;
// highest value any shire's peak-hold register captured since the last reset
temperature->minshire_high = pvt_temperature.high;
// lowest value any shire's peak-hold register captured since the last reset
temperature->minshire_low = pvt_temperature.low;
// the I/O shire's own sensor (also sent with its high and low)
temperature->ioshire_current = ...;
// the same minion-shire average, not a PMIC reading
temperature->pmic_sys = pmic_temperature;
}
Tab to (or hover) one of the five packet fields on the right to light what feeds it. The average and peak-hold boxes track whichever session (one of aifoundry2's on 20 September, or a card's reset window of the three-card check) and moment the time slider below ("What the host's temperature fields did") is set to; until you drag it, the raw-code slider follows along too, showing a 12-bit code that would truncate to the current mean. Drag it to try any other sample.
The field named pmic_sys is not a board sensor either: the firmware fills it with the same 34-shire average (), and the SP does not forward the PMIC’s own regulator temperatures. The service processor’s stats packet has a system-temperature slot meant for the PMIC, but the firmware writes 0 into it (a TODO in thermal_pwr_mgmt.c), and it read 0 in all samples of 20 September and in every committed sample from aifoundry2 and aifoundry3.
So the host never sees the current spread between shires: the low and high are peak-hold values. The firmware resets them together with the service processor’s own statistics, so both cover the same window (so the firmware source reads; on the cards, ). In aifoundry2’s telemetry of 20 September the low read °C in all samples, 1 °C under the service processor’s minimum of the average ( °C), while the average ran from °C. The high only ratchets, but it leaks one number. Every rise followed a new record of the average and settled exactly °C above the service processor’s maximum of the average (written below as maximum/high): when the log began, then during it (each after the record), and in the uncontrolled Horace session (the came s after its record). So at those peaks the hottest shire’s sensor read about °C above the 34-shire mean, in the firmware’s whole degrees. That was this evening’s figure, not a constant: . The high is a single hardware sample, so sensor noise can add to it. Which shire it was, the host cannot tell.
The host's temperature fields over time, charted: aifoundry2's sessions and a reset window on each card
The four whole-degree fields the host received, one step line each, and the band between the service processor’s maximum of the average and the high, labelled with its width. The short ticks at the top mark each rise of the high. The load step was sampled at 10 Hz; the two Horace logs were committed thinned to 2 Hz, so their step times are good to 0.5 s. The three-card check’s windows (25–26 September, 10 Hz; for each card the first window of its first pass) start at a reset of the service processor’s statistics and of the low and high, and a 7 s random-data burst begins 3 s in; each is drawn from its first sample with the statistics set. Move the time slider, or hover or tap the chart: away from a rise, the high only says what the hottest shire read at the last peak, nothing about now.
Why voltage reaches the host shire by shire, and temperature does not
The Power and temperature report (20 Sep, §3 The per-shire voltage map) reads per-shire voltages this way: with the SP’s log level at DEBUG, the firmware prints one line per shire per pass (MS %2d Voltage [mV]: VDD_MNN: %d [%d, %d] …). In one idle capture on aifoundry2 (20 September) the minion rail read 517–521 mV across the 34 shires, in whole millivolts; the monitors’ low/high captures span 513–522 mV. That line is printed because it sits inside the GET_MINION_VM macro, which the per-pass power update calls. For temperature the firmware already has the matching DEBUG line (MS %2d Temp [C]: %d [%d, %d], in pvt_print_min_shire_temperature_sampled_values()), but it is reached only through pvt_print_all(), which nothing calls, so the host is left with the chip-wide aggregate.
3. What is needed to measure spatial thermal flow
To watch heat move across the grid (say, by running multiply-add kernels on the shires in one corner and watching the heat spread to their neighbours), the following pieces are needed:
What each of the five pieces needs, table by table
| Component | Status | What is needed |
|---|---|---|
| 1. Sensor hardware | Available now | 35 Moortec temperature sensors already run continuously on the silicon, at 12-bit resolution in hardware. |
| 2. SP firmware telemetry patch | Needs code |
Change ServiceProcessorBL2 to send the individual readings:
sample.current, which is already truncated to 1 °C.
|
| 3. Firmware signing and flashing | Untested | BL1 and BL2 verify a certificate and signature against a public-key hash in OTP (crypto_rot.c / VaultIP) unless an OTP chicken bit is set, and neither the signing tool nor a test key is in the open tree. Whether these lab cards accept a rebuilt image is unknown until someone tries, and a bad image can brick the boot path, so reflashing is a lab-admin decision (the firmware caveat). |
| 4. Host telemetry extraction | Ready to build | Extend tools/ettelem (which samples libDM at 10 Hz, about 45 Hz at most, and dumps the SP trace buffer on demand with ettelem sptrace) to unpack the 35 temperature channels into a 6×6 matrix. |
| 5. 2D heat-diffusion reconstruction | Not yet fitted | The fitted thermal model is lumped: six RC stages, 1.47 °C/W in total, fitted to one day’s runs on one card (aifoundry2) in its chassis, driven by the chip-wide average. A 2D model (∂T/∂t = α∇²T + P(x, y)/C − (T − Tamb)/τ) needs lateral conductances and per-cell capacitances that only per-shire readings could fit. |
Implementation notes: the PVT controllers' register map
Implementation notes
The five PVT controllers (PVTC 0–4) are mapped in the service processor’s address space at 0x00_5400_0000 + n·0x1_0000, and all five sample continuously at 12 bits (RUN_1 mode):
| Controller | Base address | Sensors | Channels active |
|---|---|---|---|
| PVTC 0 | 0x00_5400_0000 | TS 0–7 | 8 / 8 |
| PVTC 1 | 0x00_5401_0000 | TS 8–15 | 8 / 8 |
| PVTC 2 | 0x00_5402_0000 | TS 16–23 | 8 / 8 |
| PVTC 3 | 0x00_5403_0000 | TS 24–31 | 8 / 8 |
| PVTC 4 | 0x00_5404_0000 | TS 32–34 | 3 / 8 (mask 0xF8) |
4. Current best workaround: temporal reconstruction
Until a modified SP firmware can be flashed (whether these cards require a signed image is untested), the measurements so far rely on temporal reconstruction: reading more out of the timing of the chip-wide average (see the Horace experiment, §8, A model from flips to temperature):
- Sub-degree de-quantization: the moments when the whole-degree average steps are sharp at 10 Hz, and fitting the heating power that reproduces them recovers the temperature below a degree (the Horace experiment, §2).
- Flip-to-heat modelling from the RTL: counting the toggles the operands cause in the open RTL of the multiply-add datapath predicts the board power, and from it the die temperature, with models fitted on aifoundry2 (the Horace experiment, §8).
Both sharpen the chip-wide average; neither says where on the die the heat is.
Sources
Sources, in full
- Firmware: et-platform at 353f20e,
device-bootloaders/src/ServiceProcessorBL2/:driver/pvt_controller.c(pvt_ts_conversion, the per-shire temperature print,pvt_get_and_print),include/bl2_pvt_controller.h(the conversion constants) andservices/thermal_pwr_mgmt.c(get_module_current_temperature). The cards’ own trace strings match an older build (before et-platform commit 60b40c10f, 24 Sep 2024); in that build these two PVT files differ only in their licence header, andget_module_current_temperature()is the same. - RTL: core-et’s Erbium branch at b38a1a3,
rtl/inc/pvt_defines.vh: the same Minion core lineage in a later MCU-class configuration, not the taped-out ET-SoC-1. - Manuals: ET Programmer’s Reference Manual §1.6 and §15.2.3, in et-man.
- Mesh map:
workloads/nocbench/analyze.py(marty1885’s layout and the four empty cells). - Telemetry: the three
*-telemetry.jsonlfiles indocs/reports/data/2026-09-20-power-aifoundry2/(thermal,horaceandhorace2: 3,681 samples, 20 Sep) andper-shire-voltage-idle.jsonin the same directory. The chart’s change points, the check on aifoundry2 and aifoundry3 (every committed ettelem log of 20–24 September, on either card, that records these fields), the three-card check and the grid’s voltages are the two constants thattools/ettelem/host_temp_fields.pyprints (--checkconfirms this page embeds them unchanged). - The three-card check: experiment V3-TEL of the version-3 claims check (items TEL-R, the peak-hold windows, and TEL-Q, the voltage maps;
docs/reports/data/2026-09-25-claims-v3/results/tel.json), three passes per card on aifoundry2, aifoundry3 and aifoundry1-c1 on 25–26 September. Raw data:docs/reports/data/2026-09-25-claims-v3/raw/<card>/tel/p<N>/dbg/(the 14 s reset windowsx3-w1–w4and the DEBUG trace captures); what each file holds is intools/claims-v3/tel/README.md. - Thermal model:
docs/findings/11-thermal-model.md.
Version history
- Versions (one line per date; every wording is in the file’s history). 22 September 2026: first published. 24 September: the tile size, where the firmware truncates to whole degrees, the
pmic_sysfield and the peak-hold minimum and maximum corrected. 25 September: the peak-hold high checked on both cards; in version 3, results resting on one card name it. 26 September (version 4, three cards, V3-TEL): the high 2 or 3 °C above the average at a rise (was about 3), the low 1 or 2 °C under the minimum (was always 1), the idle voltage map not confirmed. 27 September: the review’s per-card table, the sensor-to-host pipeline, detail folded into sections. 28 September: the note’s counts given by both rules.