The effect of overheating
What temperature a processor is built to run at, why its limits sit where they do, whether a limit should watch the average of the die or its hottest spot, what heat does to how fast transistors switch, and why a chip that overheats can stop working and then work again once it cools. The outside literature answers each question first; the ET-SoC-1's firmware and our own measurements then say how the chip on our cards compares.
The owner's request (28 September, lightly edited): “comprehensive outside research on the behavior of throttling and overheating … what are the typical operating requirements for processors like the Esperanto chip, why the temperature limits are what they are … the average versus highest temperature. Should we limit the highest temperature versus the average? I assumed the individual regional temperature is more important … when you're overheating, does it affect the transistor flipping speed? … [the lab lead] said in a meeting that they could get the Esperanto chip to like 120 Celsius and it stops working, but when it goes down it's fine again, so it's surprisingly resilient. I want to know more about that effect: why do things stop working when they overheat and then start working again.”
The answers
- What temperatures are processors built for, and why?
- Limits are set on the hottest point of the die, the “junction”. Consumer and data-centre chips throttle at about 95–105 °C (desktop CPUs 95–100 °C, a GPU's hottest sensor 110 °C, NVIDIA's Jetson Orin 99 °C) and switch themselves off at about 105–125 °C; industrial, military and automotive grades are rated to 105–125 °C and beyond. The throttle point is the temperature up to which the vendor guarantees the chip's timing, electrical specifications and lifetime, and the throttle keeps the die inside it; the shutdown point is where damage becomes possible. The ET-SoC-1's datasheet publishes no limit, and the temperature its timing was signed off at is not published either. On the firmware build of aifoundry2 the clock slows when the average of the sensors passes °C, and at the lowest clock nothing else acts; the other cards run other builds (section 1.3). Section 1
- Should a limit watch the average or the hottest spot?
- The hottest spot: your assumption is right, above all for wear-out, which the hottest region decides. Every hard limit documented in this survey acts on the hottest sensor (IBM's, on the hottest core); averages appear beside it, in fan control and in AMD's boost decisions. Timing is a weaker argument on this chip: at the ET's low core voltage heat barely slows its transistors and may speed them up (section 3). The ET-SoC-1 is the one chip in this survey whose throttle compares an average. On this chip that costs little: Section 2
- Does heat change how fast transistors switch?
- Yes, in both directions. Heat slows the electrons (lower mobility) but makes transistors turn on more easily (lower threshold voltage). At a high supply voltage the first wins and a hot chip is slower; at a low one the second wins and a hot chip is faster. The ET-SoC-1's cores run at about , close to where the two cancel for 7 nm in simulation, so heat should barely change them. And a chip's clock comes from an oscillator, not from its transistors: heat changes how much timing slack each path has, not how much work a cycle does. Section 3
- Why does a hot chip stop working, and work again when it cools?
- Because most of what heat does depends on the temperature now: leakage current, timing slack, how long a memory cell holds its charge, whether a clock oscillator stays in range. Past some point one of these fails, and when the temperature falls it comes back. Protective shutdowns are designed the same way: they switch the chip off before any damage. What does not come back is damage: wear-out, which accumulates slowly and invisibly with time spent hot, and fatigue of the package, which grows with every large temperature swing. A cracked joint can even fail hot and work cool, so recovering on cooling does not by itself prove that nothing was harmed. Section 4
- Is recovering from 120 °C surprising?
- No. The standard qualification test runs chips for 1,000 hours at a junction of at least 125 °C, and none of 231 may fail. Our own What stopped the lab's card near 120 °C is not recorded; the most likely candidate, untested, is power inference: Section 4.4
- So what does heat cost?
- Lifetime, fatigue and power. At the activation energies usually assumed for wear-out (0.7–0.9 eV), an hour at 117 °C ages a chip as much as hours at 65 °C; one swing from room temperature to 120 °C and back fatigues the package's joints times as much as one to 65 °C; and on these cards the leakage part of idle power doubles every °C. Section 5
Terms and labels used on this page
Junction temperature (Tj): the temperature of the transistors, on the die, as opposed to the case or the air. The ET-SoC-1 has one temperature sensor in each of its 34 minion shires (tiles of 32 cores, about 3.7 mm on a side) and one in the I/O shire. The mean is the whole-degree average of the 34 shire sensors, each first cut to a whole degree; the 0.20.0 firmware compares it with 65 °C and first acts when it reads 66. The hottest sensor is the host's peak-hold of the highest single sensor; the experiments here reset it every second, so it is the hottest sensor in that second (which shire it was is not reported). The gap is the hottest sensor minus the mean. Numbers carry their kind: measured on our cards, from source read from the ET firmware or manuals, derived arithmetic on cited numbers, external a published figure, inference reasoning no source states. Citations in brackets link to the numbered list at the end; [E1] and so on are the ET-SoC-1's own documents.
1. What the limits are, and why they sit where they do
Answer. Processors are rated by the hottest point on the die. Vendors publish a ladder rather than one number: a throttle point where the chip slows itself, then a harder clamp, then a shutdown. The throttle point is the temperature up to which the chip's timing, its electrical specifications and its lifetime budget are guaranteed; the shutdown point is where damage becomes possible.
1.1 Typical ratings
Every vendor defines the limit on the silicon itself: Intel's Tj,max is “the maximum temperature on the processor-die-active surface” . Typical figures:
- Consumer and commercial parts: a junction range of 0 to 95 °C (NXP's i.MX 8M Plus ). Industrial: −40 to 105 °C . Military: −55 to 125 °C . Automotive grades are set by the ambient temperature: up to 85, 105, 125 or 150 °C .
- Desktop CPUs throttle at 95–100 °C: AMD's Ryzen 7000 is “designed for a lifetime at 95°C” , Intel's Core i9-13900K lists 100 °C . GPUs let the hottest of many sensors reach 110 °C . NVIDIA's Jetson Orin throttles in software at 99 °C, in hardware at 103 °C, and shuts down at 104.5 and 105 °C .
- A 7 nm accelerator like the ET-SoC-1: AMD's Versal AI Core, also made at TSMC 7 nm, is rated 0 to 100 °C and may spend at most 3% of its life at 100–110 °C . TI designs its embedded processors for ten years at 105 °C, always on .
- The DRAM beside the chip has its own, lower limit: standard LPDDR4X is rated to 85 °C on its case, and above that it must be refreshed two or four times as often .
1.2 Why each rung sits where it does
- The throttle point is the guaranteed operating limit. Vendors guarantee function, timing and lifetime up to Tj,max, and the throttle keeps the die inside it. Intel's Core Duo engineers: “Functionality, electrical specifications and reliability commitments are guaranteed at maximum Tj as measured by the DTS” , and with its thermal monitor switched off “the processor operates out of specification” . Designers close timing there: Skadron and colleagues describe “the standard practice of designing the nominal operating frequency for the maximum allowed operating temperature” . The same temperature sets the lifetime budget: AMD designs for a lifetime at 95 °C , TI for ten years at 105 °C , and Versal allows 3% of its life above 100 °C . The package and cooler are sized for typical work, with the throttle catching the rest ; running cooler also saves leakage power: Intel lists the same Atom at 6.0 W with a 90 °C limit or 6.3 W at 102 °C . For the ET-SoC-1, the temperature its timing was signed off at is not published, and nothing published ties its 65 °C to either guarantee (section 1.3).
- The shutdown point is where damage becomes possible. Intel's THERMTRIP marks “a point beyond which permanent silicon damage may occur” , near 125 °C and calibrated per part . It is for failures of the cooling, not for normal work: a fan failure “will eventually reach THERMTRIP# and shut down” , and Jetson calls its hardware trip “the final failsafe” . AMD's needs “a cold reset” to leave it .
- Qualification sits above both. JEDEC's standard test runs parts powered for 1,000 hours at a junction of at least 125 °C and the maximum supply voltage, and none of 231 may fail; storage is tested at 150 °C . TSMC's 7 nm foundation IP passed the automotive grade-1 tests (ambient up to 125 °C) .
- Guard bands separate the reading from the true hot spot: sensor error (Intel specifies ±5 °C ), a sensor that is not on the hottest transistors (“the temperature observed by the sensor may be cooler by some spatial-gradient factor” ), and the heating possible between two readings (Micron's rule for its DRAM: gradient × delay ≤ 2 °C ). Calibration pushes the error to the safe side .
1.3 The ET-SoC-1's limits
Esperanto's preliminary datasheet leaves the absolute maximum ratings, the operating conditions and the package's thermal data “for a future release” . The only limits are in the service processor's firmware, and the four cards run three builds of it: What the 0.20.0 build does from source:
- Slow the clock one step at a time while the mean of the shire sensors exceeds °C, an “early indication temperature threshold” . It can only step down to the lowest operating point, 600 MHz on these cards, and stops there.
- A hardware alarm in the power-management chip (PMIC) at °C and 75 W, which reports a fatal
event . On 0.20.0 its handler would set the clock's frequency register to
MHz but, per the source, would not reprogram the clock's PLL: it does that only when the
voltage lookup before it fails (the 0.18.0 build has a real 300 MHz safe
state:
docs/findings/14-card-behaviour.md). Its temperature input is the PMIC's own “system temperature”, not the die, and the source notes that the PMIC “is currently reporting system temperature as 0” : Its 75 W threshold did not act either: - No shutdown: nothing in the service processor's firmware switches the die off on temperature. The 52 °C in its header, labelled “Expected average temperature”, is one of the default reset values that variables start from , not a limit.
So the ET's 65 °C is a product-envelope setting, not a known silicon limit. On the 0.20.0 and 0.18.0 cards, once the clock is at 600 MHz, the die's temperature is limited only by its cooling; card 0's drop to 300 MHz is the one exception on record. The highest readings on record measured:
2. Average or hottest: which temperature a limit should watch
Answer. The hottest calibrated sensor, lightly smoothed in time; not the average over the die. Wear-out is decided by the hottest region, every hard limit documented in this survey acts on the hottest sensor or core, and an average hides concentrated heat. Timing follows the hottest region only where heat slows the circuit (above the crossover voltage, and on wires); at the ET's low core voltage that argument is weak or reversed (section 3). On the ET-SoC-1 the difference is small, because its hot spots, as far as its sensors can see, are only a few degrees above the mean.
2.1 Why the hot spot decides
- Failure is local, and exponential in temperature. Reliability engineers treat a processor as a chain: “the first instance of any structure failing due to any failure mechanism causes the entire processor to fail”, and each structure's wear rate follows its own temperature . Electromigration life depends on the local wire temperature, and temperature gradients themselves “foreshorten conductor life” . JEDEC's latch-up group: the “highest Tj on the die … should be used” .
- How much the gap matters derived. Summing wear rates over many regions behaves like a “soft maximum” of their temperatures with a width of kT²/Ea, °C at 90 °C for a 0.9 eV mechanism. Gaps much smaller than that make the mean and the maximum agree; larger gaps make the hottest region decide. With one of 34 regions hotter by 3, 10, 30 or 45 °C, the effective temperature is . Wear-out is not a constant rate, though: qualification targets a small failure fraction, deep in the tail of the lifetime distribution, and there
- Hot spots form fast. Localized heating “occurs much faster than chip-wide heating”, over hundreds of microseconds to milliseconds ; in a simulated 7 nm CPU the first hot spot formed 0.2 ms after a workload started, and units above 120 °C sat 200 µm from units 30 °C cooler .
2.2 What vendors do
| Chip | What is limited or reported | Where averages are used | Source |
|---|---|---|---|
| Intel Core (12th–14th gen) | “any Digital Thermal Sensor” at TjMAX starts the thermal control circuit | a 256 ms average of the hottest sensor, for fan control; an averaged loop below TjMAX “in addition, and not instead of” it | |
| AMD Radeon RX 5700 (GPU) | “any one of the many available sensors” at 110 °C, replacing one sensor near the old diode | — | (secondary) |
| AMD Instinct MI300 (Linux driver) | reported only: the driver exposes limits on the junction (hotspot) and memory channels, and on MI300 hides the edge channel; what the throttle compares is not in these lines | — | |
| AMD Ryzen (CPU) | the hard limit: not found in an AMD document | boost decisions on “a short-duration rolling average of all temp sensors” | (secondary) |
| IBM POWER9 | firmware only, like the ET: its controller uses “the hottest core temperature”, each core's value being a weighted average of the sensors in and beside that core | a chip average is kept, but it is not the control input | |
| Intel Gaudi (accelerator) | reported only: the monitoring tool's “maximum temperature read from the four available temperature sensors”; what the throttle compares is not stated | — | |
| NVIDIA Jetson Orin | each sensor group reports “the maximum of all the sensors in the group” | a weighted average only for the fan's target | |
| ET-SoC-1 (firmware ) | the whole-degree mean of the shire sensors; no hottest-sensor path; calibration fuses not read | — |
So the hard limits documented here (Intel, the RX 5700, POWER9, Jetson) all act on the hottest sensor or core; AMD's Ryzen also uses an average of all its sensors, for boost. Where sources disagree: AMD's CPU note says its boost loop uses an average of all sensors while its GPU note throttles on any one sensor; the two fit if the average drives the boost and a maximum the hard limit, but no AMD document found says what the Ryzen hard limit compares. NVIDIA does not say what its “GPU Current Temp” is .
2.3 What goes wrong with an average
- It hides concentrated heat. If one of 34 regions heats and the rest do not, a 1 °C rise of the mean can hide up to a °C rise in that region derived, an upper bound, since heat spreads sideways.
- It forces a guard band on every workload, and fails beyond it. Intel, on its older single-diode design: the hot spot ran up to about 10 °C above the diode; a fixed offset cost 3–7% of performance, and without one “some of the workloads will run at high max Tj and therefore risk functional issues or reliability degradation” . A uniform 4 × 4 sensor grid under-read the maximum by up to 10.5 °C in simulation .
- Cooling faults show up at the hot spot first. A GPU with dried thermal paste reported 67–68 °C while its hot spot reached 107 °C and throttled; in another, a vapor chamber short of fluid drove the hot spot to 110 °C (both secondary). In both the limit that acted was the one on the hot spot.
- When an average is fine: in slow loops whose physics really is averaged (fan speed, heat-sink sizing, boost), as a per-sensor filter in time to suppress noise, and as a spatial mean only with a guard band equal to the worst gap plus sensor error .
2.4 The gap on the ET-SoC-1
The host cannot read one shire's temperature: it sees the mean, the highest and lowest single reading (anonymous peak-holds) and the I/O shire's sensor . E53 reset the peak-hold every second, so each 1 s window gives the hottest sensor in that second. OH-1 put the most concentrated load the heater can make, one shire at full load (32 cores, ), in the centre and by the I/O corner, against a central 2 × 2 block and idle; OH-2 heated the whole chip from rest to °C measured.
- At the 0.20.0 rule's decision point:
- Concentrating the load does not widen the gap:
- The gap grows a little with temperature:
- The worst cases on record:
What this means for the ET inference. For every load tried, mean > 65 behaves like
“the hottest sensor reads about °C”: the choice of the mean costs of margin, not safety
(DV2's runs on aifoundry2, the card where the rule does act, had gaps of the same size: the last item above). The chip's low power
density is the likely reason the gap is small: What the sensors cannot say is how hot it
gets inside a 3.7 mm tile: there is one sensor per tile, in one corner , and
the literature's sensor-to-hot-spot offsets are 3–10 °C on denser chips . The
bigger gaps are elsewhere: on the 0.20.0 and 0.18.0 builds nothing acts at 600 MHz, and no temperature is measured in the
memory shires or the DRAM (section 4.2).
3. Does heat change how fast transistors switch?
Answer. Yes, in both directions, and which way depends on the supply voltage. But in a clocked chip the clock does not follow: heat changes each path's timing margin, not the work done per cycle, until a margin runs out.
- Two effects pull against each other. Heat lowers the carriers' mobility (less current: slower) and lowers the threshold voltage (more overdrive: faster). “At high voltages, higher temperature increases the delay but at low voltages, higher temperature decreases the delay”; the voltage where they cancel is the crossover or zero-temperature-coefficient (ZTC) voltage . This temperature inversion “will worsen as supply voltages decrease” .
- Where the crossover lies. About 1.12 V in a 90 nm library ; in a 7 nm FinFET kit 0.49–0.55 V for single cells and about 0.53 V for a whole processor, with 0.7 V nominal (simulated) . Disagreement: another simulation of 20–10 nm FinFETs finds hot faster at every voltage from 0.3 to 0.9 V . On silicon, a 28 nm chip needed 0.48 V at 25 °C but only 0.44 V at 80 °C for the same 50 MHz , and an AMD processor used the extra margin when warm to run more than 5% lower in voltage . TSMC's N7 crossover is not public. Skadron and colleagues (2003) tie clock frequency linearly to temperature through carrier mobility and cite Garrett and Stan: “an 18% variation over the range 0-100” .
- Wires go the other way at any voltage. Copper's resistance rises 0.3–0.4% per degree , so wire-heavy paths slow when hot even where transistors speed up .
- Consequences for design. With inversion the slowest corner is the cold one (at 16 and 7 nm: slow process, low voltage, low temperature ), and the hot, fast corner is where hold violations appear: a new value arrives before the previous one was captured. A slower clock cures a setup violation but not a hold violation .
3.1 Where the ET-SoC-1 sits
Esperanto designed the chip to run at low voltage: its cores are built for about 0.4 V . On our cards at 600 MHz measured: Against the simulated 7 nm crossover inference: the cores at 600 MHz sit at it, so heat should barely change their gate delay and, if anything, shorten it; the mesh sits just below it (slightly faster hot); the SRAM rail and the 800 MHz point sit above it (slower hot), and so do the wires.
3.2 What the cards show
- The same work per cycle, hot or cool measured:
- The clock does not drift:
- Heat does not starve the supply:
- The only failures on record were at clock changes, not heat:
What could not be measured. Whether the ET's own transistors get faster or slower when hot needs one of its on-chip speed monitors. It has one ring-oscillator process detector beside each temperature sensor , but the firmware configures them as measurement-disabled and nothing reads them, and only the service processor can reach them : a firmware build is needed. The design documents say the clock-distribution delay “may change depending on voltage, temperature, process” . E53 registered its prediction before any data: on the core rail a detector's count changes by less than ±2% between 55 and 85 °C, more likely up (faster) than down; on a rail at 0.70 V or more it falls, by less than 5%.
4. Why a hot chip stops working, and works again when it cools
Answer. Almost everything heat does is a function of the temperature now. Mobility, threshold voltage, leakage, how long a memory cell keeps its charge and an oscillator's frequency all return to their old values when the temperature does. A failure caused only by their values at that moment (a parametric failure) disappears on cooling; a protective shutdown is designed to be reversible; some states need a reset. What stays is damage: wear-out, a destructive latch-up, and fatigue of the package, whose cracks can open when hot and close when cool, so a fault that comes and goes with temperature is not always harmless.
4.1 The mechanisms, and which ones reverse
| What fails | Why heat triggers it | Recovers when cooled? | On the ET-SoC-1 |
|---|---|---|---|
| Setup timing | a path gets slower than the clock period: above the crossover voltage, and on wires | yes | not predicted for the cores at ; possible on the SRAM rail and I/O (inference) |
| Hold timing | below the crossover a short path gets faster than the hold window | yes, but a slower clock does not help | hot is the cores' and mesh's fast corner (inference) |
| SRAM margins | leakage of the cell's off transistors rises exponentially | yes; a flipped bit stays wrong | no errors seen (section 4.3) |
| DRAM retention | cells leak faster: retention time −39% to −47% per +10 °C | yes; lost data stays lost | no temperature read, refresh fixed at 1× (section 4.2) |
| Power delivery | leakage current adds to the load ; a regulator or the board input could then reach its limit inference | yes, once the supply recovers (inference) | the leading candidate for a stop near 120 °C (section 4.4) |
| Clock (PLL) lock | an oscillator's frequency drifts out of its tuning range: “the PLL may lose lock” | yes, after it relocks | the service processor counts lock losses |
| Latch-up | a parasitic thyristor triggers more easily when hot | only after a power cycle; can be destructive | not seen |
| Package and joints (thermo-mechanical) | every large temperature swing fatigues the package, most at the solder joints ; temperature change, its rate and spatial gradients are possible drivers of failure beside the steady temperature | the symptom sometimes (a crack that opens hot can close cool: inference); the damage never | not known; each excursion to 110–120 °C and back is a large cycle (section 5.1) |
| Protective trip | designed shutdown before damage | yes, after cooling and a reset | only the PMIC alarm, on a reading that is not the die (section 1.3) |
| Wear-out | electromigration, oxide breakdown, the permanent part of BTI accumulate faster when hot | no; part of BTI recovers slowly | design life and qualification not published |
4.2 The DRAM: no thermometer, no derating
No temperature is measured in the memory shires or the DRAM ; the LPDDR4X's own temperature register is never read, and the memory controllers' temperature derating is commented out: “FUTURE derate is not needed for bring-up” from source. Refresh is fixed at . So if the DRAM packages ever passed 85 °C, weak cells could lose data at the fixed rate, and nothing on the card would notice inference; whether they get that hot when the die is at 110–120 °C is unknown.
4.3 Did the ET-SoC-1 compute anything wrong when hot?
4.4 The 120 °C story
The lab lead said in a meeting that an ET-SoC-1 was taken to about 120 °C, stopped working, and was fine again once it cooled. That episode is not recorded in this repository: which card, what “stopped” meant (a hang, wrong results, a lost PCIe link, a power trip) and whether a reset was needed are open questions. What the repository does record:
- Our own card that went there measured:
- Why recovering is expected. 120 °C is inside the envelope a normal part is qualified in: 1,000 hours powered at 125 °C or more with no failures allowed , military and automotive ratings to 125 °C , and below Intel's trip point near 125 °C . Stopping there and recovering when cool is most consistent with a parametric failure or a protection action; an intermittent mechanical fault, such as a cracked joint or bump that opens only when hot, cannot be excluded without the episode's record, and every such excursion is a large thermal cycle (section 5.1). The cost is lifetime: by TI's table, a processor designed for 105 °C gets 30% of its life if run continuously at 120 °C .
- What could have stopped it, ranked, none tested inference:
- Power.
- A protection action on newer firmware: card 0's drop to 300 MHz after 115–117 °C (on 1.4.1, by its idle point or its PMIC's safe state: not established), or a PMIC fatal event .
- The PCIe link.
- DRAM retention at the fixed 1× refresh, if the DRAM passed 85 °C (section 4.2).
- A hold-time or SRAM failure in the low-voltage logic, which a slower clock would not cure.
- A clock losing lock; the service processor counts these .
- A joint or bump that opens when hot (thermo-mechanical, section 4.1): the one cause on this list that is damage.
- What to record next time: board power and rail currents up to the stop, the clock, the PMIC events, the PLL lock-loss counters, the PCIe link's error counters and whether it dropped, the mean and the hottest sensor, and whether the card came back by itself or needed a reset.
5. What heat costs: lifetime, fatigue and leakage
5.1 Lifetime
Wear-out mechanisms follow the Arrhenius law: the rate grows as exp(−Ea/kT), so each mechanism's activation energy Ea sets how steeply heat shortens life . Electromigration in copper: 0.9 eV in the RAMP model and 0.8–1.2 eV in a recent review , the mechanism TI calls critical for its processors ; 0.7 eV is JEDEC's and TI's working value ; gate-oxide breakdown is driven more by voltage ; bias-temperature instability is accelerated by temperature and gate bias, and part of it recovers when the bias is removed ; hot-carrier damage has been worst cold, and JEDEC tests it at 50 °C or below , but at the reduced supply voltages of leading-edge devices “lower temperatures do not necessarily lead to accelerated degradation” , which is the ET's regime. Measured on a 16 nm FinFET FPGA, ageing slowed ring oscillators by 0–1% after 8,000 hours, a little over 2% at 115 °C and 1.15 × the nominal voltage .
TI writes the rule “with slippage at higher temperatures” , and Pecht and colleagues warn that “no simple expression can adequately describe temperature as a failure accelerator” . Far above any of this, solder melts at 217 °C and reflow peaks at 260 °C . Esperanto's design life and qualification temperature are not published.
Swings count, not only heat. Packages also fail by fatigue: “Damage accumulates every time there is a cycle in temperature”, most at the solder joints between the die and the package, and the large cycles are power-ups, power-downs and low-power modes . Pecht and colleagues name temperature change, its rate and spatial gradients beside the steady temperature as possible drivers of failure, and note that which one drives a given mechanism has generally not been quantified .
5.2 Leakage, and whether it can run away
Subthreshold leakage grows exponentially with temperature, while gate leakage barely moves ; in a 90 nm FPGA it rose five-fold from 25 to 85 °C , a doubling every °C derived. Because leakage heats the die and heat raises leakage, the loop can run away if the cooling cannot keep up .
With the heatsink removed the cooling could not keep up: Feasibility of running the ET-SoC-1 without its heatsink estimates with a Monte Carlo model that a bare 600 MHz card almost never settles at idle (1 of 3,600 draws), and that bare aifoundry2 in still air passes 90 °C 1–3 minutes after a cold power-on at idle (card 1 in 41–99 s).
6. What this means for the ET-SoC-1
| Question | The ET-SoC-1 as found | Common practice |
|---|---|---|
| What the throttle compares | on 0.20.0, the mean of 34 whole-degree sensors (costs here); it acts on aifoundry2, is latched on aifoundry3, and card 1's 0.18.0 build never moves its clock | the hottest calibrated sensor, with a time-averaged loop on top |
| Sensor calibration | generic conversion; the per-sensor calibration fuses are not read | each sensor calibrated at test, error pushed to the safe side |
| What happens at the lowest clock | on the 0.20.0 and 0.18.0 builds, nothing: a die at 600 MHz can heat without limit; card 0 (1.4.1) did drop to 300 MHz after 115–117 °C | clock modulation, harder clamps, then shutdown |
| A hardware trip | a PMIC alarm on a reading that is not the die and reads 0; on 0.20.0 its handler does not reprogram the PLL, and its 75 W threshold did not act at W | a per-part calibrated trip on the die, near 105–125 °C |
| The DRAM | no temperature, derating off, refresh fixed at 1× | read the DRAM's own sensor; refresh 2× or 4× above 85 °C |
| What the host can see | the mean, anonymous peak-holds; the driver's thermal counters never move on the 0.20.0 build | per-sensor readings and throttle counters |
What would match practice inference, not tested: a hard limit on the hottest calibrated sensor, smoothed over a few readings, with a margin for what one sensor per tile cannot see; an action below 600 MHz (clock gating or duty cycling, then shutdown); DRAM temperature polling and derating; and the mean kept as the DVFS loop's input, the role Intel and IBM give their averages. The measured gap says the first matters less on this chip than the second and third.
Asks (not answerable from the repository): the lab lead's ~120 °C episode (which card and firmware, what stopped, whether a reset was needed, whether power and the PCIe link were logged); card 0's own record from 25 September; the DRAM part and its temperature grade; what the PMIC's “system temperature” measures and whether the PMIC power-cycles the board; Esperanto's Tj max, qualification and design life; and a firmware build that enables a process-detector delay chain, if the transistor-speed question should be measured.
7. The experiments on the cards (E53)
How it ran: the freeze, two stopped attempts and their amendments, and the card time
- The freeze.
- Amendment 1.
- Amendment 2.
- Card time.
- The rules.
8. Method, data and how to reproduce
What each number is, and its limits
- Outside research. Four research tracks (limits; average against maximum; speed, failure and recovery; the ET-SoC-1's firmware, manuals and data) read the sources below on 28 September 2026. Each outside claim cites the document it is from; a source read only through a press report or a search engine's copy is flagged in the list.
- ET numbers. Every number from our cards comes from
overheat.json, whichbuild_overheat_data.pywrites from E53's reductions, from analyses of the existing record (scripts/, outputs inanalysis/) and from arithmetic on cited inputs (scripts/derived.py). Numbers read from the firmware or manuals cite the file and line. - Whole degrees. Every die temperature is the host's whole-degree reading; each sensor is truncated before the mean, so both the mean and the gap carry up to a degree of rounding.
- The illustration. The delay curves of section 3 come from a textbook model with chosen parameters, not from this chip or TSMC's process; they show the shape of temperature inversion only.
- The record.
Data. Everything is in docs/reports/data/2026-09-28-overheating:
E53's raw blocks (raw/), reductions (reductions/), the analyses of the existing record
(scripts/, analysis/), the source list (sources.json) and this page's data
(overheat.json). The tools are tools/claims-v3/oh,
with the frozen predictions in prereg/PREREG.md.
V=docs/reports/data/2026-09-28-overheating
python3 tools/claims-v3/oh/reduce.py --all --data $V/raw --out $V/reductions # E53's registered verdicts
python3 $V/extras.py # its descriptive numbers
for s in max_temps correct_vs_temp timing_vs_temp hot_minus_mean idle_vs_temp runaway events_vs_temp derived; do
python3 $V/scripts/$s.py > $V/analysis/$s.txt; done # the existing record, and arithmetic
python3 $V/build_overheat_data.py # overheat.json
python3 scripts/build-report.py effect-of-overheating $V/overheat.json docs/reports/2026-09-28-effect-of-overheating.html
Sources
Outside sources, numbered in the order the page first cites them; all fetched on 28 September 2026. Each entry names what the page uses from it.
The ET-SoC-1's own documents (firmware at the commits named, in aifoundry-org/et-platform; manuals in aifoundry-org/et-man):
Related reports
- The DVFS loop and its leakage — the governor on the cards' firmware, and leakage at 80 °C.
- Where the work sits — placement and the time to the governor's 66 °C step (E52).
- Feasibility of running the ET-SoC-1 without its heatsink — §5.2's leakage loop with the heatsink off: in its model a 600 MHz card almost never settles at idle.
- Spatial temperature — the 35 sensors, and why the host sees one mean.
- Why the chip is low power — the low-voltage design point.
- Limits of observability — the hub: every report, and what the meters cannot see.
- docs/findings/ — the knowledge base: this is experiment E53.