The effect of overheating

28 September 2026 · outside research ( sources), checked on the ET-SoC-1: its firmware source, the record of four cards, and two pre-registered experiments run the same day on aifoundry3 and aifoundry1 card 1 (E53) · part of the ET-SoC-1 measurement reports

What temperature a processor is built to run at, why its limits sit where they do, whether a limit should watch the average of the die or its hottest spot, what heat does to how fast transistors switch, and why a chip that overheats can stop working and then work again once it cools. The outside literature answers each question first; the ET-SoC-1's firmware and our own measurements then say how the chip on our cards compares.

The owner's request (28 September, lightly edited): “comprehensive outside research on the behavior of throttling and overheating … what are the typical operating requirements for processors like the Esperanto chip, why the temperature limits are what they are … the average versus highest temperature. Should we limit the highest temperature versus the average? I assumed the individual regional temperature is more important … when you're overheating, does it affect the transistor flipping speed? … [the lab lead] said in a meeting that they could get the Esperanto chip to like 120 Celsius and it stops working, but when it goes down it's fine again, so it's surprisingly resilient. I want to know more about that effect: why do things stop working when they overheat and then start working again.”

The answers

What temperatures are processors built for, and why?
Limits are set on the hottest point of the die, the “junction”. Consumer and data-centre chips throttle at about 95–105 °C (desktop CPUs 95–100 °C, a GPU's hottest sensor 110 °C, NVIDIA's Jetson Orin 99 °C) and switch themselves off at about 105–125 °C; industrial, military and automotive grades are rated to 105–125 °C and beyond. The throttle point is the temperature up to which the vendor guarantees the chip's timing, electrical specifications and lifetime, and the throttle keeps the die inside it; the shutdown point is where damage becomes possible. The ET-SoC-1's datasheet publishes no limit, and the temperature its timing was signed off at is not published either. On the firmware build of aifoundry2 the clock slows when the average of the sensors passes °C, and at the lowest clock nothing else acts; the other cards run other builds (section 1.3). Section 1
Should a limit watch the average or the hottest spot?
The hottest spot: your assumption is right, above all for wear-out, which the hottest region decides. Every hard limit documented in this survey acts on the hottest sensor (IBM's, on the hottest core); averages appear beside it, in fan control and in AMD's boost decisions. Timing is a weaker argument on this chip: at the ET's low core voltage heat barely slows its transistors and may speed them up (section 3). The ET-SoC-1 is the one chip in this survey whose throttle compares an average. On this chip that costs little: Section 2
Does heat change how fast transistors switch?
Yes, in both directions. Heat slows the electrons (lower mobility) but makes transistors turn on more easily (lower threshold voltage). At a high supply voltage the first wins and a hot chip is slower; at a low one the second wins and a hot chip is faster. The ET-SoC-1's cores run at about , close to where the two cancel for 7 nm in simulation, so heat should barely change them. And a chip's clock comes from an oscillator, not from its transistors: heat changes how much timing slack each path has, not how much work a cycle does. Section 3
Why does a hot chip stop working, and work again when it cools?
Because most of what heat does depends on the temperature now: leakage current, timing slack, how long a memory cell holds its charge, whether a clock oscillator stays in range. Past some point one of these fails, and when the temperature falls it comes back. Protective shutdowns are designed the same way: they switch the chip off before any damage. What does not come back is damage: wear-out, which accumulates slowly and invisibly with time spent hot, and fatigue of the package, which grows with every large temperature swing. A cracked joint can even fail hot and work cool, so recovering on cooling does not by itself prove that nothing was harmed. Section 4
Is recovering from 120 °C surprising?
No. The standard qualification test runs chips for 1,000 hours at a junction of at least 125 °C, and none of 231 may fail. Our own What stopped the lab's card near 120 °C is not recorded; the most likely candidate, untested, is power inference: Section 4.4
So what does heat cost?
Lifetime, fatigue and power. At the activation energies usually assumed for wear-out (0.7–0.9 eV), an hour at 117 °C ages a chip as much as hours at 65 °C; one swing from room temperature to 120 °C and back fatigues the package's joints times as much as one to 65 °C; and on these cards the leakage part of idle power doubles every °C. Section 5
The hottest sensor when the mean reads °C, the first reading the 0.20.0 rule acts on
The largest gap, hottest sensor over the mean, in any 1 s window
Wrong results among exact-checked launches, hot and cool
Work per clock cycle, hot against rest
Terms and labels used on this page

Junction temperature (Tj): the temperature of the transistors, on the die, as opposed to the case or the air. The ET-SoC-1 has one temperature sensor in each of its 34 minion shires (tiles of 32 cores, about 3.7 mm on a side) and one in the I/O shire. The mean is the whole-degree average of the 34 shire sensors, each first cut to a whole degree; the 0.20.0 firmware compares it with 65 °C and first acts when it reads 66. The hottest sensor is the host's peak-hold of the highest single sensor; the experiments here reset it every second, so it is the hottest sensor in that second (which shire it was is not reported). The gap is the hottest sensor minus the mean. Numbers carry their kind: measured on our cards, from source read from the ET firmware or manuals, derived arithmetic on cited numbers, external a published figure, inference reasoning no source states. Citations in brackets link to the numbered list at the end; [E1] and so on are the ET-SoC-1's own documents.

1. What the limits are, and why they sit where they do

Answer. Processors are rated by the hottest point on the die. Vendors publish a ladder rather than one number: a throttle point where the chip slows itself, then a harder clamp, then a shutdown. The throttle point is the temperature up to which the chip's timing, its electrical specifications and its lifetime budget are guaranteed; the shutdown point is where damage becomes possible.

1.1 Typical ratings

Every vendor defines the limit on the silicon itself: Intel's Tj,max is “the maximum temperature on the processor-die-active surface” . Typical figures:

The ladder of limits: where each design throttles, is rated and shuts down, and where the ET-SoC-1 has been

1.2 Why each rung sits where it does

1.3 The ET-SoC-1's limits

Esperanto's preliminary datasheet leaves the absolute maximum ratings, the operating conditions and the package's thermal data “for a future release” . The only limits are in the service processor's firmware, and the four cards run three builds of it: What the 0.20.0 build does from source:

So the ET's 65 °C is a product-envelope setting, not a known silicon limit. On the 0.20.0 and 0.18.0 cards, once the clock is at 600 MHz, the die's temperature is limited only by its cooling; card 0's drop to 300 MHz is the one exception on record. The highest readings on record measured:

2. Average or hottest: which temperature a limit should watch

Answer. The hottest calibrated sensor, lightly smoothed in time; not the average over the die. Wear-out is decided by the hottest region, every hard limit documented in this survey acts on the hottest sensor or core, and an average hides concentrated heat. Timing follows the hottest region only where heat slows the circuit (above the crossover voltage, and on wires); at the ET's low core voltage that argument is weak or reversed (section 3). On the ET-SoC-1 the difference is small, because its hot spots, as far as its sensors can see, are only a few degrees above the mean.

2.1 Why the hot spot decides

2.2 What vendors do

ChipWhat is limited or reportedWhere averages are usedSource
Intel Core (12th–14th gen)“any Digital Thermal Sensor” at TjMAX starts the thermal control circuita 256 ms average of the hottest sensor, for fan control; an averaged loop below TjMAX “in addition, and not instead of” it
AMD Radeon RX 5700 (GPU)“any one of the many available sensors” at 110 °C, replacing one sensor near the old diode— (secondary)
AMD Instinct MI300 (Linux driver)reported only: the driver exposes limits on the junction (hotspot) and memory channels, and on MI300 hides the edge channel; what the throttle compares is not in these lines—
AMD Ryzen (CPU)the hard limit: not found in an AMD documentboost decisions on “a short-duration rolling average of all temp sensors” (secondary)
IBM POWER9firmware only, like the ET: its controller uses “the hottest core temperature”, each core's value being a weighted average of the sensors in and beside that corea chip average is kept, but it is not the control input
Intel Gaudi (accelerator)reported only: the monitoring tool's “maximum temperature read from the four available temperature sensors”; what the throttle compares is not stated—
NVIDIA Jetson Orineach sensor group reports “the maximum of all the sensors in the group”a weighted average only for the fan's target
ET-SoC-1 (firmware )the whole-degree mean of the shire sensors; no hottest-sensor path; calibration fuses not read—

So the hard limits documented here (Intel, the RX 5700, POWER9, Jetson) all act on the hottest sensor or core; AMD's Ryzen also uses an average of all its sensors, for boost. Where sources disagree: AMD's CPU note says its boost loop uses an average of all sensors while its GPU note throttles on any one sensor; the two fit if the average drives the boost and a maximum the hard limit, but no AMD document found says what the Ryzen hard limit compares. NVIDIA does not say what its “GPU Current Temp” is .

2.3 What goes wrong with an average

2.4 The gap on the ET-SoC-1

The host cannot read one shire's temperature: it sees the mean, the highest and lowest single reading (anonymous peak-holds) and the I/O shire's sensor . E53 reset the peak-hold every second, so each 1 s window gives the hottest sensor in that second. OH-1 put the most concentrated load the heater can make, one shire at full load (32 cores, ), in the centre and by the I/O corner, against a central 2 × 2 block and idle; OH-2 heated the whole chip from rest to °C measured.

How far the hottest sensor leads the mean, by load: share of 1 s windows at each gap

The same under the whole-chip heater and its rests, by the die's mean temperature

What this means for the ET inference. For every load tried, mean > 65 behaves like “the hottest sensor reads about °C”: the choice of the mean costs of margin, not safety (DV2's runs on aifoundry2, the card where the rule does act, had gaps of the same size: the last item above). The chip's low power density is the likely reason the gap is small: What the sensors cannot say is how hot it gets inside a 3.7 mm tile: there is one sensor per tile, in one corner , and the literature's sensor-to-hot-spot offsets are 3–10 °C on denser chips . The bigger gaps are elsewhere: on the 0.20.0 and 0.18.0 builds nothing acts at 600 MHz, and no temperature is measured in the memory shires or the DRAM (section 4.2).

3. Does heat change how fast transistors switch?

Answer. Yes, in both directions, and which way depends on the supply voltage. But in a clocked chip the clock does not follow: heat changes each path's timing margin, not the work done per cycle, until a margin runs out.

How a logic gate's delay moves with temperature at the ET-SoC-1's rail voltages (an illustration)

3.1 Where the ET-SoC-1 sits

Esperanto designed the chip to run at low voltage: its cores are built for about 0.4 V . On our cards at 600 MHz measured: Against the simulated 7 nm crossover inference: the cores at 600 MHz sit at it, so heat should barely change their gate delay and, if anything, shorten it; the mesh sits just below it (slightly faster hot); the SRAM rail and the 800 MHz point sit above it (slower hot), and so do the wires.

3.2 What the cards show

What could not be measured. Whether the ET's own transistors get faster or slower when hot needs one of its on-chip speed monitors. It has one ring-oscillator process detector beside each temperature sensor , but the firmware configures them as measurement-disabled and nothing reads them, and only the service processor can reach them : a firmware build is needed. The design documents say the clock-distribution delay “may change depending on voltage, temperature, process” . E53 registered its prediction before any data: on the core rail a detector's count changes by less than ±2% between 55 and 85 °C, more likely up (faster) than down; on a rail at 0.70 V or more it falls, by less than 5%.

4. Why a hot chip stops working, and works again when it cools

Answer. Almost everything heat does is a function of the temperature now. Mobility, threshold voltage, leakage, how long a memory cell keeps its charge and an oscillator's frequency all return to their old values when the temperature does. A failure caused only by their values at that moment (a parametric failure) disappears on cooling; a protective shutdown is designed to be reversible; some states need a reset. What stays is damage: wear-out, a destructive latch-up, and fatigue of the package, whose cracks can open when hot and close when cool, so a fault that comes and goes with temperature is not always harmless.

4.1 The mechanisms, and which ones reverse

What failsWhy heat triggers itRecovers when cooled?On the ET-SoC-1
Setup timinga path gets slower than the clock period: above the crossover voltage, and on wires yesnot predicted for the cores at ; possible on the SRAM rail and I/O (inference)
Hold timingbelow the crossover a short path gets faster than the hold window yes, but a slower clock does not help hot is the cores' and mesh's fast corner (inference)
SRAM marginsleakage of the cell's off transistors rises exponentially yes; a flipped bit stays wrongno errors seen (section 4.3)
DRAM retentioncells leak faster: retention time −39% to −47% per +10 °C yes; lost data stays lostno temperature read, refresh fixed at 1× (section 4.2)
Power deliveryleakage current adds to the load ; a regulator or the board input could then reach its limit inferenceyes, once the supply recovers (inference)the leading candidate for a stop near 120 °C (section 4.4)
Clock (PLL) lockan oscillator's frequency drifts out of its tuning range: “the PLL may lose lock” yes, after it relocksthe service processor counts lock losses
Latch-upa parasitic thyristor triggers more easily when hot only after a power cycle; can be destructivenot seen
Package and joints (thermo-mechanical)every large temperature swing fatigues the package, most at the solder joints ; temperature change, its rate and spatial gradients are possible drivers of failure beside the steady temperature the symptom sometimes (a crack that opens hot can close cool: inference); the damage nevernot known; each excursion to 110–120 °C and back is a large cycle (section 5.1)
Protective tripdesigned shutdown before damage yes, after cooling and a resetonly the PMIC alarm, on a reading that is not the die (section 1.3)
Wear-outelectromigration, oxide breakdown, the permanent part of BTI accumulate faster when hot no; part of BTI recovers slowly design life and qualification not published
What heat does, by temperature, and whether it undoes itself

4.2 The DRAM: no thermometer, no derating

No temperature is measured in the memory shires or the DRAM ; the LPDDR4X's own temperature register is never read, and the memory controllers' temperature derating is commented out: “FUTURE derate is not needed for bring-up” from source. Refresh is fixed at . So if the DRAM packages ever passed 85 °C, weak cells could lose data at the fixed rate, and nothing on the card would notice inference; whether they get that hot when the die is at 110–120 °C is unknown.

4.3 Did the ET-SoC-1 compute anything wrong when hot?

Exact-checked launches by die temperature and card: the record before E53, and E53

4.4 The 120 °C story

The lab lead said in a meeting that an ET-SoC-1 was taken to about 120 °C, stopped working, and was fine again once it cooled. That episode is not recorded in this repository: which card, what “stopped” meant (a hang, wrong results, a lost PCIe link, a power trip) and whether a reset was needed are open questions. What the repository does record:

5. What heat costs: lifetime, fatigue and leakage

5.1 Lifetime

Wear-out mechanisms follow the Arrhenius law: the rate grows as exp(−Ea/kT), so each mechanism's activation energy Ea sets how steeply heat shortens life . Electromigration in copper: 0.9 eV in the RAMP model and 0.8–1.2 eV in a recent review , the mechanism TI calls critical for its processors ; 0.7 eV is JEDEC's and TI's working value ; gate-oxide breakdown is driven more by voltage ; bias-temperature instability is accelerated by temperature and gate bias, and part of it recovers when the bias is removed ; hot-carrier damage has been worst cold, and JEDEC tests it at 50 °C or below , but at the reduced supply voltages of leading-edge devices “lower temperatures do not necessarily lead to accelerated degradation” , which is the ET's regime. Measured on a 16 nm FinFET FPGA, ageing slowed ring oscillators by 0–1% after 8,000 hours, a little over 2% at 115 °C and 1.15 × the nominal voltage .

How much faster a chip wears out than at the ET-SoC-1's 65 °C threshold, by activation energy

TI writes the rule “with slippage at higher temperatures” , and Pecht and colleagues warn that “no simple expression can adequately describe temperature as a failure accelerator” . Far above any of this, solder melts at 217 °C and reflow peaks at 260 °C . Esperanto's design life and qualification temperature are not published.

Swings count, not only heat. Packages also fail by fatigue: “Damage accumulates every time there is a cycle in temperature”, most at the solder joints between the die and the package, and the large cycles are power-ups, power-downs and low-power modes . Pecht and colleagues name temperature change, its rate and spatial gradients beside the steady temperature as possible drivers of failure, and note that which one drives a given mechanism has generally not been quantified .

5.2 Leakage, and whether it can run away

Subthreshold leakage grows exponentially with temperature, while gate leakage barely moves ; in a 90 nm FPGA it rose five-fold from 25 to 85 °C , a doubling every °C derived. Because leakage heats the die and heat raises leakage, the loop can run away if the cooling cannot keep up .

Idle board power against the die's mean temperature: the September laws, E53's idle stretches, and card 0

With the heatsink removed the cooling could not keep up: Feasibility of running the ET-SoC-1 without its heatsink estimates with a Monte Carlo model that a bare 600 MHz card almost never settles at idle (1 of 3,600 draws), and that bare aifoundry2 in still air passes 90 °C 1–3 minutes after a cold power-on at idle (card 1 in 41–99 s).

6. What this means for the ET-SoC-1

QuestionThe ET-SoC-1 as foundCommon practice
What the throttle compareson 0.20.0, the mean of 34 whole-degree sensors (costs here); it acts on aifoundry2, is latched on aifoundry3, and card 1's 0.18.0 build never moves its clockthe hottest calibrated sensor, with a time-averaged loop on top
Sensor calibrationgeneric conversion; the per-sensor calibration fuses are not read each sensor calibrated at test, error pushed to the safe side
What happens at the lowest clockon the 0.20.0 and 0.18.0 builds, nothing: a die at 600 MHz can heat without limit; card 0 (1.4.1) did drop to 300 MHz after 115–117 °Cclock modulation, harder clamps, then shutdown
A hardware tripa PMIC alarm on a reading that is not the die and reads 0; on 0.20.0 its handler does not reprogram the PLL, and its 75 W threshold did not act at Wa per-part calibrated trip on the die, near 105–125 °C
The DRAMno temperature, derating off, refresh fixed at 1×read the DRAM's own sensor; refresh 2× or 4× above 85 °C
What the host can seethe mean, anonymous peak-holds; the driver's thermal counters never move on the 0.20.0 build per-sensor readings and throttle counters

What would match practice inference, not tested: a hard limit on the hottest calibrated sensor, smoothed over a few readings, with a margin for what one sensor per tile cannot see; an action below 600 MHz (clock gating or duty cycling, then shutdown); DRAM temperature polling and derating; and the mean kept as the DVFS loop's input, the role Intel and IBM give their averages. The measured gap says the first matters less on this chip than the second and third.

Asks (not answerable from the repository): the lab lead's ~120 °C episode (which card and firmware, what stopped, whether a reset was needed, whether power and the PCIe link were logged); card 0's own record from 25 September; the DRAM part and its temperature grade; what the PMIC's “system temperature” measures and whether the PMIC power-cycles the board; Esperanto's Tj max, qualification and design life; and a firmware build that enables a process-detector delay chain, if the transistor-speed question should be measured.

7. The experiments on the cards (E53)

How it ran: the freeze, two stopped attempts and their amendments, and the card time
  1. The freeze.
  2. Amendment 1.
  3. Amendment 2.
  4. Card time.
  5. The rules.

8. Method, data and how to reproduce

What each number is, and its limits

Data. Everything is in docs/reports/data/2026-09-28-overheating: E53's raw blocks (raw/), reductions (reductions/), the analyses of the existing record (scripts/, analysis/), the source list (sources.json) and this page's data (overheat.json). The tools are tools/claims-v3/oh, with the frozen predictions in prereg/PREREG.md.

V=docs/reports/data/2026-09-28-overheating
python3 tools/claims-v3/oh/reduce.py --all --data $V/raw --out $V/reductions     # E53's registered verdicts
python3 $V/extras.py                                                              # its descriptive numbers
for s in max_temps correct_vs_temp timing_vs_temp hot_minus_mean idle_vs_temp runaway events_vs_temp derived; do
  python3 $V/scripts/$s.py > $V/analysis/$s.txt; done                             # the existing record, and arithmetic
python3 $V/build_overheat_data.py                                                 # overheat.json
python3 scripts/build-report.py effect-of-overheating $V/overheat.json docs/reports/2026-09-28-effect-of-overheating.html

Sources

Outside sources, numbered in the order the page first cites them; all fetched on 28 September 2026. Each entry names what the page uses from it.

    The ET-SoC-1's own documents (firmware at the commits named, in aifoundry-org/et-platform; manuals in aifoundry-org/et-man):