HomeNAS · Hardware Validation · 28 Sep 2026 UTC

Benchmark stopped safely on CPU thermals.

A safety-gated, crash-resilient benchmark began with an idle baseline and an incremental CPU ramp. Stage CPU-2 was terminated automatically after the Ryzen 7 7700X sustained at least 90°C. The harness correctly stopped all remaining tests; it did not weaken limits or continue into RAM and dual-GPU stress.

1 / 15stages passed
1stage aborted
13stages not run
0Xid / MCE / EDAC / OOM errors
Interpretation: the measured result demonstrates that the current CPU cooling/control response is not suitable for escalating this particular sustained workload beyond the first stage under the conservative benchmark policy. It does not establish a general stability failure. RAM and dual-GPU stress results are unavailable because the stop-on-serious-error policy was honored.

System inventory

Compute

CPU: AMD Ryzen 7 7700X, 8 cores / 16 threads
Memory: 64 GiB DDR5, 2 × 32 GiB, no swap
GPU pool: 2 × NVIDIA GeForce RTX 5060 Ti 16 GiB
GPU driver: 580.173.02 · CUDA 13.0 runtime image

Platform

Board: ASRock X670E Taichi
OS: TrueNAS SCALE / Debian 12
Kernel: 6.12.105-production+truenas
Fan control: nct6687 service active throughout

Measured temperature timeline

CPU Tctl90°C sustained threshold95°C instant threshold

Stage results

StageRequested intensityDurationStatusCPU min / avg / maxDIMM maxCPU fan maxErrors
Idle baselineNormal services, GPUs unloaded60 samplesRecorded43.5 / 45.4 / 52.1°C36.0°C422 RPMNone
CPU-14 / 16 worker threads30 sPassed43.8 / 83.0 / 85.5°C35.8°C1,282 RPMNone
CPU-28 / 16 worker threads13 s observed of 45 sSafety abort46.8 / 87.2 / 94.5°C35.3°C1,486 RPMThermal watchdog
CPU-3 → CPU-512/16, 16/16, extended 16/16PlannedNot run———Stopped after serious error
RAM-1 → RAM-54, 8, 12, 16, 20 GiB verified patternsPlannedNot run———Stopped after serious error
GPU-1 → GPU-5Unified 2-rank NCCL pool, increasing BF16 GEMM + VRAMPlannedNot run———Stopped after serious error

Methodology and safeguards

Harness

  • Persistent run directory on the Fast pool
  • Timestamped JSONL stage journal and atomic status file
  • Telemetry sampled approximately once per second and flushed to storage
  • Systemd transient unit independent of the SSH session
  • Exact workload commands preserved in the stage journal

Hard stops

  • CPU: ≥95°C instant or ≥90°C sustained
  • GPU core: ≥85°C; memory/hotspot ≥100°C when exposed
  • DIMM/board: ≥80°C
  • Any Xid, MCE/EDAC, OOM, sensor loss, fan failure, workload error or instability
  • At least 6 GiB OS memory headroom

Measured facts vs. interpretation

Measured

  • CPU-1 completed successfully.
  • CPU-2 reached 94.5°C and was killed after sustained samples at or above 90°C.
  • DIMMs remained at or below 36.0°C.
  • Both GPUs stayed idle during CPU testing, peaking at 38°C and 40°C in CPU-1.
  • No benchmark-window Xid, MCE/EDAC, OOM or hardware-error kernel entries were found.

Interpretation

  • The current CPU thermal response hit the conservative policy before a 50% worker stage could finish.
  • No conclusion can be drawn about full RAM or dual-GPU stability from this run.
  • A cooling-system inspection should precede any re-run: cooler mounting/contact, pump/fan behavior, BIOS limits, airflow, and whether the software curve is appropriate.

Downloads

Complete sanitized result bundle
Telemetry, journal, harness output, inventory and error deltas
↓ TAR.GZ
Telemetry CSV
Approximately 1-second samples
↓ CSV
Stage journal
Commands, requested intensities, status and timing
↓ JSONL
Sanitized inventory
Hardware, versions, sensors and storage
↓ TXT