HomeNAS · Hardware Validation · 28 Sep 2026 UTC
Benchmark stopped safely on CPU thermals.
A safety-gated, crash-resilient benchmark began with an idle baseline and an incremental CPU ramp. Stage CPU-2 was terminated automatically after the Ryzen 7 7700X sustained at least 90°C. The harness correctly stopped all remaining tests; it did not weaken limits or continue into RAM and dual-GPU stress.
1 / 15stages passed
1stage aborted
13stages not run
0Xid / MCE / EDAC / OOM errors
Interpretation: the measured result demonstrates that the current CPU cooling/control response is not suitable for escalating this particular sustained workload beyond the first stage under the conservative benchmark policy. It does not establish a general stability failure. RAM and dual-GPU stress results are unavailable because the stop-on-serious-error policy was honored.
System inventory
Compute
CPU: AMD Ryzen 7 7700X, 8 cores / 16 threads
Memory: 64 GiB DDR5, 2 × 32 GiB, no swap
GPU pool: 2 × NVIDIA GeForce RTX 5060 Ti 16 GiB
GPU driver: 580.173.02 · CUDA 13.0 runtime image
Platform
Board: ASRock X670E Taichi
OS: TrueNAS SCALE / Debian 12
Kernel: 6.12.105-production+truenas
Fan control: nct6687 service active throughout
Measured temperature timeline
CPU Tctl90°C sustained threshold95°C instant threshold
Stage results
| Stage | Requested intensity | Duration | Status | CPU min / avg / max | DIMM max | CPU fan max | Errors |
|---|---|---|---|---|---|---|---|
| Idle baseline | Normal services, GPUs unloaded | 60 samples | Recorded | 43.5 / 45.4 / 52.1°C | 36.0°C | 422 RPM | None |
| CPU-1 | 4 / 16 worker threads | 30 s | Passed | 43.8 / 83.0 / 85.5°C | 35.8°C | 1,282 RPM | None |
| CPU-2 | 8 / 16 worker threads | 13 s observed of 45 s | Safety abort | 46.8 / 87.2 / 94.5°C | 35.3°C | 1,486 RPM | Thermal watchdog |
| CPU-3 → CPU-5 | 12/16, 16/16, extended 16/16 | Planned | Not run | — | — | — | Stopped after serious error |
| RAM-1 → RAM-5 | 4, 8, 12, 16, 20 GiB verified patterns | Planned | Not run | — | — | — | Stopped after serious error |
| GPU-1 → GPU-5 | Unified 2-rank NCCL pool, increasing BF16 GEMM + VRAM | Planned | Not run | — | — | — | Stopped after serious error |
Methodology and safeguards
Harness
- Persistent run directory on the Fast pool
- Timestamped JSONL stage journal and atomic status file
- Telemetry sampled approximately once per second and flushed to storage
- Systemd transient unit independent of the SSH session
- Exact workload commands preserved in the stage journal
Hard stops
- CPU: ≥95°C instant or ≥90°C sustained
- GPU core: ≥85°C; memory/hotspot ≥100°C when exposed
- DIMM/board: ≥80°C
- Any Xid, MCE/EDAC, OOM, sensor loss, fan failure, workload error or instability
- At least 6 GiB OS memory headroom
Measured facts vs. interpretation
Measured
- CPU-1 completed successfully.
- CPU-2 reached 94.5°C and was killed after sustained samples at or above 90°C.
- DIMMs remained at or below 36.0°C.
- Both GPUs stayed idle during CPU testing, peaking at 38°C and 40°C in CPU-1.
- No benchmark-window Xid, MCE/EDAC, OOM or hardware-error kernel entries were found.
Interpretation
- The current CPU thermal response hit the conservative policy before a 50% worker stage could finish.
- No conclusion can be drawn about full RAM or dual-GPU stability from this run.
- A cooling-system inspection should precede any re-run: cooler mounting/contact, pump/fan behavior, BIOS limits, airflow, and whether the software curve is appropriate.
Downloads
Complete sanitized result bundleTelemetry, journal, harness output, inventory and error deltas↓ TAR.GZTelemetry CSV
Approximately 1-second samples↓ CSVStage journal
Commands, requested intensities, status and timing↓ JSONLSanitized inventory
Hardware, versions, sensors and storage↓ TXT