basic-benchmark/README.md
smill 531ed18082 compare-report: median-of-reps grouping, --exclude, --results
--group NAME=t1,t2 collapses repetitions into one median column per device,
--exclude drops bad runs, and --results reads a collected results directory.
2026-09-29 23:56:54 -04:00

446 lines
20 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# basic-benchmark
A self-contained kit for benchmarking a **freshly installed Fedora 44** machine
(and comparing it against another build or an earlier baseline), using the *same*
tests on both sides, with continuous power/thermal logging so you get efficiency
comparisons — not just scores.
The kit makes **no assumptions about the host**: no baked-in CPU/GPU/user/host
names. Specs are *captured* on whatever machine runs it (`./capture-specs.sh`).
It targets a stock Fedora 44 with internet access, and can provision missing
tools itself via `dnf`.
The suite is split into two parts so it can be driven from another device:
| Part | Needs a display? | Where to run it | What it covers |
|------|------------------|-----------------|----------------|
| **headless** | no | over SSH from another device | CPU, memory, storage, thermals, power logging |
| **headed** | yes | on the desktop, in the KDE session | GPU (glmark2, vkmark) |
Both parts append their part name to the result tag, so they never overwrite
each other's CSVs.
> **PassMark PerformanceTest** is the primary cross-platform score. The suite is
> self-contained: no cloud uploads, no license keys.
---
## 1. Workflow
```
capture specs headless (from any device) headed (at the machine)
───────────── ────────────────────────── ───────────────────────
./capture-specs.sh before ssh -t <target> \
'cd basic-benchmark && \
./run-benchmarks.sh before headless'
-> results/power-before-headless.csv
results/phases-before-headless.csv
./run-benchmarks.sh before headed
-> results/power-before-headed.csv
───────────────── change hardware / reinstall ─────────────────
repeat all with "after", then:
power-report.py --tags \
before-headless before-headed \
after-headless after-headed
```
You can run the parts in any order; run both on the same machine back to back.
---
## 2. Files
| File | What it is |
|------|------------|
| `README.md` | This guide. |
| `specs.txt` | Pointer to the spec capture (no hardcoded host). |
| `specs-detailed.txt` | Pointer to the spec capture (no hardcoded host). |
| `capture-specs.sh` | Captures a portable spec sheet for the *current* host. |
| `benchmarks.txt` | The suite definition: every test, commands, ground rules, results table. |
| `run-benchmarks.sh` | Runs the suite (`headless`/`headed`/`all`), starts/stops power logging, drops phase markers. |
| `powerlog.py` | Background sampler (~1 Hz) for power, temps, clocks, GPU, fans. |
| `power-report.py` | Turns the CSVs into PNG graphs, per-phase tables and efficiency numbers. |
| `compare-report.py` | Cross-device / before-after results table + one bar chart per benchmark (`results/graphs/compare-*.png`). |
| `results/` | Created at runtime: CSVs, logs, `results/graphs/*.png`. |
| `tools/` | Bundled binaries: PassMark (`tools/pt/`). Ships with the kit — no download needed. |
| `tools/corpus/` | Fixed sample inputs for the app workloads (LibreOffice docs, a GEGL photo, Inkscape SVGs), fetched once and checksum-pinned. |
---
## 3. Prerequisites
Target: a **stock Fedora 44 with internet access** (x86-64 for the bundled
PassMark binary). Nothing host-specific is required — no fixed user, hostname, or
hardware.
- The **headless** part runs from any terminal (local or SSH); the **headed**
part must run inside a graphical session (Wayland or X11) because GPU tests
need a live display.
- `sudo` access for your user (no passwordless sudo assumed). Power logging, the
CPU governor and `turbostat` need root. Over SSH, allocate a TTY
(`ssh -t <target> ...`) so `sudo` can prompt; the script keeps the sudo
timestamp alive for the whole run so you type the password once.
- Python 3 (present on Fedora). `power-report.py` needs `matplotlib`.
- **Missing tools are installed automatically** from dnf when possible:
```bash
INSTALL_DEPS=1 ./run-benchmarks.sh before headless # auto-install, no prompt
./run-benchmarks.sh doctor # show what's missing
```
Interactive runs default to `INSTALL_DEPS=ask` (a y/N prompt). The mapping is
`7z→7zip`, `sysbench→sysbench`, `fio→fio`, `stress-ng→stress-ng`,
`sensors→lm_sensors`, `turbostat→kernel-tools`, `cpupower→cpupowerutils`,
`glmark2`, `vkmark`, `glxinfo→mesa-demos`, `vulkaninfo→vulkan-tools`,
`matplotlib→python3-matplotlib`, plus `ncurses-libs`/`ncurses-compat-libs`
for PassMark. Set `INSTALL_DEPS=0` to never install.
- The LLM test needs **no** install: `llama-bench` and its small model are
bundled under `tools/llama/` (see §4).
- The **app workloads** (LibreOffice, GEGL, Inkscape; see §4) are optional:
each step runs only if its command is installed, otherwise it is skipped and
listed by `./run-benchmarks.sh doctor`.
---
## 4. One-time tool setup
**Almost nothing to download.** PassMark, `llama-bench`, and the LLM model ship
inside `tools/`. (The optional app-workload corpora in §4 are fetched once, on
first run.) Just make sure the ncurses compat libs are present for PassMark:
```bash
sudo dnf install -y ncurses-libs ncurses-compat-libs
# if it asks for libncurses.so.5:
sudo ln -sf /usr/lib64/libncurses.so.6 /usr/lib64/libncurses.so.5
./tools/pt/PerformanceTest/PerformanceTest_Linux_x86-64 -h # help / flags
```
If the bundled binary is missing (e.g. non-x86-64 host), re-fetch it:
1. Download `https://www.passmark.com/downloads/pt_linux_x64.zip`.
2. Unpack into `tools/pt/`:
```bash
mkdir -p tools/pt && cd tools/pt
unzip ~/Downloads/pt_linux_x64.zip
# zips unpack to PerformanceTest/PerformanceTest_Linux_x86-64 (older: pt_linux_x64)
chmod +x PerformanceTest/PerformanceTest_Linux_x86-64
```
PassMark invokes `sudo dmidecode -t 17` for RAM SPD details; it's optional,
and it's silent once the script's sudo keepalive is active.
### LLM inference — llama.cpp (bundled)
**Already bundled** under `tools/llama/`: a `llama-bench` runtime (`bin/`, with its
shared libs) and a small model (`models/MiniCPM5-2B-Q8_0.gguf`, ~2.5 GB). No
download or install needed. The test runs **token generation on CPU/RAM**
(`-ngl 0`), which is memory-bandwidth-bound — so its `tg` tok/s is a direct
indication of how well the machine's RAM serves LLM inference.
```bash
tools/llama/bin/llama-bench \
-m tools/llama/models/MiniCPM5-2B-Q8_0.gguf -ngl 0 -p 0 -n 128 -r 2
```
Override with `LLAMA_MODEL=` (any GGUF), or `LLAMA_NGL=` to offload to GPU.
### App workloads — LibreOffice / GEGL / Inkscape (optional)
Stock distro applications are driven over a fixed sample corpus. Each is the
app's own CLI, timed, named after the benchmark. Install the ones you want to
measure:
```bash
sudo dnf install -y libreoffice-writer libreoffice-calc # documents -> PDF
sudo dnf install -y gegl04-tools # GEGL operations
sudo dnf install -y inkscape # SVG -> PNG
```
On first run the corpora are **fetched once** into `tools/corpus/` and
checksum-verified (`lo-sample-documents-1.zip`, `sample-photo-6000x4000-1.zip`,
`svg-test-files-1.zip`; ~12 MB total), then reused so both sides process identical
bytes. Set INKSCAPE_FILES=N to change how many SVGs Inkscape exports (default 25).
Missing apps are skipped; skip the whole group with `RUN_APPS=0`.
> `run-benchmarks.sh` finds PassMark at `tools/pt/pt_linux_x64` **or**
> `tools/pt/PerformanceTest/PerformanceTest_Linux_x86-64`. Override the location
> with `TOOLS=/path` (or set `PT=` directly).
---
## 5. Quick start
On the target machine (fresh Fedora 44, internet):
```bash
cd basic-benchmark
./run-benchmarks.sh doctor # shows what's present/missing (runs nothing)
INSTALL_DEPS=1 ./run-benchmarks.sh doctor # provision missing tools via dnf
./capture-specs.sh before # record this machine's spec sheet
```
### Part 1 — headless (run from another device over SSH)
```bash
ssh -t <target> 'cd basic-benchmark && ./run-benchmarks.sh before headless'
```
The `-t` matters: it gives `sudo` a TTY so power logging and the governor work.
If tools are missing, this prompts to install them (or set `INSTALL_DEPS=1`).
### Part 2 — headed (run at the machine, in a graphical terminal)
```bash
./run-benchmarks.sh before headed
```
This errors out (exit 2) if no display is attached, rather than silently
producing useless results. Run it when a graphical session is active.
### Compare
```bash
python3 power-report.py --tags before-headless before-headed
```
Outputs: `results/graphs/power-before-headless.png`, `power-before-headed.png`,
plus per-phase power/thermal tables and efficiency. (Comparison graph accepts
any number of tags.)
**Legacy single-shot:** `./run-benchmarks.sh before` runs `all` (headless, then
headed if a display is present) under the plain `before` tag.
**Median of 3 (recommended):** tag each repetition, then report all:
```bash
for i in 1 2 3; do
ssh -t <target> "cd basic-benchmark && ./run-benchmarks.sh before-hl$i headless"
./run-benchmarks.sh "before-hd$i" headed # at the target
done
python3 power-report.py --tags before-hl1 before-hl2 before-hl3 \
before-hd1 before-hd2 before-hd3
```
---
## 6. What the run script does (and rough times)
Headless part (`~25-30 min` per pass):
| # | Step | ~Time | Notes |
|---|------|-------|-------|
| 0 | Env capture → `results/env-<tag>.txt` | <1 min | CPU model, tools, hashes, GPU pinning |
| 0b | Idle baseline (≥15 s) | ~18 s | true idle power/thermal window |
| - | Start power sampler | 0 | runs the whole part |
| 1 | Pin CPU governor → `performance` | <1 min | sudo; **restored on exit** |
| 2 | PassMark PT Linux `-r 1` then `-r 2` | 8-12 min | CPU suite + Memory suite (separate windows) |
| 3 | `7z b` (all threads + `-mmt1`) | ~2 min | integer/cache proxy; isolated single-core |
| 4 | `openssl speed` aes-256-gcm / sha256 | 2-3 min | crypto |
| 5 | `sysbench` cpu 1-thread / all-thread / memory | 1-2 min | isolated single-core window |
| 6 | llama.cpp generation (`2B`, CPU/RAM) | ~15 s | bundled; token-generation tok/s |
| 6b | LibreOffice docs→PDF (20 docs) | ~0.5-1 min | fixed corpus |
| 6c | GEGL image ops (9 on one 6000×4000 photo) | ~0.5-1 min | fixed corpus |
| 6d | Inkscape SVG→PNG (first 25) | ~20-30 s | fixed corpus |
| 7 | `fio` 4×30 s | 3-4 min | same SSD, control; runs as your user |
| 8 | `stress-ng` 5 min + `turbostat` | ~8 min | thermals/power |
| 9 | `systemd-analyze` | <1 min | informational |
Headed part (`~1 min` per pass):
| # | Step | ~Time | Notes |
|---|------|-------|-------|
| 1 | `glmark2` (terrain, 10 s) | ~15 s | pinned to the dGPU |
| 2 | `vkmark` (texture, 10 s) | ~15 s | pinned to the dGPU |
Every step writes its own `results/<name>-<tag>.txt/json`.
Options (env):
```bash
RUN_APPS=0 ./run-benchmarks.sh before headless # skip the app workloads
CORPUS=/path ./run-benchmarks.sh before headless # app corpus cache location
GOVERNOR=powersave ./run-benchmarks.sh after headless # pin a governor (see §7)
RUN_GPU=0 ./run-benchmarks.sh before # skip GPU tests in 'all' mode
GLMARK2_DURATION=10 VKMARK_BENCH=shading ./run-benchmarks.sh before headed # GPU scenes
TOOLS=/path ./run-benchmarks.sh before headless # custom tools dir
GPU_PCI=0000:03:00.0 ./run-benchmarks.sh before headed # force a GPU
FIOFILE=/data/fio.bin ./run-benchmarks.sh before headless # fio scratch file
```
---
## 7. Ground rules for a fair BEFORE/AFTER
1. Same OS image, kernel, mesa, and **tool versions** on both runs (recorded in
`env-<tag>.txt`; tool binaries are hashed).
2. Same storage drive, same filesystem/options.
3. Same GPU, driver and displays; note which GPU drives the displays. The script
auto-detects the headless dGPU and pins GPU tests to it (see §8).
4. Machine idle: close browsers, stop docker/VMs
(`sudo systemctl stop docker docker.socket`).
5. Governor pinned to `performance` for timed runs and **restored afterwards**
(the script does both). To measure another profile, use `GOVERNOR=`:
```bash
GOVERNOR=powersave ./run-benchmarks.sh after headless # -> after-powersave-headless
GOVERNOR=performance ./run-benchmarks.sh after headless # -> after-performance-headless
./run-benchmarks.sh after headless # -> after-headless (performance)
```
An explicit `GOVERNOR` is appended to the tag, so both profiles can be captured
on the same run set and compared side by side with `compare-report.py` (the
default `performance` keeps the plain tag, so existing before/after pairs stay
comparable).
6. Warm up once, then take the **median of 3** runs. Run both parts back to back
on the same side.
7. Note ambient/room temperature next to the results.
8. **Back up** `results/` before reinstalling or re-imaging — copy them
somewhere off the machine.
---
## 8. Power & thermal logging
`powerlog.py` samples about once per second for the whole part:
- **package W** and **core W** — from RAPL energy counters
(`/sys/class/powercap/intel-rapl*`; root-only).
- **CPU temp** — `k10temp` (or `coretemp` on other platforms).
- **CPU MHz** — average `scaling_cur_freq`.
- **GPU W / GPU temp / GPU busy** — amdgpu hwmon, **tied to the dGPU** (the
first `amdgpu` hwmon exposing `power1_average`, so a hybrid iGPU+dGPU box logs
the discrete card, not the iGPU).
- **fan1, fan2** — `nct6798`.
Files:
- `results/power-<tag>.csv` — `epoch,iso,pkg_w,core_w,cpu_temp_c,cpu_mhz,gpu_w,gpu_temp_c,gpu_busy,fan1,fan2`
- `results/phases-<tag>.csv` — `epoch,phase`, one line per benchmark, so the log
can be split per benchmark.
Run it standalone for a single test (Ctrl-C to stop):
```bash
sudo python3 powerlog.py --tag before
```
### GPU pinning
The suite pins GPU tests to the GPU that owns **no connected display** (the RX
6700 XT here, while both monitors are on the Cezanne iGPU). It exports:
- `DRI_PRIME=pci-<domain:bus:dev.func>` for OpenGL (`glmark2`)
- `MESA_VK_DEVICE_SELECT` / `VK_LOADER_DEVICE_SELECT=<vendor:device>` for Vulkan
(`vkmark`)
`glxinfo -B` and `vulkaninfo --summary` are captured so you can confirm which
device actually ran.
### Graphs and efficiency
```bash
python3 power-report.py --tags before-headless
python3 power-report.py --tags before-headless before-headed
python3 power-report.py --tags before-headless before-headed \
after-headless after-headed
```
Reports per phase: **avg / p95 / peak W**, **avg / peak °C**, **energy (kJ)**,
and **score-per-watt**.
Efficiency needs a scores file `results/scores-<tag>.csv`; the name should match
a phase marker (exact or substring). Phase names:
```
env idle PassMark-cpu PassMark-mem 7zip 7zip-1t openssl
sysbench-1t sysbench-nt sysbench-mem llama
fio-seqread fio-seqwrite fio-randread fio-randwrite
thermal-load systemd libreoffice gegl inkscape
glmark2 vkmark
```
Example `results/scores-before-headless.csv`:
```
PassMark-cpu,20707
PassMark-cpu-single,3342
PassMark-mem,2450
7zip,58935
openssl-aes,30.17
sysbench-1t,4810
llama-tg,18.2
libreoffice,3.9
gegl,1.8
inkscape,12.4
fio-seqread,2010.4
glmark2,5989
```
### True wall power (optional)
The numbers above are **package** watts. For whole-system wall power, add a
smart plug (Tasmota/Shelly) or a UPS that reports load, poll it on a second
background loop, and merge on the `epoch` column.
---
## 9. Reading the results
- PassMark **CPU Mark** scales with core count; PassMark **CPU Single Threaded**
isolates IPC/clock. Read them together.
- Compare **single-thread vs all-thread** separately — thread counts differ
between platforms.
- Compare **bandwidth and latency** independently; PassMark Memory Mark also
rewards capacity.
- Components carried over unchanged (same SSD, same dGPU, same displays) should
land near-flat; a big move there means a setup fault, not a win.
- For efficiency, **score/W higher is better**, **energy kJ lower is better**.
### Comparing devices / before-after
```bash
python3 compare-report.py # every tag found in results/
python3 compare-report.py --tags t14-1 server-1-headless desktop-before-headless \
--label t14-1="ThinkPad T14" --label server-1-headless=server
```
Merges every `results/scores-<tag>.csv` into one table (also written to
`results/compare-scores.csv`) and draws **one bar chart per benchmark** to
`results/graphs/compare-<metric>.png`. Metrics missing from an older score file
are back-filled from the raw results (`fio` JSON, `openssl`, `7z`,
`sysbench-mem`, `glmark2`/`vkmark`), so existing runs compare without re-running.
Needs `matplotlib`; use `--no-graphs` for the table alone.
With **median-of-3**, collapse repetitions into one column per device and drop
any bad run:
```bash
python3 compare-report.py --results ~/Documents/results \
--group "desktop=before1-headless,before2-headless,before3-headless" \
--group "server=server1-headless,server2-headless,server3-headless" \
--group "t14=t14-2,t14-3" # t14-1 dropped (low-battery run)
```
`--group NAME=t1,t2,...` makes one median column; `--exclude t1 t2` drops tags.
`--results DIR` reads another machine's collected `results/`.
Fill the results table in `benchmarks.txt` with the medians after each side.
---
## 10. Backing up / reinstalling
Before any reinstall or re-image, copy the kit and its results somewhere safe:
```bash
cp -a ~/basic-benchmark /path/to/backup/basic-benchmark-$(date +%F)
```
After the new build: put the kit back, run `INSTALL_DEPS=1 ./run-benchmarks.sh
doctor` to provision tools, then run both parts (`headless` over SSH, `headed` at
the machine) with `after`, and `python3 power-report.py --tags before-headless
before-headed`.
---
## 11. Troubleshooting
| Symptom | Cause / fix |
|---|---|
| `powerlog: no RAPL counters found` | Not running as root. Use `sudo`; otherwise power columns stay blank. |
| Power columns blank in the CSV | Same as above, or RAPL not exposed. `turbostat` output still gives PkgWatt. |
| `WARNING: no sudo` / sampler skipped over SSH | No TTY for the password prompt. Re-run as `ssh -t <target> ...`. |
| `ERROR: no display` from `headed` | The headed part must run in the KDE session (Konsole), not over SSH. |
| `GPU: skipped` | No `DISPLAY`/`WAYLAND_DISPLAY` (e.g. over SSH) or `RUN_GPU=0`. |
| GPU tests run on the wrong card | Override with `GPU_PCI=`/`GPU_VD=`; check the renderer in `glxinfo-<tag>.txt`. |
| PassMark: `ncurses: cannot initialize terminal type` | `TERM` unset / no terminal. The script sets `TERM=xterm`; if it persists, run over `ssh -t`. |
| PassMark warns `sudo: a terminal is required` | Its internal `sudo dmidecode` (RAM SPD) is optional; it's silent once the sudo keepalive is active. |
| `skip (not found: …/pt_linux_x64)` | PassMark not installed at either expected path — see §4. |
| `report failed` / no matplotlib | `sudo dnf install -y python3-matplotlib`. |
| Scores show `no phase match` | Fix names in `scores-<tag>.csv` to match a phase name (§8). |
| Governor left on `performance` | Shouldn't happen (restored on exit); set it back with `sudo cpupower frequency-set -g powersave`. |
| App step says `skip (… missing)` | That app isn't installed; `sudo dnf install -y <pkg>`, or `RUN_APPS=0` to silence. |
| App step says `skip (corpus unavailable)` | Corpus download failed (no curl/wget or no network); re-run with network, or pre-seed `tools/corpus/`. |
---
## 12. Scope notes
- Portable by design: run it on any Fedora 44 host; specs are captured with
`./capture-specs.sh`, never hardcoded.
- The suite is intentionally split: **headless** for SSH-driven runs,
**headed** for the display-bound GPU tests.
- **PassMark** is the primary cross-platform score. The suite is
self-contained and offline: no uploads and no license keys.
- `powerlog.py`/`power-report.py` are plain Python 3; the only Python dep is
`matplotlib` (installed on demand; see §3).