basic-benchmark/benchmarks.txt

293 lines
18 KiB
Text
Raw Permalink Normal View History

############################################################################
# basic-benchmark — benchmark suite (run BEFORE and AFTER)
# host: <captured per run> suite v3 portable kit
#
# Target: a stock Fedora 44 with internet. Missing tools are installed via dnf
# (INSTALL_DEPS=1). Specs are captured with ./capture-specs.sh, not hardcoded.
#
# Purpose: measure a hardware upgrade with an identical, repeatable set of
# tests on both sides — capture the AFTER run once the new build is up.
#
# Quick path (run-benchmarks.sh is split; see README):
# headless: ./run-benchmarks.sh before headless # over SSH (ssh -t)
# headed: ./run-benchmarks.sh before headed # in the desktop session
# repeat both with "after" post-swap; tags become before-headless / before-headed
# This file is the reference for what that script runs, plus extras.
#
# Primary cross-platform score: PassMark PerformanceTest (cpubenchmark.net).
# Technical depth: the quick tools below.
############################################################################
############################################################################
# RUN LIST — what to run, in order, with rough one-pass times (idle desktop)
############################################################################
# benchmark part by ~time notes
0 env capture headless script <1 min
0b idle baseline (>=15s) headless script ~18 s true idle window
1 PassMark PT Linux -r 1 / -r 2 headless script 8-12 min CPU suite + Memory suite
2 7z b (all + mmt1) headless script ~2 min
3 openssl speed aes-256-gcm / sha256 headless script 2-3 min
4 sysbench cpu (1t + Nt) + memory headless script 1-2 min
5 llama.cpp generation (2B, CPU/RAM) headless script ~15 s bundled; tok/s
5b LibreOffice: 20 documents -> PDF headless script ~1 min fixed corpus
5c GEGL: 9 image ops (6000x4000) headless script ~1 min fixed corpus
5d Inkscape: SVG -> PNG (first 25) headless script ~30 s fixed corpus
6 fio 4x30s headless script 3-4 min runs as your user
7 stress-ng 5 min + turbostat headless script ~8 min thermals/power
8 systemd-analyze headless script <1 min
glmark2 + vkmark headed script ~1 min needs desktop session
---------- headless total ---------- ~25-30 min
---------- headed total ------------ ~1 min
=> ground rule wants MEDIAN OF 3 runs: x3 the above
* powerlog.py samples power/temp ~1/s for the WHOLE part ->
results/power-<tag>.csv + results/phases-<tag>.csv
== GROUND RULES FOR A FAIR BEFORE/AFTER ==
1. Same OS image + kernel + mesa + tool versions on both runs. Record them.
2. Same storage drive in both runs (the WD Blue SN570). Same filesystem/options.
3. Same GPU, same driver, same displays, same resolution/refresh. Record which
GPU drives the displays (iGPU vs dGPU) — it changes between platforms.
4. Kill background load: close browsers, stop docker and VMs.
sudo systemctl stop docker docker.socket
5. Pin CPU governor to performance for timed runs (BEFORE default = powersave):
sudo cpupower frequency-set -g performance
# fallback if no cpupower:
for f in /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor; do \
echo performance | sudo tee "$f" >/dev/null; done
run-benchmarks.sh does this AND restores the original governor on exit.
To measure another profile, set GOVERNOR= (the value is appended to the tag):
GOVERNOR=powersave ./run-benchmarks.sh after headless -> after-powersave-headless
GOVERNOR=performance ./run-benchmarks.sh after headless -> after-performance-headless
6. Warm up once, then take the MEDIAN of 3 runs. Throw away the first.
7. Note room/ambient temperature next to thermal results; thermals are noisy.
8. Use the SAME tool versions on both sides (with their (point) releases:
PassMark PT build). Record them in env-$TAG.txt.
9. BACK UP before any reinstall / reimage:
~/basic-benchmark/results/
copy it somewhere safe (off the machine) before reinstalling.
== 0. ENV CAPTURE (run first, each side) ==
export TAG=before # or: after
mkdir -p ~/basic-benchmark/results && cd ~/basic-benchmark/results
{ date; uname -a; lscpu; free -h; lsblk;
stress-ng --version;
glxinfo -B 2>/dev/null; vulkaninfo --summary 2>/dev/null | head -40;
sha256sum ~/basic-benchmark/tools/pt/PerformanceTest/PerformanceTest_Linux_x86-64 2>/dev/null;
} > env-$TAG.txt 2>&1
== 1. PASSMARK PERFORMANCETEST LINUX (this IS the cpubenchmark.net test) ==
# Free, CLI-only Linux build. Output is directly comparable to the numbers
# on cpubenchmark.net and cross-platform with the Windows product.
# CPU Mark, Memory Mark, and every per-test subscore incl. single-thread.
#
# -r 1 = CPU only -> results_cpu.yml
# -r 2 = Memory only-> results_memory.yml
# -r 3 = All tests -> results_all.yml <-- use this
# -d 1..4 = Short/Medium/Long/VeryLong -i <iters> -p <processes>
#
PT=~/basic-benchmark/tools/pt/PerformanceTest/PerformanceTest_Linux_x86-64
# (older PassMark zips put the binary at tools/pt/pt_linux_x64; the script
# auto-detects either; PassMark needs TERM set — the script sets TERM=xterm)
"$PT" -r 3 -d 2 | tee passmark-$TAG.txt
mv -f results_all.yml passmark-all-$TAG.yml
# KEEP the SAME -r/-d/-i/-p on both sides. Leave -p at default so it scales
# to each machine's thread count.
# Result to record: "CPU Mark", "CPU Single Threaded", "Memory Mark".
== 2. CPU — THROUGHPUT / COMPRESSION ==
# 7-Zip built-in benchmark (integer, scales with cores + cache + memory)
7z b -mmt$(nproc) | tee 7z-$TAG.txt
# zstd / xz on a FIXED corpus: create ONCE and reuse for both sides so the
# input bytes are identical (regenerating random data per run is not comparable)
[ -f /tmp/corpus-2g.bin ] || head -c 2G /dev/urandom > /tmp/corpus-2g.bin
zstd -b9 -i3 /tmp/corpus-2g.bin 2>&1 | tee zstd-$TAG.txt
/usr/bin/time -v xz -9 -T$(nproc) -k -f /tmp/corpus-2g.bin 2>&1 | tee xz-$TAG.txt
# crypto (AES/SHA throughput; exercises AVX/VAES path differently per vendor)
openssl speed -multi $(nproc) -evp aes-256-gcm 2>&1 | tail -20 | tee openssl-aes-$TAG.txt
openssl speed -multi $(nproc) -evp sha256 2>&1 | tail -20 | tee openssl-sha-$TAG.txt
== 4. CPU — SINGLE vs MULTI THREAD ==
sudo dnf install -y sysbench
sysbench cpu --threads=1 run | tee sysbench-1t-$TAG.txt # single-thread
sysbench cpu --threads=$(nproc) run | tee sysbench-nt-$TAG.txt # all-thread
# NOTE: compare single-thread AND all-thread results separately — the
# thread count will differ between the two platforms.
== 5. MEMORY — BANDWIDTH ==
sysbench memory --threads=$(nproc) run | tee sysbench-mem-$TAG.txt
# Compare bandwidth between the two builds.
== 5b. LLM INFERENCE — token generation on CPU/RAM (bundled llama.cpp) ==
# Self-contained: runtime + model live under tools/llama/. Generation (tg) is
# memory-bound, so tg t/s reflects RAM bandwidth for LLM serving.
tools/llama/bin/llama-bench \
-m tools/llama/models/MiniCPM5-2B-Q8_0.gguf -ngl 0 -p 0 -n 128 -r 2 \
| tee llama-$TAG.txt
# -ngl 0 forces CPU/RAM (no GPU offload). Record the tg t/s number.
== 5c. APP WORKLOADS (LibreOffice / GEGL / Inkscape) ==
# Stock distro apps run over a fixed sample corpus; run-benchmarks.sh fetches
# the corpus ONCE into tools/corpus/ (checksum-pinned) and times the app's own
# CLI, reporting a rate (items/s, higher is better). Missing apps are skipped.
sudo dnf install -y libreoffice-writer libreoffice-calc inkscape gegl04-tools
RUN_APPS=1 ./run-benchmarks.sh before headless # default on
RUN_APPS=0 ./run-benchmarks.sh before headless # skip the group
# Phases / scores are named after the benchmark: libreoffice gegl inkscape.
# INKSCAPE_FILES=N changes how many SVGs are exported (default 25).
== 7. GPU — DISCRETE RX 6700 XT (same card both runs; HEADED part only) ==
sudo dnf install -y glmark2 vkmark vulkan-tools mesa-demos
# Monitors hang off the Cezanne iGPU, so PIN the dGPU or these tests would
# silently measure the iGPU. run-benchmarks.sh auto-detects and exports:
# DRI_PRIME=pci-0000:03:00.0
# MESA_VK_DEVICE_SELECT=1002:73df VK_LOADER_DEVICE_SELECT=1002:73df
# (from /sys/class/drm: the dGPU is the card with no connected connector)
glxinfo -B | tee glxinfo-$TAG.txt
vulkaninfo --summary | tee vulkan-$TAG.txt
# One representative scene each (fast, comparable); tune with GLMARK2_DURATION,
# VKMARK_DURATION and VKMARK_BENCH (texture|shading|vertex).
glmark2 -b terrain:duration=10 2>&1 | tail -25 | tee glmark2-$TAG.txt
vkmark -b texture:duration=10 2>&1 | tail -30 | tee vkmark-$TAG.txt
# Confirm which GPU rendered: grep -i 'device\|renderer' glxinfo-$TAG.txt
== 8. STORAGE / IO (same drive — mostly a control; watch for regressions) ==
FIO=$HOME/.cache/basic-benchmark-fio.bin # user-owned; no sudo needed
mkdir -p "$(dirname "$FIO")"
for JOB in "seqread read 1M 16" "seqwrite write 1M 16" \
"randread randread 4k 32" "randwrite randwrite 4k 32"; do
set -- $JOB
fio --name=$1 --rw=$2 --bs=$3 --numjobs=1 --iodepth=$4 \
--size=1G --direct=1 --time_based --runtime=30 \
--filename=$FIO --group_reporting \
--output-format=json --output=fio-$1-$TAG.json
rm -f $FIO
done
# NOTE: hdparm -tT is ATA-only and does nothing useful on the NVMe SN570
# (it just returns "Inappropriate ioctl for device"), so it is dropped;
# the fio numbers above are the storage result to record.
== 9. THERMALS / POWER / SUSTAINED CLOCKS (the cooler change is measured here) ==
sensors > sensors-idle-$TAG.txt
stress-ng --cpu 0 --cpu-method matrixprod --timeout 300 --metrics-brief \
> stress-ng-$TAG.txt 2>&1 &
sleep 120 # let it reach steady state
sensors > sensors-load-$TAG.txt
cat /sys/devices/system/cpu/cpu0/cpufreq/scaling_cur_freq >> sensors-load-$TAG.txt
sudo turbostat --interval 5 --show Core,CPU,Busy%,Bzy_MHz,PkgWatt,RAMWatt,PkgTmp \
--timeout 30 | tee turbostat-load-$TAG.txt
wait
# Record: idle temp, load temp, sustained all-core clock, package watts.
# Record the fields the platform exposes (PkgWatt/RAMWatt where available).
== 9b. CONTINUOUS POWER / THERMAL LOGGING (efficiency, not just peaks) ==
# powerlog.py runs in the background for the WHOLE part and writes a ~1 Hz
# CSV. run-benchmarks.sh already starts/stops it and drops a phase marker
# before each benchmark, so the log can be split per benchmark afterwards.
# <tag> includes the part, e.g. power-before-headless.csv / power-before-headed.csv.
# GPU columns are tied to the headless dGPU (the amdgpu hwmon with power1_average).
# results/power-<tag>.csv epoch,iso,pkg_w,core_w,cpu_temp_c,cpu_mhz,
# gpu_w,gpu_temp_c,gpu_busy,fan1,fan2
# results/phases-<tag>.csv epoch,phase (one line per benchmark)
# Run it by hand for a single test (Ctrl-C to stop):
# sudo python3 powerlog.py --tag before
# Power source: RAPL package/core energy counters
# (/sys/class/powercap/intel-rapl*), root-only. GPU watts from the amdgpu
# hwmon power1_average. Thermals: k10temp/coretemp + amdgpu + nct6798 fans.
# If you add a smart plug later (Tasmota/Shelly/UPS), poll it in a second
# background loop and merge on the epoch column for true wall power.
#
# Graphs + per-phase table + efficiency:
# python3 power-report.py --tags before-headless # one part
# python3 power-report.py --tags before-headless before-headed # both parts
# python3 power-report.py --tags before-headless before-headed \
# after-headless after-headed # comparison graphs
# -> results/graphs/power-<tag>.png and power-thermal-compare.png
# (the multi-machine overlay; add --results DIR --label TAG=NAME)
# Cross-device / before-after score comparison (one bar chart per benchmark):
# python3 compare-report.py # table -> results/compare-scores.csv
# # + results/graphs/compare-*.png
# Efficiency needs numbers in results/scores-<tag>.csv; the name should match
# a phase marker (exact or substring). Phase names are:
# env PassMark 7zip openssl sysbench
# Example:
# PassMark,24500
# libreoffice,3.9 <- docs/s; energy = avg_w * seconds
# Metrics reported: avg/p95/peak W, avg/peak °C, energy (kJ) per phase,
# and score-per-watt (higher = better; lower energy = better).
== 10. DESKTOP / SYSTEM (informational — a reinstall makes these not-quite-A/B) ==
systemd-analyze | tee systemd-analyze-$TAG.txt
systemd-analyze blame | head -20 >> systemd-analyze-$TAG.txt
# Optional micro-timings with hyperfine (install if wanted):
# hyperfine --warmup 3 'gcc -O2 -o /tmp/x bench.c && rm /tmp/x'
############################################################################
# ONE-TIME TOOL SETUP (do this on BOTH sides, same versions)
############################################################################
mkdir -p ~/basic-benchmark/tools/pt
# App workloads (see == 5c): stock apps + a fixed corpus fetched once into
# tools/corpus/ (checksum-pinned). Install only the apps you want to measure.
# sudo dnf install -y libreoffice-writer libreoffice-calc inkscape gegl04-tools
# PassMark PerformanceTest Linux — free, CLI. Download the Linux x86-64
# zip (https://www.passmark.com/downloads/pt_linux_x64.zip), then:
# cd ~/basic-benchmark/tools/pt && unzip pt_linux_x64.zip
# # -> PerformanceTest/PerformanceTest_Linux_x86-64 (older zips: pt_linux_x64)
# chmod +x PerformanceTest/PerformanceTest_Linux_x86-64
# sudo dnf install -y ncurses-libs ncurses-compat-libs
# # if it complains about libncurses.so.5:
# sudo ln -sf /usr/lib64/libncurses.so.6 /usr/lib64/libncurses.so.5
# ./PerformanceTest/PerformanceTest_Linux_x86-64 -r 3 -d 1 # smoke test
# product page: https://www.passmark.com/products/pt_linux/
# Keep these binaries (or their exact versions + sha256) for the AFTER run,
# so both sides use identical builds.
############################################################################
# RESULTS TABLE — fill medians (3 runs) after each side
############################################################################
Test Unit BEFORE AFTER Δ
---------------------------------------------------------------------------
PassMark CPU Mark mark ______ ______ __%
PassMark CPU Single Threaded M ops/s ______ ______ __%
PassMark Memory Mark mark ______ ______ __%
CPU 7z (7z b) MIPS ______ ______ __%
CPU single (sysbench 1t) events/s ______ ______ __%
CPU all (sysbench Nt) events/s ______ ______ __%
Memory bandwidth (sysbench) MB/s ______ ______ __%
LLM generation (2B, CPU/RAM) tok/s ______ ______ __%
LibreOffice docs->PDF docs/s ______ ______ __%
GEGL image ops ops/s ______ ______ __%
Inkscape SVG->PNG images/s ______ ______ __%
Kernel build sec ______ ______ __%
OpenSSL aes-256-gcm GB/s ______ ______ __%
GPU glmark2 score ______ ______ __%
GPU vkmark score ______ ______ __%
SSD seqread (fio) MB/s ______ ______ __%
SSD randread 4k (fio) IOPS ______ ______ __%
CPU idle temp °C ______ ______ -
CPU load temp (all-core) °C ______ ______ -
Sustained all-core clock MHz ______ ______ -
Package power under load W ______ ______ -
Avg package power (whole pass) W ______ ______ __%
Energy — whole pass kJ ______ ______ __%
Efficiency PassMark CPU Mark mark/W ______ ______ __%
---------------------------------------------------------------------------
Ambient: BEFORE ___ °C AFTER ___ °C
Kernel: BEFORE ________ AFTER ________
Tool versions: BEFORE ____________ AFTER ____________
Notes: ______________________________________________________________
== HOW TO READ THE DELTA (not a benchmark) ==
- The BEFORE run is the control; the AFTER run is captured on the newly
installed system with this exact suite. Do not pre-fill after-specs.
- PassMark CPU Mark scales with core count; PassMark "CPU Single Threaded"
isolates IPC/clock. Read them together.
- Compare single-thread and all-thread CPU results separately: thread
counts differ between platforms, so total throughput alone can mislead.
- Compare memory bandwidth and latency independently; PassMark Memory Mark
also rewards the larger capacity.
- Components carried over unchanged (same SSD, same dGPU, same displays)
should land near-flat; a big move there means a setup fault, not a win.
- Compare idle vs load thermals and sustained clocks to judge the cooler.