Drops the optional RUN_PTS group, the pts-core phase, and every Phoronix reference from the scripts and docs. The suite is now just PassMark, the bundled llama.cpp, the app workloads, and the quick tools.
294 lines
18 KiB
Text
294 lines
18 KiB
Text
############################################################################
|
|
# basic-benchmark — benchmark suite (run BEFORE and AFTER)
|
|
# host: <captured per run> suite v3 portable kit
|
|
#
|
|
# Target: a stock Fedora 44 with internet. Missing tools are installed via dnf
|
|
# (INSTALL_DEPS=1). Specs are captured with ./capture-specs.sh, not hardcoded.
|
|
#
|
|
# Purpose: measure a hardware upgrade with an identical, repeatable set of
|
|
# tests on both sides — capture the AFTER run once the new build is up.
|
|
#
|
|
# Quick path (run-benchmarks.sh is split; see README):
|
|
# headless: ./run-benchmarks.sh before headless # over SSH (ssh -t)
|
|
# headed: ./run-benchmarks.sh before headed # in the desktop session
|
|
# repeat both with "after" post-swap; tags become before-headless / before-headed
|
|
# This file is the reference for what that script runs, plus extras.
|
|
#
|
|
# Primary cross-platform score: PassMark PerformanceTest (cpubenchmark.net).
|
|
# Technical depth: the quick tools below.
|
|
############################################################################
|
|
|
|
############################################################################
|
|
# RUN LIST — what to run, in order, with rough one-pass times (idle desktop)
|
|
############################################################################
|
|
# benchmark part by ~time notes
|
|
0 env capture headless script <1 min
|
|
0b idle baseline (>=15s) headless script ~18 s true idle window
|
|
1 PassMark PT Linux -r 1 / -r 2 headless script 8-12 min CPU suite + Memory suite
|
|
2 7z b (all + mmt1) headless script ~2 min
|
|
3 openssl speed aes-256-gcm / sha256 headless script 2-3 min
|
|
4 sysbench cpu (1t + Nt) + memory headless script 1-2 min
|
|
5 llama.cpp generation (2B, CPU/RAM) headless script ~15 s bundled; tok/s
|
|
5b LibreOffice: 20 documents -> PDF headless script ~1 min fixed corpus
|
|
5c GEGL: 9 image ops (6000x4000) headless script ~1 min fixed corpus
|
|
5d Inkscape: SVG -> PNG headless script 1-2 min fixed corpus
|
|
5e GIMP: 4 batch ops headless script 1-2 min GIMP 3, best-effort
|
|
6 fio 4x30s headless script 3-4 min runs as your user
|
|
7 stress-ng 5 min + turbostat headless script ~8 min thermals/power
|
|
8 systemd-analyze headless script <1 min
|
|
glmark2 + vkmark headed script ~1 min needs desktop session
|
|
---------- headless total ---------- ~25-30 min
|
|
---------- headed total ------------ ~1 min
|
|
=> ground rule wants MEDIAN OF 3 runs: x3 the above
|
|
* powerlog.py samples power/temp ~1/s for the WHOLE part ->
|
|
results/power-<tag>.csv + results/phases-<tag>.csv
|
|
|
|
== GROUND RULES FOR A FAIR BEFORE/AFTER ==
|
|
1. Same OS image + kernel + mesa + tool versions on both runs. Record them.
|
|
2. Same storage drive in both runs (the WD Blue SN570). Same filesystem/options.
|
|
3. Same GPU, same driver, same displays, same resolution/refresh. Record which
|
|
GPU drives the displays (iGPU vs dGPU) — it changes between platforms.
|
|
4. Kill background load: close browsers, stop docker and VMs.
|
|
sudo systemctl stop docker docker.socket
|
|
5. Pin CPU governor to performance for timed runs (BEFORE default = powersave):
|
|
sudo cpupower frequency-set -g performance
|
|
# fallback if no cpupower:
|
|
for f in /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor; do \
|
|
echo performance | sudo tee "$f" >/dev/null; done
|
|
run-benchmarks.sh does this AND restores the original governor on exit.
|
|
To measure another profile, set GOVERNOR= (the value is appended to the tag):
|
|
GOVERNOR=powersave ./run-benchmarks.sh after headless -> after-powersave-headless
|
|
GOVERNOR=performance ./run-benchmarks.sh after headless -> after-performance-headless
|
|
6. Warm up once, then take the MEDIAN of 3 runs. Throw away the first.
|
|
7. Note room/ambient temperature next to thermal results; thermals are noisy.
|
|
8. Use the SAME tool versions on both sides (with their (point) releases:
|
|
PassMark PT build). Record them in env-$TAG.txt.
|
|
9. BACK UP before any reinstall / reimage:
|
|
~/basic-benchmark/results/
|
|
copy it somewhere safe (off the machine) before reinstalling.
|
|
|
|
== 0. ENV CAPTURE (run first, each side) ==
|
|
export TAG=before # or: after
|
|
mkdir -p ~/basic-benchmark/results && cd ~/basic-benchmark/results
|
|
{ date; uname -a; lscpu; free -h; lsblk;
|
|
stress-ng --version;
|
|
glxinfo -B 2>/dev/null; vulkaninfo --summary 2>/dev/null | head -40;
|
|
sha256sum ~/basic-benchmark/tools/pt/PerformanceTest/PerformanceTest_Linux_x86-64 2>/dev/null;
|
|
} > env-$TAG.txt 2>&1
|
|
|
|
== 1. PASSMARK PERFORMANCETEST LINUX (this IS the cpubenchmark.net test) ==
|
|
# Free, CLI-only Linux build. Output is directly comparable to the numbers
|
|
# on cpubenchmark.net and cross-platform with the Windows product.
|
|
# CPU Mark, Memory Mark, and every per-test subscore incl. single-thread.
|
|
#
|
|
# -r 1 = CPU only -> results_cpu.yml
|
|
# -r 2 = Memory only-> results_memory.yml
|
|
# -r 3 = All tests -> results_all.yml <-- use this
|
|
# -d 1..4 = Short/Medium/Long/VeryLong -i <iters> -p <processes>
|
|
#
|
|
PT=~/basic-benchmark/tools/pt/PerformanceTest/PerformanceTest_Linux_x86-64
|
|
# (older PassMark zips put the binary at tools/pt/pt_linux_x64; the script
|
|
# auto-detects either; PassMark needs TERM set — the script sets TERM=xterm)
|
|
"$PT" -r 3 -d 2 | tee passmark-$TAG.txt
|
|
mv -f results_all.yml passmark-all-$TAG.yml
|
|
# KEEP the SAME -r/-d/-i/-p on both sides. Leave -p at default so it scales
|
|
# to each machine's thread count.
|
|
# Result to record: "CPU Mark", "CPU Single Threaded", "Memory Mark".
|
|
|
|
== 2. CPU — THROUGHPUT / COMPRESSION ==
|
|
# 7-Zip built-in benchmark (integer, scales with cores + cache + memory)
|
|
7z b -mmt$(nproc) | tee 7z-$TAG.txt
|
|
# zstd / xz on a FIXED corpus: create ONCE and reuse for both sides so the
|
|
# input bytes are identical (regenerating random data per run is not comparable)
|
|
[ -f /tmp/corpus-2g.bin ] || head -c 2G /dev/urandom > /tmp/corpus-2g.bin
|
|
zstd -b9 -i3 /tmp/corpus-2g.bin 2>&1 | tee zstd-$TAG.txt
|
|
/usr/bin/time -v xz -9 -T$(nproc) -k -f /tmp/corpus-2g.bin 2>&1 | tee xz-$TAG.txt
|
|
# crypto (AES/SHA throughput; exercises AVX/VAES path differently per vendor)
|
|
openssl speed -multi $(nproc) -evp aes-256-gcm 2>&1 | tail -20 | tee openssl-aes-$TAG.txt
|
|
openssl speed -multi $(nproc) -evp sha256 2>&1 | tail -20 | tee openssl-sha-$TAG.txt
|
|
|
|
== 4. CPU — SINGLE vs MULTI THREAD ==
|
|
sudo dnf install -y sysbench
|
|
sysbench cpu --threads=1 run | tee sysbench-1t-$TAG.txt # single-thread
|
|
sysbench cpu --threads=$(nproc) run | tee sysbench-nt-$TAG.txt # all-thread
|
|
# NOTE: compare single-thread AND all-thread results separately — the
|
|
# thread count will differ between the two platforms.
|
|
|
|
== 5. MEMORY — BANDWIDTH ==
|
|
sysbench memory --threads=$(nproc) run | tee sysbench-mem-$TAG.txt
|
|
# Compare bandwidth between the two builds.
|
|
|
|
== 5b. LLM INFERENCE — token generation on CPU/RAM (bundled llama.cpp) ==
|
|
# Self-contained: runtime + model live under tools/llama/. Generation (tg) is
|
|
# memory-bound, so tg t/s reflects RAM bandwidth for LLM serving.
|
|
tools/llama/bin/llama-bench \
|
|
-m tools/llama/models/MiniCPM5-2B-Q8_0.gguf -ngl 0 -p 0 -n 128 -r 2 \
|
|
| tee llama-$TAG.txt
|
|
# -ngl 0 forces CPU/RAM (no GPU offload). Record the tg t/s number.
|
|
|
|
== 5c. APP WORKLOADS (LibreOffice / GEGL / Inkscape / GIMP) ==
|
|
# Stock distro apps run over a fixed sample corpus; run-benchmarks.sh fetches
|
|
# the corpus ONCE into tools/corpus/ (checksum-pinned) and times the app's own
|
|
# CLI, reporting a rate (items/s, higher is better). Missing apps are skipped.
|
|
sudo dnf install -y libreoffice-writer libreoffice-calc inkscape gegl04-tools gimp
|
|
RUN_APPS=1 ./run-benchmarks.sh before headless # default on
|
|
RUN_APPS=0 ./run-benchmarks.sh before headless # skip the group
|
|
# Phases / scores are named after the benchmark: libreoffice gegl inkscape gimp.
|
|
# GIMP 3 reshaped Script-Fu, so the GIMP step is best-effort.
|
|
|
|
== 7. GPU — DISCRETE RX 6700 XT (same card both runs; HEADED part only) ==
|
|
sudo dnf install -y glmark2 vkmark vulkan-tools mesa-demos
|
|
# Monitors hang off the Cezanne iGPU, so PIN the dGPU or these tests would
|
|
# silently measure the iGPU. run-benchmarks.sh auto-detects and exports:
|
|
# DRI_PRIME=pci-0000:03:00.0
|
|
# MESA_VK_DEVICE_SELECT=1002:73df VK_LOADER_DEVICE_SELECT=1002:73df
|
|
# (from /sys/class/drm: the dGPU is the card with no connected connector)
|
|
glxinfo -B | tee glxinfo-$TAG.txt
|
|
vulkaninfo --summary | tee vulkan-$TAG.txt
|
|
# One representative scene each (fast, comparable); tune with GLMARK2_DURATION,
|
|
# VKMARK_DURATION and VKMARK_BENCH (texture|shading|vertex).
|
|
glmark2 -b terrain:duration=10 2>&1 | tail -25 | tee glmark2-$TAG.txt
|
|
vkmark -b texture:duration=10 2>&1 | tail -30 | tee vkmark-$TAG.txt
|
|
# Confirm which GPU rendered: grep -i 'device\|renderer' glxinfo-$TAG.txt
|
|
|
|
== 8. STORAGE / IO (same drive — mostly a control; watch for regressions) ==
|
|
FIO=$HOME/.cache/basic-benchmark-fio.bin # user-owned; no sudo needed
|
|
mkdir -p "$(dirname "$FIO")"
|
|
for JOB in "seqread read 1M 16" "seqwrite write 1M 16" \
|
|
"randread randread 4k 32" "randwrite randwrite 4k 32"; do
|
|
set -- $JOB
|
|
fio --name=$1 --rw=$2 --bs=$3 --numjobs=1 --iodepth=$4 \
|
|
--size=1G --direct=1 --time_based --runtime=30 \
|
|
--filename=$FIO --group_reporting \
|
|
--output-format=json --output=fio-$1-$TAG.json
|
|
rm -f $FIO
|
|
done
|
|
# NOTE: hdparm -tT is ATA-only and does nothing useful on the NVMe SN570
|
|
# (it just returns "Inappropriate ioctl for device"), so it is dropped;
|
|
# the fio numbers above are the storage result to record.
|
|
|
|
== 9. THERMALS / POWER / SUSTAINED CLOCKS (the cooler change is measured here) ==
|
|
sensors > sensors-idle-$TAG.txt
|
|
stress-ng --cpu 0 --cpu-method matrixprod --timeout 300 --metrics-brief \
|
|
> stress-ng-$TAG.txt 2>&1 &
|
|
sleep 120 # let it reach steady state
|
|
sensors > sensors-load-$TAG.txt
|
|
cat /sys/devices/system/cpu/cpu0/cpufreq/scaling_cur_freq >> sensors-load-$TAG.txt
|
|
sudo turbostat --interval 5 --show Core,CPU,Busy%,Bzy_MHz,PkgWatt,RAMWatt,PkgTmp \
|
|
--timeout 30 | tee turbostat-load-$TAG.txt
|
|
wait
|
|
# Record: idle temp, load temp, sustained all-core clock, package watts.
|
|
# Record the fields the platform exposes (PkgWatt/RAMWatt where available).
|
|
|
|
== 9b. CONTINUOUS POWER / THERMAL LOGGING (efficiency, not just peaks) ==
|
|
# powerlog.py runs in the background for the WHOLE part and writes a ~1 Hz
|
|
# CSV. run-benchmarks.sh already starts/stops it and drops a phase marker
|
|
# before each benchmark, so the log can be split per benchmark afterwards.
|
|
# <tag> includes the part, e.g. power-before-headless.csv / power-before-headed.csv.
|
|
# GPU columns are tied to the headless dGPU (the amdgpu hwmon with power1_average).
|
|
# results/power-<tag>.csv epoch,iso,pkg_w,core_w,cpu_temp_c,cpu_mhz,
|
|
# gpu_w,gpu_temp_c,gpu_busy,fan1,fan2
|
|
# results/phases-<tag>.csv epoch,phase (one line per benchmark)
|
|
# Run it by hand for a single test (Ctrl-C to stop):
|
|
# sudo python3 powerlog.py --tag before
|
|
# Power source: RAPL package/core energy counters
|
|
# (/sys/class/powercap/intel-rapl*), root-only. GPU watts from the amdgpu
|
|
# hwmon power1_average. Thermals: k10temp/coretemp + amdgpu + nct6798 fans.
|
|
# If you add a smart plug later (Tasmota/Shelly/UPS), poll it in a second
|
|
# background loop and merge on the epoch column for true wall power.
|
|
#
|
|
# Graphs + per-phase table + efficiency:
|
|
# python3 power-report.py --tags before-headless # one part
|
|
# python3 power-report.py --tags before-headless before-headed # both parts
|
|
# python3 power-report.py --tags before-headless before-headed \
|
|
# after-headless after-headed # comparison graphs
|
|
# -> results/graphs/power-<tag>.png and compare-before-after.png
|
|
# Cross-device / before-after score comparison (one bar chart per benchmark):
|
|
# python3 compare-report.py # table -> results/compare-scores.csv
|
|
# # + results/graphs/compare-*.png
|
|
# Efficiency needs numbers in results/scores-<tag>.csv; the name should match
|
|
# a phase marker (exact or substring). Phase names are:
|
|
# env PassMark 7zip openssl sysbench
|
|
# Example:
|
|
# PassMark,24500
|
|
# libreoffice,3.9 <- docs/s; energy = avg_w * seconds
|
|
# Metrics reported: avg/p95/peak W, avg/peak °C, energy (kJ) per phase,
|
|
# and score-per-watt (higher = better; lower energy = better).
|
|
|
|
== 10. DESKTOP / SYSTEM (informational — a reinstall makes these not-quite-A/B) ==
|
|
systemd-analyze | tee systemd-analyze-$TAG.txt
|
|
systemd-analyze blame | head -20 >> systemd-analyze-$TAG.txt
|
|
# Optional micro-timings with hyperfine (install if wanted):
|
|
# hyperfine --warmup 3 'gcc -O2 -o /tmp/x bench.c && rm /tmp/x'
|
|
|
|
############################################################################
|
|
# ONE-TIME TOOL SETUP (do this on BOTH sides, same versions)
|
|
############################################################################
|
|
mkdir -p ~/basic-benchmark/tools/pt
|
|
|
|
# App workloads (see == 5c): stock apps + a fixed corpus fetched once into
|
|
# tools/corpus/ (checksum-pinned). Install only the apps you want to measure.
|
|
# sudo dnf install -y libreoffice-writer libreoffice-calc inkscape gegl04-tools gimp
|
|
|
|
# PassMark PerformanceTest Linux — free, CLI. Download the Linux x86-64
|
|
# zip (https://www.passmark.com/downloads/pt_linux_x64.zip), then:
|
|
# cd ~/basic-benchmark/tools/pt && unzip pt_linux_x64.zip
|
|
# # -> PerformanceTest/PerformanceTest_Linux_x86-64 (older zips: pt_linux_x64)
|
|
# chmod +x PerformanceTest/PerformanceTest_Linux_x86-64
|
|
# sudo dnf install -y ncurses-libs ncurses-compat-libs
|
|
# # if it complains about libncurses.so.5:
|
|
# sudo ln -sf /usr/lib64/libncurses.so.6 /usr/lib64/libncurses.so.5
|
|
# ./PerformanceTest/PerformanceTest_Linux_x86-64 -r 3 -d 1 # smoke test
|
|
# product page: https://www.passmark.com/products/pt_linux/
|
|
|
|
# Keep these binaries (or their exact versions + sha256) for the AFTER run,
|
|
# so both sides use identical builds.
|
|
|
|
############################################################################
|
|
# RESULTS TABLE — fill medians (3 runs) after each side
|
|
############################################################################
|
|
Test Unit BEFORE AFTER Δ
|
|
---------------------------------------------------------------------------
|
|
PassMark CPU Mark mark ______ ______ __%
|
|
PassMark CPU Single Threaded M ops/s ______ ______ __%
|
|
PassMark Memory Mark mark ______ ______ __%
|
|
CPU 7z (7z b) MIPS ______ ______ __%
|
|
CPU single (sysbench 1t) events/s ______ ______ __%
|
|
CPU all (sysbench Nt) events/s ______ ______ __%
|
|
Memory bandwidth (sysbench) MB/s ______ ______ __%
|
|
LLM generation (2B, CPU/RAM) tok/s ______ ______ __%
|
|
LibreOffice docs->PDF docs/s ______ ______ __%
|
|
GEGL image ops ops/s ______ ______ __%
|
|
Inkscape SVG->PNG images/s ______ ______ __%
|
|
GIMP batch ops ops/s ______ ______ __%
|
|
Kernel build sec ______ ______ __%
|
|
OpenSSL aes-256-gcm GB/s ______ ______ __%
|
|
GPU glmark2 score ______ ______ __%
|
|
GPU vkmark score ______ ______ __%
|
|
SSD seqread (fio) MB/s ______ ______ __%
|
|
SSD randread 4k (fio) IOPS ______ ______ __%
|
|
CPU idle temp °C ______ ______ -
|
|
CPU load temp (all-core) °C ______ ______ -
|
|
Sustained all-core clock MHz ______ ______ -
|
|
Package power under load W ______ ______ -
|
|
Avg package power (whole pass) W ______ ______ __%
|
|
Energy — whole pass kJ ______ ______ __%
|
|
Efficiency PassMark CPU Mark mark/W ______ ______ __%
|
|
---------------------------------------------------------------------------
|
|
Ambient: BEFORE ___ °C AFTER ___ °C
|
|
Kernel: BEFORE ________ AFTER ________
|
|
Tool versions: BEFORE ____________ AFTER ____________
|
|
Notes: ______________________________________________________________
|
|
|
|
== HOW TO READ THE DELTA (not a benchmark) ==
|
|
- The BEFORE run is the control; the AFTER run is captured on the newly
|
|
installed system with this exact suite. Do not pre-fill after-specs.
|
|
- PassMark CPU Mark scales with core count; PassMark "CPU Single Threaded"
|
|
isolates IPC/clock. Read them together.
|
|
- Compare single-thread and all-thread CPU results separately: thread
|
|
counts differ between platforms, so total throughput alone can mislead.
|
|
- Compare memory bandwidth and latency independently; PassMark Memory Mark
|
|
also rewards the larger capacity.
|
|
- Components carried over unchanged (same SSD, same dGPU, same displays)
|
|
should land near-flat; a big move there means a setup fault, not a win.
|
|
- Compare idle vs load thermals and sustained clocks to judge the cooler.
|