basic-benchmark: portable Fedora benchmark kit
Headless + headed suites with power/thermal logging and efficiency reporting. Includes PassMark, llama.cpp, app workloads (LibreOffice/GEGL/Inkscape/GIMP), a GOVERNOR selector, compare-report.py, and captured results.
This commit is contained in:
commit
0cd63d0233
533 changed files with 43679 additions and 0 deletions
306
benchmarks.txt
Normal file
306
benchmarks.txt
Normal file
|
|
@ -0,0 +1,306 @@
|
|||
############################################################################
|
||||
# basic-benchmark — benchmark suite (run BEFORE and AFTER)
|
||||
# host: <captured per run> suite v3 portable kit
|
||||
#
|
||||
# Target: a stock Fedora 44 with internet. Missing tools are installed via dnf
|
||||
# (INSTALL_DEPS=1). Specs are captured with ./capture-specs.sh, not hardcoded.
|
||||
#
|
||||
# Purpose: measure a hardware upgrade with an identical, repeatable set of
|
||||
# tests on both sides — capture the AFTER run once the new build is up.
|
||||
#
|
||||
# Quick path (run-benchmarks.sh is split; see README):
|
||||
# headless: ./run-benchmarks.sh before headless # over SSH (ssh -t)
|
||||
# headed: ./run-benchmarks.sh before headed # in the desktop session
|
||||
# repeat both with "after" post-swap; tags become before-headless / before-headed
|
||||
# This file is the reference for what that script runs, plus extras.
|
||||
#
|
||||
# Primary cross-platform score: PassMark PerformanceTest (cpubenchmark.net).
|
||||
# Technical depth: Phoronix Test Suite + the quick tools below.
|
||||
############################################################################
|
||||
|
||||
############################################################################
|
||||
# RUN LIST — what to run, in order, with rough one-pass times (idle desktop)
|
||||
############################################################################
|
||||
# benchmark part by ~time notes
|
||||
0 env capture headless script <1 min
|
||||
0b idle baseline (>=15s) headless script ~18 s true idle window
|
||||
1 PassMark PT Linux -r 1 / -r 2 headless script 8-12 min CPU suite + Memory suite
|
||||
2 7z b (all + mmt1) headless script ~2 min
|
||||
3 openssl speed aes-256-gcm / sha256 headless script 2-3 min
|
||||
4 sysbench cpu (1t + Nt) + memory headless script 1-2 min
|
||||
5 llama.cpp generation (2B, CPU/RAM) headless script ~15 s bundled; tok/s
|
||||
5b LibreOffice: 20 documents -> PDF headless script ~1 min fixed corpus
|
||||
5c GEGL: 9 image ops (6000x4000) headless script ~1 min fixed corpus
|
||||
5d Inkscape: SVG -> PNG headless script 1-2 min fixed corpus
|
||||
5e GIMP: 4 batch ops headless script 1-2 min GIMP 3, best-effort
|
||||
6 fio 4x30s headless script 3-4 min runs as your user
|
||||
7 stress-ng 5 min + turbostat headless script ~8 min thermals/power
|
||||
8 systemd-analyze headless script <1 min
|
||||
glmark2 + vkmark headed script 3-4 min needs desktop session
|
||||
---------- headless total ---------- ~25-30 min
|
||||
---------- headed total ------------ ~4-5 min
|
||||
+ PTS core: build-linux-kernel, c-ray, stream +30-45 min RUN_PTS=1
|
||||
=> ground rule wants MEDIAN OF 3 runs: x3 the above
|
||||
* powerlog.py samples power/temp ~1/s for the WHOLE part ->
|
||||
results/power-<tag>.csv + results/phases-<tag>.csv
|
||||
|
||||
== GROUND RULES FOR A FAIR BEFORE/AFTER ==
|
||||
1. Same OS image + kernel + mesa + tool versions on both runs. Record them.
|
||||
2. Same storage drive in both runs (the WD Blue SN570). Same filesystem/options.
|
||||
3. Same GPU, same driver, same displays, same resolution/refresh. Record which
|
||||
GPU drives the displays (iGPU vs dGPU) — it changes between platforms.
|
||||
4. Kill background load: close browsers, stop docker and VMs.
|
||||
sudo systemctl stop docker docker.socket
|
||||
5. Pin CPU governor to performance for timed runs (BEFORE default = powersave):
|
||||
sudo cpupower frequency-set -g performance
|
||||
# fallback if no cpupower:
|
||||
for f in /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor; do \
|
||||
echo performance | sudo tee "$f" >/dev/null; done
|
||||
run-benchmarks.sh does this AND restores the original governor on exit.
|
||||
To measure another profile, set GOVERNOR= (the value is appended to the tag):
|
||||
GOVERNOR=powersave ./run-benchmarks.sh after headless -> after-powersave-headless
|
||||
GOVERNOR=performance ./run-benchmarks.sh after headless -> after-performance-headless
|
||||
6. Warm up once, then take the MEDIAN of 3 runs. Throw away the first.
|
||||
7. Note room/ambient temperature next to thermal results; thermals are noisy.
|
||||
8. Use the SAME tool versions on both sides (with their (point) releases:
|
||||
PassMark PT build, PTS version). Record them in env-$TAG.txt.
|
||||
9. BACK UP before any reinstall / reimage:
|
||||
~/.phoronix-test-suite/ and ~/basic-benchmark/results/
|
||||
copy both somewhere safe (off the machine) before reinstalling.
|
||||
|
||||
== 0. ENV CAPTURE (run first, each side) ==
|
||||
export TAG=before # or: after
|
||||
mkdir -p ~/basic-benchmark/results && cd ~/basic-benchmark/results
|
||||
{ date; uname -a; lscpu; free -h; lsblk;
|
||||
phoronix-test-suite version; stress-ng --version;
|
||||
glxinfo -B 2>/dev/null; vulkaninfo --summary 2>/dev/null | head -40;
|
||||
sha256sum ~/basic-benchmark/tools/pt/PerformanceTest/PerformanceTest_Linux_x86-64 2>/dev/null;
|
||||
} > env-$TAG.txt 2>&1
|
||||
|
||||
== 1. PASSMARK PERFORMANCETEST LINUX (this IS the cpubenchmark.net test) ==
|
||||
# Free, CLI-only Linux build. Output is directly comparable to the numbers
|
||||
# on cpubenchmark.net and cross-platform with the Windows product.
|
||||
# CPU Mark, Memory Mark, and every per-test subscore incl. single-thread.
|
||||
#
|
||||
# -r 1 = CPU only -> results_cpu.yml
|
||||
# -r 2 = Memory only-> results_memory.yml
|
||||
# -r 3 = All tests -> results_all.yml <-- use this
|
||||
# -d 1..4 = Short/Medium/Long/VeryLong -i <iters> -p <processes>
|
||||
#
|
||||
PT=~/basic-benchmark/tools/pt/PerformanceTest/PerformanceTest_Linux_x86-64
|
||||
# (older PassMark zips put the binary at tools/pt/pt_linux_x64; the script
|
||||
# auto-detects either; PassMark needs TERM set — the script sets TERM=xterm)
|
||||
"$PT" -r 3 -d 2 | tee passmark-$TAG.txt
|
||||
mv -f results_all.yml passmark-all-$TAG.yml
|
||||
# KEEP the SAME -r/-d/-i/-p on both sides. Leave -p at default so it scales
|
||||
# to each machine's thread count.
|
||||
# Result to record: "CPU Mark", "CPU Single Threaded", "Memory Mark".
|
||||
|
||||
== 2. CPU — THROUGHPUT / COMPRESSION ==
|
||||
# 7-Zip built-in benchmark (integer, scales with cores + cache + memory)
|
||||
7z b -mmt$(nproc) | tee 7z-$TAG.txt
|
||||
# zstd / xz on a FIXED corpus: create ONCE and reuse for both sides so the
|
||||
# input bytes are identical (regenerating random data per run is not comparable)
|
||||
[ -f /tmp/corpus-2g.bin ] || head -c 2G /dev/urandom > /tmp/corpus-2g.bin
|
||||
zstd -b9 -i3 /tmp/corpus-2g.bin 2>&1 | tee zstd-$TAG.txt
|
||||
/usr/bin/time -v xz -9 -T$(nproc) -k -f /tmp/corpus-2g.bin 2>&1 | tee xz-$TAG.txt
|
||||
# crypto (AES/SHA throughput; exercises AVX/VAES path differently per vendor)
|
||||
openssl speed -multi $(nproc) -evp aes-256-gcm 2>&1 | tail -20 | tee openssl-aes-$TAG.txt
|
||||
openssl speed -multi $(nproc) -evp sha256 2>&1 | tail -20 | tee openssl-sha-$TAG.txt
|
||||
|
||||
== 4. CPU — SINGLE vs MULTI THREAD ==
|
||||
sudo dnf install -y sysbench
|
||||
sysbench cpu --threads=1 run | tee sysbench-1t-$TAG.txt # single-thread
|
||||
sysbench cpu --threads=$(nproc) run | tee sysbench-nt-$TAG.txt # all-thread
|
||||
# NOTE: compare single-thread AND all-thread results separately — the
|
||||
# thread count will differ between the two platforms.
|
||||
|
||||
== 5. MEMORY — BANDWIDTH ==
|
||||
sysbench memory --threads=$(nproc) run | tee sysbench-mem-$TAG.txt
|
||||
# (pts/stream below gives TRIAD/COPY/SCALE/ADD MB/s)
|
||||
# Compare bandwidth between the two builds.
|
||||
|
||||
== 5b. LLM INFERENCE — token generation on CPU/RAM (bundled llama.cpp) ==
|
||||
# Self-contained: runtime + model live under tools/llama/. Generation (tg) is
|
||||
# memory-bound, so tg t/s reflects RAM bandwidth for LLM serving.
|
||||
tools/llama/bin/llama-bench \
|
||||
-m tools/llama/models/MiniCPM5-2B-Q8_0.gguf -ngl 0 -p 0 -n 128 -r 2 \
|
||||
| tee llama-$TAG.txt
|
||||
# -ngl 0 forces CPU/RAM (no GPU offload). Record the tg t/s number.
|
||||
|
||||
== 5c. APP WORKLOADS (LibreOffice / GEGL / Inkscape / GIMP — no Phoronix) ==
|
||||
# Stock distro apps run over a fixed sample corpus; run-benchmarks.sh fetches
|
||||
# the corpus ONCE into tools/corpus/ (checksum-pinned) and times the app's own
|
||||
# CLI, reporting a rate (items/s, higher is better). Missing apps are skipped.
|
||||
sudo dnf install -y libreoffice-writer libreoffice-calc inkscape gegl04-tools gimp
|
||||
RUN_APPS=1 ./run-benchmarks.sh before headless # default on
|
||||
RUN_APPS=0 ./run-benchmarks.sh before headless # skip the group
|
||||
# Phases / scores are named after the benchmark: libreoffice gegl inkscape gimp.
|
||||
# GIMP 3 reshaped Script-Fu, so the GIMP step is best-effort.
|
||||
|
||||
== 6. REAL-WORLD BUILD / RENDER (short Phoronix core) ==
|
||||
# Phoronix Test Suite pins versions/counts and stores results you can diff.
|
||||
# Installed here: 10.8.6
|
||||
sudo phoronix-test-suite make-download-cache # once, before going offline
|
||||
phoronix-test-suite benchmark \
|
||||
pts/build-linux-kernel \
|
||||
pts/c-ray \
|
||||
pts/stream
|
||||
# verify names if any differ:
|
||||
phoronix-test-suite list-available-tests | grep -Ei 'build|c-ray|stream'
|
||||
|
||||
== 7. GPU — DISCRETE RX 6700 XT (same card both runs; HEADED part only) ==
|
||||
sudo dnf install -y glmark2 vkmark vulkan-tools mesa-demos
|
||||
# Monitors hang off the Cezanne iGPU, so PIN the dGPU or these tests would
|
||||
# silently measure the iGPU. run-benchmarks.sh auto-detects and exports:
|
||||
# DRI_PRIME=pci-0000:03:00.0
|
||||
# MESA_VK_DEVICE_SELECT=1002:73df VK_LOADER_DEVICE_SELECT=1002:73df
|
||||
# (from /sys/class/drm: the dGPU is the card with no connected connector)
|
||||
glxinfo -B | tee glxinfo-$TAG.txt
|
||||
vulkaninfo --summary | tee vulkan-$TAG.txt
|
||||
glmark2 -b :duration=30 2>&1 | tail -25 | tee glmark2-$TAG.txt
|
||||
vkmark --duration 60 2>&1 | tail -30 | tee vkmark-$TAG.txt
|
||||
phoronix-test-suite benchmark pts/vkmark # if available in cache
|
||||
# Confirm which GPU rendered: grep -i 'device\|renderer' glxinfo-$TAG.txt
|
||||
|
||||
== 8. STORAGE / IO (same drive — mostly a control; watch for regressions) ==
|
||||
FIO=$HOME/.cache/basic-benchmark-fio.bin # user-owned; no sudo needed
|
||||
mkdir -p "$(dirname "$FIO")"
|
||||
for JOB in "seqread read 1M 16" "seqwrite write 1M 16" \
|
||||
"randread randread 4k 32" "randwrite randwrite 4k 32"; do
|
||||
set -- $JOB
|
||||
fio --name=$1 --rw=$2 --bs=$3 --numjobs=1 --iodepth=$4 \
|
||||
--size=1G --direct=1 --time_based --runtime=30 \
|
||||
--filename=$FIO --group_reporting \
|
||||
--output-format=json --output=fio-$1-$TAG.json
|
||||
rm -f $FIO
|
||||
done
|
||||
# NOTE: hdparm -tT is ATA-only and does nothing useful on the NVMe SN570
|
||||
# (it just returns "Inappropriate ioctl for device"), so it is dropped;
|
||||
# the fio numbers above are the storage result to record.
|
||||
|
||||
== 9. THERMALS / POWER / SUSTAINED CLOCKS (the cooler change is measured here) ==
|
||||
sensors > sensors-idle-$TAG.txt
|
||||
stress-ng --cpu 0 --cpu-method matrixprod --timeout 300 --metrics-brief \
|
||||
> stress-ng-$TAG.txt 2>&1 &
|
||||
sleep 120 # let it reach steady state
|
||||
sensors > sensors-load-$TAG.txt
|
||||
cat /sys/devices/system/cpu/cpu0/cpufreq/scaling_cur_freq >> sensors-load-$TAG.txt
|
||||
sudo turbostat --interval 5 --show Core,CPU,Busy%,Bzy_MHz,PkgWatt,RAMWatt,PkgTmp \
|
||||
--timeout 30 | tee turbostat-load-$TAG.txt
|
||||
wait
|
||||
# Record: idle temp, load temp, sustained all-core clock, package watts.
|
||||
# Record the fields the platform exposes (PkgWatt/RAMWatt where available).
|
||||
|
||||
== 9b. CONTINUOUS POWER / THERMAL LOGGING (efficiency, not just peaks) ==
|
||||
# powerlog.py runs in the background for the WHOLE part and writes a ~1 Hz
|
||||
# CSV. run-benchmarks.sh already starts/stops it and drops a phase marker
|
||||
# before each benchmark, so the log can be split per benchmark afterwards.
|
||||
# <tag> includes the part, e.g. power-before-headless.csv / power-before-headed.csv.
|
||||
# GPU columns are tied to the headless dGPU (the amdgpu hwmon with power1_average).
|
||||
# results/power-<tag>.csv epoch,iso,pkg_w,core_w,cpu_temp_c,cpu_mhz,
|
||||
# gpu_w,gpu_temp_c,gpu_busy,fan1,fan2
|
||||
# results/phases-<tag>.csv epoch,phase (one line per benchmark)
|
||||
# Run it by hand for a single test (Ctrl-C to stop):
|
||||
# sudo python3 powerlog.py --tag before
|
||||
# Power source: RAPL package/core energy counters
|
||||
# (/sys/class/powercap/intel-rapl*), root-only. GPU watts from the amdgpu
|
||||
# hwmon power1_average. Thermals: k10temp/coretemp + amdgpu + nct6798 fans.
|
||||
# If you add a smart plug later (Tasmota/Shelly/UPS), poll it in a second
|
||||
# background loop and merge on the epoch column for true wall power.
|
||||
#
|
||||
# Graphs + per-phase table + efficiency:
|
||||
# python3 power-report.py --tags before-headless # one part
|
||||
# python3 power-report.py --tags before-headless before-headed # both parts
|
||||
# python3 power-report.py --tags before-headless before-headed \
|
||||
# after-headless after-headed # comparison graphs
|
||||
# -> results/graphs/power-<tag>.png and compare-before-after.png
|
||||
# Cross-device / before-after score comparison (one bar chart per benchmark):
|
||||
# python3 compare-report.py # table -> results/compare-scores.csv
|
||||
# # + results/graphs/compare-*.png
|
||||
# Efficiency needs numbers in results/scores-<tag>.csv; the name should match
|
||||
# a phase marker (exact or substring). Phase names are:
|
||||
# env PassMark 7zip openssl sysbench
|
||||
# Example:
|
||||
# PassMark,24500
|
||||
# pts-core,420 <- kernel-build seconds; energy = avg_w * seconds
|
||||
# Metrics reported: avg/p95/peak W, avg/peak °C, energy (kJ) per phase,
|
||||
# and score-per-watt (higher = better; lower energy = better).
|
||||
|
||||
== 10. DESKTOP / SYSTEM (informational — a reinstall makes these not-quite-A/B) ==
|
||||
systemd-analyze | tee systemd-analyze-$TAG.txt
|
||||
systemd-analyze blame | head -20 >> systemd-analyze-$TAG.txt
|
||||
# Optional micro-timings with hyperfine (install if wanted):
|
||||
# hyperfine --warmup 3 'gcc -O2 -o /tmp/x bench.c && rm /tmp/x'
|
||||
|
||||
############################################################################
|
||||
# ONE-TIME TOOL SETUP (do this on BOTH sides, same versions)
|
||||
############################################################################
|
||||
mkdir -p ~/basic-benchmark/tools/pt
|
||||
|
||||
# App workloads (see == 5c): stock apps + a fixed corpus fetched once into
|
||||
# tools/corpus/ (checksum-pinned). Install only the apps you want to measure.
|
||||
# sudo dnf install -y libreoffice-writer libreoffice-calc inkscape gegl04-tools gimp
|
||||
|
||||
# PassMark PerformanceTest Linux — free, CLI. Download the Linux x86-64
|
||||
# zip (https://www.passmark.com/downloads/pt_linux_x64.zip), then:
|
||||
# cd ~/basic-benchmark/tools/pt && unzip pt_linux_x64.zip
|
||||
# # -> PerformanceTest/PerformanceTest_Linux_x86-64 (older zips: pt_linux_x64)
|
||||
# chmod +x PerformanceTest/PerformanceTest_Linux_x86-64
|
||||
# sudo dnf install -y ncurses-libs ncurses-compat-libs
|
||||
# # if it complains about libncurses.so.5:
|
||||
# sudo ln -sf /usr/lib64/libncurses.so.6 /usr/lib64/libncurses.so.5
|
||||
# ./PerformanceTest/PerformanceTest_Linux_x86-64 -r 3 -d 1 # smoke test
|
||||
# product page: https://www.passmark.com/products/pt_linux/
|
||||
|
||||
# Keep these binaries (or their exact versions + sha256) for the AFTER run,
|
||||
# so both sides use identical builds.
|
||||
|
||||
############################################################################
|
||||
# RESULTS TABLE — fill medians (3 runs) after each side
|
||||
############################################################################
|
||||
Test Unit BEFORE AFTER Δ
|
||||
---------------------------------------------------------------------------
|
||||
PassMark CPU Mark mark ______ ______ __%
|
||||
PassMark CPU Single Threaded M ops/s ______ ______ __%
|
||||
PassMark Memory Mark mark ______ ______ __%
|
||||
CPU 7z (7z b) MIPS ______ ______ __%
|
||||
CPU single (sysbench 1t) events/s ______ ______ __%
|
||||
CPU all (sysbench Nt) events/s ______ ______ __%
|
||||
Memory bandwidth (sysbench) MB/s ______ ______ __%
|
||||
LLM generation (2B, CPU/RAM) tok/s ______ ______ __%
|
||||
LibreOffice docs->PDF docs/s ______ ______ __%
|
||||
GEGL image ops ops/s ______ ______ __%
|
||||
Inkscape SVG->PNG images/s ______ ______ __%
|
||||
GIMP batch ops ops/s ______ ______ __%
|
||||
Kernel build sec ______ ______ __%
|
||||
OpenSSL aes-256-gcm GB/s ______ ______ __%
|
||||
GPU glmark2 score ______ ______ __%
|
||||
GPU vkmark score ______ ______ __%
|
||||
SSD seqread (fio) MB/s ______ ______ __%
|
||||
SSD randread 4k (fio) IOPS ______ ______ __%
|
||||
CPU idle temp °C ______ ______ -
|
||||
CPU load temp (all-core) °C ______ ______ -
|
||||
Sustained all-core clock MHz ______ ______ -
|
||||
Package power under load W ______ ______ -
|
||||
Avg package power (whole pass) W ______ ______ __%
|
||||
Energy — whole pass kJ ______ ______ __%
|
||||
Efficiency PassMark CPU Mark mark/W ______ ______ __%
|
||||
---------------------------------------------------------------------------
|
||||
Ambient: BEFORE ___ °C AFTER ___ °C
|
||||
Kernel: BEFORE ________ AFTER ________
|
||||
Tool versions: BEFORE ____________ AFTER ____________
|
||||
Notes: ______________________________________________________________
|
||||
|
||||
== HOW TO READ THE DELTA (not a benchmark) ==
|
||||
- The BEFORE run is the control; the AFTER run is captured on the newly
|
||||
installed system with this exact suite. Do not pre-fill after-specs.
|
||||
- PassMark CPU Mark scales with core count; PassMark "CPU Single Threaded"
|
||||
isolates IPC/clock. Read them together.
|
||||
- Compare single-thread and all-thread CPU results separately: thread
|
||||
counts differ between platforms, so total throughput alone can mislead.
|
||||
- Compare memory bandwidth and latency independently; PassMark Memory Mark
|
||||
also rewards the larger capacity.
|
||||
- Components carried over unchanged (same SSD, same dGPU, same displays)
|
||||
should land near-flat; a big move there means a setup fault, not a win.
|
||||
- Compare idle vs load thermals and sustained clocks to judge the cooler.
|
||||
Loading…
Reference in a new issue