Benchmark prs - #1529
Benchmark prs#1529jprendes wants to merge 9 commits into
Conversation
8ad9393 to
0e181c4
Compare
bc499ea to
a277ac1
Compare
e660953 to
9ce9ed0
Compare
Benchmark Resultskvm / amd (Linux) (❌ *1.81x slower* → 🚀 **2.41x faster**)function_call_serialization
guest_calls
guest_functions_with_large_parameters
sample_workloads
sandboxes
shared_memory
snapshots
Summary
kvm / intel (Linux) (❌ *1.88x slower* → 🚀 **7.91x faster**)function_call_serialization
guest_calls
guest_functions_with_large_parameters
sample_workloads
sandboxes
shared_memory
snapshots
Summary
mshv3 / amd (Linux) (❌ *2.19x slower* → 🚀 **2.85x faster**)function_call_serialization
guest_calls
guest_functions_with_large_parameters
sample_workloads
sandboxes
shared_memory
snapshots
Summary
mshv3 / intel (Linux) (❌ *1.56x slower* → 🚀 **1.89x faster**)function_call_serialization
guest_calls
guest_functions_with_large_parameters
sample_workloads
sandboxes
shared_memory
snapshots
Summary
hyperv-ws2025 / amd (Windows) (❌ *4.96x slower* → ✅ **1.42x faster**)function_call_serialization
guest_calls
guest_functions_with_large_parameters
sample_workloads
sandboxes
shared_memory
snapshots
Summary
hyperv-ws2025 / intel (Windows) (❌ *6.68x slower* → ✅ **1.31x faster**)function_call_serialization
guest_calls
guest_functions_with_large_parameters
sample_workloads
sandboxes
shared_memory
snapshots
Summary
|
Benchmark Resultskvm / amd (Linux) (❌ *1.92x slower* → 🚀 **2.44x faster**)function_call_serialization
guest_calls
guest_functions_with_large_parameters
sample_workloads
sandboxes
shared_memory
snapshots
Summary
kvm / intel (Linux) (❌ *1.70x slower* → 🚀 **5.39x faster**)function_call_serialization
guest_calls
guest_functions_with_large_parameters
sample_workloads
sandboxes
shared_memory
snapshots
Summary
mshv3 / amd (Linux) (❌ *1.98x slower* → 🚀 **2.60x faster**)function_call_serialization
guest_calls
guest_functions_with_large_parameters
sample_workloads
sandboxes
shared_memory
snapshots
Summary
mshv3 / intel (Linux) (❌ *2.33x slower* → 🚀 **1.94x faster**)function_call_serialization
guest_calls
guest_functions_with_large_parameters
sample_workloads
sandboxes
shared_memory
snapshots
Summary
hyperv-ws2025 / amd (Windows) (❌ *8.66x slower* → ✅ **1.11x faster**)function_call_serialization
guest_calls
guest_functions_with_large_parameters
sample_workloads
sandboxes
shared_memory
snapshots
Summary
hyperv-ws2025 / intel (Windows) (❌ *3.32x slower* → ✅ **1.24x faster**)function_call_serialization
guest_calls
guest_functions_with_large_parameters
sample_workloads
sandboxes
shared_memory
snapshots
Summary
|
Benchmark Resultskvm / amd (Linux) (❌ *1.35x slower* → 🚀 **2.81x faster**)function_call_serialization
guest_calls
guest_functions_with_large_parameters
sample_workloads
sandboxes
shared_memory
snapshots
Summary
kvm / intel (Linux) (✅ **1.01x slower** → 🚀 **13.07x faster**)function_call_serialization
guest_calls
guest_functions_with_large_parameters
sample_workloads
sandboxes
shared_memory
snapshots
Summary
mshv3 / amd (Linux) (❌ *1.44x slower* → 🚀 **4.04x faster**)function_call_serialization
guest_calls
guest_functions_with_large_parameters
sample_workloads
sandboxes
shared_memory
snapshots
Summary
mshv3 / intel (Linux) (✅ **1.10x slower** → 🚀 **2.89x faster**)function_call_serialization
guest_calls
guest_functions_with_large_parameters
sample_workloads
sandboxes
shared_memory
snapshots
Summary
hyperv-ws2025 / amd (Windows) (❌ *3.62x slower* → 🚀 **1.82x faster**)function_call_serialization
guest_calls
guest_functions_with_large_parameters
sample_workloads
sandboxes
shared_memory
snapshots
Summary
hyperv-ws2025 / intel (Windows) (❌ *1.99x slower* → ✅ **1.38x faster**)function_call_serialization
guest_calls
guest_functions_with_large_parameters
sample_workloads
sandboxes
shared_memory
snapshots
Summary
|
Benchmark Resultskvm / amd (Linux) (❌ *1.81x slower* → 🚀 **2.37x faster**)function_call_serialization
guest_calls
guest_functions_with_large_parameters
sample_workloads
sandboxes
shared_memory
snapshots
Summary
kvm / intel (Linux) (❌ *1.92x slower* → 🚀 **8.06x faster**)function_call_serialization
guest_calls
guest_functions_with_large_parameters
sample_workloads
sandboxes
shared_memory
snapshots
Summary
mshv3 / amd (Linux) (❌ *2.29x slower* → 🚀 **2.72x faster**)function_call_serialization
guest_calls
guest_functions_with_large_parameters
sample_workloads
sandboxes
shared_memory
snapshots
Summary
mshv3 / intel (Linux) (❌ *2.15x slower* → ✅ **1.70x faster**)function_call_serialization
guest_calls
guest_functions_with_large_parameters
sample_workloads
sandboxes
shared_memory
snapshots
Summary
hyperv-ws2025 / amd (Windows) (❌ *6.07x slower* → ✅ **1.44x faster**)function_call_serialization
guest_calls
guest_functions_with_large_parameters
sample_workloads
sandboxes
shared_memory
snapshots
Summary
hyperv-ws2025 / intel (Windows) (❌ *5.82x slower* → ✅ **1.29x faster**)function_call_serialization
guest_calls
guest_functions_with_large_parameters
sample_workloads
sandboxes
shared_memory
snapshots
Summary
|
Benchmark Resultskvm / amd (Linux) (❌ *1.77x slower* → 🚀 **2.48x faster**)function_call_serialization
guest_calls
guest_functions_with_large_parameters
sample_workloads
sandboxes
shared_memory
snapshots
Summary
kvm / intel (Linux) (❌ *2.04x slower* → 🚀 **12.49x faster**)function_call_serialization
guest_calls
guest_functions_with_large_parameters
sample_workloads
sandboxes
shared_memory
snapshots
Summary
mshv3 / amd (Linux) (❌ *1.87x slower* → 🚀 **2.73x faster**)function_call_serialization
guest_calls
guest_functions_with_large_parameters
sample_workloads
sandboxes
shared_memory
snapshots
Summary
mshv3 / intel (Linux) (❌ *1.43x slower* → 🚀 **2.26x faster**)function_call_serialization
guest_calls
guest_functions_with_large_parameters
sample_workloads
sandboxes
shared_memory
snapshots
Summary
hyperv-ws2025 / amd (Windows) (❌ *7.91x slower* → ✅ **1.21x faster**)function_call_serialization
guest_calls
guest_functions_with_large_parameters
sample_workloads
sandboxes
shared_memory
snapshots
Summary
hyperv-ws2025 / intel (Windows) (❌ *10.73x slower* → ✅ **1.46x faster**)function_call_serialization
guest_calls
guest_functions_with_large_parameters
sample_workloads
sandboxes
shared_memory
snapshots
Summary
|
Benchmark Resultskvm / amd (Linux) (❌ *1.15x slower* → 🚀 **2.76x faster**)alloc_fragmented
alloc_lifo
alloc_single
free
free_list_reuse
function_call_serialization
guest_calls
guest_functions_with_large_parameters
recycle_pool
sample_workloads
sandboxes
segmented_payload
shared_memory
snapshots
virtq_readonly_allocator_strategy
virtq_readwrite_allocator_strategy
Summary
kvm / intel (Linux) (✅ **1.07x slower** → 🚀 **11.85x faster**)alloc_fragmented
alloc_lifo
alloc_single
free
free_list_reuse
function_call_serialization
guest_calls
guest_functions_with_large_parameters
recycle_pool
sample_workloads
sandboxes
segmented_payload
shared_memory
snapshots
virtq_readonly_allocator_strategy
virtq_readwrite_allocator_strategy
Summary
mshv3 / amd (Linux) (❌ *1.18x slower* → 🚀 **8.04x faster**)alloc_fragmented
alloc_lifo
alloc_single
free
free_list_reuse
function_call_serialization
guest_calls
guest_functions_with_large_parameters
recycle_pool
sample_workloads
sandboxes
segmented_payload
shared_memory
snapshots
virtq_readonly_allocator_strategy
virtq_readwrite_allocator_strategy
Summary
mshv3 / intel (Linux) (✅ **1.11x slower** → 🚀 **2.85x faster**)alloc_fragmented
alloc_lifo
alloc_single
free
free_list_reuse
function_call_serialization
guest_calls
guest_functions_with_large_parameters
recycle_pool
sample_workloads
sandboxes
segmented_payload
shared_memory
snapshots
virtq_readonly_allocator_strategy
virtq_readwrite_allocator_strategy
Summary
hyperv-ws2025 / amd (Windows) (❌ *1.29x slower* → ✅ **1.46x faster**)alloc_fragmented
alloc_lifo
alloc_single
free
free_list_reuse
function_call_serialization
guest_calls
guest_functions_with_large_parameters
recycle_pool
sample_workloads
sandboxes
segmented_payload
shared_memory
snapshots
virtq_readonly_allocator_strategy
virtq_readwrite_allocator_strategy
Summary
|
Benchmark Resultskvm / amd (Linux) (❌ *1.13x slower* → 🚀 **2.73x faster**)alloc_fragmented
alloc_lifo
alloc_single
free
free_list_reuse
function_call_serialization
guest_calls
guest_functions_with_large_parameters
recycle_pool
sample_workloads
sandboxes
segmented_payload
shared_memory
snapshots
virtq_readonly_allocator_strategy
virtq_readwrite_allocator_strategy
Summary
kvm / intel (Linux) (✅ **1.08x slower** → 🚀 **11.46x faster**)alloc_fragmented
alloc_lifo
alloc_single
free
free_list_reuse
function_call_serialization
guest_calls
guest_functions_with_large_parameters
recycle_pool
sample_workloads
sandboxes
segmented_payload
shared_memory
snapshots
virtq_readonly_allocator_strategy
virtq_readwrite_allocator_strategy
Summary
mshv3 / amd (Linux) (✅ **1.11x slower** → 🚀 **4.25x faster**)alloc_fragmented
alloc_lifo
alloc_single
free
free_list_reuse
function_call_serialization
guest_calls
guest_functions_with_large_parameters
recycle_pool
sample_workloads
sandboxes
segmented_payload
shared_memory
snapshots
virtq_readonly_allocator_strategy
virtq_readwrite_allocator_strategy
Summary
mshv3 / intel (Linux) (❌ *1.16x slower* → 🚀 **3.13x faster**)alloc_fragmented
alloc_lifo
alloc_single
free
free_list_reuse
function_call_serialization
guest_calls
guest_functions_with_large_parameters
recycle_pool
sample_workloads
sandboxes
segmented_payload
shared_memory
snapshots
virtq_readonly_allocator_strategy
virtq_readwrite_allocator_strategy
Summary
hyperv-ws2025 / amd (Windows) (✅ **1.11x slower** → 🚀 **1.95x faster**)alloc_fragmented
alloc_lifo
alloc_single
free
free_list_reuse
function_call_serialization
guest_calls
guest_functions_with_large_parameters
recycle_pool
sample_workloads
sandboxes
segmented_payload
shared_memory
snapshots
virtq_readonly_allocator_strategy
virtq_readwrite_allocator_strategy
Summary
hyperv-ws2025 / intel (Windows) (❌ *1.23x slower* → ✅ **1.76x faster**)alloc_fragmented
alloc_lifo
alloc_single
free
free_list_reuse
function_call_serialization
guest_calls
guest_functions_with_large_parameters
recycle_pool
sample_workloads
sandboxes
segmented_payload
shared_memory
snapshots
virtq_readonly_allocator_strategy
virtq_readwrite_allocator_strategy
Summary
|
No, some of the benchmarks were slower, and some were faster, e.g.:
For that platform / hypervisor combination there was at least one benchmark that was 1.71x slower, and at least one that was 2.54x faster. The values between brackets show the "spread" in benchmark results for that platform. I guess it could be confused as <before> → <after>. I'm open to ideas on how to improve it :-) |
Benchmark Resultskvm / amd (Linux) (❌ *1.83x slower* → 🚀 **2.45x faster**)alloc_fragmented
alloc_lifo
alloc_single
free
free_list_reuse
function_call_serialization
guest_calls
guest_functions_with_large_parameters
recycle_pool
sample_workloads
sandboxes
segmented_payload
shared_memory
snapshot_files
snapshots
virtq_readonly_allocator_strategy
virtq_readwrite_allocator_strategy
Summary
kvm / intel (Linux) (❌ *1.74x slower* → 🚀 **12.06x faster**)alloc_fragmented
alloc_lifo
alloc_single
free
free_list_reuse
function_call_serialization
guest_calls
guest_functions_with_large_parameters
recycle_pool
sample_workloads
sandboxes
segmented_payload
shared_memory
snapshot_files
snapshots
virtq_readonly_allocator_strategy
virtq_readwrite_allocator_strategy
Summary
mshv3 / amd (Linux) (❌ *2.04x slower* → 🚀 **2.96x faster**)alloc_fragmented
alloc_lifo
alloc_single
free
free_list_reuse
function_call_serialization
guest_calls
guest_functions_with_large_parameters
recycle_pool
sample_workloads
sandboxes
segmented_payload
shared_memory
snapshot_files
snapshots
virtq_readonly_allocator_strategy
virtq_readwrite_allocator_strategy
Summary
mshv3 / intel (Linux) (❌ *1.37x slower* → 🚀 **2.27x faster**)alloc_fragmented
alloc_lifo
alloc_single
free
free_list_reuse
function_call_serialization
guest_calls
guest_functions_with_large_parameters
recycle_pool
sample_workloads
sandboxes
segmented_payload
shared_memory
snapshot_files
snapshots
virtq_readonly_allocator_strategy
virtq_readwrite_allocator_strategy
Summary
hyperv-ws2025 / amd (Windows) (❌ *5.42x slower* → ✅ **1.45x faster**)alloc_fragmented
alloc_lifo
alloc_single
free
free_list_reuse
function_call_serialization
guest_calls
guest_functions_with_large_parameters
recycle_pool
sample_workloads
sandboxes
segmented_payload
shared_memory
snapshot_files
snapshots
virtq_readonly_allocator_strategy
virtq_readwrite_allocator_strategy
Summary
hyperv-ws2025 / intel (Windows) (❌ *4.82x slower* → ✅ **1.32x faster**)alloc_fragmented
alloc_lifo
alloc_single
free
free_list_reuse
function_call_serialization
guest_calls
guest_functions_with_large_parameters
recycle_pool
sample_workloads
sandboxes
segmented_payload
shared_memory
snapshot_files
snapshots
virtq_readonly_allocator_strategy
virtq_readwrite_allocator_strategy
Summary
|
|
We discuss this and decided we should not run on every PR due to variance but we will make this triggerable and should run it on the PR to create a Release. We can determine if that is valuable after running it there multiple time. |
|
Would it be possible to run just benchmarks that are stable to start? that way we get some signal and are able to move forward with this? |
786230d to
3d40580
Compare
Benchmark Resultskvm / amd (Linux) (❌ *2.57x slower* → 🚀 **5.71x faster**)alloc_fragmented
alloc_lifo
alloc_single
free
free_list_reuse
guest_calls
recycle_pool
sample_workloads
segmented_payload
shared_memory
snapshot_files
virtq_readonly_allocator_strategy
virtq_readwrite_allocator_strategy
Summary
kvm / intel (Linux) (❌ *3.05x slower* → 🚀 **4.78x faster**)alloc_fragmented
alloc_lifo
alloc_single
free
free_list_reuse
guest_calls
recycle_pool
sample_workloads
segmented_payload
shared_memory
snapshot_files
virtq_readonly_allocator_strategy
virtq_readwrite_allocator_strategy
Summary
mshv3 / amd (Linux) (❌ *1.99x slower* → ✅ **1.60x faster**)alloc_fragmented
alloc_lifo
alloc_single
free
free_list_reuse
guest_calls
recycle_pool
sample_workloads
segmented_payload
shared_memory
snapshot_files
virtq_readonly_allocator_strategy
virtq_readwrite_allocator_strategy
Summary
mshv3 / intel (Linux) (❌ *1.15x slower* → ✅ **1.28x faster**)alloc_fragmented
alloc_lifo
alloc_single
free
free_list_reuse
guest_calls
recycle_pool
sample_workloads
segmented_payload
shared_memory
snapshot_files
virtq_readonly_allocator_strategy
virtq_readwrite_allocator_strategy
Summary
hyperv-ws2025 / amd (Windows) (❌ *1.57x slower* → 🚀 **5.56x faster**)alloc_fragmented
alloc_lifo
alloc_single
free
free_list_reuse
guest_calls
recycle_pool
sample_workloads
segmented_payload
shared_memory
snapshot_files
virtq_readonly_allocator_strategy
virtq_readwrite_allocator_strategy
Summary
hyperv-ws2025 / intel (Windows) (❌ *1.94x slower* → 🚀 **6.97x faster**)alloc_fragmented
alloc_lifo
alloc_single
free
free_list_reuse
guest_calls
recycle_pool
sample_workloads
segmented_payload
shared_memory
snapshot_files
virtq_readonly_allocator_strategy
virtq_readwrite_allocator_strategy
Summary
|
1e53cd1 to
3c5e059
Compare
Benchmark Resultskvm / amd (Linux) (❌ *1.88x slower* → 🚀 **5.73x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
kvm / intel (Linux) (❌ *2.65x slower* → 🚀 **4.84x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
mshv3 / amd (Linux) (❌ *1.74x slower* → ✅ **1.58x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
mshv3 / intel (Linux) (❌ *1.29x slower* → ✅ **1.28x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
hyperv-ws2025 / amd (Windows) (❌ *1.63x slower* → 🚀 **1.88x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
hyperv-ws2025 / intel (Windows) (❌ *1.44x slower* → 🚀 **6.44x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
|
Introduce a new internal tooling crate (hyperlight-ci) that provides: - bench subcommand: Runs criterion benchmarks in parallel via criterion-swarm. Features include: - Configurable parallelism (-j N, defaults to all P-cores) - Configurable output modes (spinner, stream, summary) - Support for pre-built binaries (--binary) to skip rebuilds - Trailing args forwarded to criterion (filter, --exact, etc.) - bench-report subcommand: Generates markdown comparison tables from criterion's target/criterion/ JSON output via criterion-markdown. Features include: - Benchmark discovery via criterion-swarm - Optional allowlist filtering via --binary or trailing args - Output to stdout This replaces ad-hoc benchmark scripting with a unified tool suitable for both local development and CI report generation. Signed-off-by: Jorge Prendes <jorge.prendes@gmail.com>
- Add cargo alias (`cargo ci`) for convenient hyperlight-ci invocation - Update dep_benchmarks workflow to use `cargo ci bench` and generate a markdown report via `cargo ci bench-report`, posting results as a PR comment per hypervisor/cpu matrix entry - Add benchmarks job to ValidatePullRequest workflow with hypervisor and cpu matrix, gated behind docs-only and build-guests checks - Grant pull-requests: write permission for PR comment posting - Simplify Justfile bench recipes to delegate to `cargo ci bench` - Update benchmarking docs to reflect the new workflow Signed-off-by: Jorge Prendes <jorge.prendes@gmail.com>
The reusable benchmark workflow declares `cpu_vendor` and a required `arch`. The pull request matrix supplies both, and the report label and artifact name use `cpu_vendor`. Signed-off-by: Jorge Prendes <jorge.prendes@gmail.com>
3c5e059 to
1dea3cb
Compare
`bench_report.toml` lists regular expressions matched against criterion benchmark ids. A benchmark is selected when it matches `allowlist` and no `denylist` entry, and an empty `allowlist` keeps everything the denylist does not exclude. Omitting both lists disables filtering. `cargo ci bench-report --config-file` reports the selection, and `cargo ci bench --config-file` runs it. CI keeps running every benchmark and filters at report time. Both subcommands apply the selection through `CriterionSwarm::retain`, added in criterion-swarm 0.2.1, so running and reporting cannot drift apart. An allowlist pattern that matches no benchmark fails, so a rename surfaces instead of dropping out of the comment silently. A denylist pattern matching nothing is accepted, because a benchmark may be absent on some platforms. Unknown keys are rejected so a stale key cannot silently disable filtering. The initial lists keep 47 of 116 benchmarks: those whose median drifted by at most 5% across five back-to-back runs on an idle machine, less the snapshot cold start and restore families. Signed-off-by: Jorge Prendes <jorge.prendes@gmail.com>
0.2.2 stops pinning benchmarks to CPU 0, which hosts the boot CPU's timers, RCU kthreads and unbound kworkers. A benchmark pinned there waited 327us per timeslice to be scheduled against 47us elsewhere, running up to 4x slower. Which benchmark drew that core varied per run, so it surfaced as drift. Measured over the full suite at equal parallelism, p90 run-to-run drift falls from 257% to 6% for the sandbox benchmarks, and from 14% to 5% for the rest. One core of parallelism is given up, costing about 17% wall time per run. Signed-off-by: Jorge Prendes <jorge.prendes@gmail.com>
A vCPU created without an in-kernel LAPIC bumps the kernel's `kvm_has_noapic_vcpu` static key, and teardown drops it again. Each transition through zero rewrites kernel text and IPIs every core. Benchmarks that create and drop sandboxes leave no resident VM between iterations, so they cross that boundary constantly: a full suite run issues 176k broadcast IPIs, against 560 with one sandbox held. `cargo ci bench` now keeps one sandbox alive for the duration of a run, so that cost lands on neither the benchmark that triggers it nor its neighbours. Excluding benchmarks that ran on CPU 0, which has its own much larger effect, median drift for the sandbox group falls from 6.5% to 1.7%. Pass `--no-ballast` to measure the cold path instead. The helper is an example rather than a dependency of this crate, keeping hyperlight-host out of the CI tool's build. It exits on stdin EOF, so it cannot outlive the run even when this process is killed without unwinding. Signed-off-by: Jorge Prendes <jorge.prendes@gmail.com>
1dea3cb to
1481413
Compare
Benchmark Resultskvm / amd (Linux) (❌ *1.83x slower* → 🚀 **5.28x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
kvm / intel (Linux) (❌ *2.91x slower* → 🚀 **4.65x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
mshv3 / amd (Linux) (❌ *1.29x slower* → ✅ **1.58x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
mshv3 / intel (Linux) (❌ *1.14x slower* → ✅ **1.29x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
hyperv-ws2025 / amd (Windows) (❌ *1.49x slower* → 🚀 **2.47x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
hyperv-ws2025 / intel (Windows) (❌ *1.25x slower* → 🚀 **7.29x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
|
1 similar comment
Benchmark Resultskvm / amd (Linux) (❌ *1.83x slower* → 🚀 **5.28x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
kvm / intel (Linux) (❌ *2.91x slower* → 🚀 **4.65x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
mshv3 / amd (Linux) (❌ *1.29x slower* → ✅ **1.58x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
mshv3 / intel (Linux) (❌ *1.14x slower* → ✅ **1.29x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
hyperv-ws2025 / amd (Windows) (❌ *1.49x slower* → 🚀 **2.47x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
hyperv-ws2025 / intel (Windows) (❌ *1.25x slower* → 🚀 **7.29x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
|
Benchmark Resultskvm / amd (Linux) (❌ *1.88x slower* → 🚀 **5.76x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
kvm / intel (Linux) (❌ *3.02x slower* → 🚀 **4.81x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
mshv3 / amd (Linux) (❌ *1.28x slower* → ✅ **1.59x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
mshv3 / intel (Linux) (✅ **1.07x slower** → 🚀 **1.84x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
hyperv-ws2025 / amd (Windows) (❌ *1.45x slower* → 🚀 **2.40x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
hyperv-ws2025 / intel (Windows) (❌ *2.05x slower* → 🚀 **7.05x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
|
Criterion stores an id per result directory but nothing about the run as a whole, and results accumulate: a directory keeps benchmarks that no longer exist, indistinguishable from the ones just measured. CI compounds this by unpacking a baseline into the same directory before running, so its uploaded artifact holds 155 result directories for a suite of 87. `cargo ci bench` now writes `benchmarks.json` next to the results, listing the benchmarks the run covers along with a timestamp and the host it ran on. Reading an archived run no longer needs the binaries that produced it, which is the only other way to tell current results from leftovers. The list is the post-filter set, so it describes what was measured rather than what was discovered. Signed-off-by: Jorge Prendes <jorge.prendes@gmail.com>
Benchmark Resultskvm / amd (Linux) (❌ *2.37x slower* → 🚀 **5.76x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
kvm / intel (Linux) (❌ *2.56x slower* → 🚀 **4.21x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
mshv3 / amd (Linux) (❌ *1.30x slower* → ✅ **1.59x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
mshv3 / intel (Linux) (❌ *1.17x slower* → ✅ **1.29x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
hyperv-ws2025 / amd (Windows) (❌ *1.33x slower* → 🚀 **2.19x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
hyperv-ws2025 / intel (Windows) (❌ *2.17x slower* → 🚀 **6.92x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
|
Naming the benchmarks by listing the binaries builds them first, and describes the current checkout rather than the results being reported. Those are the same thing only while reporting a local run of the current tree. `bench-report` now takes the ids from `benchmarks.json` when the results carry one, so a criterion directory from elsewhere, a CI artifact say, reads without a toolchain and without the checkout matching. Rendering the downloaded results of a full suite takes 11ms rather than a build. Explicit `--binary` or trailing bench args still list the binaries, since both name benchmarks the manifest cannot filter, as do results from before the manifest existed. Signed-off-by: Jorge Prendes <jorge.prendes@gmail.com>
Benchmark Resultskvm / amd (Linux) (❌ *1.83x slower* → 🚀 **4.97x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
kvm / intel (Linux) (❌ *2.44x slower* → 🚀 **4.94x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
mshv3 / amd (Linux) (❌ *1.16x slower* → ✅ **1.44x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
mshv3 / intel (Linux) (❌ *1.17x slower* → ✅ **1.29x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
hyperv-ws2025 / amd (Windows) (❌ *1.51x slower* → 🚀 **2.31x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
hyperv-ws2025 / intel (Windows) (❌ *1.36x slower* → 🚀 **7.13x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
|
Comparing what CI measured meant downloading six artifacts by hand, unpacking each somewhere, and running the report once per configuration. The results are richer than the comment CI posts, holding every benchmark rather than the reported subset and the samples behind each estimate, so reaching for them is worth making cheap. `bench-report --from-run <id>` renders a whole run, one section per hypervisor and cpu vendor, the way the pull request comment reads. `--pr <number>` finds the run, taking the most recent one that still has its artifacts: the newest is often a label check, or a run whose benchmarks have not finished. Runs land under `target/ci-runs` and are reused, artifacts being immutable. A run is about 400MB unpacked, and the first report of one waits on the download. This leans on the manifest, so the ids come from what the run measured rather than from building this checkout, which need not even be the same commit. Signed-off-by: Jorge Prendes <jorge.prendes@gmail.com>
Benchmark Resultskvm / amd (Linux) (❌ *1.86x slower* → 🚀 **5.75x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
kvm / intel (Linux) (❌ *2.65x slower* → 🚀 **4.97x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
mshv3 / amd (Linux) (❌ *1.26x slower* → ✅ **1.59x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
mshv3 / intel (Linux) (❌ *1.17x slower* → ✅ **1.28x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
hyperv-ws2025 / amd (Windows) (❌ *1.61x slower* → 🚀 **2.12x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
hyperv-ws2025 / intel (Windows) (❌ *1.99x slower* → 🚀 **6.67x faster**)guest_calls
payload_allocation
sample_workloads
shared_memory
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Summary
|
No description provided.