Skip to content

Performance Baseline Tests Performance

github-actions[bot] edited this page Oct 3, 2026 · 6 revisions

Performance Baseline Tests

These tests serve as automated CI regression guards. They verify that critical operations complete within acceptable time bounds, detecting performance regressions before they reach production.

Baseline Philosophy

Baselines are set generously (2-3x expected typical performance) to account for CI environment variability while still catching significant regressions. A test failure indicates a performance regression that needs investigation.

Test Categories

  • Spatial Trees: QuadTree2D, KdTree2D, KdTree3D, OctTree3D, RTree2D construction and query performance
  • PRNG: Random number generation throughput for PcgRandom, XoroShiroRandom, SplitMix64, RomuDuo
  • Pooling: Collection pool rent/return overhead for List, HashSet, Dictionary, StringBuilder, SystemArrayPool
  • Serialization: JSON and Protobuf serialization/deserialization throughput

Current benchmark measurements

No successful complete Unity Benchmarks run is available for publication yet. The committed measurements below are historical and are not results from the current candidate.

Publishing current measurements

Every complete successful Unity Benchmarks execution updates the current measurements section above and commits the guide alongside its raw NUnit XML and comparison report. A detected comparison regression still publishes the measured values and is labeled in the comparison report. An unseeded acceptance baseline does not prevent publication. If a seeded baseline cannot be compared completely, the current measurements still publish; stale comparison files are removed and the workflow reports the comparison contract failure after committing the data. Partial, failed, empty, malformed, duplicate, or mismatched result sets preserve the last complete measurement report.

Each publication records its workflow run and attempt, measured candidate commit, Unity version, test mode, and UTC publication time. Raw XML links target the same repository and publication ref as the run; measurements and evidence are committed together. The candidate SHA identifies the measured code, not the later documentation commit. Single Stopwatch aggregates are labeled recorded values; reported distributions retain their median and sample count. These current measurements do not seed or promote the separate canonical acceptance baseline.

Benchmarks run weekly on Wednesday at 10:29 UTC through a separate hosted scheduler, which requests an eligible dispatch on protected main. The scheduler performs no licensed Unity work and holds no Unity credentials. Direct scheduled execution remains excluded from the licensed workflow by the enrollment contract; that workflow accepts controlled dispatches. Manual dispatch remains available. Every complete successful execution uses the same committed measurement publication path.

Historical performance baseline report

Generated: 2026-01-12 01:36:55 UTC

Spatial Trees

Test Iterations Time (ms) Baseline (ms) % of Baseline Status
QuadTree2DRangeQuery 1K 27 200 13.5% Pass
QuadTree2DBoundsQuery 1K 29 200 14.5% Pass
KdTree2DRangeQuery 1K 27 200 13.5% Pass
KdTree2DNearestNeighbor 1K 32 200 16.0% Pass
RTree2DRangeQuery 1K 2479 200 1239.5% FAIL
OctTree3DRangeQuery 1K 15 200 7.5% Pass
KdTree3DRangeQuery 1K 33 200 16.5% Pass
QuadTree2DConstruction 1 2 500 0.4% Pass
KdTree2DConstruction 1 2 500 0.4% Pass
RTree2DConstruction 1 1 500 0.2% Pass

PRNG

Test Iterations Time (ms) Baseline (ms) % of Baseline Status
PcgRandomNextInt 1M 1 500 0.2% Pass
PcgRandomNextFloat 1M 5 500 1.0% Pass
XoroShiroRandomNextInt 1M 1 500 0.2% Pass
SplitMix64NextInt 1M 1 500 0.2% Pass
RomuDuoNextInt 1M 1 500 0.2% Pass

Pooling

Test Iterations Time (ms) Baseline (ms) % of Baseline Status
ListPooling 100K 239504 200 119752.0% FAIL
HashSetPooling 100K 16503 200 8251.5% FAIL
DictionaryPooling 100K 16997 200 8498.5% FAIL
SystemArrayPool 100K 8 200 4.0% Pass
StringBuilderPooling 100K 16456 200 8228.0% FAIL

Serialization

Test Iterations Time (ms) Baseline (ms) % of Baseline Status
JsonSerialize 10K 43 500 8.6% Pass
JsonDeserialize 10K 64 500 12.8% Pass
JsonRoundTrip 10K 113 1000 11.3% Pass
ProtobufSerialize 10K 1169 500 233.8% FAIL
ProtobufDeserialize 10K 12 500 2.4% Pass
ProtobufRoundTrip 10K 1728 1000 172.8% FAIL

Summary

19 passed, 7 failed out of 26 tests.

Running the Tests

Dispatch Unity Benchmarks to execute the current benchmark profile and publish its complete measurements to this guide. To inspect the historical baseline test report locally:

  1. Open Unity Test Runner
  2. Navigate to PerformanceBaselineTests
  3. Run GeneratePerformanceBaselineReport explicitly (it is marked [Explicit])
  4. Results are output to the console; they do not replace the committed current measurements

Interpreting Results

  • Time (ms): Actual measured time for the operation
  • Baseline (ms): Maximum allowed time before test failure
  • % of Baseline: How much of the baseline budget was used (lower is better)
  • Status: Pass if within baseline, Fail if exceeded

Refreshing these numbers

Run PerformanceBaselineTests.GeneratePerformanceBaselineReport from Unity's Test Runner.

Calibrated paired evidence

The generous baseline budgets above detect large regressions. They do not establish that an optimization meets the acceptance criteria in issue #636. Unity Benchmarks publishes current measured aggregates in this guide. Comparative aggregate reports are not calibrated acceptance evidence. Their renderer never writes the canonical baseline, including after a regression, and refuses empty, malformed, duplicate, or incomplete metric sets. A positive cost against a zero baseline is a regression. The previous automatic baseline update option is rejected.

BenchmarkProtocol.MeasureCalibrated now provides the first part of the stronger protocol:

  • Correctness runs before warmup and again after the retained timing samples.
  • Each arm warms separately for at least 100 ms and three executions. Calibration doubles the common iteration count until both arms take at least 20 ms; inability to calibrate is inconclusive.
  • Eight predeclared ABBABAAB batches retain 32 observations per arm. Each slot begins with heap settling, consumes a checksum, and must take at least 10 ms. There are no acceptance retries.
  • Immutable raw milliseconds, paired log ratios, geometric throughput ratio, median, p95, MAD, and within-arm spread accompany a 95% interval from 2,000 deterministic bootstrap repetitions. Bootstrap resamples whole counterbalanced batches to preserve within-batch correlation.
  • Environment metadata records commit/base commit, Unity version, backend, build configuration, OS, CPU, process width, platform, and whether execution is inside the Editor. Unavailable code generation and optimization settings remain unknown. Seed, iteration count, and measured warmup durations/executions accompany the samples.

IntMapPerformanceTests emits the full raw record as INTMAP_PAIRED_SAMPLES JSON in its NUnit output through an explicit JSON writer that needs no runtime serializer generation. Incomplete counterbalanced batches cannot claim timing acceptance and report an undefined interval as JSON null. Existing fixtures that still call MeasurePaired retain their older advisory protocol; this change does not make their tables calibrated evidence.

For the IntMap player experiment, manually dispatch Unity Tests with acceptance=intmap and a supported unity-version. Normal selected tests remain mandatory. The extra test runs in a fresh Release IL2CPP player against Dictionary compiled in the same candidate, so its commit and reference commit metadata are identical. The artifact retains all four workload records. The verifier derives ratios, spreads and the bootstrap interval from the raw samples. It reports inconclusive for unstable arms, meets-hit-margin for stable results with at least 1.3× at both hit-only sizes and a favorable interval, or below-hit-margin. This is the timing decision for #578; it is not full #636 acceptance or proof of unmeasured allocation and code-size properties. The current corrected-map player campaign has rejected unstable workloads; those observations establish no speed claim.

Set UH_PERF_DIAGNOSTICS=1 only for a predeclared investigation run, or select intmap-diagnostics=true when manually dispatching acceptance=intmap. The workflow enables diagnostics only in that additional player step. Calibrated records then include SlotDiagnostics for every chronological slot, retaining the same clock boundaries, sample counts, calibration and stability limit. Windows diagnostics capture native thread identity, thread CPU accounting, raw thread cycle counts and the processor observed at each boundary. Unavailable APIs remain explicit. The player verifier rejects any record containing diagnostics as acceptance evidence, even when its timings appear stable.

Before correctness checks, warmup and calibration, diagnostics query the host's Windows CPU sets once using a fixed 64 KiB buffer. CpuSetTopology retains the machine name, processor group, logical processor index, core index, efficiency class and flags. Unsupported APIs, malformed data and insufficient buffer space remain explicit; the query does not retry or change scheduling. Endpoint mapping requires one active group, unique logical indices and no unknown record types. Efficiency classes are reported OS metadata, not inferred processor labels.

Thread CPU accounting brackets the timed work and includes diagnostic-call overhead and checksum publication; its resolution limits comparisons with elapsed time. Processor numbers are relative to the current processor group. Endpoints reveal some migrations but cannot exclude movement during a slot or between groups. Raw cycle counts are never converted to elapsed time or frequency, as Microsoft's API contract prohibits that inference. These diagnostics narrow scheduling hypotheses; they do not by themselves prove a cause or justify changing affinity, priority, samples or acceptance thresholds.

The encoded timing improvement predicate requires at least 5% less runtime (throughput ratio at least 1 / 0.95), an entirely favorable interval, and stable arms. The allocation-change timing predicate requires the entire runtime interval to stay within 5% of the reference. These are timing predicates only: they cannot accept a change without green instrument controls, identical semantics, and verified allocation, retention, and code-size evidence. Unsupported counters are explicitly labeled; an absent measurement never becomes zero.

Issue #636 remains open. Required follow-up includes calibrated allocating/non-allocating and fast/slow player canaries, workload-specific retention probes and build-size measurements, the full acceptance policy including declared tradeoffs, and explicit post-merge promotion after 20 clean floor/latest Mono/IL2CPP player calibration repetitions. There is currently no automatic promotion command. Analyzer findings, changed-branch coverage, mutation, replay/minimization, and touched-group CI tiers also remain separate requirements. Current measured numbers are published only after a complete successful benchmark run; this infrastructure change does not claim new measurements or a nonempty acceptance baseline.

Clone this wiki locally