Performance Baseline Tests¶
These tests serve as automated CI regression guards. They verify that critical operations complete within acceptable time bounds, detecting performance regressions before they reach production.
Baseline Philosophy¶
Baselines are set generously (2-3x expected typical performance) to account for CI environment variability while still catching significant regressions. A test failure indicates a performance regression that needs investigation.
Test Categories¶
- Spatial Trees: QuadTree2D, KdTree2D, KdTree3D, OctTree3D, RTree2D construction and query performance
- PRNG: Random number generation throughput for PcgRandom, XoroShiroRandom, SplitMix64, RomuDuo
- Pooling: Collection pool rent/return overhead for List, HashSet, Dictionary, StringBuilder, SystemArrayPool
- Serialization: JSON and Protobuf serialization/deserialization throughput
Performance Baseline Report¶
Generated: 2026-01-12 01:36:55 UTC
Spatial Trees¶
| Test | Iterations | Time (ms) | Baseline (ms) | % of Baseline | Status |
|---|---|---|---|---|---|
| QuadTree2DRangeQuery | 1K | 27 | 200 | 13.5% | Pass |
| QuadTree2DBoundsQuery | 1K | 29 | 200 | 14.5% | Pass |
| KdTree2DRangeQuery | 1K | 27 | 200 | 13.5% | Pass |
| KdTree2DNearestNeighbor | 1K | 32 | 200 | 16.0% | Pass |
| RTree2DRangeQuery | 1K | 2479 | 200 | 1239.5% | FAIL |
| OctTree3DRangeQuery | 1K | 15 | 200 | 7.5% | Pass |
| KdTree3DRangeQuery | 1K | 33 | 200 | 16.5% | Pass |
| QuadTree2DConstruction | 1 | 2 | 500 | 0.4% | Pass |
| KdTree2DConstruction | 1 | 2 | 500 | 0.4% | Pass |
| RTree2DConstruction | 1 | 1 | 500 | 0.2% | Pass |
PRNG¶
| Test | Iterations | Time (ms) | Baseline (ms) | % of Baseline | Status |
|---|---|---|---|---|---|
| PcgRandomNextInt | 1M | 1 | 500 | 0.2% | Pass |
| PcgRandomNextFloat | 1M | 5 | 500 | 1.0% | Pass |
| XoroShiroRandomNextInt | 1M | 1 | 500 | 0.2% | Pass |
| SplitMix64NextInt | 1M | 1 | 500 | 0.2% | Pass |
| RomuDuoNextInt | 1M | 1 | 500 | 0.2% | Pass |
Pooling¶
| Test | Iterations | Time (ms) | Baseline (ms) | % of Baseline | Status |
|---|---|---|---|---|---|
| ListPooling | 100K | 239504 | 200 | 119752.0% | FAIL |
| HashSetPooling | 100K | 16503 | 200 | 8251.5% | FAIL |
| DictionaryPooling | 100K | 16997 | 200 | 8498.5% | FAIL |
| SystemArrayPool | 100K | 8 | 200 | 4.0% | Pass |
| StringBuilderPooling | 100K | 16456 | 200 | 8228.0% | FAIL |
Serialization¶
| Test | Iterations | Time (ms) | Baseline (ms) | % of Baseline | Status |
|---|---|---|---|---|---|
| JsonSerialize | 10K | 43 | 500 | 8.6% | Pass |
| JsonDeserialize | 10K | 64 | 500 | 12.8% | Pass |
| JsonRoundTrip | 10K | 113 | 1000 | 11.3% | Pass |
| ProtobufSerialize | 10K | 1169 | 500 | 233.8% | FAIL |
| ProtobufDeserialize | 10K | 12 | 500 | 2.4% | Pass |
| ProtobufRoundTrip | 10K | 1728 | 1000 | 172.8% | FAIL |
Summary¶
19 passed, 7 failed out of 26 tests.
Running the Tests¶
These tests run automatically during CI to catch regressions. To generate fresh benchmark results:
- Open Unity Test Runner
- Navigate to
PerformanceBaselineTests - Run
GeneratePerformanceBaselineReportexplicitly (it is marked[Explicit]) - Results will be output to the console and can be copied to this document
Interpreting Results¶
- Time (ms): Actual measured time for the operation
- Baseline (ms): Maximum allowed time before test failure
- % of Baseline: How much of the baseline budget was used (lower is better)
- Status: Pass if within baseline, Fail if exceeded
Refreshing these numbers¶
Run PerformanceBaselineTests.GeneratePerformanceBaselineReport from Unity's Test Runner.
Calibrated paired evidence¶
The generous baseline budgets above detect large regressions. They do not establish that an optimization meets the acceptance criteria in issue #636. Scheduled aggregate reports are advisory. Their renderer never writes the canonical baseline, including after a regression, and refuses empty, malformed, duplicate, or incomplete metric sets. A positive cost against a zero baseline is a regression. The previous automatic baseline update option is rejected.
BenchmarkProtocol.MeasureCalibrated now provides the first part of the stronger protocol:
- Correctness runs before warmup and again after the retained timing samples.
- Each arm warms separately for at least 100 ms and three executions. Calibration doubles the common iteration count until both arms take at least 20 ms; inability to calibrate is inconclusive.
- Eight predeclared
ABBABAABbatches retain 32 observations per arm. Each slot begins with heap settling, consumes a checksum, and must take at least 10 ms. There are no acceptance retries. - Immutable raw milliseconds, paired log ratios, geometric throughput ratio, median, p95, MAD, and within-arm spread accompany a 95% interval from 2,000 deterministic bootstrap repetitions. Bootstrap resamples whole counterbalanced batches to preserve within-batch correlation.
- Environment metadata records commit/base commit, Unity version, backend, build configuration, OS, CPU, process width, platform, and whether execution is inside the Editor. Unavailable code generation and optimization settings remain
unknown. Seed, iteration count, and measured warmup durations/executions accompany the samples.
IntMapPerformanceTests emits the full raw record as INTMAP_PAIRED_SAMPLES JSON in its NUnit output through an explicit JSON writer that needs no runtime serializer generation. Incomplete counterbalanced batches cannot claim timing acceptance and report an undefined interval as JSON null. Existing fixtures that still call MeasurePaired retain their older advisory protocol; this change does not make their tables calibrated evidence.
For the IntMap player experiment, manually dispatch Unity Tests with acceptance=intmap and a supported unity-version. Normal selected tests remain mandatory. The extra test runs in a fresh Release IL2CPP player against Dictionary compiled in the same candidate, so its commit and reference commit metadata are identical. The artifact retains all four workload records. The verifier derives ratios, spreads and the bootstrap interval from the raw samples. It reports inconclusive for unstable arms, meets-hit-margin for stable results with at least 1.3× at both hit-only sizes and a favorable interval, or below-hit-margin. This is the timing decision for #578; it is not full #636 acceptance or proof of unmeasured allocation and code-size properties. No IntMap player result has been claimed before this workflow actually runs.
The encoded timing improvement predicate requires at least 5% less runtime (throughput ratio at least 1 / 0.95), an entirely favorable interval, and stable arms. The allocation-change timing predicate requires the entire runtime interval to stay within 5% of the reference. These are timing predicates only: they cannot accept a change without green instrument controls, identical semantics, and verified allocation, retention, and code-size evidence. Unsupported counters are explicitly labeled; an absent measurement never becomes zero.
Issue #636 remains open. Required follow-up includes calibrated allocating/non-allocating and fast/slow player canaries, workload-specific retention probes and build-size measurements, the full acceptance policy including declared tradeoffs, and explicit post-merge promotion after 20 clean floor/latest Mono/IL2CPP player calibration repetitions. There is currently no automatic promotion command. Analyzer findings, changed-branch coverage, mutation, replay/minimization, and touched-group CI tiers also remain separate requirements. No new measured performance numbers or nonempty baseline are claimed by this infrastructure change.