Repository navigation
Benchmark Framework
The Benchmark Framework provides a structured way to define, run, and compare performance benchmarks against saved baselines. It detects regressions by comparing measured metrics against stored thresholds, making it suitable for CI integration and automated performance regression testing.
BenchmarkFramework::Initialize() is called during engine startup (in GameplayLifecycleShared.cpp) and starts empty by design — scenarios are registered by whoever wants to benchmark. Engine startup then calls RegisterBuiltinScenarios(), which adds two dependency-free, deterministic throughput canaries (cpu.sort_1m, mem.churn_4k_x10k) so benchmark.run has real regression signal out of the box; feature code registers domain-specific scenarios via RegisterScenario(). Runs are triggered on demand via the benchmark.* console commands (see Console Commands).
Source: SparkEngine/Source/Utils/BenchmarkFramework.h
| Class | Responsibility |
|---|---|
IBenchmarkScenario |
Abstract interface for defining a repeatable performance test scenario |
BenchmarkFramework |
Singleton that registers scenarios, runs them, saves/loads baselines, and detects regressions |
BenchmarkMetric |
A single named measurement from a benchmark run (value + unit + direction) |
BenchmarkResult |
Aggregated results from running one scenario across all iterations |
BenchmarkBaseline |
Stored reference values and tolerance for a scenario |
BenchmarkComparison |
Pass/fail result of comparing a scenario run against its baseline |
RegressionDetail |
Detailed information about a single metric that regressed |
A single named metric captured during a benchmark iteration:
struct BenchmarkMetric
{
std::string name; // Metric name (e.g. "FrameTime")
double value = 0.0; // Measured value
std::string unit; // Unit string (e.g. "ms", "MB", "count")
bool lowerIsBetter = true; // Whether lower values indicate better performance
};Results from running a complete benchmark scenario (averaged across iterations):
struct BenchmarkResult
{
std::string scenarioName; // Name of the scenario that was run
std::vector<BenchmarkMetric> metrics; // All collected metrics (averaged)
uint32_t iterations = 0; // Number of iterations executed
std::string timestamp; // ISO-8601 timestamp of the run
};Stored reference values loaded from a baseline file:
struct BenchmarkBaseline
{
std::string scenarioName; // Scenario this baseline belongs to
std::map<std::string, double> metrics; // Metric name -> baseline value
float tolerancePercent = 5.0f; // Allowed deviation before flagging regression
};Comparison output with details on any regressions detected:
struct RegressionDetail
{
std::string metricName; // Which metric regressed
double baseline = 0.0; // Expected baseline value
double measured = 0.0; // Actually measured value
double percentChange = 0.0; // Percentage change from baseline
double threshold = 0.0; // Allowed threshold percentage
};
struct BenchmarkComparison
{
std::string scenarioName; // Scenario that was compared
bool passed = true; // True if no regressions detected
std::vector<RegressionDetail> regressions; // Details of any regressions found
};Implement IBenchmarkScenario to define a repeatable performance test. The framework calls Setup() once, then Run() for GetIterationCount() iterations, then TearDown() once. Metrics are averaged across all iterations.
#include "Utils/BenchmarkFramework.h"
class PhysicsBenchmark : public Spark::IBenchmarkScenario
{
public:
std::string_view GetName() const override { return "Physics_1000Bodies"; }
void Setup() override
{
// Create 1000 physics bodies for the test
m_world = CreatePhysicsWorld();
for (int i = 0; i < 1000; ++i)
m_world->AddBody(RandomBody());
}
std::vector<Spark::BenchmarkMetric> Run() override
{
auto start = std::chrono::high_resolution_clock::now();
m_world->Step(1.0f / 60.0f);
auto end = std::chrono::high_resolution_clock::now();
double ms = std::chrono::duration<double, std::milli>(end - start).count();
return {
{"StepTime", ms, "ms", true}, // lower is better
{"ActiveBodies", 1000.0, "count", false} // higher is better
};
}
void TearDown() override
{
m_world.reset();
}
uint32_t GetIterationCount() const override { return 20; }
private:
std::unique_ptr<PhysicsWorld> m_world;
};auto& bench = Spark::BenchmarkFramework::GetInstance();
bench.Initialize();
// Register one or more scenarios
bench.RegisterScenario(std::make_unique<PhysicsBenchmark>());
bench.RegisterScenario(std::make_unique<RenderBenchmark>());
// Run all registered scenarios
auto results = bench.RunAll();
// Or run a single scenario by name
auto singleResult = bench.RunScenario("Physics_1000Bodies");
for (const auto& metric : singleResult.metrics)
{
std::print(" {}: {} {}\n", metric.name, metric.value, metric.unit);
}Baselines are stored as simple JSON files. Save after a known-good run, then compare future runs against the baseline.
// Save baseline after a reference run
bench.SaveBaseline("baselines.json", results);
// Later: load baseline and compare
auto baselines = bench.LoadBaseline("baselines.json");
auto comparisons = bench.CompareWithBaseline(results, baselines);
if (bench.HasRegressions(comparisons))
{
for (const auto& comp : comparisons)
{
for (const auto& reg : comp.regressions)
{
std::print("REGRESSION: {} in {} -- baseline: {:.2f}, measured: {:.2f} ({:+.1f}%)\n",
reg.metricName, comp.scenarioName,
reg.baseline, reg.measured, reg.percentChange);
}
}
}Each BenchmarkBaseline has a tolerancePercent field (default 5.0%) that controls how much deviation is allowed before a regression is flagged:
- For lower-is-better metrics (e.g., frame time): a positive percent change exceeding the tolerance triggers a regression.
- For higher-is-better metrics (e.g., throughput): a negative percent change exceeding the tolerance triggers a regression.
// Custom tolerance for a specific baseline
BenchmarkBaseline baseline;
baseline.scenarioName = "CriticalPath";
baseline.metrics["FrameTime"] = 16.0;
baseline.tolerancePercent = 2.0f; // Stricter: only 2% allowedOverride GetIterationCount() in your scenario to control averaging. More iterations reduce noise but increase benchmark duration:
uint32_t GetIterationCount() const override { return 100; } // Default is 10Baselines are stored in a straightforward JSON structure:
{
"baselines": [
{
"scenario": "Physics_1000Bodies",
"metrics": {
"StepTime": 2.45,
"ActiveBodies": 1000.0
},
"tolerance": 5.0
},
{
"scenario": "Render_ShadowMaps",
"metrics": {
"FrameTime": 8.12,
"DrawCalls": 150.0
},
"tolerance": 10.0
}
]
}The framework exposes a console status method for runtime debugging:
auto status = bench.Console_GetStatus();
// Output: "[BenchmarkFramework] initialized=true, scenarios=3"Use the benchmark framework in CI pipelines to catch performance regressions before merge:
int main()
{
auto& bench = Spark::BenchmarkFramework::GetInstance();
bench.Initialize();
bench.RegisterScenario(std::make_unique<PhysicsBenchmark>());
bench.RegisterScenario(std::make_unique<RenderBenchmark>());
bench.RegisterScenario(std::make_unique<ECSBenchmark>());
auto results = bench.RunAll();
auto baselines = bench.LoadBaseline("Tests/baselines.json");
auto comparisons = bench.CompareWithBaseline(results, baselines);
if (bench.HasRegressions(comparisons))
{
for (const auto& comp : comparisons)
{
for (const auto& reg : comp.regressions)
{
std::print(stderr, "FAIL: {} {} {:+.1f}% (limit {}%)\n",
comp.scenarioName, reg.metricName,
reg.percentChange, reg.threshold);
}
}
return 1; // Non-zero exit code fails CI
}
std::println("All benchmarks passed.");
return 0;
}After intentional performance changes (e.g., adding a new rendering pass), update baselines:
auto results = bench.RunAll();
bench.SaveBaseline("Tests/baselines.json", results);The Benchmark Framework is a standalone utility that does not depend on other engine systems at runtime. It integrates with:
- CTest / Unit Tests -- Benchmark scenarios can be wrapped in CTest cases for automated CI runs.
-
Engine Console --
Console_GetStatus()reports framework state toSimpleConsole. -
Golden Image Testing -- Use alongside
GoldenImageTestRunnerfor combined performance and visual regression testing. -
Profiler -- Benchmark scenarios can use
Spark::Profilerinternally to collect fine-grained timing data.
// Example: wrapping a benchmark in a CTest test case
TEST_CASE("Performance regression check")
{
auto& bench = Spark::BenchmarkFramework::GetInstance();
bench.Initialize();
bench.RegisterScenario(std::make_unique<PhysicsBenchmark>());
auto results = bench.RunAll();
auto baselines = bench.LoadBaseline("TestData/baselines.json");
auto comparisons = bench.CompareWithBaseline(results, baselines);
REQUIRE_FALSE(bench.HasRegressions(comparisons));
bench.Shutdown();
}| Method | Description |
|---|---|
GetName() -> string_view |
Unique name for this scenario |
Setup() |
One-time setup before iterations begin (default: no-op) |
Run() -> vector<BenchmarkMetric> |
Execute one iteration and return metrics |
TearDown() |
One-time teardown after all iterations (default: no-op) |
GetIterationCount() -> uint32_t |
Number of iterations to run (default: 10) |
| Method | Description |
|---|---|
GetInstance() -> BenchmarkFramework& |
Access the singleton |
Initialize() |
Clear scenarios and mark as initialized |
Shutdown() |
Release all scenarios and reset state |
RegisterScenario(unique_ptr<IBenchmarkScenario>) |
Register a benchmark scenario |
RunAll() -> vector<BenchmarkResult> |
Run all registered scenarios |
RunScenario(string_view) -> BenchmarkResult |
Run a single scenario by name |
SaveBaseline(string_view, vector<BenchmarkResult>) |
Save results as baseline JSON |
LoadBaseline(string_view) -> vector<BenchmarkBaseline> |
Load baselines from JSON |
CompareWithBaseline(results, baselines) -> vector<BenchmarkComparison> |
Compare results against baselines |
HasRegressions(vector<BenchmarkComparison>) -> bool |
Check if any comparisons failed |
Console_GetStatus() -> string |
Human-readable status string |
-
BenchmarkFrameworkis not thread-safe. All calls toRegisterScenario(),RunAll(),RunScenario(), and baseline I/O must happen on the same thread. - Individual
IBenchmarkScenario::Run()implementations may use threads internally (e.g., multithreaded physics simulation), but the framework itself is single-threaded. -
GetInstance()uses a function-local static and is safe for concurrent first-access under C++11 magic-statics guarantees.
Registered during engine startup (BenchmarkFramework::RegisterConsoleCommands()):
| Command | Description |
|---|---|
benchmark.status |
Show initialization state and registered scenario count |
benchmark.run |
Run all registered scenarios and print averaged metrics per scenario |
The built-in canaries (cpu.sort_1m, mem.churn_4k_x10k) are covered by Tests/TestBenchmarkFramework.cpp (BenchmarkFramework_BuiltinScenariosRunAndProduceMetrics).
- Golden-Image-Testing -- Visual regression testing framework
- Profiler -- Fine-grained performance profiling
- Console-System -- Engine console integration
- Testing -- Unit test infrastructure and CTest setup
Published from 2b03dc797148. Edit the canonical source in wiki/.
- Documentation
- Docs route
- Wiki index
- Guides
- Tutorials
- Samples
- Examples
- API Reference
- API route
- Reference
- Build Guide
- Dependencies
- FAQ
- Changelog
- Roadmap
- Contributing
- Code of Conduct
- Home
- FAQ
- Getting Started
- Quick-Start Tutorial
- Making Your First Game
- Making Your First Multiplayer Game
- Artist Workflow Guide
- Editor Walkthrough
- Migration Guide
- How SparkEngine Works
- Architecture Overview
- Engine Architecture Flowchart
- Creating a Game Module
- Game Modules (catalog)
- Entity Component System
- Rendering and Graphics
- Physics
- Cloth Simulation
- Audio
- Input System
- Camera System
- Scripting with AngelScript
- Visual Scripting
- AI and Navigation
- Animation
- 2D Systems
- Networking
- Dedicated Server
- Multiplayer Quick Start
- Area Server Architecture
- Scene Management
- Large World Support
- Collaborative Editing
- Coroutine System
- Event System
- Event Response System
- Job System
- UI System
- UI Layout Extensions
- Localization
- Dialogue System
- Destruction System
- Replay System
- Achievement System
- Loading System
- Mod System
- Content Delivery
- Tween System
- Memory Integrity
- Gameplay Systems
- Terrain and Procedural Generation
- Save System
- Persistence System
- Day Night Cycle and Weather
- Cinematic Sequencer
- Runtime Prefabs
- SparkEditor
- Editor Tutorials
- SparkConsole
- SparkDaemon
- Shader Pipeline
- Asset Pipeline
- Asset Validation
- Asset Migration
- Game Packaging
- Online Services
- DataTable System
- Loot and Crafting System
- CSG System
- Font System
- Timer Manager
- Movie Render Pipeline
- HLOD and World Partition
- Remote Debug System
- Selection Manager
- Asset Dependency Graph
- Editor Automation
- File Watcher
- Project Templates
- System Requirements
- VR Support
- Mobile Platform
- Accessibility
- Platform Input
- Platform Certification
- Cross-Compilation: Wine Testing
- RHI Abstraction Layer
- D3D11 Backend
- D3D12 Backend
- Vulkan Backend
- OpenGL Backend
- Metal Backend
- DXR Raytracing
- Hybrid Ray Tracing
- Upscaling (DLSS/FSR)
- Render Graph
- Shader Graph
- GPU Particles
- GPU-Driven Rendering
- Volumetric Fog
- Volumetric Clouds
- Global Illumination
- Virtual Texturing
- Water Rendering
- Clustered Lighting
- Material System
- Post-Processing
- Shadow System
- Particle System
- Decal System
- Sky and Atmosphere
- Foliage System
- Mesh Shaders
- Neural Rendering
- Configuration Reference
- Performance Tips
- Benchmark Framework
- Threading Model
- Fuzz Policy and Parser Security
- Memory Safety
- Memory Management Patterns
- Build System and CMake Modules
- Profiler and Debugging
- Performance Profiling Guide
- Telemetry System
- Crash Reporting
- Golden Image Testing
- Utilities
- Testing
- Fuzz Policy and Parser Security
- Codebase Statistics
- Codebase Health
- Error Handling Patterns
- Hot Reload Overview
- Troubleshooting
- Contributing
- Workflow Patterns
- Build Optimizations
- CI Reproducible Builds
- GitHub API and PR Checks
- Git Rebase Conflicts
- Clang-Format
- Code Quality Violations
- AI Bloat Pattern
- MinGW + Wine Cross-Compilation
- Live Editor Testing
- Engine & Renderer Landscape
- DuetOS Portability Catalog
- Five-Engine Analysis
- Eleven-Engine Analysis
- ThorVG / Unity Graphics Analysis
- Advanced Techniques Catalog
- Third-Party Library Evaluation
- Engine Viability Evaluation
- Engine Feature Recommendations
- Project Recommendations
- Mac Compatibility Analysis
- Codebase Observations
- Codebase Bloat Audit
- Test Suite Audit
- Documentation Coverage Audit
- ThirdParty Dependencies Audit
- Load Test Baseline
- Gameplay Systems Status
- SparkGame Module Status
- Stub and Abandoned Features
- Memory Integrity System
- Memory Safety Evaluation
- Hardware Acceleration Systems
- Jolt Physics Integration
- GPU/CPU Separation Plan
- Daemon Services Architecture
- Reflection & Polymorphism Refactoring Plan
- SparkBuild In-Tree
- Wine No-JobSystem Breakthrough
- Wine Role and Fallback Tiers