Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -3,3 +3,6 @@
*.jl.mem
/Manifest.toml
/docs/build/
.vscode
.perftest_logs
.perftests
37 changes: 21 additions & 16 deletions Project.toml
Original file line number Diff line number Diff line change
Expand Up @@ -4,25 +4,27 @@ authors = ["Dvegrod <dvegrodu@gmail.com>, Samuel Omlin <samuel.omlin@cscs.ch>, a
version = "0.2.1"

[deps]
BandwidthBenchmark = "68eb07c1-04fd-4e62-9736-d6127c4c03c6"
BenchmarkTools = "6e4b80f9-dd63-53aa-95a3-0cdb28fa8baf"
Configurations = "5218b696-f38b-4ac9-8b61-a12ec717816d"
CountFlops = "1db9610d-79e1-487a-8d40-77f3295c7593"
CpuId = "adafc99b-e345-5852-983c-f28acb93d879"
DataFrames = "a93c6f00-e57d-5684-b7b6-d8193f3e46c0"
Dates = "ade2ca70-3891-5945-98fb-dc099432e06a"
HTTP = "cd3eb016-35fb-5094-929b-558a96fad6f3"
Hwloc = "0e44f5e4-bd66-52a0-8798-143a42290a1d"
JLD2 = "033835bb-8acc-5ee8-8aae-3f567f8a3819"
JSON = "682c06a0-de6a-54ab-a142-c8b1cf79cde6"
LibGit2 = "76f85450-5226-5b5a-8eaa-529ad045b433"
LinearAlgebra = "37e2e46d-f89d-539d-b4ee-838fcccc9c8e"
MLStyle = "d8e11817-5142-5d16-987a-aa16d5891078"
MacroTools = "1914dd2f-81c6-5fcd-8719-6d5c9610ff09"
Pkg = "44cfe95a-1eb2-52ea-b672-e2afdf69b78f"
PrecompileTools = "aea7be01-6a6a-4083-8856-8a6e6704d82a"
Printf = "de0858da-6303-5e67-8744-51eddeeeb8d7"
Revise = "295af30f-e4ad-537b-8983-00126c2a3abe"
STREAMBenchmark = "05e9033e-e298-417a-adae-495536c11ad4"
Suppressor = "fd094767-a336-5f1f-9728-57cf17d0bbfb"
TOML = "fa267f1f-6049-4f14-aa54-33bafae1ed76"
Test = "8dfed614-e22c-5e08-85e1-65c5234f0b40"
ThreadPinningCore = "6f48bc29-05ce-4cc8-baad-4adcba581a18"
UnicodePlots = "b8865327-cd53-5732-bb35-84acbb429228"

[weakdeps]
Expand All @@ -32,16 +34,19 @@ MPI = "da04e1cc-30fd-572f-bb4f-1f8673147195"
PerfTest_MPIExt = "MPI"

[compat]
BenchmarkTools="1"
Configurations="0"
CountFlops="0"
CpuId="0"
HTTP="1"
JLD2="0"
JSON="0"
MLStyle="0"
MacroTools="0"
Revise="3"
STREAMBenchmark="0"
Suppressor="0"
UnicodePlots="3"
julia = "1, 1.11"
BandwidthBenchmark = "0.2.0"
BenchmarkTools = "1"
Configurations = "0"
CountFlops = "0"
DataFrames = "1"
HTTP = "1"
Hwloc = "3.3"
JLD2 = "0"
JSON = "0"
MLStyle = "0"
MacroTools = "0"
PrecompileTools = "1.2"
Suppressor = "0"
ThreadPinningCore = "0.4"
UnicodePlots = "3"
3 changes: 2 additions & 1 deletion docs/make.jl
Original file line number Diff line number Diff line change
Expand Up @@ -38,13 +38,14 @@ makedocs(;
warnonly = [:missing_docs],
pages = [
"Introduction" => "index.md",
"Usage" => "usage.md",
"Quickstart" => "usage.md",
"Macros" => "macros.md",
"Examples" => [hide("..." => "examples.md"),
"examples/mock2-memorythroughput.md",
"examples/mock3-roofline.md",
"examples/mock4-recursive.md",
],
"Configuration" => "configuration_params.md",
"Internals" => "internals.md",
"Limitations" => "limitations.md",
"API reference" => "api.md",
Expand Down
104 changes: 104 additions & 0 deletions docs/src/configuration_params.MD
Original file line number Diff line number Diff line change
@@ -0,0 +1,104 @@
# Configuration Parameters

This document describes all configuration parameters available. Parameters are organized by section, they are specified in TOML format.

---

## `[general]`

General settings that control the overall behavior of the package.

| Parameter | Type | Default | Description |
|---|---|---|---|
| `autoflops` | `Bool` | `true` | Whether the autoflop counter is enabled and accessible during testing. |
| `numas` | `String \| Integer \| Float64` | `"single"` | Amount of NUMAS to use when pinning threads. |
| `threads_per_numa` | `String \| Integer \| Float64` | `"single"` | Amount of threads to pin per NUMA. |
| `save_results` | `Bool` | `true` | Whether to record the results of executing the performance suite. |
| `logs_enabled` | `Bool` | `true` | Enable or disable the log subsystem. If false, this will prevent PerfTest from creating a log folder and log files.|
| `save_folder` | `String` | `".perftests"` | Folder where performance suite results shall be stored. |
| `max_saved_results` | `Int` | `20` | Maximum amount of performance test suite execution results to be saved in the result file for each suite. |
| `plotting` | `Bool` | `true` | Whether to have plots in the test output of methodologies that support it. |
| `verbose` | `Int` | `0` | Whether to output the collected logs in the standard output. |
| `recursive` | `Bool` | `true` | If enabled, whenever a file is included inside a recipe that is being transformed, it will be transformed as well. |
| `safe_formulas` | `Bool` | `false` | Deprecated. |
| `suppress_output` | `Bool` | `true` | Whether to hide the output of the test targets. If false the output of each execution will be shown, which will be probably a long output. |

---

## `[regression]`

Settings for regression testing.

| Parameter | Type | Default | Description |
|---|---|---|---|
| `enabled` | `Bool` | `true` | Whether to use regression testing in this suite. |
| `dedicated_reference_file` | `String` | `""` | If not an empty string, a succesful suite execution will be saved into this file, as well as in the (result file see `[general] save_folder`). The last saved execution will be used as reference for regression testing. If empty the package will look at the result file instead, and find the latest sucessful execution as reference. |
| `default_threshold` | `Number` | `1.1` | The default threshold of the `@regression` macro if no threshold is specified. |
| `use_bencher` | `Bool` | `false` | Whether to enable Bencher to upload results to an online CI/CD performance benchmark platform. THIS FEATURE IS EXPERIMENTAL and likely unstable. |

---

## `[roofline]`

Settings for roofline analysis.

| Parameter | Type | Default | Description |
|---|---|---|---|
| `enabled` | `Bool` | `true` | Whether to use roofline methodologies in this suite. |
| `default_threshold` | `Number` | `0.5` | Default threshold of the `@roofline` macro if no threshold is specified. |

---

## `[memory_bandwidth]`

Settings for memory bandwidth analysis.

| Parameter | Type | Default | Description |
|---|---|---|---|
| `enabled` | `Bool` | `true` | Whether to use effective memory throughput methodology in this suite. |
| `default_threshold` | `Number` | `0.5` | Default thresholds of the `@define_eff_mem_throughput` macro if no threshold is specified. |

---

## `[perfcompare]`

Settings for performance comparison.

| Parameter | Type | Default | Description |
|---|---|---|---|
| `enabled` | `Bool` | `true` | Whether to use the `@perfcompare` macro in this suite. |

---

## `[machine_benchmarking]`

Settings for benchmarking the underlying machine.

| Parameter | Type | Default | Description |
|---|---|---|---|
| `memory_bandwidth_test_buffer_size` | `Int` | `0` | If 0 a buffer size for bandwith benchmarks is used so its at least 4 times bigger than the biggest cache level of the machine. If non-zero this value is used to set the buffer size instead. |

---

## `[MPI]`

Settings for MPI-based execution.

| Parameter | Type | Default | Description |
|---|---|---|---|
| `enabled` | `Bool` | `false` | Whether to enable the MPI aware performance suite generation, this will mainly make sure tests are measured in all ranks but evaluated on the main rank only, and that machine benchmarks take into account all ranks.|
| `mode` | `String` | `"reduce"` | This has no function as of now, but it will be used in future versions of perftest. |

---

## `[bencher]`

Settings for [Bencher](https://bencher.dev) integration. This feature is experimental, and likely prone to bugs.

| Parameter | Type | Default | Description |
|---|---|---|---|
| `api_key` | `String` | `""` | The Bencher platform key to use to connect to bencher. |
| `api_url` | `String` | `"https://api.bencher.dev"` | The API url. |
| `project_name` | `String` | `""` | The name of the project where metrics and testbeds shall be posted. |
| `organization` | `String` | `""` | The organization that holds the project. |
| `custom_testbed_name` | `String` | `""` | The customized identification of this machine (a.k.a testbed for Bencher) when uploading posting suite execution results. |
5 changes: 3 additions & 2 deletions docs/src/index.MD
Original file line number Diff line number Diff line change
Expand Up @@ -4,10 +4,11 @@ The package `PerfTest` provides the user with a performance regression unit test
## Dependencies
`PerfTest` relies on:
- BenchmarkTools
- BandwidthBenchmark
- Configurations
- CountFlops
- CpuId
- Dates
- Hwloc
- HTTP
- JLD2
- JSON
Expand All @@ -16,9 +17,9 @@ The package `PerfTest` provides the user with a performance regression unit test
- MLStyle
- MacroTools
- Pkg
- PrecompileTools
- Printf
- Revise
- STREAMBenchmark
- Suppressor
- TOML
- Test
Expand Down
91 changes: 86 additions & 5 deletions docs/src/usage.MD
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
# Usage
# Quickstart

`PerfTest` provides a set of macros to instrument ordinary Julia test files with performance tests. The idea is to have the posibility of having a functional and a performance suite all in the same place.

Expand All @@ -8,7 +8,7 @@ The underlying idea of declaring performance tests can be boiled down the follow
2. Tell PerfTest what is the target to be tested by using the macro @perftest
3. Tell PerfTest how the target shall be tested, which metrics are interesting, which of those metrics values would be considered a failure, this can be declared using the metric and methodology macros (see Macros)

The following dummy example embodies the paradigm of the package:
The following dummy example presents how a recipe file looks like:

```julia
using ExampleModule : innerProduct, Test, PerfTest # Importing the target and test libraries
Expand All @@ -29,14 +29,95 @@ The following things can be appreciated in this example:
2. The target of the perftest is the innerProduct function
3. The performance test methodology is a roofline model, the developer expects innerProduct to perform at least at 50% of the maximum flop performance set by the roofline. The operational intensity is defined on the main block of the macro. :autoflop is a symbol that enables the use of an automatic flop count feature.

## Execution
## How to use, a first PerfTest.jl recipe

To execute the functional test, simply run the file.
Lets assume we are a developer that wants to track performance regressions on the components of a package in development. The files discussed here can be accessed in `examples/example_quickstart`. We have the following module:

```julia
# module.jl
module MyPackage

# Add [a[1],a[2],...a[end]] and [b[end], b[end-1],...,b[1]] elementwise
function addReversed(A :: Vector{<: Number}, B:: Vector{<: Number}) :: Vector{<:Number}
return [a + b for (a,b) in zip(A, reverse(B))]
end

end
```

We want to track the performance of this package in order to detect performance regressions. To do so we build the following test recipe:

```julia
# testfile.jl
using Test,PerfTest

include("module.jl")

@perftest_config "
[general]
verbose = 3
[regression]
dedicated_reference_file='reference.JLD2'
"

@testset "addReversed tests" begin
N = 10
# We want the size to be bigger on the performance test
@on_perftest_exec begin
N = 1_000_000
end
# We set the regression checker, we dont specify a metric therefore the default (median time elapsed) is used
# low_is_bad=false time elapsed metrics are considered worse the bigger they are
# threshold = 1.05 the test will fail if the time is 105% of the reference or greater, in other words: @test time_elapsed < 1.05 * reference
@regression threshold=1.05 low_is_bad=false

A = [i for i in 1:N]
B = [N-i for i in 1:N]

result = @perftest MyPackage.addReversed(A, B)
@test sum(result) == N*N
end
```

### Running the functional test side of the recipe:

```julia
include("testfile.jl")
```

Or if the module is setup as its own package (not this example):
```julia
using Pkg; Pkg.test()
```

### Running the performance test side of the recipe:

We are doing regression, so in case no reference has been made we will need to execute it at least twice. Once to record the reference, and after that whenever the developer wants to check for a performance regression.

The first time PerfTest is run on a specific directory, a configuration file will be created with a set of default values. Configuration parameters can be set as well by the `@perftest_config` macro, as seen above. The macro takes the highest priority over any other configuration source.
```julia
# This will record a successful performance suite
using PerfTest; runperftests("testfile.jl")
```

In between, the developer may add new changes to the implementation or checks out a different git branch.

```julia
# This will compare the suite results against the reference, if the tests are sucessful the results become the reference.
using PerfTest; runperftests("testfile.jl")
```

To execute the performance test, pass the file path to `PerfTest.transform` and evaluate the resulting test suite. The result of transform can be also saved as a file (see `PerfTest.saveExprAsFile`) for later execution.

For more information have a look at the [Examples](@ref) and see the [API reference](@ref) for details on the usage of `PerfTest`.

## Main use cases of this package:

This package supports both performance testing by regression checks and by performance methodologies, e.g roofline models. It is meant to cover the following two situations:

1. A developer wants to track potential regressions over the course of package development.
2. A developer wants to set machine-agnostic performance tests by using methodologies or custom metrics that are normalized by machine parameters, with the purpose of verifying that the package profits from the capabilities of the machines its been executed on.
2.1. During development.
2.2. During the whole software lifecycle, users can check if the package is properly set up in their machine.

## Installation

Expand Down
11 changes: 4 additions & 7 deletions examples/example-paper-implicitglobalgrid/EXP_test_halo_thr.jl
Original file line number Diff line number Diff line change
@@ -1,10 +1,7 @@
# IMPORTANT: To replicate the same results as in the original paper, divide the GB/s values by 3
# this difference is due to the STREAM benchmark policy, which counts each byte transferred thrice (write-allocate cache) and thus the target does as write_allocate
# IMPORTANT: To replicate the same results as in the original paper, divide the GB/s values by 2
# this difference is due to the benchmark policy, which counts each byte transferred twice (no write-allocate in copy in PerfTest 1.2.3 in PerfTest 1.2.1 this would be thrice instead)
# in our example the policy is to measure throughput as the amount of bytes transferred on an execution which is a different policy
# Nevertheless, the percentages remain constant and the test is valid regardless of the used criterion
# NOTE: All tests of this file can be run with any number of processes.
# Nearly all of the functionality can however be verified with one single process
# (thanks to the usage of periodic boundaries in most of the full halo update tests).

push!(LOAD_PATH, "../src")
using Test
Expand All @@ -21,7 +18,7 @@ NTHREADS = 16
# 256M Elements on STREAM Benchmark (2GB)
@perftest_config "
[general]
verbose = true
verbose = 3
autoflops = false

[regression]
Expand Down Expand Up @@ -74,7 +71,7 @@ dz = 1.0
i = 0
# For info in the x3 multiplier see beggining of file
@define_eff_memory_throughput ratio=0.9 begin
(nx * ny * 8) * 3 * MPI.Comm_size(MPI.COMM_WORLD) / :median_time
(nx * ny * 8) * 2 * MPI.Comm_size(MPI.COMM_WORLD) / :median_time
end
@auxiliary_metric name="Time" units="s" begin
:median_time
Expand Down
Loading
Loading