Skip to content

Commit cdfe51c

Browse files
Merge branch 'master' into dependabot/pip/distributed-lt-2021.8
2 parents 59314c9 + 78d3237 commit cdfe51c

22 files changed

Lines changed: 185 additions & 1197 deletions

File tree

‎benchmarks/user/README.md‎

Lines changed: 19 additions & 67 deletions
Original file line numberDiff line numberDiff line change
@@ -98,17 +98,12 @@ below.
9898

9999
## The optimization level
100100

101-
In Devito, an Operator has two preset optimization levels: `noop` and
102-
`advanced`. With `noop`, no performance optimizations are introduced by the
103-
compiler. With `advanced`, several flop-reducing and data locality
104-
optimizations are applied. Examples of flop-reducing optimizations are common
105-
sub-expressions elimination and factorization; examples of data locality
106-
optimizations are loop fusion and cache blocking. SIMD vectorization is also
107-
applied through compiler auto-vectorization.
108-
109-
`benchmark.py` has two preset optimization modes, that for historical reasons
110-
are called `O1` and `O2`. Basically, `O1` corresponds to `noop`, while `O2`
111-
corresponds to `advanced`.
101+
`benchmark.py` allows to set optimization mode, as well as several optimization
102+
options, via the `--opt` argument. Please refer to
103+
[this](https://github.com/devitocodes/devito/blob/master/examples/performance/00_overview.ipynb)
104+
notebook for a comprehensive list of all optimization modes and options
105+
available in Devito. You may also want to take a look at the example command
106+
lines a few sections below.
112107

113108
## Auto-tuning
114109

@@ -195,6 +190,12 @@ grid:
195190
```
196191
python benchmark.py run -P tti -d 512 402 890 -so 12 -a basic --tn 100
197192
```
193+
Same as before, but telling devito not to use temporaries to store the
194+
intermediate values which stem from mixed derivatives:
195+
```
196+
python benchmark.py run -P tti -d 512 402 890 -so 12 -a basic --tn 100 --opt
197+
"('advanced', {'cire-mingain: 1000000'})"
198+
```
198199
Do not forget to pin processes, especially on NUMA systems; below, we use
199200
`numactl` to pin processes and threads to one specific NUMA domain.
200201
```
@@ -212,8 +213,9 @@ watch numastat -m
212213

213214
## The run-jit-backdoor mode
214215

215-
As of Devito v3.5 it is possible to customize the code generated by Devito. This
216-
is often referred to as the ["JIT backdoor" mode](https://github.com/devitocodes/devito/wiki/FAQ#can-i-manually-modify-the-c-code-generated-by-devito-and-test-these-modifications).
216+
As of Devito v3.5 it is possible to customize the code generated by Devito.
217+
This is often referred to as the ["JIT backdoor"
218+
mode](https://github.com/devitocodes/devito/wiki/FAQ#can-i-manually-modify-the-c-code-generated-by-devito-and-test-these-modifications).
217219
With ``benchmark.py`` we can exploit this feature to manually hack and test the
218220
code generated for a given benchmark. So, we first run a problem, for example
219221
```
@@ -248,57 +250,7 @@ experiments.
248250
## Benchmark output
249251

250252
The GFlops/s and GPoints/s performance, Operational Intensity (OI) and
251-
execution time are emitted to standard output at the end of each run.
252-
Further, when running in `bench` mode, a `.json` file is produced
253-
(see `python benchmark.py bench --help` for more info) in a folder named
254-
`results` except if otherwise specified with the `-r` option.
255-
256-
## Generating a roofline model
257-
258-
To generate a roofline model from the results obtained in `bench` mode,
259-
one can execute `benchmark.py` in `plot` mode. For example, the command
260-
261-
```
262-
python benchmark.py plot -P acoustic -d 512 512 512 -so 12 --tn 100 -a aggressive --max-bw 12.8 --flop-ceil 80 linpack
263-
```
264-
265-
will generate a roofline model for the results obtained from
266-
267-
```
268-
python benchmark.py bench -P acoustic -d 512 512 512 -so 12 --tn 100 -a
269-
```
270-
271-
The `plot` mode expects the same arguments used in `bench` mode plus
272-
two additional arguments to generate the roofline:
273-
274-
* `--max-bw <float>`: DRAM bandwidth (GB/s).
275-
* `--flop-ceil <float, str>`: CPU machine peak. The CPU performance ceil
276-
(GFlops/s) and how the ceil was obtained (ideal peak, linpack, ...).
277-
278-
There also are two optional arguments:
279-
280-
* `--point-runtime` (bool switch): Annotate points with the runtime value.
281-
* `--section <str>`: The code section for which the roofline is produced.
282-
An Operator consists of multiple sections. Each section typically
283-
comprises a loop nest and a sequence of equations. Different sections
284-
are created for logically-distinct parts of the computation
285-
(finite-difference stencils, boundary conditions, interpolation, etc.).
286-
The naming convention is `sectionX`, where `X` is a progress id (`section0`,
287-
`section1`, ...). In the generated code the beginning and the end
288-
of a section are marked with suitable comments. Currently, there is
289-
no way other than looking at the generated code to understand which
290-
section the user-provided equations belong to.
291-
292-
To obtain the DRAM bandwidth of a system, we advise to use
293-
[STREAM](http://www.cs.virginia.edu/stream/ref.html).
294-
295-
To obtain the ideal CPU peak, one should instantiate this formula
296-
297-
#[cores] · #[avx units] · #[vector lanes] · #[FMA ports] · [ISA base frequency]
298-
299-
More details in this [paper](https://arxiv.org/pdf/1807.03032.pdf).
300-
301-
## Do not hesitate to contact us
302-
303-
Should you encounter any issues, do not hesitate to
304-
[get in touch with the development team](https://join.slack.com/t/devitocodes/shared_invite/zt-gtd2yxj9-Y31YKk_7lr9AwfXeL2iMFg)
253+
execution time are emitted to standard output at the end of each run. You may
254+
find this
255+
[FAQ](https://github.com/devitocodes/devito/wiki/FAQ#how-does-devito-compute-the-performance-of-an-operator)
256+
useful.

0 commit comments

Comments
 (0)