Single owner for (n1, n2, n3) scratch arrays. Every column-
local kernel slot (continuity, coriolis_adv, pressure_force,
lateral-mix vorticity scratch, KPP solve scratch) declares its
per-step workspace as a scratch_3d_buffer_t rather than as a
bare real(wp), allocatable :: foo(:, :, :).
Why centralise this:
1. One allocation policy. The strided hybrid loop
pattern (see src/core/ocean/README.md) sizes every
column-local scratch at (stride, stride, nz_ml) with
stride = ocean_state%column_stride. Putting that policy
in one type means a new slot follows the convention by
construction, not by remembering to copy-paste it.
2. One GPU-mapping site.enter_data / exit_data
live on the type, so every slot that uses this type
inherits the mapping for free — no per-slot directive
copies to keep in sync.
3. One resize site. If a kernel discovers it needs a
different shape mid-run (rare — only happens if stride
becomes dynamic), ensure_size reallocates and remaps in
one place.
4. Cheap finite-check. Phase 6 debug guards walk every
scratch_3d_buffer_t instance after each slot kernel and
verify no NaN / out-of-range leaked into scratch.
Nodes of different colours represent the following:
Solid arrows point from a submodule to the (sub)module which it is
descended from. Dashed arrows point from a module or program unit to
modules which it uses.
Where possible, edges connecting nodes are
given different colours to make them easier to distinguish in
large graphs.
Nodes of different colours represent the following:
Solid arrows point from a submodule to the (sub)module which it is
descended from. Dashed arrows point from a module or program unit to
modules which it uses.
Where possible, edges connecting nodes are
given different colours to make them easier to distinguish in
large graphs.
The scratch storage. Access in a kernel as
slot%scratch%data(i, j, k). Note: this is a single
derived type (not array-of-derived-types), so the
!$acc indirection bug from the array-of-DT pattern
does not apply.
logical,
public
::
is_init
=
.false.
True between init and destroy. Prefer this to
allocated(this%data) — tracks GPU device attachment too.
integer,
public
::
n1
=
0
First-dim size, typically stride (or stride+1 on
face-located buffers).
integer,
public
::
n2
=
0
Second-dim size, typically stride.
integer,
public
::
n3
=
0
Third-dim size, typically nz_ml (or nz_ml+1 for
layer-interface buffers).
character(len=32),
public
::
name
=
""
Human-readable tag for diagnostics + debug guard
messages (“kpp_bulk_ri_scratch”, “ppm_h_face_left_x”, …).
Attach the scratch payload to the device holding the HOST
payload: zero from init (allocate(..., source = 0.0_wp)), or
whatever a warm-restart read restored into it before the map.
The zero half is a contract, not an implementation detail: a
consumer whose producer was SKIPPED this step is entitled to read
zero on a cold start, and it must read zero on BOTH toolchains
(see the inline note below for the pred_corr predictor that
does exactly that). type(...) (not class) dummy on purpose:
a by-reference non-polymorphic dummy aliases the heap object, so
copyin(this%data) attaches against a heap base — no polymorphic
stack box for AMD to reject.
Release the scratch payload from the device. No copyout
— scratch contents are per-step intermediates with no
host-side meaning. No-op when never allocated (mirrors
scratch_3d_buffer_enter_data_impl’s gate).
Type-bound wrapper — keeps buf%enter_data() call sites working.
Delegates to the non-polymorphic impl so the OpenMP map base is
the heap object (see the impl + the public-decl comment above).
Allocate the host-side storage at (n1, n2, n3), zero-filled.
GPU attachment is a separate step via enter_data —
init runs before the device exists in the typical
init-then-enter-data sequence.