Convenience wrapper: wrap h_layer, u/v layer faces, and every
registered tracer. Called at stage entry (before
derive_bt_from_layers) and after continuity (before hdiff).
Per-tracer loop is OUTSIDE the DC kernels (outer-shim pattern: array-of-derived-types cannot be dereferenced on device).
No-op when neither bc%periodic_x nor bc%periodic_y is set.
D4: skip_x / skip_y (optional, default .false.): when .true.,
the local wrap on that axis is suppressed. The caller passes
ocean_halo_is_decomposed_x/y() so that a multi-rank periodic-x
decomposition does not double-wrap the seam (the MPI halo already
owns those ghost columns). Single-rank ⇒ caller passes .false.
(or omits) ⇒ bit-identical.
NOTE: rdb_ocean_periodic must NOT import rdb_ocean_halo (USE cycle —
the halo backends import this module’s wrap kernels). The skip is
passed in by the caller.
| Type | Intent | Optional | Attributes | Name | ||
|---|---|---|---|---|---|---|
| type(hgrid_t), | intent(in) | :: | grid | |||
| type(ocean_bc_state_t), | intent(in) | :: | bc | |||
| type(multilayer_state_t), | intent(inout) | :: | ms | |||
| logical, | intent(in), | optional | :: | skip_x |
When .true., suppress the local x-axis wrap (D4 multi-rank). |
|
| logical, | intent(in), | optional | :: | skip_y |
When .true., suppress the local y-axis wrap (D4 multi-rank). |
| Type | Visibility | Attributes | Name | Initial | |||
|---|---|---|---|---|---|---|---|
| integer, | private | :: | it | ||||
| integer, | private | :: | nghost | ||||
| integer, | private | :: | nx | ||||
| integer, | private | :: | nx_phys | ||||
| integer, | private | :: | ny | ||||
| integer, | private | :: | ny_phys | ||||
| integer, | private | :: | nz | ||||
| logical, | private | :: | per_x | ||||
| logical, | private | :: | per_y |
subroutine ocean_periodic_wrap_state(grid, bc, ms, skip_x, skip_y) !! Convenience wrapper: wrap h_layer, u/v layer faces, and every !! registered tracer. Called at stage entry (before !! `derive_bt_from_layers`) and after continuity (before hdiff). !! !! Per-tracer loop is OUTSIDE the DC kernels (outer-shim pattern: !! array-of-derived-types cannot be dereferenced on device). !! !! No-op when neither `bc%periodic_x` nor `bc%periodic_y` is set. !! !! D4: `skip_x` / `skip_y` (optional, default .false.): when .true., !! the local wrap on that axis is suppressed. The caller passes !! `ocean_halo_is_decomposed_x/y()` so that a multi-rank periodic-x !! decomposition does not double-wrap the seam (the MPI halo already !! owns those ghost columns). Single-rank ⇒ caller passes .false. !! (or omits) ⇒ bit-identical. !! NOTE: rdb_ocean_periodic must NOT import rdb_ocean_halo (USE cycle — !! the halo backends import this module's wrap kernels). The skip is !! passed in by the caller. type(hgrid_t), intent(in) :: grid type(ocean_bc_state_t), intent(in) :: bc type(multilayer_state_t), intent(inout) :: ms logical, intent(in), optional :: skip_x !! When .true., suppress the local x-axis wrap (D4 multi-rank). logical, intent(in), optional :: skip_y !! When .true., suppress the local y-axis wrap (D4 multi-rank). integer :: it integer :: nx, ny, nz, nx_phys, ny_phys, nghost logical :: per_x, per_y per_x = bc%periodic_x per_y = bc%periodic_y if (.not. per_x .and. .not. per_y) return nx = grid%nx_total ny = grid%ny_total nz = ms%nz_ml nx_phys = grid%nx_phys ny_phys = grid%ny_phys nghost = grid%nghost ! D4: skip the LOCAL periodic wrap on any axis split across MPI ranks. ! On a decomposed axis the ghost columns are owned by a NEIGHBOUR rank — ! the local wrap kernel copies the wrong (local-subdomain) columns into ! those ghosts, overwriting what ocean_halo_* just filled correctly. ! The MPI halo already ran before this call, so the ghosts are correct; ! just leave them alone. Single-rank ⇒ skip_x absent/.false. ! ⇒ local wrap governs as before (bit-identical). if (present(skip_x)) per_x = per_x .and. .not. skip_x if (present(skip_y)) per_y = per_y .and. .not. skip_y if (.not. per_x .and. .not. per_y) return ! Batched async wrap: every per-field wrap issues on queue 1 ! (no_wait=.true.); the whole batch is synced ONCE below. Pipelines the ! tiny launch-latency-bound ghost-slab kernels; same queue ⇒ ordered. call ocean_periodic_wrap_centre_3d(ms%h_layer, nx, ny, nz, & nx_phys, ny_phys, nghost, per_x, per_y, no_wait=.true.) call ocean_periodic_wrap_face_x_3d(ms%u_face_x_layer, nx + 1, ny, nz, & nx_phys, ny_phys, nghost, per_x, per_y, no_wait=.true.) call ocean_periodic_wrap_face_y_3d(ms%v_face_y_layer, nx, ny + 1, nz, & nx_phys, ny_phys, nghost, per_x, per_y, no_wait=.true.) ! Per-tracer loop outside DCs — outer-shim pattern for array-of-DTs. if (allocated(ms%tracers)) then do it = 1, size(ms%tracers) if (.not. allocated(ms%tracers(it)%hTr)) cycle call ocean_periodic_wrap_centre_3d(ms%tracers(it)%hTr, nx, ny, nz, & nx_phys, ny_phys, nghost, per_x, per_y, no_wait=.true.) end do end if ! Single sync for the whole batch (no-op on non-OpenACC builds). !$acc wait(1) end subroutine ocean_periodic_wrap_state