Exchanges ghost-cell strips between neighbouring MPI ranks
Uses pic_mpi_lib for non-blocking sends/receives. Plus-shaped stencil: 4 exchanges (N/S/E/W), no corner exchanges.
Three exchange modes: halo_exchange_2d — host-staged, blocking (caller copies device<->host) halo_exchange_2d_device — GPU-direct, blocking (CUDA-aware MPI) halo_exchange_begin/end — async split for comm/compute overlap
| Type | Visibility | Attributes | Name | Initial | |||
|---|---|---|---|---|---|---|---|
| integer, | private, | parameter | :: | MAX_REQS | = | 8 |
Max MPI requests: 4 directions x 2 (send + recv) |
| integer, | private, | parameter | :: | NFIELDS | = | 4 |
Number of fields exchanged (h, hu, hv, b) |
| real(kind=wp), | private, | allocatable | :: | ha_buf_recv_east(:) | |||
| real(kind=wp), | private, | allocatable | :: | ha_buf_recv_north(:) | |||
| real(kind=wp), | private, | allocatable | :: | ha_buf_recv_south(:) | |||
| real(kind=wp), | private, | allocatable | :: | ha_buf_recv_west(:) | |||
| real(kind=wp), | private, | allocatable | :: | ha_buf_send_east(:) | |||
| real(kind=wp), | private, | allocatable | :: | ha_buf_send_north(:) | |||
| real(kind=wp), | private, | allocatable | :: | ha_buf_send_south(:) | |||
| real(kind=wp), | private, | allocatable | :: | ha_buf_send_west(:) | |||
| real(kind=wp), | private, | allocatable | :: | hs3_buf_recv_east(:) | |||
| real(kind=wp), | private, | allocatable | :: | hs3_buf_recv_north(:) | |||
| real(kind=wp), | private, | allocatable | :: | hs3_buf_recv_south(:) | |||
| real(kind=wp), | private, | allocatable | :: | hs3_buf_recv_west(:) | |||
| real(kind=wp), | private, | allocatable | :: | hs3_buf_send_east(:) | |||
| real(kind=wp), | private, | allocatable | :: | hs3_buf_send_north(:) | |||
| real(kind=wp), | private, | allocatable | :: | hs3_buf_send_south(:) | |||
| real(kind=wp), | private, | allocatable | :: | hs3_buf_send_west(:) | |||
| integer, | private | :: | hs3_nghost | = | 0 | ||
| integer, | private | :: | hs3_nx_total | = | 0 | ||
| integer, | private | :: | hs3_ny_total | = | 0 | ||
| integer, | private | :: | hs3_nz_capacity | = | 0 | ||
| real(kind=wp), | private, | allocatable | :: | hs_buf_recv_east(:) | |||
| real(kind=wp), | private, | allocatable | :: | hs_buf_recv_north(:) | |||
| real(kind=wp), | private, | allocatable | :: | hs_buf_recv_south(:) | |||
| real(kind=wp), | private, | allocatable | :: | hs_buf_recv_west(:) | |||
| real(kind=wp), | private, | allocatable | :: | hs_buf_send_east(:) | |||
| real(kind=wp), | private, | allocatable | :: | hs_buf_send_north(:) | |||
| real(kind=wp), | private, | allocatable | :: | hs_buf_send_south(:) | |||
| real(kind=wp), | private, | allocatable | :: | hs_buf_send_west(:) | |||
| integer, | private | :: | hs_nghost | = | 0 | ||
| integer, | private | :: | hs_nx_total | = | 0 | ||
| integer, | private | :: | hs_ny_total | = | 0 |
Persistent state for split begin/end halo exchange
| Type | Visibility | Attributes | Name | Initial | |||
|---|---|---|---|---|---|---|---|
| type(decomp_t), | public | :: | decomp | ||||
| logical, | public | :: | initialised | = | .false. | ||
| integer, | public | :: | nghost | = | 0 | ||
| integer, | public | :: | nreq | = | 0 | ||
| integer, | public | :: | nx_local | = | 0 | ||
| integer, | public | :: | nx_total | = | 0 | ||
| integer, | public | :: | ny_local | = | 0 | ||
| integer, | public | :: | ny_total | = | 0 | ||
| integer, | public | :: | rank_east | = | -1 | ||
| integer, | public | :: | rank_north | = | -1 | ||
| integer, | public | :: | rank_south | = | -1 | ||
| integer, | public | :: | rank_west | = | -1 | ||
| type(request_t), | public | :: | reqs(MAX_REQS) | ||||
| integer, | public | :: | strip_ew | = | 0 | ||
| integer, | public | :: | strip_sn | = | 0 |
Order-invariant EXACT cross-rank combine of nval EFP values in
ONE collective (PR-32). Replaces N separate scalar
halo_allreduce_sum calls with one packed allreduce.
| Type | Intent | Optional | Attributes | Name | ||
|---|---|---|---|---|---|---|
| type(efp_t), | intent(in) | :: | local_list(:) | |||
| type(efp_t), | intent(out) | :: | global_list(:) | |||
| integer, | intent(in) | :: | nval |
MPI_Allreduce with MPI_MAX — max-type reductions are exact in FP, so a global max stays layout-reproducible (ocean-MPI plan D5). Used for auto_n_inner’s global gravity-wave CFL (shared n_inner).
| Type | Intent | Optional | Attributes | Name | ||
|---|---|---|---|---|---|---|
| real(kind=wp), | intent(in) | :: | local_val | |||
| real(kind=wp), | intent(out) | :: | global_val |
MPI_Allreduce with MPI_MIN for global timestep
| Type | Intent | Optional | Attributes | Name | ||
|---|---|---|---|---|---|---|
| real(kind=wp), | intent(in) | :: | local_val | |||
| real(kind=wp), | intent(out) | :: | global_val |
MPI_Allreduce with MPI_SUM for CG dot products
| Type | Intent | Optional | Attributes | Name | ||
|---|---|---|---|---|---|---|
| real(kind=wp), | intent(in) | :: | local_val | |||
| real(kind=wp), | intent(out) | :: | global_val |
Cross-rank int64 sum for the decomposition-invariant chksum
bitcount (rdb_ocean_chksum). pic_mpi_lib exposes no
integer(int64) allreduce overload (MPI_INTEGER8 reaches only
send/recv), so the value rides the EXACT real64 allreduce: a
per-field POPCNT sum is bounded by (#elements x 64), which stays
FAR below the double-mantissa bound 253 for any realistic grid
(253/64 ~ 1.4e14 cells), so both the local->double cast and
every partial MPI_SUM are exact — the invariant survives.
| Type | Intent | Optional | Attributes | Name | ||
|---|---|---|---|---|---|---|
| integer(kind=int64), | intent(in) | :: | local_val | |||
| integer(kind=int64), | intent(out) | :: | global_val |
Free pre-allocated halo buffers
| Type | Intent | Optional | Attributes | Name | ||
|---|---|---|---|---|---|---|
| type(halo_async_t), | intent(inout) | :: | ha |
Pre-allocate halo buffers for all 4 fields
| Type | Intent | Optional | Attributes | Name | ||
|---|---|---|---|---|---|---|
| type(halo_async_t), | intent(out) | :: | ha | |||
| type(decomp_t), | intent(in) | :: | decomp | |||
| integer, | intent(in) | :: | nghost | |||
| integer, | intent(in) | :: | nx_local | |||
| integer, | intent(in) | :: | ny_local |
Exchange ghost-cell halos for a single 2D field
| Type | Intent | Optional | Attributes | Name | ||
|---|---|---|---|---|---|---|
| real(kind=wp), | intent(inout) | :: | fld(:,:) |
2D field with ghost cells |
||
| type(decomp_t), | intent(in) | :: | decomp |
Domain decomposition descriptor |
||
| integer, | intent(in) | :: | nghost |
Ghost cell width |
||
| integer, | intent(in) | :: | nx_local |
Local physical cells in x |
||
| integer, | intent(in) | :: | ny_local |
Local physical cells in y |
GPU-direct halo exchange via CUDA-aware MPI
| Type | Intent | Optional | Attributes | Name | ||
|---|---|---|---|---|---|---|
| real(kind=wp), | intent(inout) | :: | fld(:,:) |
2D field with ghost cells (present on device) |
||
| type(decomp_t), | intent(in) | :: | decomp | |||
| integer, | intent(in) | :: | nghost | |||
| integer, | intent(in) | :: | nx_local | |||
| integer, | intent(in) | :: | ny_local |
Exchange ghost-cell halos for a 3D field (all nz layers packed per direction)
| Type | Intent | Optional | Attributes | Name | ||
|---|---|---|---|---|---|---|
| real(kind=wp), | intent(inout) | :: | fld(:,:,:) | |||
| type(decomp_t), | intent(in) | :: | decomp | |||
| integer, | intent(in) | :: | nghost | |||
| integer, | intent(in) | :: | nx_local | |||
| integer, | intent(in) | :: | ny_local | |||
| integer, | intent(in) | :: | nz |
GPU-direct halo exchange for a 3D field, batched across layers.
| Type | Intent | Optional | Attributes | Name | ||
|---|---|---|---|---|---|---|
| real(kind=wp), | intent(inout) | :: | fld(nx_local+2*nghost,ny_local+2*nghost,nz) | |||
| type(decomp_t), | intent(in) | :: | decomp | |||
| integer, | intent(in) | :: | nghost | |||
| integer, | intent(in) | :: | nx_local | |||
| integer, | intent(in) | :: | ny_local | |||
| integer, | intent(in) | :: | nz |
Pack and post non-blocking MPI sends/recvs for all 4 fields
| Type | Intent | Optional | Attributes | Name | ||
|---|---|---|---|---|---|---|
| type(halo_async_t), | intent(inout) | :: | ha | |||
| real(kind=wp), | intent(in) | :: | h(:,:) | |||
| real(kind=wp), | intent(in) | :: | hu(:,:) | |||
| real(kind=wp), | intent(in) | :: | hv(:,:) | |||
| real(kind=wp), | intent(in) | :: | b_fld(:,:) |
Wait for MPI to complete and unpack received ghost cells
| Type | Intent | Optional | Attributes | Name | ||
|---|---|---|---|---|---|---|
| type(halo_async_t), | intent(inout) | :: | ha | |||
| real(kind=wp), | intent(inout) | :: | h(:,:) | |||
| real(kind=wp), | intent(inout) | :: | hu(:,:) | |||
| real(kind=wp), | intent(inout) | :: | hv(:,:) | |||
| real(kind=wp), | intent(inout) | :: | b_fld(:,:) |
Release the persistent halo buffers. Idempotent.
Release the persistent 3D halo buffers. Idempotent.
Lazy-allocate the persistent send/recv buffers used by
halo_exchange_2d_device (and halo_exchange_2d via the same
pool). Sized to the per-rank subdomain; resizes on grid change.
| Type | Intent | Optional | Attributes | Name | ||
|---|---|---|---|---|---|---|
| integer, | intent(in) | :: | nghost | |||
| integer, | intent(in) | :: | nx_total | |||
| integer, | intent(in) | :: | ny_total |
Lazy-allocate the persistent 3D send/recv buffers used by
halo_exchange_3d_device. Sized to the per-rank subdomain
and the largest nz seen so far – subsequent calls with a
smaller nz reuse the existing buffer (just write fewer
elements), grow-only on nz so an ML run that later does an
NH w halo (nz_ml+1) doesn’t free/reallocate. Resizes in
full if (nx_total, ny_total, nghost) change.
| Type | Intent | Optional | Attributes | Name | ||
|---|---|---|---|---|---|---|
| integer, | intent(in) | :: | nghost | |||
| integer, | intent(in) | :: | nx_total | |||
| integer, | intent(in) | :: | ny_total | |||
| integer, | intent(in) | :: | nz |