When use_io_server is enabled, the last rank on each node becomes
a dedicated I/O server (no GPU, no solver). All solver/decomp/halo
operations use compute_comm (excludes I/O ranks). When disabled,
compute_comm = comm_world (current behavior, no overhead).
comm_env_push_compute_comm / comm_env_pop_compute_comm are a
TEST/TOOLING-ONLY seam: a one-level override of what
comm_env_compute_comm() returns, for tests that need a collective
to run over something other than the real job communicator (see
their docstrings below). Production code never calls them.
Nodes of different colours represent the following:
Solid arrows point from a submodule to the (sub)module which it is
descended from. Dashed arrows point from a module or program unit to
modules which it uses.
Where possible, edges connecting nodes are
given different colours to make them easier to distinguish in
large graphs.
Nodes of different colours represent the following:
Solid arrows point from a submodule to the (sub)module which it is
descended from. Dashed arrows point from a module or program unit to
modules which it uses.
Where possible, edges connecting nodes are
given different colours to make them easier to distinguish in
large graphs.
Variables
Type
Visibility
Attributes
Name
Initial
integer,
private,
parameter
::
COMPUTE_COMM_OVERRIDE_MAX_DEPTH
=
1
TEST/TOOLING-ONLY. One-level stack depth for
comm_env_push_compute_comm / comm_env_pop_compute_comm — see
their docstrings. Kept at 1 deliberately: nothing in production
code nests the override, and a deeper stack would hide a
forgotten pop instead of failing loud on the very next push.
integer,
private,
save
::
cached_compute_rank
=
-1
integer,
private,
save
::
cached_compute_size
=
-1
integer,
private,
save
::
cached_io_server_rank
=
-1
World rank of this node’s I/O server (-1 if none)
logical,
private,
save
::
cached_is_io_server
=
.false.
integer,
private,
save, allocatable
::
cached_node_compute_ranks(:)
World ranks of compute processes on this node
integer,
private,
save
::
cached_node_n_compute
=
0
Number of compute ranks on this node
integer,
private,
save
::
cached_rank
=
-1
integer,
private,
save
::
cached_size
=
-1
type(comm_t),
private,
save
::
comm_compute
Compute-only communicator (excludes I/O ranks)
type(comm_t),
private,
save
::
comm_global
Global MPI communicator (cached after init)
type(comm_t),
private,
save
::
compute_comm_override
TEST/TOOLING-ONLY override payload (valid only while
compute_comm_override_depth == 1).
integer,
private,
save
::
compute_comm_override_depth
=
0
TEST/TOOLING-ONLY override stack depth, 0 or 1.
logical,
private,
save
::
env_initialised
=
.false.
True between comm_env_init and comm_env_finalize. Makes init
idempotent and finalize a no-op when MPI was never initialised
(e.g. a test binary that never inits the comm-env), so the shared
per-test main can call finalize unconditionally at exit.
Finalise MPI. No-op if the comm-env was never initialised, so it
is safe to call unconditionally at program exit – the shared
per-test main does this for the few tests that init MPI, and it
costs nothing for the rest (and on serial builds via the stub).
Phase 1: Initialise MPI, cache world rank/size
Call this before reading config. Idempotent: a second call is a
no-op, so multiple entry points / testdrive cases can call it.
TEST/TOOLING-ONLY. Pops the override pushed by
comm_env_push_compute_comm, restoring comm_env_compute_comm()
to the real compute communicator. Fails loud on pop-without-push
(a bug in the caller, not a state this seam should absorb
silently). Does NOT finalize the popped communicator – the
pusher owns that (e.g. freeing a self-comm built for the push).
Per-rank one-shot diagnostic: world rank -> requested device num /
actual device num reported by the OpenACC runtime. Lets us
confirm that mpirun is binding each rank to its own GPU rather
than serialising N ranks on device 0.