Gather-scatter communication using an MPI neighbourhood collective. The whole halo exchange is carried out by a single MPI_Ineighbor_alltoallv over a distributed-graph communicator built once from the send/recv peer lists. Compared to the point-to-point gs_mpi_t this trades N Isend/Irecv pairs for one collective call: a single entry into the (serialised, on Fugaku) MPI runtime, leaving the library free to stripe the transfers across all network interfaces internally.
More...
|
| procedure, pass(this) | init (this, send_pe, recv_pe) |
| | Initialise the neighbourhood-collective communication method See gs_comm.f90 for details.
|
| |
| procedure, pass(this) | free (this) |
| | Deallocate the neighbourhood-collective communication method.
|
| |
| procedure, pass(this) | nbsend (this, u, n, tag, deps, strm) |
| | See gs_comm.f90 for more details on these routines.
|
| |
| procedure, pass(this) | nbrecv (this, tag) |
| | No-op: the neighbourhood collective issued in nbsend handles both the send and receive directions, so there is nothing to pre-post here.
|
| |
| procedure, pass(this) | nbwait (this, u, n, op, strm) |
| | Wait for the neighbourhood collective and reduce the received slabs.
|
| |
| procedure, pass(this) | init_vec (this) |
| | Allocate the fused vector exchange buffers and the nc-scaled collective descriptors, sized for GS_VEC_NC components. Deferred to the first fused exchange, see gs_comm_t. The graph communicator built in init is shared with the scalar exchange and is not touched here, so this stays rank local.
|
| |
| procedure, pass(this) | nbsend_vec (this, u, n, nc, tag, deps, strm) |
| | Pack the nc components and launch a single neighbourhood collective.
|
| |
| procedure, pass(this) | nbrecv_vec (this, tag, nc) |
| | No-op: the collective is launched in nbsend_vec.
|
| |
| procedure, pass(this) | nbwait_vec (this, u, n, nc, op, strm) |
| | Wait for the vector collective and reduce each received slab into u. Peer i's received slab holds nc consecutive blocks of recvcounts(i) values, matching the sender's packing convention.
|
| |
| procedure(gs_comm_init), deferred, pass | init gs_comm_init |
| |
| procedure(gs_comm_free), deferred, pass | free gs_comm_free |
| |
| procedure(gs_nbsend), deferred, pass | nbsend gs_nbsend |
| |
| procedure(gs_nbrecv), deferred, pass | nbrecv gs_nbrecv |
| |
| procedure(gs_nbwait), deferred, pass | nbwait gs_nbwait |
| |
| procedure, pass(this) | init_dofs (this) |
| |
| procedure, pass(this) | free_dofs (this) |
| |
| procedure, pass(this) | init_order (this, send_pe, recv_pe) |
| | Obtains which ranks to send and receive data from.
|
| |
| procedure, pass(this) | free_order (this) |
| |
| procedure, pass(this) | take_schedule (this, src) |
| | Take over the gather-scatter schedule (dof lists and peer order) of src, avoiding a second (expensive) pass over the connectivity. The data is moved rather than copied, so src is left without a schedule and must not be used for communication afterwards (it can still be freed). No communication resources are set up here; complete the handover with init_schedule once src has been freed, so that the two backends never hold their resources at the same time.
|
| |
| procedure, pass(this) | init_schedule (this) |
| | Set up this communication method for the schedule taken over by take_schedule. Collective, as init is.
|
| |
| procedure, pass(this) | init_vec (this) |
| | Fused vector halo exchange. Default implementations abort; backends that set vec_supported = .true. override them.
|
| |
| procedure, pass(this) | nbsend_vec (this, u, n, nc, tag, deps, strm) |
| | Default fused vector send. Abort unless a backend overrides it.
|
| |
| procedure, pass(this) | nbrecv_vec (this, tag, nc) |
| | Default fused vector receive. Abort unless a backend overrides it.
|
| |
| procedure, pass(this) | nbwait_vec (this, u, n, nc, op, strm) |
| | Default fused vector wait/reduce. Abort unless a backend overrides it.
|
| |
|
| real(kind=rp), dimension(:), allocatable | send_buf |
| | Concatenated send slabs, one per peer, in send_pe order.
|
| |
| real(kind=rp), dimension(:), allocatable | recv_buf |
| | Concatenated recv slabs, one per peer, in recv_pe order.
|
| |
| integer, dimension(:), allocatable | sendcounts |
| | Per-peer counts and 0-based element displacements for the neighbourhood collective. send_* follow send_pe (the graph destinations), recv_* follow recv_pe (the graph sources).
|
| |
| integer, dimension(:), allocatable | sdispls |
| |
| integer, dimension(:), allocatable | recvcounts |
| |
| integer, dimension(:), allocatable | rdispls |
| |
| type(mpi_comm) | neigh_comm |
| | Distributed-graph communicator over the halo neighbourhood.
|
| |
| type(mpi_request) | request |
| | In-flight non-blocking neighbourhood collective.
|
| |
| real(kind=rp), dimension(:), allocatable | send_buf_v |
| | Fused vector (multi-component) exchange buffers, sized for up to GS_VEC_NC components. Each peer slab holds nc consecutive component blocks. sendcounts_v/sdispls_v etc. are the nc-scaled collective descriptors, filled per call.
|
| |
| real(kind=rp), dimension(:), allocatable | recv_buf_v |
| |
| integer, dimension(:), allocatable | sendcounts_v |
| |
| integer, dimension(:), allocatable | sdispls_v |
| |
| integer, dimension(:), allocatable | recvcounts_v |
| |
| integer, dimension(:), allocatable | rdispls_v |
| |
| type(stack_i4_t), dimension(:), allocatable | send_dof |
| | A list of stacks of dof indices local to this process to send to rank_i.
|
| |
| type(stack_i4_t), dimension(:), allocatable | recv_dof |
| | recv_dof(rank_i) is a stack of dof indices local to this process to receive from rank_i. size(recv_dof) == pe_size
|
| |
| integer, dimension(:), allocatable | send_pe |
| | Array of ranks that this process should send to.
|
| |
| integer, dimension(:), allocatable | recv_pe |
| | array of ranks that this process will receive messages from
|
| |
| logical | vec_supported = .false. |
| | Whether this backend implements the fused vector (multi-component) halo exchange (nbsend_vec/nbrecv_vec/nbwait_vec). When .false., the gs_op_r3 caller falls back to nc independent scalar exchanges.
|
| |
| logical | vec_ready = .false. |
| | Whether the buffers the fused vector exchange needs are in place. They are sized GS_VEC_NC times the halo, quadrupling what the backend holds, and only gs_op_r3 ever touches them, so a backend that can allocate them on its own defers that to the first fused exchange (see init_vec) rather than paying for it in every run. A backend whose vector buffers are part of an allocation the whole run has to agree on – symmetric memory, coarrays, registered memory, an RMA window – cannot defer, since a rank with no shared dofs never reaches the first fused exchange; those allocate in init and set this there.
|
| |
Definition at line 56 of file gs_neighbour.f90.