Gather-scatter communication using MPI one-sided puts into a passive target window, with per-rank signalling for completion.
More...
|
| procedure, pass(this) | init (this, send_pe, recv_pe) |
| | Initialise MPI RMA based communication method.
|
| |
| procedure, pass(this) | free (this) |
| | Deallocate MPI RMA based communication method.
|
| |
| procedure, pass(this) | nbsend (this, u, n, tag, deps, strm) |
| | Pack the gathered shared dofs and put them into each neighbour's receive window, then announce the puts. See the type comment for why the puts, the flush and the signals are separate phases.
|
| |
| procedure, pass(this) | nbrecv (this, tag) |
| | No-op: receives are completed by the remote put and its signal.
|
| |
| procedure, pass(this) | nbwait (this, u, n, op, strm) |
| | Wait per neighbour for the signal that its data has landed, reduce the slab into u, and ack the sender so it may overwrite the slab next round.
|
| |
| procedure, pass(this) | nbsend_vec (this, u, n, nc, tag, deps, strm) |
| | Fused nc-component send: pack nc contiguous component blocks per peer slab and put nc*ndofs reals. Buffer indexing and put sizes scale by nc; the signalling is unchanged.
|
| |
| procedure, pass(this) | nbrecv_vec (this, tag, nc) |
| | No-op: receives are completed by the remote put and its signal.
|
| |
| procedure, pass(this) | nbwait_vec (this, u, n, nc, op, strm) |
| | Fused nc-component wait/reduce: per peer, wait on the data signal and reduce nc component blocks into u, then ack the sender.
|
| |
| procedure(gs_comm_init), deferred, pass | init gs_comm_init |
| |
| procedure(gs_comm_free), deferred, pass | free gs_comm_free |
| |
| procedure(gs_nbsend), deferred, pass | nbsend gs_nbsend |
| |
| procedure(gs_nbrecv), deferred, pass | nbrecv gs_nbrecv |
| |
| procedure(gs_nbwait), deferred, pass | nbwait gs_nbwait |
| |
| procedure, pass(this) | init_dofs (this) |
| |
| procedure, pass(this) | free_dofs (this) |
| |
| procedure, pass(this) | init_order (this, send_pe, recv_pe) |
| | Obtains which ranks to send and receive data from.
|
| |
| procedure, pass(this) | free_order (this) |
| |
| procedure, pass(this) | take_schedule (this, src) |
| | Take over the gather-scatter schedule (dof lists and peer order) of src, avoiding a second (expensive) pass over the connectivity. The data is moved rather than copied, so src is left without a schedule and must not be used for communication afterwards (it can still be freed). No communication resources are set up here; complete the handover with init_schedule once src has been freed, so that the two backends never hold their resources at the same time.
|
| |
| procedure, pass(this) | init_schedule (this) |
| | Set up this communication method for the schedule taken over by take_schedule. Collective, as init is.
|
| |
| procedure, pass(this) | nbsend_vec (this, u, n, nc, tag, deps, strm) |
| | Fused vector halo exchange. Default implementations abort; backends that set vec_supported = .true. override them.
|
| |
| procedure, pass(this) | nbrecv_vec (this, tag, nc) |
| | Default fused vector receive. Abort unless a backend overrides it.
|
| |
| procedure, pass(this) | nbwait_vec (this, u, n, nc, op, strm) |
| | Default fused vector wait/reduce. Abort unless a backend overrides it.
|
| |
|
| real(kind=rp), dimension(:), allocatable | send_buf |
| | Origin buffer for the puts. Plain local memory, no window needed.
|
| |
| integer, dimension(:), allocatable | send_ndofs |
| | Number of dofs to send to each peer, in send_pe order.
|
| |
| integer, dimension(:), allocatable | send_offset |
| | Offset of each peer's slab in send_buf (scalar path units)
|
| |
| integer, dimension(:), allocatable | send_rdisp |
| | Offset of our slab in each peer's receive window (scalar path units), exchanged at init.
|
| |
| type(mpi_win) | win_data |
| | Receive window, holding GS_VEC_NC*recv_total reals.
|
| |
| type(c_ptr) | recv_ptr = C_NULL_PTR |
| | Base of the receive window.
|
| |
| integer, dimension(:), allocatable | recv_ndofs |
| | Number of dofs received from each peer, in recv_pe order.
|
| |
| integer, dimension(:), allocatable | recv_offset |
| | Offset of each peer's slab in the receive window (scalar path units)
|
| |
| integer | recv_total = 0 |
| | Total number of received dofs on this rank.
|
| |
| type(mpi_win) | win_sig |
| | Signal window, 2*pe_size counters (data signals then acks)
|
| |
| type(c_ptr) | sig_ptr = C_NULL_PTR |
| | Base of the signal window.
|
| |
| integer(kind=i8) | iter = 0 |
| | Monotonically increasing round counter. The sender replaces the receiver's data signal with it; the receiver replaces the sender's ack with the same value once consumed.
|
| |
| logical | unified = .true. |
| | Whether the windows use MPI_WIN_UNIFIED. In the separate memory model a remote update is only guaranteed visible to a local load after MPI_Win_sync, so the spin waits call it every iteration.
|
| |
| logical | win_alloc = .false. |
| | Whether the windows have been allocated, so free can tell a live backend from one that never got past init.
|
| |
| type(stack_i4_t), dimension(:), allocatable | send_dof |
| | A list of stacks of dof indices local to this process to send to rank_i.
|
| |
| type(stack_i4_t), dimension(:), allocatable | recv_dof |
| | recv_dof(rank_i) is a stack of dof indices local to this process to receive from rank_i. size(recv_dof) == pe_size
|
| |
| integer, dimension(:), allocatable | send_pe |
| | Array of ranks that this process should send to.
|
| |
| integer, dimension(:), allocatable | recv_pe |
| | array of ranks that this process will receive messages from
|
| |
| logical | vec_supported = .false. |
| | Whether this backend implements the fused vector (multi-component) halo exchange (nbsend_vec/nbrecv_vec/nbwait_vec). When .false., the gs_op_r3 caller falls back to nc independent scalar exchanges.
|
| |
Definition at line 73 of file gs_mpi_rma.f90.