Skip to content

Xo_res initialization on CPU for OpenACC version and on GPU for CUDAFortran #153

Description

@bellenlau

Line 69 in X_irredux_residuals.F fills Xo_res with zeros

  • on the GPU for CUDAFortran version (as being a device variable)
  • on the CPU for OpenACC version (as being a host variable)

The time taken for this operation on the CPU and on the GPU differs significantly and becomes evident when distributing the simulation in the OpenACC version, as being done on the CPU and not on the GPU.

In the following pictures, a small test running on 4 gpus is traced with nsys; nvtx range labelled "issue3" wraps line 68-69.

OpenACC small test

openacc-case

CUDAFortran small test

cudaf-case

The time becomes detrimental when running larger systems maximally distributed (e.g. GrCo-7k on 16 nodes in the following picture). This simulation takes 3 minutes for CUDAFortran version, but overcomes the walltime for OpenACC version.

image

If zeroing Xo_res is actually needed, a possible fix is kernels construct: bellenlau@8efaf61, but I am not sure if a counterpart in OpenMP offload exists.

With kernels, small test on 4 gpus, OpenACC version:

openacc-kernels-fix

The data: issue.tar.gz
software stack: nvhpc/23.1, no present clauses, openmpi/4.1.4 on Leonardo

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions