Skip to content

Collective NetCDF IO to improve performance in MOM5 (ESM 1.6) #25

Description

@manodeep

Back in the day, @rxy900 and @marshallward implemented a collective io patch for doing parallel netcdf read/write (paper is here, patch is here) . That patch never got merged in to our FMS fork. The biggest advantage is that MOM5 can then create a single output file, while still writing in parallel. This improves throughput (i.e., no longer serial io in MOM5), reduces the number of intermediate files created, and reduces an extra job step (i.e, no collation phase required).

The paper also said that the performance gradient was higher with higher values of io_ny.

Background

Parallel writes into multiple files (i.e., with io_layout = io_nx,io_ny, where at least one of io_nx, io_ny is not 1) was abandoned because the collation step would crash intermittently. However, I had noticed back in Dec, that with io_layout = 4,3, the throughput of ESM 1.6 increased by ~5% (say ~1-1.5 year/day) but the upstream development had moved sufficiently along that it was going to take a fair bit of time to incorporate the changes into our fork. That has changed now with the rapid advancement of the AI agents.

#24 needs to be resolved first.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions