Back in the day, @rxy900 and @marshallward implemented a collective io patch for doing parallel netcdf read/write (paper is here, patch is here) . That patch never got merged in to our FMS fork. The biggest advantage is that MOM5 can then create a single output file, while still writing in parallel. This improves throughput (i.e., no longer serial io in MOM5), reduces the number of intermediate files created, and reduces an extra job step (i.e, no collation phase required).
The paper also said that the performance gradient was higher with higher values of io_ny.
Background
Parallel writes into multiple files (i.e., with io_layout = io_nx,io_ny, where at least one of io_nx, io_ny is not 1) was abandoned because the collation step would crash intermittently. However, I had noticed back in Dec, that with io_layout = 4,3, the throughput of ESM 1.6 increased by ~5% (say ~1-1.5 year/day) but the upstream development had moved sufficiently along that it was going to take a fair bit of time to incorporate the changes into our fork. That has changed now with the rapid advancement of the AI agents.
#24 needs to be resolved first.
Back in the day, @rxy900 and @marshallward implemented a collective io patch for doing parallel netcdf read/write (paper is here, patch is here) . That patch never got merged in to our FMS fork. The biggest advantage is that MOM5 can then create a single output file, while still writing in parallel. This improves throughput (i.e., no longer serial io in MOM5), reduces the number of intermediate files created, and reduces an extra job step (i.e, no collation phase required).
The paper also said that the performance gradient was higher with higher values of
io_ny.Background
Parallel writes into multiple files (i.e., with
io_layout = io_nx,io_ny, where at least one ofio_nx, io_nyis not 1) was abandoned because the collation step would crash intermittently. However, I had noticed back in Dec, that withio_layout = 4,3, the throughput of ESM 1.6 increased by ~5% (say ~1-1.5 year/day) but the upstream development had moved sufficiently along that it was going to take a fair bit of time to incorporate the changes into our fork. That has changed now with the rapid advancement of the AI agents.#24 needs to be resolved first.