Skip to content

Significant difference in pickle file sizes containing CV trajectories, depending on source of CV functions #1171

Description

@Bernadette-Mohr

I have made an interesting observation:

I have to convert coordinate trajectories to CV trajectories and store them separately before generating path histograms (memory concerns, cluster walltime concerns). I pickle the resulting lists of [CV-Trajectory, weight] pairs using the currently highest pickle protocol 5.

The pickled CV trajectories generated from the CV functions used during the TPS runs, and therefore loaded from storage, result in double the file size on disk compared to CV functions defined later on. The generated list objects are the same size in memory for all tested CV combinations, independent of source, as reported by sys.getsizeof().

Is there maybe some metadata stored and returned with the CV objects in the simstore? Or are there other potential reasons I could investigate?

Activity

  1. dwhswenson commented on Feb 1, 2025

    @dwhswenson
    Member

    (Sorry for the delay; this sat in a drafted response from the weekend I received the message -- my apologies.)

    Thanks for the report, @Bernadette-Mohr. I don't have a direct answer, but maybe I can point you in some useful directions.

    First, let me double-check that I'm understanding what's going on (basically, that we're using the same terms). What I understood is that you're taking OPS-generated trajectories, and using OPS (SimStore) CVs to generate values from those trajectories, and storing those CV results (this is what I understand a "CV trajectory" to be).

    Then (and this is where I'm not 100% clear), it sounds like you add some other CVs of interest, which you calculate after the sampling. And the weirdness is that file size (in pickle) for CVs calculated during sampling is about double the size of CVs calculated after sampling. This is, indeed, weird.

    So first, let me know if that isn't accurate. The rest is based on that interpretation.

    My first guess is that this might have to do with the StorableFunctionResult object. I added a StorableFunctionResult object between the main SimStore StorableFunction (generic version of a CV) and actual results because I could use it to ensure storage of new results when running multiple trajectories in parallel. It was supposed to be ephemeral (just a carrier that allowed the results you can about to get stored) but maybe something in your approach is storing both the CV results and the ephemeral container?

    To the broader question:

    Is there maybe some metadata stored and returned with the CV objects in the simstore?

    That's a definite possibility. OPS objects carry a lot of metadata, which if not handled in the way OPS expects, might lead to unexpected results. Pickle is usually pretty good about most of these (e.g., it also knows not to store a new copy of the engine with every snapshot), but maybe you're hitting an edge case I don't know.

    My biggest question here is, what OPS objects are being stored in the larger file? Assuming that you pickle a single object, when you load that from pickle, try the function openpathsampling.experimental.simstore.serialization_helpers.find_all_uuids (yeah, that's a little hidden). It will return a dict mapping UUID to object with that UUID. I would recommend using something like collections.Counter([obj.__class__.__name__ for obj in all_uuids.values()]) to see what classes you have in there (not tested; but something similar should work).

    No guarantee that will help, but it is the first step that I would try!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions