Skip to content

Add path ensemble theory properties to ensembles #205

Description

@dwhswenson

I think we need to add the path ensemble theory properties to our ensembles. In some cases, these properties ensure that certain algorithms are correct (so they serve as a safeguard against user error). In other cases, they enable faster algorithms in common situations.

There are several approaches to this, but I think the best is to have a dictionary of properties, and then each "construction" ensemble (logical combinations, sequential ensembles, etc) should have a path_ensemble_theory function which builds that dictionary using stuff from path ensemble theory.

The most urgent example where we need this is for frame-by-frame ensembles. Ensembles are functions of trajectories, so when we check an ensemble, we need to check all frames in it. The exception is frame-by-frame ensembles (AllInX and AllOutX being prime examples) for which (assuming we've previously passed checks on all previous frames) we only need to check the most recently added frame to know whether we're (still) in the ensemble. Because of this, the caching I'm doing in #204 won't really solve the problem: it has to cache on an ensemble level, not a frame level. We need a way to implement the faster algorithm if our ensemble allows it, while still falling back on the slower algorithm in other cases.

The work I've been doing on path ensemble theory gives us most of the tools we need, although I think I need to find a better description than frame-by-frame for the exact concept I'm getting at.

Activity

  1. dwhswenson commented on Apr 5, 2015

    @dwhswenson
    MemberAuthor

    On further thought, I think the best way to solve the immediate problem might be with the lazy flag we have on Ensemble.__call__ (and adding it to can_append and can_prepend). I don't think we currently use this flag, but wasn't something like this its original intention? I would change the name to trusted, though (indicating that the trajectory, except for the last frame, is trusted to have been checked frame-by-frame). For ensembles where this allows a faster algorithm (AllInX, AllOutX) this will allow us to speed things up by having the SequentialEnsemble call the trusted version in its _find_subtraj_first/final methods, which I think are the major culprits.

    I still think we'll eventually want the path ensemble theory metadata (for useful sanity checks and preventing users from doing things that aren't well-defined), but that's a lower priority if this speed issue can be solved another way.

  2. jhprinz commented on Apr 6, 2015

    @jhprinz
    Contributor

    Yes, I used the lazy property to allow for "recursive" testing if can_append makes sense. Usually this is what we want since if it has not been in the expected ensemble before we would not have run a next step. Still, for later checking we have to be able to check the whole thing at once.
    I thought lazy is a commonly used term to describe that not everything is checked but in a loose, and less complete sense, hene lazy. Trusted is more precise, but still misses the recursive idea and is more general. It would be good if the name indicated that we assume that .can_append(traj, trusted=True) just assumes __call__(traj) is True

    For the other properties we could have a class like EnsembleProperties that knows all the properties we are interested in. This way we can also use auto-complete in IPython and IDEs.

  3. dwhswenson commented on Apr 6, 2015

    @dwhswenson
    MemberAuthor

    I thought lazy is a commonly used term to describe that not everything is checked but in a loose, and less complete sense, hene lazy.

    It's already too commonplace in our code. We had at least two different meanings of lazy in ensemble.py alone. So I tried to sort out the ones that were about the evaluation of __call__ (there was no such argument in recent versions of can_append; I re-added that) and I renamed those as trusted. We also have that as an ensemble attribute: that is incorrect. It is not a property of the ensemble, but a property of the context in which the function is called. I haven't removed the ensemble property yet, but we should.

    It would be good if the name indicated that we assume that .can_append(traj, trusted=True) just assumes __call__(traj) is True

    For the sake of precision, this statement is not quite correct. can_append(traj, trusted=True) assumes that can_append(traj[:-1]) is True. It does not relate to __call__ except insofar as, by coincidence, AllInXEnsemble.__call__(traj) == AllInXEnsemble.can_append(traj) for all traj. It also, of course, depends on traj[:-1], not traj.

    I thought trusted fit the description, in the sense that "At this point, when I call this function, I trust that the trajectory has been built in such a way that all previous frames were checked." Again, it's not a property of the ensemble, but a property of when you call the function. If you have a different word to use, I'd be happy to use it instead, but I think that we need to avoid re-using lazy in too many different contexts.

  4. jhprinz commented on Apr 7, 2015

    @jhprinz
    Contributor

    I see, then I understood this a little wrong. Okay so then trusted is indeed a good name. The old code just used this recursive idea and so I assumed this is here the same.

    Actually, I also think that lazy is not a good name and too broad in general and we might try to avoid it everywhere it is used so far. I think trusted is good in this context and is short enough.

    Will merge this, no further objections.

  5. added this to the Future milestone on May 14, 2015
  6. dwhswenson commented on Mar 28, 2016

    @dwhswenson
    MemberAuthor

    About adding Path Ensemble Theory ideas to our ensembles: I mentioned this again in #455, I thought I'd add to this (older) issue to clarify what I meant in there.

    Currently, the approach we use is very much like path ensemble theory (well, more accurately, path ensemble theory was developed to match the approach we use). As a specific example, think about what happens when we do a shooting move with a TIS ensemble. First, we have to generate the candidate trajectory. Then we take that candidate trajectory and check whether it is in the ensemble. In particular, the code asks the following questions:

    1. Does this start with one frame in stateA?
    2. Is that then followed by frames that are neither in stateA nor stateB?
    3. In the bunch of frames from the previous question, is at least one outside of interface?
    4. After those frames, is there exactly 1 frame in either stateA or stateB?

    That's 4 calls to check volumes (technically more, since stateA | stateB calls other checks) for each frame.

    However, by the nature of the candidate trajectory, the only question that isn't already guaranteed to be true is (3). So we could avoid asking all the other questions (or take some shortcut to say they were already implicitly answered), and speed up our calculation. This speed is actually important when dealing with toy models like the 2D stuff (though probably not significant if studying solvated kinases....)

    Now, this is all easy enough to see in the case of TIS. But the beauty of OPS is that it is designed to work with any ensemble we might create. And that's where Path Ensemble Theory comes in. We define the ensemble of candidate trajectories and the ensemble of acceptable trajectories. For some ensembles, all candidates are also acceptable. This property is propagated in certain ways if the ensemble is combined in intersections with other ensembles, or if it is combined in sequences with other ensembles. As a result, we should be able to give a general mathematical justification, based on the properties of an ensemble, for not needing to check whether a candidate trajectory satisfies that ensemble.

    In other words, that little PET project of mine will also lead to much faster code, once we put everything together properly. There are many other things to consider adding, but since I just thought of this one today, I'm writing it down now so that I don't forget about it.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions