Skip to content

Organize by ensemble strategy - #320

Merged
jhprinz merged 60 commits into
openpathsampling:masterfrom
dwhswenson:organize_by_ensemble_strategy
Oct 14, 2015
Merged

jhprinz merged 60 commits into
openpathsampling:masterfrom
dwhswenson:organize_by_ensemble_strategy

Conversation

@dwhswenson

Copy link
Copy Markdown
Member

What this PR does: Makes it possible to rearrange the “global” structure of the move decision process. Standard behavior is to select a move type, then a specific move within that. Now we can select an ensemble first, then a move within that. We can also switch between these.

Why we should have this: Ensemble-first organization is a step toward SRTIS. The ability to conserve move probabilities when switching selection order makes comparison between SRTIS and normal RETIS more straightforward, and makes it easy for users to switch from one approach to the other.

What is hard about this: My intuition expects different behaviors between ensemble-first and move group-first organization. Details below, but figuring our that there was a difference, and then figuring our exactly what the difference was took some effort (and a major rewrite after my first draft of this).

Tasks:

  • EnsembleDictionaryMover
  • OrganizeByEnsembleStrategy
  • OrganizeByMoveGroupStrategy (rename of Default)
  • Check that we get the correct weights regardless of top-level organization strategy
  • Update docstrings

Background

Before getting into the details of what this PR contains, let me refresh a few essential background points. A few of these came from #289. The big thing is that movers can be uniquely defined by (groupname, ensemble_signature), and that the main thing that comes out of the highest-level organization strategy is the choice_probabilities, which is a dictionary of {PathMover : probability} .

Goals

  • code should allow different intuitive behavior for move-type organization and for ensemble organization
  • scheme.choice_probabilities should be preserved (unless explicitly overridden)
  • general schemes should have a reasonable default weights

Intuitive behavior differs by global organization type

This is one of the weirder parts of this PR, but it basically boils down to this: when I set the move-type weights in a move-type-first organization scheme, I think of a dictionary like this:

group_weights = {
    ‘shooting’ : 1.0,
    ‘repex’ : 0.5,
    ‘pathreversal’ : 0.5,
    ‘minus’ : 0.2
}

The tricky bit is that there may be 5 shooting movers for every minus mover. If the weights represented “probability of choosing this move type”, then we’d do 5 shooting moves for every minus move. But since there are 5 different shooting movers, that would mean each individual shooting mover was used as frequently as the minus mover.

That’s not what I want. I want the minus move to be done less frequently than shooting: about 1/5 as often. So rather than make the user do the math to get that to work (as my old code did), I make it so that the code figures it out automatically. This means that (assuming all ensembles have the same weight) each shooting move will be 5 times as likely to occur as a minus move. That’s what we want.

On the other hand, ensemble-first organization does just what you’d expect: first you pick an ensemble with the random probability given by ensemble_weights, then you pick a move which has that ensemble as one of its initial ensembles.

Ensemble-first organization does not have a unique solution

The quantity that is conserved is the total choice probability for each ensemble. However, some ensembles (e.g., replica exchange) have more than one input ensemble. This means that they show up in more than one place in the ensemble-first decision process. There is not a unique way to split these weights, so the approach I’ve used is to say that each ensemble contributes the same to the total probability of choosing that mover. Details on this problem, and the approach I used to solve it, are in this gist.

Reasonable defaults

On one hand, it would make sense to normalize most of these results: so the probabilities should be normalized probabilities. On the other hand, that’s harder as a way to think about it -- and obviously doesn’t work with the group_weights described above.

So the approach I’ve taken is the following:

  1. choice_probability really is a normalized probability
  2. group_weights is rescaled such that ’shooting’ is 1.0 (we usually think in terms of shooting movers); if there is no group called ’shooting’, then we rescale so that the most common value in the group_weights dictionary is 1.0.
  3. mover_weights are rescaled so that the most common number within each second-level decision is 1.0.

These are so that, if you say “I want this twice as much shooting in ensembleA as ensembleB,” you can just change the default 1.0 to 2.0.

Caveats when overriding

There are two levels of weights to override -- the weight that is sorted by (either move group or input ensemble) and the weight within that sorting group.

For the sorting groups, overriding the weights are easy: the input dictionaries take the group name or the ensemble and map it to a weight (group_weight or ensemble_weight).

For the second level, this is a little more complicated. First, note that these weights should only be used by pretty advanced users -- typical needs are met without them. These weights are referred to as mover_weights, and the shape of the dictionary keys is different between the two types of organization strategy. For OrganizeByMoveGroupStrategy, the keys are (groupname, ensemble_signature); for OrganizeByEnsembleStrategy, they are (groupname, ensemble_signature, ensemble).

In both cases, the mover_weights default to 1.0, as does ensemble_weights. However, group_weights defaults to the version shown above. This means that to get the standard default behavior with ensemble-first organization, your scheme needs an OrganizeByMoveGroupStrategy, followed by an OrganizeByEnsembleStrategy. (This will eventually be hidden to SRTIS users, who will just ask for an SRTISScheme that handles all of that automatically.)

Still needs completed tests
In particular, there's a problem that the make_choosers was only working if
there was actually a mover of each type for each ensemble. Of course, this
isn't the case: for the minus ensemble, there's only one legal type of move!

Still need to do some serious cleanup to get this working.
Should be ready to add the new choosers and root next -- looks like the
default weight numbers are reasonable.
Runs, but still needs extensive testing.
@dwhswenson dwhswenson added the ref docs issues/PRs with discussion that might be added to docs label Oct 6, 2015
@dwhswenson dwhswenson self-assigned this Oct 7, 2015
@dwhswenson dwhswenson added this to the 1.0 milestone Oct 8, 2015
@dwhswenson dwhswenson changed the title [WIP] Organize by ensemble strategy Organize by ensemble strategy Oct 12, 2015
@dwhswenson

Copy link
Copy Markdown
Member Author

Finally! I'm calling this one ready for review and merge. I've only been working on this since, oh, June.

@dwhswenson dwhswenson assigned jhprinz and unassigned dwhswenson Oct 12, 2015

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What would you call this then? The underlying idea was that since we accept samples we check, if all samples can be accepted and otherwise we multiply the bias which I think is something like the forward proposal probability.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This should not be a method. That's my point. It breaks the ability to maintain detailed balance for arbitrary moves. The overall acceptance is not of a sample. It is of an entire active sample set. In most moves, most of the sample set doesn't change, so this doesn't matter. However, the current structure of path movers will not work in the general case. See #251, #294.

@jhprinz

jhprinz commented Oct 13, 2015

Copy link
Copy Markdown
Contributor

Hmm, my post disappeared... Anyway just a quick question to confirm.
Choosing a different organization will completely rearrange the move tree and have first a random choice mover that selects the ensemble or the move group and then have a next set of random choice below that. So the move tree will look very different from what it used to if you choose ensemble first, right?

Anyway, I think this is ready to be merged, yes?

@dwhswenson

Copy link
Copy Markdown
Member Author

Choosing a different organization will completely rearrange the move tree and have first a random choice mover that selects the ensemble or the move group and then have a next set of random choice below that. So the move tree will look very different from what it used to if you choose ensemble first, right?

Precisely. Single replica TIS needs to be based on a "choose-the-ensemble-first" strategy (because the ensemble is given, so there's no "choice" to be had). This lets us switch between the two approaches.

The important thing is that we maintain the same probabilities for each mover. The original approach was "choose a move type, then choose a mover." From this we can get an absolute probability of choosing each individual mover. The ensemble-first strategy is "choose an ensemble, then choose a mover." Again, we can get the absolute probability of each individual mover. The smart part of this PR is that when we change from one organization strategy to the other, those absolute probabilities remain the same.

Anyway, I think this is ready to be merged, yes?

As soon as the newest commit passes tests, yes, it should be.

@jhprinz

jhprinz commented Oct 14, 2015

Copy link
Copy Markdown
Contributor

Great, this is really cool. Looking forward to seeing this in the move tree ...

Merging...

jhprinz added a commit that referenced this pull request Oct 14, 2015
@jhprinz
jhprinz merged commit 8b36fdf into openpathsampling:master Oct 14, 2015
@dwhswenson
dwhswenson deleted the organize_by_ensemble_strategy branch October 26, 2015 00:50
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ref docs issues/PRs with discussion that might be added to docs

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants