Repository navigation
Major performance improvements: Ensembles, Volumes, Splitting - #455
Merged
Merged
Conversation
Member
Author
Member
Author
|
Passed tests. Ready for review. |
Contributor
|
Very nice! Really. This is how I like improvements... I think with all the changes in storage lately we should change the chaindict approach also a little. But that is different PR. Maybe cython will do some tricks. I think we might remove some convenience and trade it for speed. Especially the flexibility in storing and caching etc. is possible not necessary. Looks good and will merge. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.


Performance improvements. Five-fold speedup in the DNA analysis I was doing; probably less (but still significant) in our examples.
logger.isEnabledFor(logging.DEBUG)scopes.splitis using an algorithm that does not scale linearly (not usingtrusted=True)LengthbeforeVolume) in flux calculationTimings (in order of my implementations, so these are cumulative improvements) to use$N=1$ ).
SingleTrajectoryAnalysis.analyze_flux()on a DNA trajectory of 10k frames. Inexact, since I just ran the test once on my laptop, but when you see a calculation drop from ~5 minutes to ~1 minute, that’s well beyond the margin of error (even for my sample ofmasterbranch:264.98s user 1.95s system 89% cpu 4:56.82 total127.29s user 1.64s system 89% cpu 2:24.61 total96.98s user 1.45s system 90% cpu 1:49.07 total62.83s user 1.35s system 80% cpu 1:19.98 total57.13s user 1.20s system 85% cpu 1:08.24 totalOf course, getting such improvement from the short-circuit volumes requires defining your volumes to take the most advantage of short-circuit logic. But that was easy to do with this particular system.
There’s some room to improve
SequentialEnsemble.can_append, especially in its use of_find_subtraj_final. This should give us a significant speed boost. I’ll do that at a later date.Other notes for future speedups:
chaindictis always going to be inner-loop stuff. It looks like a good target for cythonization -- it does a lot of python function calls deep in inner loops, and Cython can massively decrease the overhead cost of those function calls.