Conversation
|
The docs for this PR live here. All of your documentation changes will be reflected on that endpoint. The docs are available until 30 days after the last update. |
sayakpaul
left a comment
There was a problem hiding this comment.
Didn't review it fully but some changes are too intrusive, IMO. Hopefully, the existing comments will be helpful to get an idea of what I am talking about and help you make changes to the files I didn't review yet.
|
|
||
| ## Optimization | ||
|
|
||
| You can optimize the pipeline's runtime and memory consumption with torch.compile and feed-forward chunking. To learn about other optimization methods, check out the [Speed up inference](../../optimization/fp16) and [Reduce memory usage](../../optimization/memory) guides. |
There was a problem hiding this comment.
I think precision and compilation as a title is less exciting as a reader in the sense that it doesn't tell me that it's related to speeding up of inference.
There was a problem hiding this comment.
i think it's a little strange to broadly name it "Accelerate inference" but then it doesn't include things like caching or the attention backends, which may cause users to overlook those last two options. so i'd rather keep it as "Precision and compilation"
| <hfoption id="memory"> | ||
|
|
||
| Refer to the [Reduce memory usage](../../optimization/memory) guide for more details about the various memory saving techniques. | ||
| Refer to the [Memory and offloading](../../optimization/memory) guide for more details about the various memory saving techniques. |
There was a problem hiding this comment.
I think this change makes it less obvious to the reader the page is about something they would actually care about (i.e., reducing memory usage).
There was a problem hiding this comment.
i agree here that "Reduce memory usage" is better since it has all the memory techniques on one page :)
Softened some overly absolute claims, hopefully this is what you were referring to 😄 |
sayakpaul
left a comment
There was a problem hiding this comment.
Left some further comments. Thanks for iterating!
|
|
||
| | Method | Use when | Tradeoff | | ||
| |--------|----------|----------| | ||
| | Text KV Cache | NucleusMoE image only, need exact text K/V reuse across steps | Lossless | |
There was a problem hiding this comment.
Hmm, it was added alongside NucleusMoE but I don't see anything that would prevent it from being used in other pipelines.
Trade-offs are usually brittle because quality is subjective. So, maybe we should suggest to the readers that they should inspect the quality of the outputs because each caching technique has its own trade-off (better memory consumption, better speed, better preservation of the original outputs, etc.). Not sure how to present this faithfully in the table.
The current table and wording have most of it. My biggest worry is about terms like "loseless". Open to ideas.
There was a problem hiding this comment.
it seems like apply_text_kv_cache only applies its hooks to NucleusMoEImageTransformerBlock so on other models it doesn't do anything?
| | Method | Use when | Tradeoff | | ||
| |--------|----------|----------| | ||
| | Text KV Cache | NucleusMoE image only, need exact text K/V reuse across steps | Lossless | | ||
| | SeaCache | Video transformers that already have a SeaCache path | Approximate, settings often do not transfer across models | |
There was a problem hiding this comment.
As per the Cosmos3 documentation page, it seems like it's supposed to only work with Cosmos3. @yzhautouskay can probably confirm.
| attention_weight_callback=lambda _: 0.3, | ||
| unconditional_batch_skip_range=5, | ||
| unconditional_batch_timestep_skip_range=(-1, 781), | ||
| unconditional_batch_timestep_skip_range=(-1, 641), |
There was a problem hiding this comment.
i used the default value from below, but happy to change it back to 781 if thats better
diffusers/src/diffusers/hooks/faster_cache.py
Line 151 in c60830e
| ``` | ||
|
|
||
| To learn more, take a look at the [Reduce memory usage](../optimization/memory) and [Accelerate inference](../optimization/fp16) guides. | ||
| To learn more, take a look at [Optimize and scale](../stable_diffusion#optimization-techniques). |
There was a problem hiding this comment.
../stable_diffusion#optimization-techniques
Do we still have a reason keep these under "stable_diffusion". I don't see any reason to.
There was a problem hiding this comment.
stable_diffusion.md is actually the overview, so it may be nice to point users toward an overview of optimization techniques for a particular task and then they can choose what they want
fedf28e to
342b599
Compare
Refreshes the Optimize and scale section:
TextKVCache(can remove if this is too niche, but good to document for completeness)