Skip to content

[docs] Optimize and scale pt. 1 - #14867

Open
stevhliu wants to merge 9 commits into
huggingface:mainfrom
stevhliu:optimize-scale
Open

stevhliu wants to merge 9 commits into
huggingface:mainfrom
stevhliu:optimize-scale

Conversation

@stevhliu

Copy link
Copy Markdown
Member

Refreshes the Optimize and scale section:

  • add an Overview with a starter path (dtype + device → offload) and links to other optimization techniques such as kernels and quantization
  • more accurate doc titles like "Precision and compilation" that reflect what they actually teach vs a more vague promise like "Reduce memory usage"
  • adds a section to guide cache method selection and also adds TextKVCache (can remove if this is too niche, but good to document for completeness)
  • reduce noisy prose (tighter intros, fewer callouts, etc.) and more consistent navigation
  • make sure hardware/API claims (FA3/Hopper, kernels requirements) matches code

@github-actions github-actions Bot added documentation Improvements or additions to documentation size/L PR with diff > 200 LOC labels Sep 24, 2026
@HuggingFaceDocBuilderDev

Copy link
Copy Markdown

The docs for this PR live here. All of your documentation changes will be reflected on that endpoint. The docs are available until 30 days after the last update.

@stevhliu
stevhliu requested a review from sayakpaul September 24, 2026 21:57

@sayakpaul sayakpaul left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Didn't review it fully but some changes are too intrusive, IMO. Hopefully, the existing comments will be helpful to get an idea of what I am talking about and help you make changes to the files I didn't review yet.


## Optimization

You can optimize the pipeline's runtime and memory consumption with torch.compile and feed-forward chunking. To learn about other optimization methods, check out the [Speed up inference](../../optimization/fp16) and [Reduce memory usage](../../optimization/memory) guides.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think precision and compilation as a title is less exciting as a reader in the sense that it doesn't tell me that it's related to speeding up of inference.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

i think it's a little strange to broadly name it "Accelerate inference" but then it doesn't include things like caching or the attention backends, which may cause users to overlook those last two options. so i'd rather keep it as "Precision and compilation"

<hfoption id="memory">

Refer to the [Reduce memory usage](../../optimization/memory) guide for more details about the various memory saving techniques.
Refer to the [Memory and offloading](../../optimization/memory) guide for more details about the various memory saving techniques.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this change makes it less obvious to the reader the page is about something they would actually care about (i.e., reducing memory usage).

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

i agree here that "Reduce memory usage" is better since it has all the memory techniques on one page :)

Comment thread docs/source/en/optimization/attention_backends.md Outdated
Comment thread docs/source/en/optimization/attention_backends.md Outdated
Comment thread docs/source/en/optimization/attention_backends.md Outdated
Comment thread docs/source/en/optimization/attention_backends.md Outdated
Comment thread docs/source/en/optimization/attention_backends.md Outdated
@stevhliu

Copy link
Copy Markdown
Member Author

Hopefully, the existing comments will be helpful to get an idea of what I am talking about and help you make changes to the files I didn't review yet.

Softened some overly absolute claims, hopefully this is what you were referring to 😄

@sayakpaul sayakpaul left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Left some further comments. Thanks for iterating!

Comment thread docs/source/en/optimization/attention_backends.md Outdated
Comment thread docs/source/en/optimization/attention_backends.md
Comment thread docs/source/en/optimization/cache.md Outdated

| Method | Use when | Tradeoff |
|--------|----------|----------|
| Text KV Cache | NucleusMoE image only, need exact text K/V reuse across steps | Lossless |

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm, it was added alongside NucleusMoE but I don't see anything that would prevent it from being used in other pipelines.

Trade-offs are usually brittle because quality is subjective. So, maybe we should suggest to the readers that they should inspect the quality of the outputs because each caching technique has its own trade-off (better memory consumption, better speed, better preservation of the original outputs, etc.). Not sure how to present this faithfully in the table.

The current table and wording have most of it. My biggest worry is about terms like "loseless". Open to ideas.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it seems like apply_text_kv_cache only applies its hooks to NucleusMoEImageTransformerBlock so on other models it doesn't do anything?

Comment thread docs/source/en/optimization/cache.md Outdated
| Method | Use when | Tradeoff |
|--------|----------|----------|
| Text KV Cache | NucleusMoE image only, need exact text K/V reuse across steps | Lossless |
| SeaCache | Video transformers that already have a SeaCache path | Approximate, settings often do not transfer across models |

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

As per the Cosmos3 documentation page, it seems like it's supposed to only work with Cosmos3. @yzhautouskay can probably confirm.

attention_weight_callback=lambda _: 0.3,
unconditional_batch_skip_range=5,
unconditional_batch_timestep_skip_range=(-1, 781),
unconditional_batch_timestep_skip_range=(-1, 641),

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What happened here? 👀

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

i used the default value from below, but happy to change it back to 781 if thats better

unconditional_batch_timestep_skip_range: tuple[int, int] = (-1, 641)

Comment thread docs/source/en/optimization/cache.md Outdated
Comment thread docs/source/en/optimization/fp16.md Outdated
Comment thread docs/source/en/optimization/fp16.md Outdated
```

To learn more, take a look at the [Reduce memory usage](../optimization/memory) and [Accelerate inference](../optimization/fp16) guides.
To learn more, take a look at [Optimize and scale](../stable_diffusion#optimization-techniques).

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

../stable_diffusion#optimization-techniques

Do we still have a reason keep these under "stable_diffusion". I don't see any reason to.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

stable_diffusion.md is actually the overview, so it may be nice to point users toward an overview of optimization techniques for a particular task and then they can choose what they want

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation size/L PR with diff > 200 LOC

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants