Skip to content

[BugFix] Reject round on a vectorized cast from a non-f32 source - #3350

Open
173787247 wants to merge 1 commit into
tile-ai:mainfrom
173787247:contrib/3301-vectorized-cast-round
Open

173787247 wants to merge 1 commit into
tile-ai:mainfrom
173787247:contrib/3301-vectorized-cast-round

Conversation

@173787247

@173787247 173787247 commented Sep 30, 2026 •

Copy link
Copy Markdown

Refs #3301 (F318).

Why

round and its rbits operand have a PTX lowering only for an f32 source, in the four codegen groups that consult cast_round themselves. Every other source dtype reaches a vectorized branch that emits the plain conversion helper and returns, before the rejection at the end of the cast lowering is reached. The requested rounding mode was dropped silently:

for i in T.vectorized(8):
    B[i] = T.cast(A[i], "float8_e4m3fn", round="rs", rbits=T.uint32(seed))

with A in float16 emitted __tl_cvt_half2_to_fp8x2, the round-to-nearest helper, and no rbits operand anywhere in the generated CUDA. The output was bit-identical across seed values 0x0, 0xFFFFFFFF and 0x12345678, and identical to the same cast with no rounding argument at all. The scalar form of the same cast was rejected:

Fatal: round 'rs' is not supported for cast from float16 to float8_e4m3fn
       (only supported f32 packed stochastic conversions are available)

So one request had two outcomes, decided only by whether the loop was vectorized.

What changed

A non-f32 source with a non-empty round is rejected before the vectorized branches, with the message the scalar path already produces. An f32 source is untouched and still takes the stochastic path through the groups that handle it.

Verification

  • Two cases added to testing/python/language/test_tilelang_cast_rounding.py, the test module for this feature: the rejection of a vectorized rs cast from float16, and a control that the same vectorized cast with no rounding mode still lowers and still emits __tl_cvt_half2_to_fp8x2. The rejection fails on an unmodified tree with DID NOT RAISE Exception; the control passes in both states.

  • test_tilelang_cast_rounding.py, test_tilelang_language_fp8.py, testing/python/quantize/ and test_tilelang_language_vectorized_cast.py: 134 passed, 1 skipped, 19 failed. Those 19 fail identically on an unmodified tree and are an environment limit, not a regression:

    tl_templates/cuda/cuda_fp4.h(313): static assertion failed with
      "Stochastic rounding f32-to-FP4 requires sm_100a or sm_103a"
    

    This machine is sm_120, so the tests that need sm_100a cannot build there. The count is 19 with and without the change.

  • The kernel cache was cleared before each run.

  • pre-commit run --files <changed files>: all 13 hooks pass.

Summary

Reject vectorized casts that request rounding when the source dtype is not float32. The check uses the existing scalar-path error message. float32 conversion paths remain unchanged.

Add tests that check rejection of vectorized float16 casts with round="rs" and verify that casts without a rounding mode still lower to the half-to-FP8 conversion helper.

Tests

The author reports 134 passed, 1 skipped, and 19 failed across the listed test suites. The author attributes the failures to the test machine's sm_120 architecture, which cannot build tests requiring sm_100a or sm_103a; the author reports the same failures on an unmodified tree. The author also reports that all 13 pre-commit hooks passed. These results were not independently verified.

C++ style / lint notes

The change touches C++ source covered by docs/developer_guide/cpp_style.md, but it does not change the documented style rules. The CI configuration includes the “C++ API Style Audit (warning only)” step. No audit result or current style findings were supplied.


Verified end to end, together with the other contributions from this batch: VERIFICATION.md

`round` and its `rbits` operand have a PTX lowering only for an f32 source, in
the four codegen groups that consult `cast_round` themselves. Every other source
dtype reaches a vectorized branch that emits the plain conversion helper and
returns, so it never reaches the rejection at the end of the cast lowering. The
requested rounding mode was dropped silently:

    for i in T.vectorized(8):
        B[i] = T.cast(A[i], "float8_e4m3fn", round="rs", rbits=T.uint32(seed))

emitted `__tl_cvt_half2_to_fp8x2`, the round-to-nearest helper, with no `rbits`
operand anywhere; the output was bit-identical across seeds and identical to a
cast with no rounding argument at all. The same cast written in a `T.serial`
loop was rejected with "round 'rs' is not supported for cast from float16 to
float8_e4m3fn", so one request had two outcomes decided only by the loop form.

Refuse it for a non-f32 source before the vectorized branches, with the message
the scalar path already produces. An f32 source is untouched and still takes the
stochastic path.
@github-actions

Copy link
Copy Markdown

👋 Hi! Thank you for contributing to the TileLang project.

Please remember to run pre-commit run --all-files in the root directory of the project to ensure your changes are properly linted and formatted. This will help ensure your contribution passes the format check.

We appreciate you taking this step! Our team will review your contribution, and we look forward to your awesome work! 🚀

@coderabbitai

coderabbitai Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: tile-ai/tilelang/.coderabbit.yaml

Review profile: CHILL

Plan: Advanced

Run ID: 039fc29a-ba03-486f-b854-c20f9f8b1b83

📥 Commits

Reviewing files that changed from the base of the PR and between 994b44e and 36c5aaa.

📒 Files selected for processing (2)
  • src/cuda/codegen/codegen_cuda.cc
  • testing/python/language/test_tilelang_cast_rounding.py

Included review availability: This review used your included allowance. Your plan provides up to 8 included reviews per hour; 3 remain after this review.


📝 Walkthrough

Walkthrough

The CUDA cast code generator now rejects annotated rounding modes when the source type is not FP32. New tests cover rejection of stochastic rounding and lowering of vectorized FP16-to-FP8 casts without a rounding mode.

Changes

Cast rounding validation

Layer / File(s) Summary
Rounding guard and cast tests
src/cuda/codegen/codegen_cuda.cc, testing/python/language/test_tilelang_cast_rounding.py
The code generator rejects annotated rounding modes for non-FP32 source types. CUDA tests verify rejection for vectorized FP16-to-FP8 stochastic rounding and verify lowering without a rounding mode.

Priority: ⬇️ Low

Estimated code review effort: 2 (Simple) | ~10 minutes

Change: Bug fix

Suggested reviewers: ljc00118

Merge Risk: ⚪ Minimal · up to 36c5a

Unsupported rounded non-f32 casts now fail instead of silently ignoring the rounding request, while unrounded vectorized casts remain supported. No concrete merge-blocking risk is established.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 50.00% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 6 functions across 2 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the main change: rejecting a rounding mode on vectorized casts from non-f32 sources.
  • Fix all pre-merge checks with AI
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant