Skip to content

docs(dead-ends): DE-58 — censoring-aware CLint retrain is null; assay-limit clipping recorded - #109

Open
jam-sudo wants to merge 2 commits into
mainfrom
docs/de-58-clint-censoring
Open

jam-sudo wants to merge 2 commits into
mainfrom
docs/de-58-clint-censoring

Conversation

@jam-sudo

@jam-sudo jam-sudo commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner

Summary

  • Finding: Hepatocyte_AZ CLint labels are clipped at the assay limits (16.1% exactly at the 3.0 LLOQ, 11.3% at the 150 cap), so the model never predicts below ~3 µL/min/10⁶ cells. For high-fu, low-clearance drugs this sets a hepatic-CL floor: against PK-DB open-licence plasma timecourses, engine curves are 28× (caffeine) and 22× (acetaminophen) off; midazolam ~2×.
  • Fix tested: XGBoost survival:aft interval loss vs a matched plain-regression control (same data, exclusion, features, folds, hyperparameters).
  • Result: null. 107-holdout, public-clone, 4 seeds: ΔMeta −0.065 / −0.004 / −0.009 / −0.089 (mean −0.042), ΔEngine mixed (mean −0.013), within seed noise and the ~0.42 CI half-width. Exact-point scaffold-CV R² gets worse (0.08 → −0.16). The binding limit is structural discrimination, not the clipping.
  • Side check: production xgboost_clint.json is holdout-clean (median |log err| 0.076 on training rows vs 0.297 on the 24 holdout rows in Hep_AZ; a holdout-excluded retrain gives 0.327).
  • Open (correctness, not headline): renal CL = GFR·fup has no tubular reabsorption (~5–6 L/h excess for caffeine/acetaminophen).

No production model, source, or headline number changes. Local stack reproduces the canonical cache bit-identically (Meta 2.74279) before any arm was run.

Files

  • docs/research/dead-ends.md — DE-58
  • docs/research/diagnosis.md — §1 bullet
  • scripts/clint_censored_ab.py — CV A/B + --benchmark holdout arm (runtime patch, no tracked-model change)
  • data/validation/clint_censored_ab_2026-09-22.json — all numbers

…-limit clipping recorded

PK-DB open-licence timecourse overlay showed engine curves 28x (caffeine) and
22x (acetaminophen) off, traced to Hepatocyte_AZ labels clipped at the 3.0
LLOQ (16.1%) and 150 cap (11.3%). An AFT interval-loss retrain vs a matched
plain control is null on the 107-holdout (dMeta -0.042 mean over 4 seeds,
engine mixed) and worsens exact-point scaffold-CV R2 (0.08 -> -0.16): the
binding limit is structural discrimination, not the clipping. Production
CLint model verified holdout-clean. Renal reabsorption gap noted as open.
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 23, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-23T03:11:31.423475Z 66f027b PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 66f027b130

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +178 to +180
m = fit_aft(X, y, a.sigma)
m.set_attr(aft="1") # predict_log10 converts AFT's original-scale output
m.save_model(a.save_model)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Make saved AFT models compatible with production inference

When --save-model is used to produce a replacement for models/adme/xgboost_clint.json, the saved AFT booster predicts CLint on the original scale, but src/sisyphus/predict/adme.py::_predict_clint always treats model output as log10(CLint) and applies 10**. The custom aft attribute is only checked by this experiment script, so deploying the generated model would yield exponentially inflated clearance values; either save an inference-compatible model or update the production loader to honor the objective/attribute.

Useful? React with 👍 / 👎.

Comment on lines +144 to +146
adme._predict_clint = _predict_clint
sys.argv = ["run_engine_benchmark.py", "--save-json", out]
runpy.run_path(str(ROOT / "scripts" / "run_engine_benchmark.py"), run_name="__main__")

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Allow benchmarks to run without an output file

When a caller supplies --benchmark A or --benchmark B without --out, a.out is None and this constructs sys.argv with a non-string element; the downstream argparse parser raises a TypeError before running the benchmark. Since --out is currently optional and run_engine_benchmark.py itself supports printing results without saving, omit the --save-json pair when no path is provided or make --out conditionally required.

Useful? React with 👍 / 👎.

…dated

The header count had drifted since DE-44; entries run DE-01..DE-58 with no gaps.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant