Repository navigation
Conversation
…-limit clipping recorded PK-DB open-licence timecourse overlay showed engine curves 28x (caffeine) and 22x (acetaminophen) off, traced to Hepatocyte_AZ labels clipped at the 3.0 LLOQ (16.1%) and 150 cap (11.3%). An AFT interval-loss retrain vs a matched plain control is null on the 107-holdout (dMeta -0.042 mean over 4 seeds, engine mixed) and worsens exact-point scaffold-CV R2 (0.08 -> -0.16): the binding limit is structural discrimination, not the clipping. Production CLint model verified holdout-clean. Renal reabsorption gap noted as open.
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 66f027b130
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| m = fit_aft(X, y, a.sigma) | ||
| m.set_attr(aft="1") # predict_log10 converts AFT's original-scale output | ||
| m.save_model(a.save_model) |
There was a problem hiding this comment.
Make saved AFT models compatible with production inference
When --save-model is used to produce a replacement for models/adme/xgboost_clint.json, the saved AFT booster predicts CLint on the original scale, but src/sisyphus/predict/adme.py::_predict_clint always treats model output as log10(CLint) and applies 10**. The custom aft attribute is only checked by this experiment script, so deploying the generated model would yield exponentially inflated clearance values; either save an inference-compatible model or update the production loader to honor the objective/attribute.
Useful? React with 👍 / 👎.
| adme._predict_clint = _predict_clint | ||
| sys.argv = ["run_engine_benchmark.py", "--save-json", out] | ||
| runpy.run_path(str(ROOT / "scripts" / "run_engine_benchmark.py"), run_name="__main__") |
There was a problem hiding this comment.
Allow benchmarks to run without an output file
When a caller supplies --benchmark A or --benchmark B without --out, a.out is None and this constructs sys.argv with a non-string element; the downstream argparse parser raises a TypeError before running the benchmark. Since --out is currently optional and run_engine_benchmark.py itself supports printing results without saving, omit the --save-json pair when no path is provided or make --out conditionally required.
Useful? React with 👍 / 👎.
…dated The header count had drifted since DE-44; entries run DE-01..DE-58 with no gaps.
Summary
survival:aftinterval loss vs a matched plain-regression control (same data, exclusion, features, folds, hyperparameters).xgboost_clint.jsonis holdout-clean (median |log err| 0.076 on training rows vs 0.297 on the 24 holdout rows in Hep_AZ; a holdout-excluded retrain gives 0.327).No production model, source, or headline number changes. Local stack reproduces the canonical cache bit-identically (Meta 2.74279) before any arm was run.
Files
docs/research/dead-ends.md— DE-58docs/research/diagnosis.md— §1 bulletscripts/clint_censored_ab.py— CV A/B +--benchmarkholdout arm (runtime patch, no tracked-model change)data/validation/clint_censored_ab_2026-09-22.json— all numbers