Skip to content

feat(oauth): route Grok Build image, video and audio - #15305

Open
Kizuno18 wants to merge 2 commits into
diegosouzapw:release/v3.8.52from
Kizuno18:kz/grok-oauth-media
Open

Kizuno18 wants to merge 2 commits into
diegosouzapw:release/v3.8.52from
Kizuno18:kz/grok-oauth-media

Conversation

@Kizuno18

@Kizuno18 Kizuno18 commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Grok Build logins can already call these media endpoints, but OmniRoute only exposes their LLM routes. This adds image, video, TTS, and batch transcription through the existing grok-cli OAuth connection pool.

  • grok-cli/ and gc/ now resolve the four media models. Imagine uses the CLI proxy; TTS and STT use api.x.ai. Existing xai API-key routes and bare Imagine model resolution stay unchanged. This does not register the separate xai-oauth provider.
  • Media reuses the existing refresh, rotation persistence, and concurrent-refresh deduplication. API-key-only credentials are rejected rather than silently changing billing identity.
  • Video submits once and polls with the creating account's token. Failed submissions and polling failures keep their upstream status, errors are sanitized, and timeouts do not resubmit. Fractional timeout options are normalized before creating abort signals.
  • The existing multipart and image-result helpers were extracted with their original exports preserved, avoiding new import cycles and growth in frozen handlers. The provider reference was regenerated from source.

Validation

Added tests/unit/grok-oauth-media.test.ts and updated tests/unit/video-xai-grok-imagine.test.ts. The new suite covers all four dispatch families, payloads and hosts, UTF-8, codec validation, API-key rejection, refresh deduplication and persisted rotation in an isolated SQLite database, sticky video polling, timeout normalization, and error redaction. The timeout regressions were reproduced with failing tests before their fixes.

The final focused regression run passed 127/127 tests:

node --import tsx/esm --test \
  tests/unit/grok-oauth-media.test.ts \
  tests/unit/video-xai-grok-imagine.test.ts \
  tests/unit/video-registry-dispatch-parity.test.ts \
  tests/unit/audio-speech-handler.test.ts \
  tests/unit/audio-transcription-handler.test.ts \
  tests/unit/audio-transcription-opus-filename.test.ts \
  tests/unit/audio-alias-prefix-10586.test.ts \
  tests/unit/antigravity-image-credential-retry.test.ts \
  tests/unit/image-credential-retry-429.test.ts \
  tests/unit/specialty-model-catalog-routes.test.ts \
  tests/unit/grok-cli-proactive-refresh-7610.test.ts \
  tests/unit/grok-cli-oauth.test.ts

Also passed: direct tsc --noEmit -p open-sse/tsconfig.json, npm run typecheck:core, provider consistency and assets, base-relative file-size checks, error-helper checks, the standard cycle gate, and each docs gate. On Windows the env/docs gate needed Git Bash as the child shell. The wider handlers/utils cycle scan reports the same five existing components on both the base and this branch; neither new helper nor the Grok adapter adds a cycle. Changed-file ESLint and the normal commit hooks passed without disabling hooks or using a stash.

An isolated live bundle of these handlers returned HTTP 200 for all four operations using an existing Grok OAuth access token:

  • Image: a decoded 1024×1024 JPEG.
  • Video: one submission followed by polling, producing a decoded 480×480 H.264/AAC MP4, 2.04 seconds.
  • Speech: a decoded 24 kHz MP3, 3.24 seconds.
  • Transcription: nonempty Portuguese text with intact accents and eight word entries. This is a smoke test, not an accuracy benchmark.

The live harness deliberately stubbed token refresh, usage logging, feature flags, and log-payload normalization to avoid changing production state. It tests real auth headers, payload translation, HTTP calls, and polling with a fresh token, not a deployed gateway or end-to-end account selection. Refresh and persistence were exercised separately against the real refresh service with a mocked upstream and a disposable database. The probe required verified quota headroom and a zero on-demand cap; reported on-demand usage stayed zero afterward. No API-key fallback, browser session, deployment, or production token refresh was used.

Remaining checks

Full npm run lint stops on unused suppression entries in six untouched base files. Running it with --pass-on-unpruned-suppressions exits zero; no suppressions were pruned or widened. The provider translate-path golden fails locally only on six Copilot User-Agent values changing from linux to win32; the affected provider code, test, and snapshot are unchanged from the base, and the snapshot was not regenerated. The stock open-sse typecheck launcher also fails on Windows with spawnSync npx.cmd EINVAL; the compiler itself passed via the direct command above.

The base is release/v3.8.52 at 228d72db. Its active-release validation was still running, and its multi-arch manifest publishing check had already failed before this PR; no open base-red tracking issue was found. Full-repository unit/Vitest suites, aggregate coverage, and the production build were not run locally. Those results remain unverified until CI/review; the focused tests and live smokes are not a claim that the full matrix is green.

Media availability and billing remain upstream/account-dependent. This adds batch STT and TTS, not realtime voice, audio translation, or a guarantee of free usage.

@Kizuno18
Kizuno18 requested a review from diegosouzapw as a code owner October 1, 2026 21:29
Update the image error-log contract for the extracted result leaf
and verify credential redaction in persisted generation errors.
Split media payload validation and video error formatting to stay
within the complexity gate, with option-mapping regressions.

Add the feature fragment missing from the original PR. The expanded
regression run passes 142 tests. Scoped ESLint and complexity checks,
changelog integrity, and direct open-sse typechecking also pass.
@Kizuno18

Kizuno18 commented Oct 2, 2026

Copy link
Copy Markdown
Contributor Author

Pushed 57df992f to fix the failures this PR introduced in the first CI run.

  • The image-error source-contract test now follows the extraction into imageResult.ts, while still checking the shared stringifier and preserved exports. A behavioral regression also verifies that generation call logs redact credentials from a null-prototype error and preserve status/retryability.
  • Image/speech payload validation and video error formatting are split into small helpers. The three complexity violations (20, 18, 17 against a maximum of 15) now pass the same scoped complexity configuration without raising baselines or adding suppressions.
  • Added the missing feature changelog fragment and regressions for invalid image options, native image overrides, speech format/voice precedence, malformed speeds, and array-format defaults.

The expanded thirteen-file regression run passed 142/142 tests. Scoped ESLint, the complexity configuration, direct open-sse TypeScript checking, changelog integrity, and the normal pre-commit gates also passed. The full lint command still fails on obsolete suppressions in the six untouched files noted in the PR description; a diagnostic pass found no remaining code errors.

The old CI run also reports env/docs drift for GITHUB_OUTPUT and 25 missing mutation test registrations outside this PR. The fetched release branch already excludes GITHUB_OUTPUT from the env contract. Those failures are distinct from the PR-owned defects above; the new CI run still needs to establish the updated merge result. I have not merged or deployed this change.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant