Skip to content

Add Turkish document retrieval evaluation - #3144

Draft
Cemilcanoz wants to merge 2 commits into
openai:mainfrom
Cemilcanoz:add-turkish-retrieval-evaluation
Draft

Cemilcanoz wants to merge 2 commits into
openai:mainfrom
Cemilcanoz:add-turkish-retrieval-evaluation

Conversation

@Cemilcanoz

@Cemilcanoz Cemilcanoz commented Sep 30, 2026 •

Copy link
Copy Markdown

Summary

Add a runnable Turkish retrieval evaluation that extends the existing semantic text search example with labeled queries and per-category metrics. It compares Unicode lowercasing, Turkish casing, ASCII folding, and an optional OpenAI embeddings path using Recall@3 and MRR@3.

The fictional fixture contains 40 documents, 88 answerable queries, and 20 separately stored unanswerable queries. Intent groups keep nearby query variants in the same split. Evaluation defaults to 56 development queries; the 32 held-out queries have not been benchmarked.

An offline Markdown failure report lists development queries with missing relevant documents, including partial misses. It consumes saved rankings and bundled fixtures without API calls and rejects held-out input.

Motivation

Turkish I/ı and İ/i casing, ASCII spelling, paraphrases, and inflection can affect retrieval. This example makes those effects measurable instead of demonstrating only a single semantic-search query. These synthetic development results are illustrative, not general Turkish retrieval-quality claims.

Validation

  • 15 offline unittest checks passed, including grouped-split validation, multi-document metrics, mocked embedding request/response ordering, and full/partial/empty retrieval failure reports with invalid-input checks.
  • Ruff and git diff --check passed.
  • New registry entry validated against the repository JSON Schema; path, unique slug, and author checked.
  • Development evaluation reproduced Recall@3 / MRR@3: Unicode 0.866 / 0.833; Turkish and ASCII 0.920 / 0.887.
  • README reviewed with the repository docs-editor skill; no notebooks changed.

Remaining before ready for review

This PR is a draft. The live embedding path has not been executed because a local API key was unavailable; it has mocked coverage only. Labels were authored and reviewed by agents, not independently adjudicated by humans. Unanswerable queries are included for future evaluation, but the current evaluator does not score abstention or false positives.

For new content

  • Added a registry.yaml entry. Author attribution uses the GitHub profile; no custom authors.yaml entry is needed.
  • Relevance: OpenAI embeddings comparison is included as an opt-in path.
  • Uniqueness: Searched existing PRs and related examples; this adds Turkish categories, relevance labels, and grouped evaluation splits to the existing semantic search example.
  • Spelling and Grammar: README reviewed.
  • Clarity: Setup, commands, methods, metrics, and limitations documented.
  • Correctness: Offline code is verified; live embedding execution remains pending.
  • Completeness: References, dependencies, environment variable, API cost boundary, and current limitations documented.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant