You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
A maintainer has triaged this issue and applied the ready label
This issue has no assignee
No duplicate PR exists
PRs not meeting these requirements may be automatically closed.
Willingness to contribute
Yes. I would be willing to contribute this feature with guidance from the MLflow community.
Proposal Summary
I would like maintainer feedback on a narrowly scoped, deterministic trace-evidence scorer for memory-defense regression evaluation in MLflow GenAI Evaluation.
This is a personal-capacity, nonconfidential proposal. It does not represent an employer and does not claim or imply endorsement, adoption, support, or a commitment by MLflow, Databricks, OWASP, or OWASP Agent Memory Guard.
The first increment would be an optional, trace-aware scorer (for example, AgentMemoryGuardScorer) that consumes an MLflow Trace and scans only explicitly mapped SpanType.MEMORY evidence. It would not infer memory activity from a final model output or from arbitrary tool/LLM spans. This is deliberately not a claim of runtime protection, backend integrity, lifecycle conformance, production monitoring, or certification of a deployed memory system.
Motivation
What is the use case for this feature?
Evaluate agent-memory defenses against synthetic poisoned or sensitive memory payloads using deterministic evidence recorded in MLflow traces. For each explicitly mapped, evaluation-permitted memory payload, the scorer would run the deterministic text scanner exposed at agent_memory_guard.scan and emit separately named code-assessment feedback such as amg_memory_trace_scan and amg_memory_evidence_coverage.
Why is this use case valuable to support for MLflow users in general?
It would let users regression-test memory-defense behavior without treating a final model output or arbitrary tool/LLM span as memory evidence. The scorer would normalize results into stable, non-secret fields and would not place raw payloads, raw keys, arbitrary span metadata, timestamps, UUIDs/event IDs, or measured latency in feedback metadata or rationales.
Absence of evidence must never become a clean result. If a selected memory payload is missing, redacted, hashed, encrypted/unavailable to the scorer, malformed, or otherwise not permitted for evaluation, the scorer should return an explicit named not_evaluable evidence result (and use error feedback for invalid configuration or scorer failure). It must not silently scan a substitute field or report a pass because payload content is unavailable.
Why is this use case valuable to support for your project(s) or organization?
For the OWASP Agent Memory Guard project and potential adopters, this would provide a reproducible, offline MLflow evaluation surface for content-level memory-trace scanning while keeping the claim boundary explicit. This is a personal-capacity proposal and does not represent an employer or any organization.
Why is it currently difficult to achieve this use case?
MLflow custom code scorers provide the general extension surface, but there is no maintainer-approved memory-evidence mapping or AMG-specific feedback contract. Generic traces also do not provide enough information to reconstruct AMG's stateful write/read/integrity lifecycle. A lifecycle or conformance scorer would require a canonical, versioned memory trace schema; until that exists, content scanning and lifecycle conformance should remain separate outputs.
For each mapped, evaluation-permitted memory payload, the scorer would run agent_memory_guard.scan and emit separate code-assessment feedback:
amg_memory_trace_scan — normalized scanner findings for the observed payload; and
amg_memory_evidence_coverage — whether the mapped evidence was available and evaluable.
The scorer should be optional-dependency based, use MLflow code assessment feedback, and normalize results into stable, non-secret fields such as detector, severity, action, operation, opaque key identifier, count, AMG version, and trace-schema version. It should not place raw payloads, raw keys, arbitrary span metadata, timestamps, UUIDs/event IDs, or measured latency in feedback metadata or rationales.
The trace mapping should be opt-in and versioned. Missing, redacted, hashed, encrypted/unavailable, malformed, or evaluation-disallowed payloads must produce an explicit not_evaluable result rather than a false clean result.
Synthetic fixtures and determinism
Any contribution should use only synthetic, nonconfidential traces and fixtures:
a clean mapped memory payload;
a synthetic prompt-injection-like memory payload;
a synthetic sensitive-data-like payload using placeholders, not real secrets;
missing, redacted, and hashed/unavailable cases asserting not_evaluable;
a malicious-looking final output outside a mapped memory span, asserting no memory-scan result; and
repeat-run tests proving identical score, rationale, and stable metadata for the same trace.
Tests should also confirm that feedback never echoes the synthetic sensitive payload. An end-to-end mlflow.genai.evaluate() example could log the separate scan and evidence-coverage feedbacks for a trace-backed synthetic dataset.
I do not propose reconstructing AMG's stateful write/read/integrity lifecycle from generic traces in the first increment. A later lifecycle or conformance scorer would need a maintainer-approved canonical memory trace schema with ordered operation (write/read/delete), stable opaque key identifier, payload or explicit unevaluable-redaction state, source/provenance class, policy version, and baseline/snapshot semantics. Until such a contract exists, content scanning and lifecycle conformance should remain separate outputs.
Decision requested
Before any implementation work, could maintainers choose one of these paths (or suggest another supported path)?
First-party optional scorer: an MLflow-owned optional scorer module, potentially under mlflow/genai/scorers/agent_memory_guard/, with tests and documentation;
Maintained external package: a versioned external scorer package that MLflow maintainers explicitly choose to document and maintain as an integration; or
Wait for a canonical memory trace schema: defer both packaging and lifecycle work until the required memory-operation trace contract is agreed.
The content-level scorer could be considered independently of a future lifecycle schema, but I would wait for explicit maintainer agreement on placement, dependency policy, feedback semantics, and the opt-in trace mapping before coding or opening a pull request.
Warning
Before submitting a PR, please make sure that:
readylabelPRs not meeting these requirements may be automatically closed.
Willingness to contribute
Yes. I would be willing to contribute this feature with guidance from the MLflow community.
Proposal Summary
I would like maintainer feedback on a narrowly scoped, deterministic trace-evidence scorer for memory-defense regression evaluation in MLflow GenAI Evaluation.
This is a personal-capacity, nonconfidential proposal. It does not represent an employer and does not claim or imply endorsement, adoption, support, or a commitment by MLflow, Databricks, OWASP, or OWASP Agent Memory Guard.
The first increment would be an optional, trace-aware scorer (for example,
AgentMemoryGuardScorer) that consumes an MLflowTraceand scans only explicitly mappedSpanType.MEMORYevidence. It would not infer memory activity from a final model output or from arbitrary tool/LLM spans. This is deliberately not a claim of runtime protection, backend integrity, lifecycle conformance, production monitoring, or certification of a deployed memory system.Motivation
What is the use case for this feature?
Evaluate agent-memory defenses against synthetic poisoned or sensitive memory payloads using deterministic evidence recorded in MLflow traces. For each explicitly mapped, evaluation-permitted memory payload, the scorer would run the deterministic text scanner exposed at
agent_memory_guard.scanand emit separately named code-assessment feedback such asamg_memory_trace_scanandamg_memory_evidence_coverage.Why is this use case valuable to support for MLflow users in general?
It would let users regression-test memory-defense behavior without treating a final model output or arbitrary tool/LLM span as memory evidence. The scorer would normalize results into stable, non-secret fields and would not place raw payloads, raw keys, arbitrary span metadata, timestamps, UUIDs/event IDs, or measured latency in feedback metadata or rationales.
Absence of evidence must never become a clean result. If a selected memory payload is missing, redacted, hashed, encrypted/unavailable to the scorer, malformed, or otherwise not permitted for evaluation, the scorer should return an explicit named
not_evaluableevidence result (and use error feedback for invalid configuration or scorer failure). It must not silently scan a substitute field or report a pass because payload content is unavailable.Why is this use case valuable to support for your project(s) or organization?
For the OWASP Agent Memory Guard project and potential adopters, this would provide a reproducible, offline MLflow evaluation surface for content-level memory-trace scanning while keeping the claim boundary explicit. This is a personal-capacity proposal and does not represent an employer or any organization.
Why is it currently difficult to achieve this use case?
MLflow custom code scorers provide the general extension surface, but there is no maintainer-approved memory-evidence mapping or AMG-specific feedback contract. Generic traces also do not provide enough information to reconstruct AMG's stateful write/read/integrity lifecycle. A lifecycle or conformance scorer would require a canonical, versioned memory trace schema; until that exists, content scanning and lifecycle conformance should remain separate outputs.
Details
Proposed initial scope: content-level trace scanning
For each mapped, evaluation-permitted memory payload, the scorer would run
agent_memory_guard.scanand emit separate code-assessment feedback:amg_memory_trace_scan— normalized scanner findings for the observed payload; andamg_memory_evidence_coverage— whether the mapped evidence was available and evaluable.The scorer should be optional-dependency based, use MLflow code assessment feedback, and normalize results into stable, non-secret fields such as detector, severity, action, operation, opaque key identifier, count, AMG version, and trace-schema version. It should not place raw payloads, raw keys, arbitrary span metadata, timestamps, UUIDs/event IDs, or measured latency in feedback metadata or rationales.
The trace mapping should be opt-in and versioned. Missing, redacted, hashed, encrypted/unavailable, malformed, or evaluation-disallowed payloads must produce an explicit
not_evaluableresult rather than a false clean result.Synthetic fixtures and determinism
Any contribution should use only synthetic, nonconfidential traces and fixtures:
not_evaluable;Tests should also confirm that feedback never echoes the synthetic sensitive payload. An end-to-end
mlflow.genai.evaluate()example could log the separate scan and evidence-coverage feedbacks for a trace-backed synthetic dataset.Deliberately deferred lifecycle/conformance scoring
I do not propose reconstructing AMG's stateful write/read/integrity lifecycle from generic traces in the first increment. A later lifecycle or conformance scorer would need a maintainer-approved canonical memory trace schema with ordered operation (
write/read/delete), stable opaque key identifier, payload or explicit unevaluable-redaction state, source/provenance class, policy version, and baseline/snapshot semantics. Until such a contract exists, content scanning and lifecycle conformance should remain separate outputs.Decision requested
Before any implementation work, could maintainers choose one of these paths (or suggest another supported path)?
mlflow/genai/scorers/agent_memory_guard/, with tests and documentation;The content-level scorer could be considered independently of a future lifecycle schema, but I would wait for explicit maintainer agreement on placement, dependency policy, feedback semantics, and the opt-in trace mapping before coding or opening a pull request.
Public references
MEMORYspansWhat machine learning domain(s) is this feature request about?
domain/genai: LLMs, Agents, and other GenAI-related use casesdomain/classical-ml: Traditional machine learning, such as linear regression.domain/deep-learning: Deep learning and neural networks.domain/platform: MLflow platform foundation, not specific to a particular machine learning domain.What area(s) of MLflow is this feature request about?
area/tracking: Tracking Service, tracking client APIs, autologgingarea/model-registry: Model Registry service, APIs, and the fluent client calls for Model Registryarea/scoring: MLflow model serving, deployment tools, Spark UDFsarea/evaluation: MLflow model evaluation features, evaluation metrics, and evaluation workflowsarea/prompt: MLflow prompt engineering features, prompt templates, and prompt managementarea/tracing: MLflow Tracing features, tracing APIs, and LLM tracing functionalityarea/gateway: MLflow AI Gateway client APIs, server, and third-party integrationsarea/projects: MLproject format, project running backendsarea/uiux: Front-end, user experience, plottingarea/docs: MLflow documentation pages