Skip to content

Python: [Bug]: Automatic memory consolidation clears sessions still extracting memories #9218

Description

@ktz03

Observed Behavior

Automatic memory consolidation can clear a session's maintenance marker while that session is still extracting a new memory. The newly extracted fact is saved correctly, but it no longer counts toward consolidation_min_sessions even though it was absent from the completed consolidation's topic snapshot.

With consolidation_min_sessions=2 and a zero interval, the reproduction compares the same four sessions in two schedules:

Schedule Pending sessions after B Topics sent to consolidation after C
Serial ["B"] Old, then New and Old
Interleaved [] only Old

In the interleaved case, C leaves ["C"] pending and does not trigger consolidation. New still contains B new fact in both schedules. This concerns consolidation cadence, not lost memory facts.

Expected Behavior

Completing consolidation should retire only the work covered by that pass. A session whose extraction completes after the topic snapshot should remain eligible for the subsequent consolidation window.

Steps to Reproduce

  1. Share one MemoryContextProvider and real temporary MemoryFileStore between distinct sessions belonging to the same owner. Set the minimum to two sessions and the interval to zero.
  2. Seed an Old topic and finish the first session. Start a second Agent.run; its automatic consolidation snapshots Old and awaits the client.
  3. Start B's Agent.run. Its after_run registers B, then awaits extraction of a New topic.
  4. Let the Old consolidation finish, then allow B's extraction to finish.
  5. Read the persisted state and run C. Compare with the serial control.

Minimal Reproduction

Run this against the current Python core source. Both clients are deterministic and offline; no model or service credentials are required.

import asyncio
import json
import tempfile
from datetime import timedelta

from agent_framework import Agent, AgentSession, ChatResponse, MemoryContextProvider, MemoryFileStore, Message, SessionContext


def reply(payload):
    return ChatResponse(messages=[Message(role="assistant", contents=[json.dumps(payload)])])


class Client:
    def __init__(self, interleaved):
        self.additional_properties = {}
        self.interleaved = interleaved
        self.consolidating = asyncio.Event()
        self.extracting_b = asyncio.Event()
        self.finish_consolidation = asyncio.Event()
        self.finish_extraction = asyncio.Event()
        self.consolidated_topics = []

    async def get_response(self, messages, **kwargs):
        system = messages[0].text if messages[0].role == "system" else ""
        if "consolidate one topic memory file" in system.lower():
            payload = json.loads(messages[-1].text)
            self.consolidated_topics.append(payload["topic"])
            self.consolidating.set()
            if self.interleaved:
                await self.finish_consolidation.wait()
            return reply({"summary": payload["summary"], "memories": payload["memories"]})
        if "extract durable memory candidates" in system.lower():
            if "B marker" not in messages[-1].text:
                return reply({"memories": []})
            self.extracting_b.set()
            if self.interleaved:
                await self.finish_extraction.wait()
            return reply({"memories": [{"topic": "New", "memory": "B new fact"}]})
        return reply("Reply")


async def reproduce(interleaved):
    with tempfile.TemporaryDirectory() as tmp:
        store = MemoryFileStore(tmp, owner_state_key="owner")
        client = Client(interleaved)
        provider = MemoryContextProvider(store=store, recent_turns=0, max_extractions=1,
            consolidation_min_sessions=2, consolidation_interval=timedelta(0), consolidation_client=client)
        agent = Agent(client=client, context_providers=[provider], default_options={"store": False})
        sessions = [AgentSession(session_id=sid) for sid in ["Seed", "A", "B", "C"]]
        for session in sessions:
            session.state["owner"] = "same-owner"
        seed, a, b, c = sessions
        context = SessionContext(session_id=seed.session_id, input_messages=[])
        await provider.before_run(agent=agent, session=seed, context=context, state={})
        write = next(tool for tool in context.tools if tool.name == "write_memory")
        await write.invoke(arguments={"topic": "Old", "memory": "Old fact"}, skip_parsing=True)
        await agent.run("ordinary turn", session=seed)
        if interleaved:
            first = asyncio.create_task(agent.run("ordinary turn", session=a))
            await asyncio.wait_for(client.consolidating.wait(), 5)
            second = asyncio.create_task(agent.run("B marker", session=b))
            await asyncio.wait_for(client.extracting_b.wait(), 5)
            client.finish_consolidation.set()
            await asyncio.wait_for(first, 5)
            client.finish_extraction.set()
            await asyncio.wait_for(second, 5)
        else:
            await agent.run("ordinary turn", session=a)
            await agent.run("B marker", session=b)
        state_after_b = store.read_state(b, source_id="memory")
        fact = store.get_topic(b, source_id="memory", topic="New").memories
        await agent.run("ordinary turn", session=c)
        return {"interleaved": interleaved, "after_b": state_after_b,
            "after_c": store.read_state(c, source_id="memory"),
            "new_fact": fact, "consolidated_topics": client.consolidated_topics}


async def main():
    print(json.dumps([await reproduce(False), await reproduce(True)], indent=2))


asyncio.run(main())

Error Messages and Stack Traces

No product exception is raised. The harness emits its existing ExperimentalWarning. Both Agent.run calls complete, and the newly extracted fact can be loaded from the real Markdown store.

Package Versions

Executed core package source is byte-identical to main 91ab44faa4824a30247d00c341f3498ba46b1ca3 (core source advertised as 1.21.0). The cached environment's installed agent-framework-core metadata is 1.20.0; the source was explicitly selected through PYTHONPATH. This was not a fresh 1.21.0 lock installation.

Python Version

Python 3.14.3

Operating System

Windows

Regression

Unknown

Additional Context

The state-lock release between the topic snapshot and consolidation completion was introduced in #5613 to allow concurrent before_run/after_run calls during model requests. See the original review request and the author's response. This report preserves that concurrency goal.

Six local controls cover serial/interleaved success, transient failure, and different owners. Failure and owner-isolation controls retain B's marker; all preserve the new fact. An additional independent public Agent.run check with two delayed extraction sessions reproduces the same marker loss. These overlap in scope and are not additional product defects. No live model, remote store, cross-process or performance behavior is claimed.

A possible direction is to advance the maintenance window using coverage captured by a pass while preserving work that arrived or completed later. The treatment of repeated runs with the same session ID, overlapping consolidation passes, partial failures, and existing persisted state should be agreed before choosing an implementation. Merely holding the state lock through model calls would undo the original concurrency goal. No implementation is included; I can take this on after maintainer agreement on the direction.

AI Assistance

AI-assisted analysis, reproduction, and writing.

Acknowledgements

  • I searched existing issues and did not find a duplicate.
  • I personally verified this behavior and the reproduction details are authentic.
  • I will wait for explicit maintainer agreement before starting implementation of a non-trivial change.

Activity

  1. added
    pythonUsage: [Issues, PRs], Target: Python
    triageUsage: [Issues], Target: All issues that still need to be triaged
    on Oct 8, 2026
  2. added
    reproducedUsage: [Issues], Target: all issues that can be reproduced by the triage workflow
    on Oct 8, 2026
  3. github-actions commented on Oct 8, 2026

    @github-actions
    Contributor

    🤖 Automated triage reproduction notes (agent-authored — trust but verify)

    Agent analysis

    The bug reproduces in python/packages/core/agent_framework/_harness/_memory.py::MemoryContextProvider._run_consolidation, particularly the unconditional marker reset at line 1593. It occurs when same-owner runs overlap so session B registers while session A's successful consolidation is awaiting its client. Minimal repro: use a shared MemoryFileStore, consolidation_min_sessions=2, zero interval, and event-gated consolidation/extraction calls; B's fact persists but its pending marker becomes empty.

    • Failing test: python/packages/core/tests/core/test_harness_memory.py::test_automatic_consolidation_preserves_session_registered_during_pass
    • Files examined: python/packages/core/agent_framework/_harness/_memory.py, python/packages/core/tests/core/test_harness_memory.py, python/packages/core/pyproject.toml, python/pyproject.toml
    • Tests run: test_memory_consolidation_transient_failure_preserves_state, test_automatic_consolidation_preserves_session_registered_during_pass
    • Reported version: 1.21.0
    • Current version: 1.21.0
  4. added
    harness[Issues, PRs], Target: harness-level items
    and removed
    triageUsage: [Issues], Target: All issues that still need to be triaged
    on Oct 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

harness[Issues, PRs], Target: harness-level itemspythonUsage: [Issues, PRs], Target: PythonreproducedUsage: [Issues], Target: all issues that can be reproduced by the triage workflow

Type

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions