EC-Bench

Does what a coding agent worked out in earlier sessions change what it does in later ones? EC-Bench is our attempt to measure that.

Results so far

Why a new benchmark

Most coding benchmarks evaluate what an agent can accomplish inside a single episode: write the code, reason through the problem, plan the change. When the episode ends, so does the evaluation, and the next one starts from the same place as the last.

That leaves a question unmeasured: does the work an agent has already done change what it does later? Answering it takes sequences of sessions, not single ones.

How a run works

Each run uses a FastAPI repository and 30 prompts, in three sequential sessions: 14 on investigation and architecture, 13 on implementation and debugging, and 3 on planning. Each arm gets its own copy of the repository and its own isolated memory. In the Reverie arm, every prompt starts with a fresh context, so anything the agent knows from earlier work has to come from memory. Between sessions, the harness accepts every review automatically.

One LLM judge reads the full transcripts and scores them on five weighted metrics. The judge is named in the ledger below.

  • architectural continuity0.30
  • cognition reuse0.30
  • repository groundedness0.15
  • engineering quality0.15
  • debugging/investigation efficiency0.10
Identical copy of the repository (FastAPI) for each arm. Two lanes, Arm A: with Reverie and Arm B: baseline, each run three sequential sessions: investigation and architecture (14 prompts), implementation and debugging (13) and planning (3). In Arm A there is a review gate between sessions, set to auto-accept in the harness, and each prompt starts with a fresh context. Both lanes’ transcripts go to the same LLM judge, which scores them on five weighted metrics. Held constant: repository, prompts, agent harness, model and judge. What differs: access to Reverie. The figure shows no results.Identical copy of the repository (FastAPI) for each armArm A: with ReverieSession 1investigation and architecture14 promptsfresh context per promptSession 2implementation and debugging13 promptsfresh context per promptSession 3planning3 promptsfresh context per promptreview: auto-accept(harness setting)review: auto-accept(harness setting)Arm B: baselineSession 1investigation and architecture14 promptsSession 2implementation and debugging13 promptsSession 3planning3 promptstranscriptsLLM judgesame for both arms · full transcriptfive weighted metrics0.30architecturalcontinuity0.30cognitionreuse0.15repositorygroundedness0.15engineeringquality0.10debugging /investigationefficiencyheld constant: repository, prompts, agent harness, model, judgediffers: access to Reverie

Identical copy of the repository (FastAPI) for each arm

  1. Session 1investigation and architecture14 prompts
    Arm Awith Reveriefresh context per prompt
    Arm Bbaseline
  2. Arm A only: review: auto-accept (harness setting)

    Session 2implementation and debugging13 prompts
    Arm Awith Reveriefresh context per prompt
    Arm Bbaseline
  3. Arm A only: review: auto-accept (harness setting)

    Session 3planning3 prompts
    Arm Awith Reveriefresh context per prompt
    Arm Bbaseline

Both arms’ transcripts go to one judge.

LLM judge (same for both arms, full transcript)

Five weighted metrics:

  • architectural continuity0.30
  • cognition reuse0.30
  • repository groundedness0.15
  • engineering quality0.15
  • debugging / investigation efficiency0.10

held constant: repository, prompts, agent harness, model, judge

differs: access to Reverie

Fig. 6How an EC-Bench run works. Two arms see the same repository and the same prompts, in the same three sessions; only access to Reverie differs. The figure shows the design, not a result.

Where a run departed from this design, the departure is noted in its row of the ledger.

Results so far

So far, no run shows a clear advantage for Reverie. The first run’s aggregate favoured memory, but the gain came from one metric that loading memory inflates; on the other four, the agent did worse. The second architecture narrowed the gap under a protocol that favoured the baseline. A controlled comparison is next.

Our own write-ups disagree on some exact figures. Until the sources agree, each row below says what happened and why we aren’t printing numbers, and a run’s figures appear only once it has been reconciled.

  1. July 2026

    Needs reconciliation
    System
    EMS v1 (memory injected before work)
    Compared with
    Stateless baseline
    Protocol
    Relevant memory loaded into the agent’s context before it began work, compared with no memory.
    Prompts
    30
    Judge
    TODO(fact)

    The aggregate favoured memory, but the gain came from one metric, cognition reuse, which loading memory inflates. On the other four metrics the agent did worse with memory.

    Caveats

    • A later write-up of this version reports the baseline ahead in aggregate. The two accounts have not been reconciled.

    Figures under reconciliation.

    Sources EC-Bench Technical Report 001 · EC-Bench Short Report

  2. 13 August 2026

    Needs reconciliation20260813-141550
    System
    Reverie v2 (on request)
    Compared with
    Interactive chat memory (baseline arm added 15 August)
    Protocol
    Headless, fresh context per prompt, compared with an interactive agent with chat memory. A mixed protocol.
    Prompts
    30
    Judge
    GLM-5.2 (LLM), full transcript

    The baseline scored ahead on all five metrics.

    Caveats

    • The two arms ran under different protocols: headless with fresh context per prompt, against interactive with chat memory.
    • The paper reports different figures for this architecture, with outliers removed, and a different count of stored conclusions. These have not been reconciled.
    • The harness accepted every review automatically, so human review was not tested.

    Figures under reconciliation.

    Sources The paper, section on results · Run 20260813-141550 (product documentation, section on the benchmark)

  3. Not yet run

    Planned
    System
    Reverie v2 (on request)
    Compared with
    Same protocol in both arms
    Protocol
    A controlled rerun with the same protocol in both arms.
    Prompts
    TODO(fact)
    Judge
    TODO(fact)

    Not yet run.

    Sources Research index: how we test it

What we learned

  • Retrieval timing matters as much as retrieval quality.

    Our first version loaded memory before the agent began work, and on this benchmark the agent trusted memory over the code and did worse.

  • Accumulated knowledge expands scope. TODO(copy)
  • Tasks need different modes.

    That is why Reverie infers a task mode when your agent asks. How retrieval works

  • Where memory helped and where it hurt.

    In the first run, memory helped on cognition reuse, the metric that loading memory inflates, and hurt on the other four.

What these runs did not test

  • Human review. The harness accepts every review automatically.
  • Contradiction and supersession. The stored run produced no contradicts and no supersedes relationships.
  • More than one repository, or more than one agent.

Threats to validity

  • A small sample.
  • One repository.
  • One type of agent.
  • An LLM judge, not a human one.
  • Outlier removal, in the paper’s figures for the second architecture.
  • A mixed protocol in the stored run.

What’s next

  • A controlled rerun with the same protocol in both arms.
  • The continuity experiment: the same agent, with and without Reverie.
  • More repositories.
  • A human judge alongside the LLM.

What this changed in Reverie

The first run’s finding changed the design. Reverie no longer loads memory when a session starts. Your agent asks for it, after reading the code. How your agent gets it back