---
title: "D-Mem: A Dual-Process Memory System for LLM Agents"
authors: ["Zhixing You", "Einstein Institute of Mathematics", "The Hebrew University of Jerusalem, Israel", "Jiachen Yuan", "Independent Researcher", "Corresponding author", "Jason Cai", "AWS AI"]
url: "https://arxiv.org/abs/2603.18631"
sections: 55
estimated_tokens: "18.4k"
---

## Contents
- 1 Introduction
- 2 Related Work
  - 2.1 Retrieval-Augmented Generation (RAG)
  - 2.2 Agentic Memory for Autonomous Agents
- 3 Methodology
  - 3.1 Mem0 ∗ : The System 1 Retrieval Foundation
    - Robust Query Resolution.
  - 3.2 Gated Deliberation Policies
    - Policy 1: Majority Voting.
    - Policy 2: Consensus.
    - Policy 3: Quality Gating (ours).
  - 3.3 Full Deliberation
    - Stage 1: Chunked Fact Extraction.
    - Stage 2: Multi-stage Filtering.
    - Stage 3: Answer Generation.
- 4 Experiments
  - 4.1 Setup
    - Datasets.
    - Metrics.
    - LLMs.
    - Baselines.
  - 4.2 Overall Performance
    - Comparison with Baselines.
    - Effectiveness of Dual-Process Routing.
    - Generalization and Robustness.
  - 4.3 In-depth Analysis of Quality Gating
    - Per-Category Analysis
    - Impact of Filtering and Model-Specific Variations.
  - 4.4 Fallback Mechanism Analysis
    - Adaptability to Dataset Complexity
    - Selection Bias in Fallback Routing.
    - Model-Specific Sensitivity in Quality Gating.
  - 4.5 Case Study: Why Static Retrieval Fails
- 5 Conclusion
- 6 Limitations
- References
- Appendix A Prompt Templates
  - A.1 Answer Generation Prompt (Mem0)
  - A.2 Filter Prompt (Ours)
  - A.3 Fact Extraction Prompt (Ours)
  - A.4 Fact Filtering System Prompt (Ours)
  - A.5 Consensus Prompt (Ours)
  - A.6 Majority Voting Prompt (Ours)
  - A.7 Quality Gating Prompt (Ours)
  - A.8 LLM-as-a-Judge Prompt (Mem0)
- Appendix B Implementation Details
  - B.1 Hyperparameter Settings
  - B.2 System Configuration
- Appendix C Comprehensive Evaluation Results
- Appendix D Computational Resources and Reproducibility
- Appendix E Ethics Statement
  - Privacy and Data Security.
  - Environmental and Computational Impact.
  - Memory-Induced Bias and Safety.
- Appendix F Artifact Licenses and Terms of Use

## Abstract

Abstract Driven by the development of persistent, self-adapting autonomous agents,
equipping these systems with high-fidelity memory access for long-horizon reasoning has emerged as a critical requirement.
However, prevalent retrieval-based memory frameworks
often follow an incremental processing paradigm that
continuously extracts and
updates conversational memories into vector databases,
relying on semantic retrieval when queried.
While this approach is fast,
it inherently relies on lossy abstraction,
frequently missing contextually critical information
and struggling to resolve queries that rely on fine-grained contextual understanding.
To address this, we introduce D-Mem, a dual-process
memory system. It retains lightweight vector retrieval
for routine queries while establishing an exhaustive Full
Deliberation module as a high-fidelity fallback.
To achieve cognitive economy without sacrificing
accuracy, D-Mem employs a Multi-dimensional Quality Gating policy to
dynamically bridge these two processes.
Experiments on the LoCoMo and RealTalk benchmarks using
GPT-4o-mini and Qwen3-235B-Instruct demonstrate the
efficacy of our approach. Notably, our Multi-dimensional Quality Gating
policy achieves an F1 score of 53.5 on LoCoMo with GPT-4o-mini.
This outperforms our static
retrieval baseline, Mem0 ∗ (51.2), and recovers 96.7% of the
Full Deliberation’s performance (55.3), while incurring
significantly lower computational costs. 1 1 1 Code will be released upon acceptance of the paper.

## 1 Introduction

The development of LLM-based autonomous agents marks a significant evolution, shifting the application focus from stateless text generators to persistent, self-adapting entities deployed in dynamic, long-horizon environments
Xi et al. (2025); Wang et al. (2024, 2026). As agents operate
in increasingly complex scenarios, the capability to
accumulate experiences and self-evolve across extended interactions
becomes a critical necessity. However, this continuous adaptation is fundamentally
hindered by the static nature of deployed LLM parameters.
Despite the advent of ever-expanding context windows,
simply feeding massive dialogue histories into the model is computationally expensive and frequently exacerbates the "Lost-in-the-Middle" phenomenon
Liu et al. (2024). To address this,
recent frameworks like MemoryBank, Zep and Mem0
developed an incremental memory processing paradigm Zhong et al. (2024); Rasmussen et al. (2025); Chhikara et al. (2025).
These systems
continuously extract, compress, and update
conversational memories into vector databases to
enable efficient, scalable semantic retrieval.

While highly effective for routine, explicit
queries, this paradigm introduces a critical
vulnerability for deep reasoning: lossy abstraction Zhu et al. (2026).
By aggressively performing query-agnostic
compression—condensing continuous dialogues into
semantic snippets independent of future
queries—the system strips away potentially crucial
contextual nuances. Consequently, when faced
with queries
requiring rigorous deduction, it frequently misses
contextually critical information—such as unstated
temporal logic (e.g., relative time calculations) or
multi-hop dependencies. This leaves agents struggling with
a severe performance bottleneck: static retrieval
simply cannot reconstruct the precise logical
chains lost during compression. To address the fundamental
limitations of query-agnostic memory compression, we
introduce D-Mem, a dual-process memory system that
emulates the human cognitive process of metacognitive
monitoring Kahneman (2011); Evans and Stanovich (2013).
Rather than abandoning efficient vector
retrieval, D-Mem leverages it as a rapid System 1 baseline (Mem0^∗), while
introducing a robust System 2 fallback mechanism. D-Mem operates
on two distinct cognitive levels:

- •
System 1 (Mem0^∗):
Acting as an enhanced incremental memory processing module built upon the Mem0 architecture, this system
performs a
rapid, low-cost static retrieval (Top-K) based on
surface similarity within the compressed memory space.
- •
System 2 (Full Deliberation): When
System 1 fails to resolve the query, the system
escalates to an exhaustive deliberate reading mode. Bypassing the
compressed memory, this module directly processes the raw dialogue
history. It executes a query-guided temporal scan,
utilizing the original question as a discriminative
anchor to systematically extract query-relevant facts
chunk-by-chunk via a quantitative scoring mechanism (0-10).
By applying strict
multi-stage filtering, it synthesizes a highly grounded, high-fidelity
answer-closely mimicking the human cognitive
process of purposeful reading.

Crucially, executing
System 2 indiscriminately incurs massive computational overhead. On the LoCoMo dataset,
it increases input tokens and inference time by
over 10$\times$ compared to the fast path (see Table [1](#S4.T1) for details). To achieve
cognitive economy, D-Mem incorporates a metacognitive
assessment step. Acting as a rigorous analytical
gatekeeper, this module evaluates the initial System 1
output against a strict multi-dimensional pass/fail
rubric—encompassing relevance, faithfulness & consistency,
and completeness. The architecture escalates to the
computationally intensive System 2 exclusively when
the initial output fails to satisfy any of these criteria,
ensuring that Full Deliberation is
triggered only by critical information deficits.

In summary,
our main contributions are as follows:

- •
The dual-process D-Mem Framework.
To overcome the fundamental limitations of
query-agnostic memory compression, we propose D-Mem,
a novel architecture that integrates efficient
vector retrieval (System 1) with an exhaustive
deliberate reading mode (System 2). To
bridge these processes and balance cognitive economy
with reasoning accuracy, we introduce a
Multi-dimensional Quality Gating policy (hereafter, Quality Gating).
Serving as a metacognitive checkpoint, it offers an
accuracy-efficiency trade-off.
- •
Full Deliberation as a robust baseline. We
establish a Full Deliberation method that exhaustively
processes raw dialogue history chunk-by-chunk. Because
this extraction mechanism is strictly query-guided,
it ensures that explicit, nuanced details are meticulously
preserved. Furthermore, its chunk-by-chunk processing
effectively mitigates the "Lost-in-the-Middle" phenomenon.
By establishing this high-fidelity upper bound,
we isolate a critical bottleneck shift: to surpass
this baseline, future long memory architectures must
evolve beyond extracting explicit facts to actively
capturing and synthesizing implicit information across
longitudinal history.
- •
High Performance with Computational Efficiency. Through
comprehensive evaluations on the LoCoMo Maharana et al. (2024) and RealTalk Lee et al. (2025)
benchmarks using GPT-4o-mini and Qwen3-235B-Instruct,
we demonstrate the efficacy of our framework. Notably,
our Quality Gating achieves an F1 score of 53.5
on LoCoMo and exhibits consistent substantial gains on RealTalk.
This outperforms our Mem0^∗ (51.2) and
recovers 96.7% of the
Full Deliberation’s performance (55.3)
with significantly fewer tokens and inference time.

## 2 Related Work

### 2.1 Retrieval-Augmented Generation (RAG)

Standard RAG retrieves documents via dense similarity
and feeds them to a generator. Subsequent work has
improved this pipeline through query
expansion (Zhang et al., 2024),
re-ranking (Yu et al., 2024), relevance
filtering (Rossi et al., 2024), and
query-transformation (Zheng et al., 2024).
These techniques refine how to
retrieve.
Recent Agentic RAG paradigm
actively assess their own information needs,
refines whether retrieval is sufficient
(Asai et al., 2024; Xu et al., 2026).

### 2.2 Agentic Memory for Autonomous Agents

Early frameworks like MemGPT (Packer et al., 2024)
introduced operating system-inspired virtual context
management via hierarchical paging.
Other memory systems
like Mem0 (Chhikara et al., 2025) and MemoryBank
(Zhong et al., 2024) employ an incremental
processing paradigm that continuously extracts and
updates conversational memories into vector databases.
Such aggressive vectorization inherently fragments semantic context.
To preserve relational integrity, recent works have explored structural
optimizations. For instance,
G-Memory (Zhang et al., 2025) and Mem0^g(Chhikara et al., 2025)
construct graph-based hierarchies,
and A-Mem (Xu et al., 2025) utilizes interconnected notes.
Despite these structural enhancements, these methods remain
fundamentally query-agnostic. Because the memory representation
is aggressively compressed and pre-fixed during the storage phase,
they inevitably suffer from lossy abstraction—stripping
away nuanced temporal and causal logic.

To circumvent the pitfalls of lossy abstraction,
recent advancements like GAM (Yan et al., 2025)
and E-mem (Wang et al., 2026) entirely abandon query-agnostic compression.
Instead, they pivot towards episodic context reconstruction and multi-agent
deliberative paradigms, prioritizing reasoning fidelity by exhaustively
processing raw, uncompressed historical context.
However, discarding lightweight retrieval to apply such heavy
deliberation indiscriminately incurs a massive, inflexible computational
overhead. This approach ignores the "cognitive economy" of System 1,
where the majority of routine queries can be efficiently resolved
via rapid semantic recall, obviating the need for deep episodic reconstruction.

## 3 Methodology

Figure: Figure 1: Architectural overview of D-Mem. Components of Part A illustrates the memory mechanism of Mem0^∗. The elements highlighted in soft pale pink denote our key modifications over the original Mem0 framework: (1) utilizing the top-10 similar memories rather than a general summary during the Extraction Phase, and (2) incorporating an additional relevance filtering step prior to memory updates. Part B illustrates the Quality Gating mechanism, which evaluates the initial answer from Mem0^∗ and dynamically determines whether to trigger the Full Deliberation fallback.
Refer to caption: 2603.18631v1/x1.png

To circumvent the pitfalls of lossy abstraction
inherent in prior designs while maintaining
computational efficiency, we introduce D-Mem,
a dual-process memory architecture. The framework
is operationalized through three coupled components:
(1) Mem0^∗ serves as the foundational
lightweight retrieval module. It improves upon
the standard Mem0 paradigm by efficiently
extracting and updating salient conversational
memories within a vector database, designed to
rapidly resolve the majority of routine queries.
(2) The Quality Gating acts as a
dynamic evaluative router. It evaluates the initially retrieved context from
Mem0^∗ against a multi-dimensional quality
rubric.
Specifically, the gate assesses the retrieved
memory along three orthogonal dimensions:
Relevance, Faithfulness & Consistency, and Completeness.
If the retrieved context falls short on any of these axes,
the gate triggers the fallback mechanism.
(3) The Full
Deliberation mechanism provides a high-fidelity
fallback. Triggered exclusively when the quality
gate deems the lightweight retrieval inadequate,
this module bypasses compressed representations
to exhaustively process the raw, uncompressed
historical context. By structurally decoupling
routine semantic recall from
resource-intensive exhaustive full deliberation,
D-Mem enables high-fidelity reasoning for
complex queries while preserving strict
cognitive economy.

### 3.1 Mem0 ∗ : The System 1 Retrieval Foundation

The Mem0^∗ module follows an incremental processing paradigm
and serves as our foundational System 1, enabling fast associative retrieval.
As illustrated in Figure [1](#S3.F1) Part A,
this module comprises two phases: extraction phase and update phase.

Formally, to reflect the dynamic nature of
incremental processing, we model the
conversation as a continuous stream. Let
$\mathcal{H}_{t-1}=(m_{1},m_{2},\dots,m_{2t-2})$
denote the accumulated dialogue history prior to
the $t$-th interaction round. Each message $m_{i}$
represents an individual utterance augmented with
timestamp and speaker metadata. The maintenance
pipeline dynamically triggers upon the ingestion
of the new interaction round at step $t$, defined
as the message pair $(m_{2t-1},m_{2t})$.

- •
Extraction Phase: Mem0^∗
conditions the LLM on two complementary
sources: (1) the top-$10$ most semantically similar
existing memories $\mathcal{F}$ retrieved from the vector database,
and (2) recent messages $(m_{2t-10},\dots,m_{2t-2})$
from $\mathcal{H}_{t-1}$. Here $\mathcal{F}$ serves as a
dynamic background context and the recent
message sequence offers granular temporal
context. This dual contextual information,
combined with the new message pair, forms
a comprehensive input for the LLM to extract
salient memories $\Omega=\{\omega_{1},\dots,\omega_{n}\}$.
- •
Update Phase: Once the set of
candidate memories $\Omega$ is extracted, the system
retrieves the top-5 semantically similar historical
memories for each $\omega_{i}\in\Omega$ via vector
embeddings. To minimize cognitive load and mitigate
noise, low-relevance items are explicitly discarded
using a strict cosine similarity threshold ($>0.8$).
The retained historical memories constitute the
contextual background $\mathcal{B}$.
Subsequently, an LLM cross-references the newly
extracted memories $\Omega$ against the contextual
background $\mathcal{B}$ to evaluate potential
contradictions and redundancies.
This dictates precise memory operations
for memories in $\mathcal{B}$ and $\Omega$:
ADD, UPDATE, DELETE, or
NOOP.

#### Robust Query Resolution.

During inference, given a user query $q$,
Mem0^∗ retrieves the top-30 most similar memories $C$
and generates an initial answer $A_{init}=\text{LLM}(q,C)$.

To further mitigate hallucination,
we also employ a rigorous filtering pipeline to
derive a highly refined context $C^{\prime}$:

To filter out noise memories in $C$, an LLM
evaluates the initial context $C$ against the query $q$.
It retains only strictly necessary memories $C_{filter}$ to answer
$q$. In the edge case where aggressive filtering
eliminates all candidates (i.e., $C_{filter}=\emptyset$), the system conservatively falls back
to the original unfiltered context $C$ to
preserve recall.

### 3.2 Gated Deliberation Policies

To operationalize the routing between the Mem0^∗
and the Full Deliberation fallback (as mentioned in Subsection [3.3](#S3.SS3)),
we formalize a Gating Policy $\mathcal{G}$.

Formally, given the initial answer $A_{init}$,
the user query $q$, and the retrieved context
$C$, we define the gating function $\mathcal{G}$:

$$ $\mathcal{G}:(A_{init},q,C)\rightarrow\{0,1\}.$ $$

Here, the system outputs $A_{init}$ if $\mathcal{G}(\dots)=0$,
and it executes the
full deliberation if $\mathcal{G}(A_{init},q,C)=1$.

To comprehensively evaluate this dual-process
routing, we investigate three distinct
gated deliberation policies:

#### Policy 1: Majority Voting.

The system generates three candidate answers from
the same Top-30 context (with temperature $>0$
for diversity). A judge determines whether a
majority consensus exists—i.e., at
least two of the three answers are semantically
equivalent. If a majority exists, one answer is
selected from the majority group (e.g., by
preferring clarity and verbatim extraction);
otherwise $\mathcal{G}(A_{init},q,C)=1$ and
the system falls back to Full Deliberation.

#### Policy 2: Consensus.

As in Majority Voting, three answers are generated
from Top-30 retrieval. The trigger requires
full semantic agreement: all three
answers must be equivalent. If any answer differs,
$\mathcal{G}(A_{init},q,C)=1$ and the system
invokes Full Deliberation.
This policy is stricter than Majority Voting and
triggers fallback more often.

#### Policy 3: Quality Gating (ours).

An LLM checks the quality of the initial answer $A_{init}$
against a quality rubric with three
dimensions: Relevance, Faithfulness & Consistency,
and Completeness.
If the answer successfully passes all three dimensions,
then the system returns $A_{init}$; otherwise,
$\mathcal{G}(A_{init},q,C)=1$
and the system falls back to Full Deliberation.

### 3.3 Full Deliberation

The Full Deliberation mechanism serves
both as a new baseline and as the
fallback path for all Gated Deliberation policies.
Instead of relying on semantic search, it
processes the complete conversation
history through three stages:

#### Stage 1: Chunked Fact Extraction.

The conversation is partitioned into chunks of 60
messages each. For each chunk, an LLM
extracts query-relevant facts and assigns a
relevance score (0–10). A sliding context
window of 4 preceding messages maintains
continuity across chunks.

#### Stage 2: Multi-stage Filtering.

Extracted facts are sorted by relevance score.
A preliminary threshold ($>6$) filters out
weakly relevant facts; if more than 6 facts
remain, an additional LLM-based filter
selects the most pertinent subset.

#### Stage 3: Answer Generation.

The filtered facts replace the Top-30 memories $C$
in the answer generation prompt, producing
the final response.

This method is computationally expensive, which
motivates gated deliberation that
invokes it only when needed.

## 4 Experiments

### 4.1 Setup

#### Datasets.

We evaluate on two benchmark datasets:

(1) LoCoMo Maharana et al. (2024).
It contains 10 long-term dialogues with an average length of 24K tokens.
We utilize 1,540 questions spanning four core reasoning
categories: Single-hop, Multi-hop, Temporal,
and Open-domain (the adversarial question category
is explicitly excluded from our evaluation).

(2) RealTalk Lee et al. (2025).
A real-world dialogue dataset containing 10 conversations,
each averaging over 16,000 words.
It features 728 questions evaluated across three categories: Multi-hop,
Temporal, and Open-domain.

These datasets primarily focus on English-language
dialogue and are designed to evaluate the agent’s ability
to maintain coherence over long-term dialogue history.
For comprehensive details regarding the demographic distribution and specific linguistic phenomena of these datasets, we refer readers to their original documentation.

#### Metrics.

We report three complementary metrics:
the F1 score for lexical overlap,
the BLEU score for n-gram fidelity, and an
LLM-as-a-Judge (hereafter, LLM) score for semantic
equivalence. All scores are reported as
percentages.
For the LLM-as-a-Judge evaluation,
we follow the binary accuracy protocol established
by Chhikara et al. (2025).
Specifically, this approach employs
GPT-4o-mini to evaluate whether the generated
response is semantically consistent with the ground truth.
Rather than penalizing minor formatting differences, it provides a
robust assessment that better correlates with human judgment by tolerating generative
variations and relative temporal expressions.
See Appendix [A.8](#A1.SS8) for the complete prompt template.

#### LLMs.

All methods are evaluated with GPT-4o-mini
as the primary backbone.
We use OpenAI text-embedding-3-small as our embedding model.
We additionally report
results with Qwen3-235B-Instruct to assess
generalization.

#### Baselines.

We compare our results against the following baselines:
Full Context, which feeds the
entire conversation history to the model;
standard RAG, a standard retrieval-augmented generation
approach that chunks
dialogues into 4096-token segments for dense
retrieval;
LangMem Chase (2022), Mem0 Chhikara et al. (2025),
Zep Rasmussen et al. (2025), Nemori Nan et al. (2025),
EMem-G Zhou and Han (2025). See Nan et al. (2025)
for more details on these baselines.

### 4.2 Overall Performance

Table [1](#S4.T1) summarizes the overall
F1, LLM, and BLEU scores plus response time and token
usage for all methods across the four experimental
settings (see Tables [3](#A3.T3)–[6](#A3.T6) for details).

**Table 1: Summary of overall scores and efficiency across four settings. The baseline performance comes from Nan et al. (2025) and Zhou and Han (2025).**
|  | LoCoMo | RealTalk |  |  |  |  |  |  |  |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| Method | F1 | LLM | BLEU | Time(s) | Tokens | F1 | LLM | BLEU | Time(s) | Tokens |
| GPT-4o-mini |  |  |  |  |  |  |  |  |  |  |
| Full Context | 46.2 | 72.3 | 37.8 | – | – | – | – | – | – | – |
| RAG | 20.8 | 30.2 | 16.4 | – | – | – | – | – | – | – |
| LangMem | 35.8 | 51.3 | 29.4 | – | – | – | – | – | – | – |
| Mem0 | 41.5 | 61.3 | 34.2 | – | – | – | – | – | – | – |
| Zep | 37.5 | 58.5 | 30.9 | – | – | – | – | – | – | – |
| Nemori | 49.5 | 74.4 | 38.5 | – | – | – | – | – | – | – |
| Mem0^∗ | 51.2 | 72.7 | 41.0 | 1.28 | 2191 | 37.3 | 59.1 | 23.7 | 3.27 | 2303 |
| Filter | 51.6 | 74.0 | 41.6 | 2.67 | 3190 | 38.4 | 60.7 | 24.8 | 4.14 | 3376 |
| Majority Voting | 51.3 | 73.1 | 41.1 | 3.32 | 7534 | 37.5 | 60.4 | 23.7 | 5.40 | 8794 |
| Consensus | 53.5 | 76.1 | 43.0 | 9.55 | 15757 | 39.1 | 62.4 | 25.0 | 15.45 | 21949 |
| Quality Gating(ours) | 53.5 | 76.3 | 43.1 | 8.03 | 12681 | 39.4 | 62.5 | 25.4 | 13.00 | 16786 |
| Full Deliberation | 55.3 | 78.4 | 44.2 | 23.73 | 35435 | 40.6 | 62.8 | 27.4 | 27.95 | 48772 |
| Qwen3-235B-Instruct |  |  |  |  |  |  |  |  |  |  |
| Mem0^∗ | 48.1 | 74.7 | 39.7 | 1.52 | 2443 | 35.4 | 62.5 | 21.3 | 2.00 | 2777 |
| Filter | 50.3 | 76.8 | 42.0 | 2.71 | 3554 | 35.2 | 60.0 | 21.2 | 3.22 | 4119 |
| Majority Voting | 49.2 | 75.4 | 40.8 | 4.41 | 9149 | 35.8 | 60.6 | 21.4 | 5.45 | 11009 |
| Consensus | 51.1 | 76.4 | 42.4 | 7.84 | 15407 | 36.3 | 62.4 | 21.8 | 12.35 | 23673 |
| Quality Gating(ours) | 51.0 | 78.6 | 42.6 | 8.56 | 15574 | 35.5 | 63.1 | 22.0 | 16.51 | 26417 |
| Full Deliberation | 53.7 | 78.6 | 45.1 | 17.03 | 39101 | 37.5 | 64.4 | 24.8 | 27.79 | 57956 |

#### Comparison with Baselines.

Our enhanced System 1 baseline
(Mem0^∗) achieves an F1 score of 51.2
on the LoCoMo dataset utilizing GPT-4o-mini,
significantly outperforming both the original Mem0 (41.5)
and outperforming the recently proposed Nemori (49.5).

Furthermore, by establishing an exhaustive deliberate
reading process, the Full Deliberation mechanism
boosts the F1 score to 55.3
and the LLM score to 78.4 on the LoCoMo dataset
with GPT-4o-mini. This approach effectively mitigates
the "Lost-in-the-Middle" phenomenon,
demonstrating substantial gains over the Full Context (F1: 46.2, LLM: 72.3).
It also significantly outperforms Nemori (LLM 74.4).

In summary, while our System 1 baseline (Mem0^∗) already
achieves competitive performance, the Full
Deliberation establishes a new upper bound.

#### Effectiveness of Dual-Process Routing.

The core objective of D-Mem is to approximate the Full Deliberation without incurring its massive computational overhead.
The Majority Voting policy achieves a modest F1
improvement over Mem0^∗ (51.3 vs. 51.2) but at the
cost of an approximate 3$\times$ token overhead (7,534 vs. 2,191 tokens).
This suggests that Majority Voting
is prone to premature
convergence on incorrect semantic snippets.
Conversely, the
Consensus policy proves overly conservative
compared to Quality Gating:
while it successfully elevates the F1 score to 53.5,
it forces excessive fallbacks, resulting in heavy
computational costs (15,757 tokens).

#### Generalization and Robustness.

The efficacy of the Quality Gating mechanism transcends
specific datasets and model architectures.

On the Real-world Dialogue (RealTalk) dataset with GPT-4o-mini, the
structural advantages of D-Mem remain substantial:
Quality Gating achieves the better F1 score (39.4) and
LLM-as-a-judge score (62.5) than Consensus (39.1 and 62.4).
Mirroring the LoCoMo results,
it successfully recovers 97.0% of the Full Deliberation’s
F1 score (40.6) while consuming merely 34.4% of its
token cost (16,786 vs. 48,772).

With the Qwen3-235B-Instruct backbone, we observe a notable
divergence between evaluation metrics: while Quality
Gating outperforms Consensus in LLM-as-a-judge scores
(78.6 vs. 76.4 on LoCoMo; 63.1 vs. 62.4 on RealTalk),
it yields slightly lower F1 scores (e.g., 51.0 vs. 51.1 on LoCoMo).
This discrepancy is primarily attributed to the
distinct alignment styles of the underlying model.
Qualitative inspection reveals that Qwen3-235B-Instruct is
inclined toward explanatory generation,
explicitly outputting a “step-by-step” prior
to its final answer in 12 separate instances with Quality Gating.
In contrast, the Consensus policy—which derives the
final response via extraction from multiple
candidate drafts—effectively prunes these
deliberative steps.

### 4.3 In-depth Analysis of Quality Gating

Figure: Figure 2: F1 Improvement of Quality Gating over Mem0^∗ by Question Category on the LoCoMo Dataset.
Refer to caption: 2603.18631v1/x2.png

#### Per-Category Analysis

To further understand where the Quality Gating
mechanism yields its gains, we break down the
F1 improvement of Quality Gating over
Mem0^∗ by question category on the LoCoMo
dataset (Figure [2](#S4.F2)). A clear positive correlation emerges:
the improvement scales with question
difficulty. For Single-Hop questions—simple
fact lookups where the top-$K$ retrieval
typically provides sufficient evidence—Quality
Gating produces the smallest absolute gain
(+1.9 F1 with GPT-4o-mini, +2.5 with
Qwen3-235B-Instruct). For Multi-Hop questions, which
require chaining facts scattered across
multiple memories, the gain increases notably
(+2.6 and +3.0, respectively), as the quality
rubric’s Completeness dimension effectively
detects when retrieved context covers only
a subset of the required reasoning chain.
Temporal and Open-Domain categories benefit
the most (+2.7/+4.1 with GPT-4o-mini; +4.1/+3.2
with Qwen3-235B-Instruct), precisely because these
categories suffer most severely from lossy
abstraction: temporal queries lose relative
time anchors during memory compression, while
open-domain queries demand broader contextual
grounding that top-$K$ retrieval rarely captures.
This pronounced upward trend between question
difficulty and Quality Gating improvement
confirms that the gating mechanism selectively
escalates queries where static retrieval is
fundamentally insufficient, rather than
applying uniform overhead across all question
types.

#### Impact of Filtering and Model-Specific Variations.

We observe that the explicit filtering step in
Mem0^∗ yields a consistent improvement for
GPT-4o-mini across datasets, successfully
mitigating the “Lost-in-the-Middle” phenomenon.
However, this benefit does not generalize uniformly.
When deploying Qwen3-235B-Instruct on the highly
noisy RealTalk dataset, the addition of explicit
filtering counterintuitively degrades performance compared to
the unfiltered Mem0^∗ (e.g., the LLM score drops
from 62.5 to 60.0). We hypothesize that this
performance inversion occurs because
Qwen3-235B-Instruct struggles to recognize unstated
temporal anchors or implicit logical bridges
during the relevance assessment.
In such noise-heavy contexts, this can lead to the
erroneous pruning of critical context.
This indicates that the efficacy of aggressive
context filtering
is dependent on the underlying model’s alignment style
and inherent discriminative capabilities.

Crucially, our proposed Quality Gating policy
achieves a superior balance between performance and efficiency. On the LoCoMo
dataset with GPT-4o-mini, it achieves a strict
improvement over the Consensus policy: it matches Consensus F1 score (53.5) and
yields a higher LLM-as-a-judge score (76.3),
while simultaneously reducing token consumption by
nearly 19.5% (12,681 vs. 15,757 tokens). Compared to the
exhaustive fallback, Quality Gating effectively
recovers nearly 96.7% of the Full Deliberation’s F1
performance. Remarkably, it accomplishes this while
utilizing only 35.8% of the Full Deliberation’s tokens
(12,681 vs. 35,435) and approximately one-third of the
inference latency.

### 4.4 Fallback Mechanism Analysis

**Table 2: Impact of Fallback Mechanism in Gated Deliberation across Datasets. W/O FB: Without Fallback; W/ FB: With Fallback; Rate: Percentage of queries falling into this routing status.**
| Method | Status | LoCoMo | RealTalk |  |  |  |  |  |  |  |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| Rate | F1 | LLM | BLEU | Tokens | Rate | F1 | LLM | BLEU | Tokens |  |  |
| GPT-4o-mini |  |  |  |  |  |  |  |  |  |  |  |
| Majority Voting | W/O FB | 98.2% | 51.6 | 73.4 | 41.4 | 6902.1 | 97.1% | 38.0 | 61.1 | 23.9 | 7262.1 |
| W/ FB | 1.8% | 35.1 | 59.3 | 28.7 | 41082.7 | 2.9% | 19.7 | 38.1 | 17.3 | 58445.0 |  |
| Consensus | W/O FB | 75.5% | 57.4 | 78.1 | 46.7 | 6834.7 | 70.5% | 42.1 | 65.7 | 25.9 | 7148.4 |
| W/ FB | 24.5% | 41.6 | 69.8 | 31.5 | 42315.1 | 29.5% | 32.0 | 54.4 | 22.9 | 55959.9 |  |
| Quality Gating | W/O FB | 75.9% | 53.8 | 78.1 | 43.4 | 4254.6 | 75.3% | 41.6 | 66.8 | 26.4 | 4512.6 |
| W/ FB | 24.1% | 52.7 | 70.6 | 42.1 | 38582.9 | 24.7% | 32.8 | 49.4 | 22.4 | 53218.3 |  |
| Qwen3-235B-Instruct |  |  |  |  |  |  |  |  |  |  |  |
| Majority Voting | W/O FB | 96.4% | 49.8 | 76.2 | 41.3 | 7610.5 | 96.0% | 36.2 | 61.2 | 21.4 | 8609.1 |
| W/ FB | 3.6% | 33.3 | 52.7 | 26.4 | 48192.8 | 4.0% | 26.6 | 44.8 | 21.1 | 65075.4 |  |
| Consensus | W/O FB | 80.8% | 53.4 | 77.6 | 44.6 | 7548.9 | 74.2% | 38.3 | 63.9 | 21.6 | 8557.2 |
| W/ FB | 19.2% | 41.5 | 71.6 | 33.1 | 47288.6 | 25.8% | 30.6 | 58.0 | 22.5 | 65111.9 |  |
| Quality Gating | W/O FB | 72.1% | 52.7 | 81.8 | 44.6 | 4728.4 | 65.0% | 39.1 | 67.0 | 22.9 | 5431.0 |
| W/ FB | 27.9% | 46.6 | 71.6 | 37.1 | 42751.3 | 35.0% | 28.9 | 55.7 | 20.3 | 63345.1 |  |

Table [2](#S4.T2) breaks down
the performance and resource consumption based
on whether the fallback mechanism was triggered.

#### Adaptability to Dataset Complexity

Consistently across all settings,
the fallback rate consistently scales up from LoCoMo dataset to RealTalk dataset.
For example, on the LoCoMo dataset with Quality Gating, Qwen3-235B-Instruct triggers fallback 27.9% of the time;
on the substantially more challenging RealTalk dataset,
this rate rises to 35.0%. This dynamic behavior
demonstrates that Quality Gating successfully
calibrates to the varying complexity of the
dialogue history.

#### Selection Bias in Fallback Routing.

Consistently across all settings, queries requiring
the fallback mechanism yield lower scores than those
resolved in the fast path (e.g., F1 drops from 41.6 to
32.8 on RealTalk using GPT-4o-mini). This reflects a
natural selection bias: the quality gate effectively
isolates the most inherently difficult queries,
thus demonstrating the effectiveness of the gating mechanism in identifying queries that are likely to fail under static retrieval.

#### Model-Specific Sensitivity in Quality Gating.

Notably, Qwen3-235B-Instruct exhibits significantly higher fallback rates under the Quality Gating policy compared to GPT-4o-mini (e.g., 35.0% vs. 24.7% on RealTalk).
We attribute this to Qwen3-235B-Instruct’s advanced reasoning
capabilities, which facilitate a more rigorous assessment
against the quality rubric. As a stricter adjudicator,
the model more readily identifies subtle logical gaps
or evidentiary deficiencies, ensuring the fallback
mechanism is proactively engaged whenever static
retrieval is insufficient.

### 4.5 Case Study: Why Static Retrieval Fails

To better understand the underlying failure modes
of static semantic retrieval—and why our
D-Mem framework is necessary—we examine a
representative error from the Temporal reasoning
category.

> Query: When did Caroline go to the LGBTQ support group? Ground Truth: 7 May 2023 Raw Context ($\mathcal{H}$): 1:56 pm on 8 May, 2023, Caroline: I went to a LGBTQ support group yesterday and it was so powerful. Baseline Output: 8 May 2023 (Retrieved Memory: “1:56 pm on 8 May, 2023: Caroline found the transgender stories at the LGBTQ support group inspiring”) Full Deliberation Output: 7 May 2023

This failure illustrates how standard
incremental memory acts as a lossy
abstraction. Because background updates
occur without a guiding query, the LLM
over-compresses the dialogue. It retains
the thematic essence and the message’s
absolute timestamp (8 May) but
permanently discards the relative
temporal arithmetic needed to resolve
“yesterday”, losing crucial
context-dependent dependencies before
retrieval even begins.

## 5 Conclusion

In this work, we first introduced Full Deliberation,
a mechanism designed to mitigate the lossy
abstraction of standard retrieval-based memory.
It also mitigates the "Lost in the Middle" phenomenon
inherent in full-context processing
and achieves strong performance.
To address the prohibitive computational costs of
Full Deliberation, we further proposed D-Mem,
a dual-process memory architecture. Through Quality
Gating, D-Mem dynamically routes queries between
rapid associative recall and Full Deliberation.
Empirical evaluations demonstrate that D-Mem achieves
near-Full Deliberation performance on the LoCoMo and
RealTalk benchmarks, while successfully reducing both
token consumption and inference latency.

## 6 Limitations

While Full Deliberation is a powerful mechanism for mitigating the limitations of retrieval-based memory, it inherently consumes a massive amount of input tokens by exhaustively processing the entire dialogue history. Although our Quality Gating mechanism effectively reduces these token costs, this exhaustive context scanning still limits the architecture’s scalability for lifelong deployments with infinite context horizons.

Furthermore, because the current Full Deliberation mechanism splits the conversation history into isolated chunks for parallel fact extraction and relies solely on the LLM’s native self-attention over the extracted isolated facts, the system lacks explicit logical chaining. This poses a significant bottleneck for long-horizon reasoning that requires global context and cross-chunk dependencies. Consequently, a state-tracking architecture that explicitly synthesizes causal and temporal logic across chunks is needed.

## References

- A. Asai, Z. Wu, Y. Wang, A. Sil, and H. Hajishirzi (2024)
Self-RAG: learning to retrieve, generate, and critique through self-reflection.
In The Twelfth International Conference on Learning Representations,
External Links: [Link](https://openreview.net/forum?id=hSyW5go0v8)
Cited by: [§2.1](#S2.SS1.p1.1).
- H. Chase (2022)
LangChain.
GitHub.
Note: [https://github.com/langchain-ai/langchain](https://github.com/langchain-ai/langchain)Accessed: 2025-07-20
Cited by: [§4.1](#S4.SS1.SSS0.Px4.p1.1).
- P. Chhikara, D. Khant, S. Aryan, T. Singh, and D. Yadav (2025)
Mem0: building production-ready ai agents with scalable long-term memory.
External Links: 2504.19413,
[Link](https://arxiv.org/abs/2504.19413)
Cited by: [Appendix A](#A1.p1.1),
[§1](#S1.p1.1),
[§2.2](#S2.SS2.p1.1),
[§4.1](#S4.SS1.SSS0.Px2.p1.1),
[§4.1](#S4.SS1.SSS0.Px4.p1.1).
- J. St. B. T. Evans and K. E. Stanovich (2013)
Dual-process theories of higher cognition: advancing the debate.
Perspectives on Psychological Science 8 (3), pp. 223–241.
Note: PMID: 26172965
External Links: [Document](https://dx.doi.org/10.1177/1745691612460685),
[Link](https://doi.org/10.1177/1745691612460685),
https://doi.org/10.1177/1745691612460685
Cited by: [§1](#S1.p2.1).
- D. Kahneman (2011)
Thinking, fast and slow.
Farrar, Straus and Giroux, New York.
Cited by: [§1](#S1.p2.1).
- D. Lee, A. Maharana, J. Pujara, X. Ren, and F. Barbieri (2025)
REALTALK: a 21-day real-world dataset for long-term conversation.
External Links: 2502.13270,
[Link](https://arxiv.org/abs/2502.13270)
Cited by: [3rd item](#S1.I2.i3.p1.1),
[§4.1](#S4.SS1.SSS0.Px1.p3.1).
- N. F. Liu, K. Lin, J. Hewitt, A. Paranjape, M. Bevilacqua, F. Petroni, and P. Liang (2024)
Lost in the middle: how language models use long contexts.
Transactions of the Association for Computational Linguistics 12, pp. 157–173.
External Links: [Link](https://aclanthology.org/2024.tacl-1.9/),
[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00638)
Cited by: [§1](#S1.p1.1).
- A. Maharana, D. Lee, S. Tulyakov, M. Bansal, F. Barbieri, and Y. Fang (2024)
Evaluating very long-term conversational memory of LLM agents.
In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.),
Bangkok, Thailand, pp. 13851–13870.
External Links: [Link](https://aclanthology.org/2024.acl-long.747/),
[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.747)
Cited by: [3rd item](#S1.I2.i3.p1.1),
[§4.1](#S4.SS1.SSS0.Px1.p2.1).
- J. Nan, W. Ma, W. Wu, and Y. Chen (2025)
Nemori: self-organizing agent memory inspired by cognitive science.
External Links: 2508.03341,
[Link](https://arxiv.org/abs/2508.03341)
Cited by: [§4.1](#S4.SS1.SSS0.Px4.p1.1),
[Table 1](#S4.T1).
- C. Packer, S. Wooders, K. Lin, V. Fang, S. G. Patil, I. Stoica, and J. E. Gonzalez (2024)
MemGPT: towards llms as operating systems.
External Links: 2310.08560,
[Link](https://arxiv.org/abs/2310.08560)
Cited by: [§2.2](#S2.SS2.p1.1).
- P. Rasmussen, P. Paliychuk, T. Beauvais, J. Ryan, and D. Chalef (2025)
Zep: a temporal knowledge graph architecture for agent memory.
External Links: 2501.13956,
[Link](https://arxiv.org/abs/2501.13956)
Cited by: [§1](#S1.p1.1),
[§4.1](#S4.SS1.SSS0.Px4.p1.1).
- N. Rossi, J. Lin, F. Liu, Z. Yang, T. Lee, A. Magnani, and C. Liao (2024)
Relevance filtering for embedding-based retrieval.
In Proceedings of the 33rd ACM International Conference on Information and Knowledge Management,
CIKM ’24, New York, NY, USA, pp. 4828–4835.
External Links: ISBN 9798400704369,
[Link](https://doi.org/10.1145/3627673.3680095),
[Document](https://dx.doi.org/10.1145/3627673.3680095)
Cited by: [§2.1](#S2.SS1.p1.1).
- K. Wang, Y. Lin, J. Lou, Z. Zhou, B. Suvonov, and J. Li (2026)
E-mem: multi-agent based episodic context reconstruction for llm agent memory.
External Links: 2601.21714,
[Link](https://arxiv.org/abs/2601.21714)
Cited by: [§1](#S1.p1.1),
[§2.2](#S2.SS2.p2.1).
- L. Wang, C. Ma, X. Feng, Z. Zhang, H. Yang, J. Zhang, Z. Chen, J. Tang, X. Chen, Y. Lin, et al. (2024)
A survey on large language model based autonomous agents.
Frontiers of Computer Science 18 (6), pp. 186345.
Cited by: [§1](#S1.p1.1).
- Z. Xi, W. Chen, X. Guo, W. He, Y. Ding, B. Hong, M. Zhang, J. Wang, S. Jin, E. Zhou, et al. (2025)
The rise and potential of large language model based agents: a survey.
Science China Information Sciences 68 (2), pp. 121101.
Cited by: [§1](#S1.p1.1).
- W. Xu, Z. Liang, K. Mei, H. Gao, J. Tan, and Y. Zhang (2025)
A-mem: agentic memory for LLM agents.
In The Thirty-ninth Annual Conference on Neural Information Processing Systems,
External Links: [Link](https://openreview.net/forum?id=FiM0M8gcct)
Cited by: [§2.2](#S2.SS2.p1.1).
- Z. Xu, Z. Liu, Y. Yan, S. Wang, S. Yu, Z. Zeng, C. Xiao, Z. Liu, G. Yu, and C. Xiong (2026)
ThinkNote: enhancing knowledge integration and utilization of large language models via constructivist cognition modeling.
In Findings of the Association for Computational Linguistics: EACL 2026,
Note: To appear
Cited by: [§2.1](#S2.SS1.p1.1).
- B. Y. Yan, C. Li, H. Qian, S. Lu, and Z. Liu (2025)
General agentic memory via deep research.
External Links: 2511.18423,
[Link](https://arxiv.org/abs/2511.18423)
Cited by: [§2.2](#S2.SS2.p2.1).
- Y. Yu, W. Ping, Z. Liu, B. Wang, J. You, C. Zhang, M. Shoeybi, and B. Catanzaro (2024)
RankRAG: unifying context ranking with retrieval-augmented generation in LLMs.
In The Thirty-eighth Annual Conference on Neural Information Processing Systems,
External Links: [Link](https://openreview.net/forum?id=S1fc92uemC)
Cited by: [§2.1](#S2.SS1.p1.1).
- G. Zhang, M. Fu, K. Wang, G. Wan, M. Yu, and S. YAN (2025)
G-memory: tracing hierarchical memory for multi-agent systems.
In The Thirty-ninth Annual Conference on Neural Information Processing Systems,
External Links: [Link](https://openreview.net/forum?id=mmIAp3cVS0)
Cited by: [§2.2](#S2.SS2.p1.1).
- W. Zhang, Z. Liu, K. Wang, and S. Lian (2024)
Query expansion and verification with large language model for information retrieval.
In Advanced Intelligent Computing Technology and Applications: 20th International Conference, ICIC 2024, Tianjin, China, August 5–8, 2024, Proceedings, Part IV,
Berlin, Heidelberg, pp. 341–351.
External Links: ISBN 978-981-97-5671-1,
[Link](https://doi.org/10.1007/978-981-97-5672-8_29),
[Document](https://dx.doi.org/10.1007/978-981-97-5672-8%5F29)
Cited by: [§2.1](#S2.SS1.p1.1).
- H. S. Zheng, S. Mishra, X. Chen, H. Cheng, E. H. Chi, Q. V. Le, and D. Zhou (2024)
Take a step back: evoking reasoning via abstraction in large language models.
In The Twelfth International Conference on Learning Representations,
External Links: [Link](https://openreview.net/forum?id=3bq3jsvcQ1)
Cited by: [§2.1](#S2.SS1.p1.1).
- W. Zhong, L. Guo, Q. Gao, H. Ye, and Y. Wang (2024)
MemoryBank: enhancing large language models with long-term memory.
In Proceedings of the Thirty-Eighth AAAI Conference on Artificial Intelligence and Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence and Fourteenth Symposium on Educational Advances in Artificial Intelligence,
AAAI’24/IAAI’24/EAAI’24.
External Links: ISBN 978-1-57735-887-9,
[Link](https://doi.org/10.1609/aaai.v38i17.29946),
[Document](https://dx.doi.org/10.1609/aaai.v38i17.29946)
Cited by: [§1](#S1.p1.1),
[§2.2](#S2.SS2.p1.1).
- S. Zhou and J. Han (2025)
A simple yet strong baseline for long-term conversational memory of llm agents.
External Links: 2511.17208,
[Link](https://arxiv.org/abs/2511.17208)
Cited by: [§4.1](#S4.SS1.SSS0.Px4.p1.1),
[Table 1](#S4.T1).
- Q. Zhu, S. Chen, R. Yu, Z. Wu, and B. Wang (2026)
From lossy to verified: a provenance-aware tiered memory for agents.
External Links: 2602.17913,
[Link](https://arxiv.org/abs/2602.17913)
Cited by: [§1](#S1.p2.1).

## Appendix A Prompt Templates

In this section, we provide the complete prompt templates used in our experiments.
For the sake of reproducibility and fair comparison,
the prompts utilized for the static retrieval baseline and the evaluation judge are sourced from the mem0 open-source repository Chhikara et al. (2025). Templates specifically designed for our Multi-dimensional Quality Gating and Full Deliberation modules are denoted with “(Ours)”.

### A.1 Answer Generation Prompt (Mem0)

The following prompt is used for generating answers from retrieved memories in all methods:

### A.2 Filter Prompt (Ours)

This prompt is used in the Filter method to filter irrelevant memories:

### A.3 Fact Extraction Prompt (Ours)

This prompt is used in the Full Deliberation method to extract relevant facts from conversation history:

### A.4 Fact Filtering System Prompt (Ours)

This prompt is used in the Full Deliberation method to further filter extracted facts:

### A.5 Consensus Prompt (Ours)

This prompt is used in the Consensus policy to check if all three generated answers are semantically equivalent:

### A.6 Majority Voting Prompt (Ours)

This prompt is used in the Majority Voting policy
to check if at least 2 out of 3 generated answers are
semantically equivalent:

### A.7 Quality Gating Prompt (Ours)

This prompt is used in the Quality Gating policy
to evaluate answer quality and trigger
Full Deliberation:

### A.8 LLM-as-a-Judge Prompt (Mem0)

This prompt is used to evaluate the quality of generated answers using an LLM as a judge:

## Appendix B Implementation Details

### B.1 Hyperparameter Settings

Key hyperparameters used in our experiments:

- •
TopK: 30 (number of memories retrieved per speaker)
- •
Temperature for Answer Generation: 0.0 (deterministic)
- •
Temperature for Multiple Answers: 0.7 (for diversity in Majority Voting and Consensus methods)
- •
MESSAGES_CHUNK_SIZE: 60 (messages per chunk in Full Deliberation method)
- •
HISTORY_SIZE: 4 (previous messages as context in Full Deliberation method)
- •
Preliminary Score Threshold: 6 (minimum relevance score for filtering in Full Deliberation method)
- •
LLM Filter Threshold: 6 (minimum facts to trigger LLM-based filtering)
- •
Fact Extraction Score Threshold: 5 (minimum relevance score for extraction in Full Deliberation method)

### B.2 System Configuration

- •
Vector Store: Qdrant (local instance)
- •
Embedding Model: OpenAI’s text-embedding-3-small
- •
API Providers: OpenAI (GPT-4o-mini), Qwen (Qwen3-235B-Instruct)

## Appendix C Comprehensive Evaluation Results

In the main text, due to space constraints, we
present the aggregated performance and key
analytical visualizations of our proposed D-Mem
framework. In this section, we provide the
complete, fine-grained evaluation results across all tested dimensions, datasets, and base models.

- •
Table [3](#A3.T3): Full results on the LoCoMo dataset using GPT-4o-mini. This table also includes results from prior baseline frameworks (LangMem, Mem0, RAG, Zep, Nemori) for a comprehensive historical comparison.
- •
Table [4](#A3.T4): Full results on the RealTalk dataset using GPT-4o-mini. Note that RealTalk inherently excludes Single-Hop questions.
- •
Table [5](#A3.T5): Full results on the LoCoMo dataset using the Qwen3-235B-Instruct model.
- •
Table [6](#A3.T6): Full results on the RealTalk dataset using the Qwen3-235B-Instruct model.

These detailed tables substantiate the core claims made in Section [4](#S4) (Main Paper), particularly demonstrating the consistent superiority of our Multi-dimensional Quality Gating policy in balancing high-fidelity reasoning with cognitive economy.

**Table 3: Overall Performance Comparison (LoCoMo, GPT-4o-mini)**
| Method | F1 Score | LLM-as-a-Judge | BLEU | Efficiency |  |  |  |  |  |  |  |  |  |  |  |  |  |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| S-H | M-H | Temp | O-D | Avg | S-H | M-H | Temp | O-D | Avg | S-H | M-H | Temp | O-D | Avg | In | Out | Time |  |
| Full Context | 53.1 | 35.4 | 44.1 | 24.5 | 46.2 | 83.0 | 66.8 | 56.2 | 48.6 | 72.3 | 44.7 | 26.1 | 36.1 | 17.2 | 37.8 | – | – | – |
| LangMem | 38.8 | 33.5 | 31.9 | 29.4 | 35.8 | 61.4 | 52.4 | 24.9 | 47.6 | 51.3 | 33.1 | 23.9 | 26.2 | 23.5 | 29.4 | – | – | – |
| Mem0 | 44.4 | 34.3 | 44.4 | 27.1 | 41.5 | 68.1 | 60.3 | 50.4 | 40.6 | 61.3 | 37.7 | 25.2 | 37.6 | 19.4 | 34.2 | – | – | – |
| RAG | 22.2 | 18.6 | 19.5 | 19.0 | 20.8 | 32.0 | 31.3 | 23.7 | 32.6 | 30.2 | 18.6 | 11.7 | 15.7 | 13.5 | 16.4 | – | – | – |
| Zep | 39.7 | 27.5 | 44.8 | 22.9 | 37.5 | 63.2 | 50.5 | 58.9 | 39.6 | 58.5 | 33.7 | 19.3 | 38.1 | 15.7 | 30.9 | – | – | – |
| Nemori | 54.4 | 36.5 | 56.7 | 20.8 | 49.5 | 82.1 | 65.3 | 71.0 | 44.8 | 74.4 | 43.2 | 25.6 | 46.6 | 15.1 | 38.5 | – | – | – |
| Basic Methods |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |
| Mem0^∗ | 55.2 | 38.8 | 59.3 | 25.8 | 51.2 | 79.2 | 63.8 | 73.2 | 40.6 | 72.7 | 45.5 | 27.0 | 48.3 | 19.4 | 41.1 | 2186.2 | 5.2 | 1.278 |
| Filter | 54.3 | 41.0 | 61.0 | 27.9 | 51.6 | 78.4 | 67.4 | 75.7 | 50.0 | 74.0 | 45.2 | 29.4 | 48.8 | 20.8 | 41.6 | 3161.7 | 28.5 | 2.673 |
| Gated Deliberation |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |
| Majority Voting | 54.5 | 40.2 | 60.2 | 25.1 | 51.3 | 79.1 | 65.6 | 72.0 | 46.9 | 73.1 | 45.1 | 27.9 | 49.2 | 18.4 | 41.1 | 7501.4 | 32.6 | 3.318 |
| Consensus | 57.5 | 42.2 | 61.5 | 25.3 | 53.5 | 83.5 | 68.4 | 72.9 | 44.8 | 76.1 | 47.3 | 30.3 | 49.6 | 19.9 | 43.0 | 15543.6 | 213.0 | 9.549 |
| Quality Gating (ours) | 57.1 | 41.4 | 62.0 | 29.9 | 53.5 | 83.0 | 68.8 | 73.2 | 50.0 | 76.3 | 47.2 | 29.8 | 50.3 | 22.4 | 43.1 | 12524.6 | 156.6 | 8.03 |
| Full Deliberation | 58.9 | 44.2 | 63.0 | 31.1 | 55.3 | 83.1 | 73.4 | 77.9 | 54.2 | 78.4 | 48.2 | 32.0 | 50.6 | 24.0 | 44.2 | 34805.0 | 629.9 | 23.725 |

**Table 4: Overall Performance Comparison (RealTalk, GPT-4o-mini)**
| Method | F1 Score | LLM-as-a-Judge | BLEU | Efficiency |  |  |  |  |  |  |  |  |  |  |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| M-H | Temp | O-D | Avg | M-H | Temp | O-D | Avg | M-H | Temp | O-D | Avg | In | Out | Time |  |
| Basic Methods |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |
| Mem0^∗ | 32.0 | 47.5 | 21.7 | 37.3 | 51.2 | 68.0 | 54.6 | 59.1 | 23.5 | 26.1 | 17.0 | 23.7 | 2296.5 | 6.0 | 3.272 |
| Filter | 33.2 | 48.3 | 24.0 | 38.4 | 54.8 | 68.7 | 53.7 | 60.7 | 24.6 | 27.3 | 17.8 | 24.8 | 3345.1 | 31.3 | 4.140 |
| Gated Deliberation |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |
| Majority Voting | 31.6 | 48.2 | 22.2 | 37.5 | 53.2 | 69.6 | 53.7 | 60.4 | 22.6 | 27.0 | 17.3 | 23.7 | 8738.5 | 55.2 | 5.404 |
| Consensus | 34.4 | 49.1 | 22.8 | 39.1 | 57.5 | 69.6 | 54.6 | 62.4 | 24.4 | 28.2 | 17.3 | 25.0 | 21563.9 | 385.2 | 15.446 |
| Quality Gating (ours) | 33.6 | 50.6 | 22.7 | 39.4 | 55.2 | 70.9 | 58.3 | 62.5 | 24.3 | 29.0 | 17.7 | 25.4 | 16555.2 | 231.1 | 13.003 |
| Full Deliberation | 36.4 | 50.1 | 24.4 | 40.6 | 59.5 | 68.0 | 56.5 | 62.8 | 27.9 | 30.0 | 18.3 | 27.4 | 47955.0 | 816.5 | 27.951 |

**Table 5: Overall Performance Comparison (LoCoMo, Qwen3-235B-Instruct)**
| Method | F1 Score | LLM-as-a-Judge | BLEU | Efficiency |  |  |  |  |  |  |  |  |  |  |  |  |  |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| S-H | M-H | Temp | O-D | Avg | S-H | M-H | Temp | O-D | Avg | S-H | M-H | Temp | O-D | Avg | In | Out | Time |  |
| Mem0 (Baseline) | 38.0 | 29.5 | 42.5 | 16.3 | 36.0 | 58.0 | 52.5 | 49.8 | 37.5 | 54.0 | 31.7 | 20.1 | 35.1 | 13.1 | 29.1 | 1977.9 | 5.7 | 0.702 |
| Basic Methods |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |
| Mem0^∗ | 51.7 | 41.5 | 51.8 | 22.6 | 48.1 | 79.9 | 74.8 | 65.7 | 59.4 | 74.7 | 45.4 | 32.1 | 38.1 | 18.0 | 39.7 | 2431.2 | 11.6 | 1.518 |
| Filter | 52.8 | 41.6 | 59.1 | 24.5 | 50.3 | 79.8 | 74.5 | 77.0 | 57.3 | 76.8 | 46.7 | 32.5 | 44.6 | 20.3 | 42.0 | 3521.8 | 32.5 | 2.713 |
| Gated Deliberation |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |
| Majority Voting | 52.6 | 41.3 | 54.9 | 23.8 | 49.2 | 80.5 | 75.9 | 67.0 | 57.3 | 75.4 | 46.2 | 32.4 | 40.6 | 18.2 | 40.8 | 9059.9 | 89.0 | 4.407 |
| Consensus | 54.6 | 43.4 | 56.8 | 23.4 | 51.1 | 81.9 | 75.9 | 67.9 | 58.3 | 76.4 | 48.0 | 33.8 | 42.6 | 17.3 | 42.4 | 15187.2 | 219.3 | 7.840 |
| Quality Gating (ours) | 54.2 | 44.5 | 55.9 | 25.8 | 51.0 | 83.7 | 78.7 | 70.1 | 61.5 | 78.6 | 47.8 | 34.6 | 42.4 | 21.7 | 42.6 | 15319.4 | 254.7 | 8.560 |
| Full Deliberation | 56.6 | 47.7 | 58.8 | 28.6 | 53.7 | 83.6 | 78.7 | 71.7 | 57.3 | 78.6 | 49.6 | 38.2 | 45.3 | 25.0 | 45.1 | 38488.9 | 611.6 | 17.032 |

**Table 6: Overall Performance Comparison (RealTalk, Qwen3-235B-Instruct)**
| Method | F1 Score | LLM-as-a-Judge | BLEU | Efficiency |  |  |  |  |  |  |  |  |  |  |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| M-H | Temp | O-D | Avg | M-H | Temp | O-D | Avg | M-H | Temp | O-D | Avg | In | Out | Time |  |
| Basic Methods |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |
| Mem0^∗ | 32.1 | 44.1 | 18.6 | 35.4 | 59.1 | 68.0 | 55.6 | 62.5 | 26.1 | 18.8 | 15.4 | 21.3 | 2765.5 | 11.0 | 1.995 |
| Filter | 32.1 | 43.3 | 19.8 | 35.2 | 56.8 | 65.8 | 51.8 | 60.0 | 24.9 | 19.7 | 15.4 | 21.2 | 4075.5 | 43.9 | 3.215 |
| Gated Deliberation |  |  |  |  |  |  |  |  |  |  |  |  |  |  |  |
| Majority Voting | 32.8 | 44.8 | 17.9 | 35.8 | 55.8 | 67.7 | 52.8 | 60.6 | 26.5 | 18.9 | 14.8 | 21.4 | 10858.4 | 150.7 | 5.446 |
| Consensus | 33.6 | 44.4 | 20.1 | 36.3 | 59.8 | 68.0 | 52.8 | 62.4 | 26.5 | 19.2 | 16.5 | 21.8 | 23162.0 | 510.5 | 12.353 |
| Quality Gating (ours) | 33.3 | 43.1 | 19.2 | 35.5 | 61.5 | 67.7 | 53.7 | 63.1 | 27.5 | 19.3 | 14.8 | 22.0 | 25716.8 | 700.2 | 16.515 |
| Full Deliberation | 35.4 | 44.6 | 22.6 | 37.5 | 61.8 | 68.7 | 59.3 | 64.4 | 29.1 | 23.4 | 17.3 | 24.8 | 56703.4 | 1253.0 | 27.792 |

## Appendix D Computational Resources and Reproducibility

To ensure the reproducibility of our experimental results, we provide details regarding the computational environment and the models utilized in this study:

- •
Infrastructure: All experimental orchestration, data pre-processing, and local evaluation scripts were executed on a Lenovo laptop equipped with an NVIDIA GeForce RTX 4070 Laptop GPU (8GB VRAM), and an Intel Core i9-13900HX CPU.
- •
Model Access: We utilized proprietary models via official API endpoints to ensure consistency:
–
GPT-4o-mini: Accessed via the OpenAI API. Its exact parameter count remains proprietary and has not been disclosed by the provider.
–
Qwen-235B-Instruct: Accessed via the Alibaba DashScope API. It is a Mixture-of-Experts (MoE) model with a total of 235 billion parameters and 22 billion active parameters.
- •
Computational Budget: The entire evaluation process for D-Mem is highly efficient; all experimental runs were completed within 24 wall-clock hours through parallel API invocations, involving a total consumption of approximately 400 million tokens.
- •
Software and Evaluation Packages: Our implementation and evaluation framework was built upon the following technical stack:
–
Orchestration: Python 3.10 with the openai (v2.7.1) library for model inference and API orchestration. We also utilized langchain (v1.2.7) for auxiliary memory management tasks.
–
Tokenization: tiktoken (v0.12.0) with the cl100k_base encoding was employed for precise token counting and to ensure compliance with model-specific context window constraints.
–
Evaluation Metrics: Lexical overlap metrics were computed using nltk (v3.9.2) for F1 and BLEU scores, and rouge-score (v0.1.2) for ROUGE-L. Semantic evaluations (LLM-as-a-Judge) were executed by GPT-4o-mini, following the multi-dimensional rubric detailed in Appendix [A](#A1). Statistical analysis and visualization were performed using scikit-learn (v1.7.2), matplotlib (v3.10.7), and pandas (v2.3.3).
- •
Statistical Transparency: Due to the substantial computational costs associated with processing the extensive contexts in the LoCoMo and RealTalk benchmarks, all performance metrics reported in this paper are derived from a single, exhaustive execution of the evaluation pipeline. To ensure the reliability and reproducibility of these results, we utilized deterministic decoding (e.g., setting $\texttt{temperature}=0$) for memory retrieval, except for the Majority Voting and Consensus.

## Appendix E Ethics Statement

In accordance with the ACL Code of Ethics, we acknowledge and discuss the potential risks and broader impacts associated with the deployment of long-term memory systems for LLM agents like D-Mem.

#### Privacy and Data Security.

The core capability of D-Mem involves persistently storing and retrieving extensive user interaction histories. This inherently introduces risks related to data privacy in real-world applications, especially if the conversational context contains Personally Identifiable Information (PII) or sensitive operational data. To mitigate these concerns during our research phase, we strictly evaluated our framework on publicly available benchmark datasets (e.g., LoCoMo and RealTalk). We verified that these standard benchmarks have been appropriately pre-processed and anonymized by their creators to remove PII and mitigate offensive content. However, for future real-world deployment, practitioners must implement strict data encryption and allow users to actively manage or delete their memory states.

#### Environmental and Computational Impact.

While our Quality Gating mechanism successfully mitigates redundant compute for simple queries (System 1), the Full Deliberation module (System 2) requires exhaustive context processing. As demonstrated in our efficiency metrics, this exhaustive nature increases token consumption and inference latency. Large-scale deployment of such dual-process systems could lead to a substantial carbon footprint. Future work should explore more eco-friendly deliberation alternatives, such as deploying smaller, specialized language models for the gating functions and Full Deliberation.

#### Memory-Induced Bias and Safety.

A highly retentive memory system runs the risk of perpetuating or amplifying historical biases. If an agent ingests toxic or factually incorrect statements from a user, these “poisoned” memories could be retrieved during future multi-hop reasoning, leading to unsafe or hallucinated outputs over time. We urge developers to pair D-Mem with robust safety guardrails and memory-sanitization protocols before user-facing deployment.

## Appendix F Artifact Licenses and Terms of Use

To ensure responsible NLP research and compliance with intellectual property guidelines, we outline the licenses of the scientific artifacts used and created in this work:

- •
Utilized Datasets: The LoCoMo and RealTalk datasets are used strictly for academic evaluation purposes, adhering to their respective open-source distribution terms (e.g., CC BY 4.0).
- •
Utilized Models and Frameworks: We accessed GPT-4o-mini via the official OpenAI API under their terms of service. The Qwen3-235B-Instruct model via the official Qwen API under their terms of service. The baseline memory framework, Mem0, is distributed under the Apache License 2.0.
- •
Created Artifacts: The source code for our D-Mem framework, along with all evaluation scripts, is distributed under the MIT License. The intended use of our created artifacts is to facilitate reproducibility and future academic research, which is entirely compatible with the original licenses of the utilized data and frameworks.