MemUse: Moving Memory Evaluation from Direct QA to Natural Integration in Long-Term Human-AI Conversation
Abstract
Direct QA benchmarks for conversational memory do not predict user satisfaction, whereas natural integration of prior context does, revealing a large gap between elicited recall and conversational use.
Memory systems for conversational LLMs are conventionally evaluated by direct, fact-seeking questions about prior dialogue (Direct QA): can the model recall fact X from a prior conversation? We tested whether higher Direct QA accuracy correlates with higher user satisfaction in a 4-month deployment (40 users, 1,872 sessions, 7 memory conditions). Existing-benchmark Direct QA varies from 19.7% to 70.1% across the 7 conditions, but satisfaction does not change. We hypothesize that existing benchmarks and user satisfaction are tracking different capabilities: benchmarks measure elicited retrieval (recall when asked), while conversation requires natural integration (detecting relevance and naturally weaving prior context into a response). To examine this, we introduce MemUse, a set of real user-cued memory moments drawn from the deployment, scored by an integration-aware judgment of the natural conversational response. Holding the model and context fixed, the same system that scores 78.8% on Direct QA references only 7.9% of those facts in conversation -- a 71-point gap. Within these moments, Natural Integration is associated with satisfaction, whereas Direct QA is not. We release the deployment corpus and MemUse together with all judgments and scoring prompts at https://github.com/ryuichi-sumida/memuse.
Community
Accepted to EMNLP 2026 Main.
We ran a 4-month real-world deployment of a memory-augmented conversational AI and found that Direct QA accuracy - the standard way to evaluate long-term memory - doesn't predict user satisfaction: a system scoring 78.8% on Direct QA spontaneously referenced only 7.9% of those facts in actual conversation.
We introduce MemUse, a benchmark built from real user interactions that instead measures whether models naturally integrate remembered context into responses.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- When Your Agent Opens the Chat App: Agent-Controlled Search over Raw Chat Logs Rivals Structured Memory (2026)
- Don't Scroll Back: Missing-Evidence Memory for Streaming Dialogue Summarization (2026)
- MemDelta: Controlled Baselines and Hidden Confounds in Agent Memory Evaluation (2026)
- Total Recall at What Cost? Benchmarking the Serving Cost of Agentic Memory Systems (2026)
- FTA-Mem: Fact-Time-Affect Anchored Memory for Low-Density Long-Term Dialogue (2026)
- RUMBA: Russian User Memory Benchmark (2026)
- TrajWiki: Source-Grounded Memory Trajectories for Long-Horizon Dialogue Agents (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.24189 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper