SpeakerMem-R1: Solving the Who-Said-What Dilemma in Multi-Party AI Memory
- Conversational AI often struggles with message attribution and tracking relational dynamics in multi-party group chats.
- SpeakerMem-R1 solves this using dual-track memory: verbatim speaker-labeled logs paired with structured person and group views.
- Reinforcement Learning (GRPO) training boosts memory construction accuracy from 57.38% to 68.20% in controlled tests.
- Achieves a state-of-the-art score of 62.33% on EverMemBench and 70.85% on the LoCoMo long-term conversation benchmark.
The Challenge of Multi-Party Conversations
When Large Language Models (LLMs) participate in long multi-party conversations, they face a unique challenge: keeping track of social context. In a busy group chat, it isn't enough to simply remember what was said. An AI needs to know who said it, whom it concerns, how individual perceptions differ, and what information is shared by the whole group over time.
Existing LLM memory systems frequently fail in group settings due to two main bottlenecks:
- Message Attribution & Relational Understanding: Mixing up speaker identities or failing to connect clues spread across multiple members.
- State Reconstruction: Rebuilding a coherent timeline from tangled, interleaved group chat logs.
Introducing SpeakerMem-R1
To address these hurdles, SpeakerMem-R1 introduces a speaker-centered dual-track memory framework. Instead of treating dialogue as a uniform stream of text, SpeakerMem-R1 splits memory into two distinct pathways:
- Verbatim Track: Retains raw, speaker-labeled messages to preserve exact dialogue details.
- Structured Track: Extracts derived states organized into dedicated person-level and group-level perspectives.
When a user asks a question, the system dynamically retrieves and combines evidence from both tracks using entity, event, and temporal indexing.
Smarter Memory Writing via RL
Building structured memories without introducing errors is notoriously difficult. To solve this locally, the research introduces Writer-R1, trained with a combination of SpeakerLevenshtein distance loss and speaker-conditioned Group Relative Policy Optimization (GRPO).
In controlled tests involving 305 complex questions, Reinforcement Learning (RL) increased the memory writer's mean accuracy from 57.38% (Supervised Fine-Tuning) to 68.20% (RL)—proving that targeted RL substantially reduces attribution and update errors.
Benchmark Performance
SpeakerMem-R1 was evaluated across several key benchmarks, setting top-tier performance standard across the board:
- EverMemBench: Achieved 62.33% on the public leaderboard, taking the top spot among state-of-the-art frameworks.
- LoCoMo Benchmark: Reached 70.85% across 1,986 long-term conversation questions.
- SocialMemBench & GroupMemBench: Recorded 69.2% and 47.9% binary accuracy, respectively.
Ablation studies confirmed that neither verbatim logs nor structured views are sufficient on their own; their combination is what enables reliable, context-aware AI memory in complex social dialogues.