MMLAv4

在推理中,
更新策略与记忆。
Adapt the policy.
Update the memory.

用一次尝试的反馈,调整后续尝试。
Φ 改变生成方式,M 保存可供读取的记录。
Feedback from one attempt can inform the next.
Φ adapts generation. M retains records for later reasoning.

两种状态,分别更新Two states. Separate updates.Φ / M
已到达的反馈Feedback received
Φ数值状态Policy values
有界更新 / 保持原状Bounded update / no-op
已完成的片段Completed segment
M记忆表Memory table
整行提交 / NULLWhole-row commit / NULL
Φ+M后续推理Later reasoning读取兼容版本Read compatible versions

同一道题仍在进行时,反馈可调整 Φ,参与下一次生成。While a problem is still open, feedback can adjust Φ for the next generation.

Junyi Zou · Avrova Donz222 页 · 2026 年 9 月 14 日222 pages · 14 Sep 2026实验统计截至 2026.09.13Results through 13 Sep 2026

01 / 运行机制Mechanisms

策略状态与记忆状态Policy state and memory state

Φ 和 M 都能影响下一次尝试,但保存的内容、更新的时机和写入方式各不相同。Both Φ and M can affect the next attempt. They store different things and have different update rules and timing.

POLICY STATE

一组可更新的数值An adjustable numerical state

存什么Stores
例如小型适配器的参数,大小有上限。For example, a small adapter's parameters, with a fixed size limit.
怎么改Updates
用已收到的反馈执行受限数值更新,或保持不变。A bounded numerical update from received feedback, or no change.
怎么用Used for
参与后续计算,改变模型生成候选的方式。Participating in later computation to change how candidates are generated.

AUTHORITATIVE MEMORY

一张容量有限的记忆表A memory table with limited capacity

存什么Stores
有类型的完整记录,包含内容、来源、版本等字段。Complete typed records with content, provenance, version, and other fields.
怎么改Updates
片段结束后,每个事件替换一行,或选择 NULL 不写入。After a segment ends, each event replaces one row or takes NULL.
怎么用Used for
向后续推理提供已经保存、仍然有效的信息。Providing retained, still-valid information to later reasoning.

两条路径,各有触发时机Two paths, with different triggers

这些在线更新中,基础参数 θ 保持固定Base parameters θ stay fixed during these online updates
策略路径 ΦPolicy path Φ一次尝试结束,收到反馈An attempt ends; feedback arrives更新 Φ,或不更新Update Φ, or leave it unchanged后续尝试读取当前的 ΦLater attempts read the resulting Φ
记忆路径 MMemory path M片段结束,整理已完成事件A segment ends; completed events are collected逐事件生成候选,写入或 NULLFor each event: propose, then write or take NULL全部处理完,后续推理读取 MAfter all events are processed, later reasoning reads M

“尝试”是一次候选生成;“片段”是系统划定的一段记录,结束后才能回看整合。两条路径可以分别发生,无需每次同时更新。An attempt generates a candidate. A segment is a delimited record that becomes available for consolidation after it closes. Either update path can run independently.

贯穿示例 · 非实验结果Worked example · Not an experimental result

一道尚未结束的整数约束题An integer-constraint problem still in progress

题目先给出 x + y = 10。第一次尝试提交 (4, 6) 后,校验器才返回一条额外约束:x > y。The problem initially gives x + y = 10. After the first attempt submits (4, 6), the checker returns an additional constraint: x > y.

一次尝试开始时At the start of an attempt

用当前信息,
生成一个候选。
Use what is known
to generate a candidate.

模型把题目、当前 Φ 和 M 中可读取的信息用于生成。一次尝试固定读取这组版本,途中不会悄悄换成刚更新的状态。The model generates using the problem, current Φ, and information read from M. The attempt uses a fixed pair of versions throughout.

此时只知道 x + y = 10,可以提出 (4, 6)。反馈 x > y 尚未到达,不能提前使用。Only x + y = 10 is known, so (4, 6) is a possible candidate. The feedback x > y has not arrived yet.

本阶段At this stage读取 Φ、M,生成候选;不改写这两种状态。Read Φ and M and generate a candidate; neither state is rewritten.
01 / 04
MMLAv4 / 状态流STATE FLOWΦM
Φi策略版本Policy version
Mj记忆版本Memory version
模型计算 · 基础参数 θ 固定MODEL COMPUTATION · FIXED BASE θ
已知题目Known constraintx + y = 10
一次候选One candidate(4, 6)

x > y 尚未返回,不能作为本次输入。x > y has not been returned and cannot be an input yet.

两条更新路径可分别发生,不要求每次都更新 Φ 或写入 M。The two update paths are independent. Neither must run on every attempt.v4 第 2–3 节v4 §§2–3

预测式记忆准入Predictive memory admission

凭什么决定
写入还是 NULL?
What makes a write
worth choosing?

“准入”就是决定一条候选能否进入 M,以及替换哪个槽位。容量固定,保存新信息可能挤掉仍有用的旧信息;有价值的候选也未必值得写入。Admission decides whether a candidate enters M and which slot it replaces. With fixed capacity, new information can displace useful old information, so even a useful candidate may not be worth writing.

每个可行动作的目标Objective for each feasible action预期后续误差 + 旧记忆损失 + 动作成本Expected future error + damage to retained memory + action cost在这些后果之间比较;NULL 也要参与,越低越好。Compare these consequences, including NULL. Lower is better.

在例子中,保存 x > y 可能方便以后读取;但若它已被可靠保留,或覆盖会丢掉更重要的约束,NULL 可能更合适。Retaining x > y may help later reads. If it is already reliably retained, or a write would remove a more important constraint, NULL may be preferable.

v4 §3.4 · 训练目标v4 §3.4 · Training objective
离线训练目标OFFLINE TRAINING OBJECTIVE以单次写入决策为例A single-write decision
  1. 从同一起点试各动作Try each action from the same state复制写入前的 M,分别执行各可行槽位写入和 NULL。Copy the pre-write M and apply each feasible slot write or NULL separately.
  2. 重新运行相同的一组后续Replay the same set of continuations让后续计算实际读取各自处理后的 M,比较平均误差、旧记忆损失和成本。Later computation reads the M produced by each action. Compare mean error, damage to retained memory, and cost.
  3. 把比较结果作为监督Use the comparison as supervision训练准入器仅凭决策时已有的信息,预测各动作的预期风险。Train admission to predict each action's expected risk using only information available at decision time.

比较的是多种可能后续的平均表现,不能看过某个实际未来,再为它挑选动作。The target averages over possible continuations. An action cannot be chosen after seeing which future actually occurred.

未来分支与答案不进入线上输入Future branches and answers are excluded from deployment inputs
线上决策DEPLOYMENT预测器看不到未来Future outcomes are unavailable
有界历史摘要
当前状态与候选
Bounded history summary
Current state and candidates
估计各动作风险
比较可行选择
Estimate action risks
Compare feasible choices
槽位 / NULLSlot / NULL

“预测式”指用已有信息估计未来后果。部署时没有未来分支、标准答案或事后评估结果作为输入。“Predictive” means estimating future consequences from available information. Future branches, answer keys, and later evaluations are absent from deployment inputs.

机制与训练目标依据 v4 第 2–3 节。各模块的定义、实现条件和评估方法见下方论文详解。Mechanisms and training objectives follow Sections 2–3 of v4. The paper guide develops each module's definitions, implementation conditions, and evaluation methods.

论文机制The papers

v4 与五篇配套论文v4 and the five companion papers

完整机制详解Full mechanism guide

02 / 系统架构Architecture

MMLAv4 系统架构The MMLAv4 architecture

策略和记忆各自维护版本、写入权限与回滚记录。一次推理读取兼容的版本组合,更新后的状态用于之后的推理。Policy and memory have separate versions, write permissions, and rollback records. Each attempt reads a compatible pair; updates become available to later attempts.

MMLAv4 架构MMLAv4 architecture

PNG SVG
根据 v4 第 2–3 节及图 3 绘制。Based on v4 Sections 2–3 and Figure 3.论文图 3 · 第 10 页Paper Fig. 3 · p. 10
ht

因果 token 状态Causal token state

当前生成过程的工作状态,随 token 和尝试变化。它与跨片段保留的记忆 M 有不同的生命周期。The working state of generation changes across tokens and attempts. Its lifetime differs from memory M, which persists across segments.

Bn

片段工作区Segment workspace

汇总片段内已经发生的信息,容量有界。片段结束后,系统据此构造记忆候选。A bounded summary of information observed within a segment. After the segment ends, it supports the construction of memory candidates.

θ

慢变基础参数Slow base parameters

在预训练或离线训练中更新,与推理时可变的 Φ、M 分开管理。Updated during pretraining or offline training, and managed separately from Φ and M, which can change during reasoning.

03 / 实验Experiments

组件实验结果Component results

选自 v4 第 5–8 节的记忆管理、类型化传输与扩展上下文问答实验。Selected memory management, typed transport, and extended-context QA experiments from Sections 5–8 of v4.

01 / 记忆生命周期Memory lifecycle300/300

三个随机种子下,各 300 条留出记录均正确执行了生命周期操作。任务使用合成数据。All 300 held-out records passed the lifecycle checks for each of three seeds in synthetic tasks.

表 7 · 第 20 页Table 7 · p. 20
02 / 类型化传输Typed transport240/240

受控 anchor-filler 任务中,三个随机种子下各 240 条留出记录均完整匹配。All 240 held-out records matched exactly for each of three seeds in the controlled anchor-filler task.

表 9 · 第 25 页Table 9 · p. 25
03 / 多跳问答 F1Multi-hop QA F10.813–0.830

Qwen2.5-14B 在扩展上下文实验中的结果范围,覆盖三个选择器种子。The range across three selector seeds for Qwen2.5-14B in the extended-context experiment.

表 8 · 第 22 页Table 8 · p. 22

多跳问答 · 论文表 8Multi-hop QA · Table 8

问答 F1 与输入长度F1 and prompt length

查看原表View Table 8

Qwen2.5-14B · 扩展上下文(约 8,200 词)

方法MethodF1平均输入 tokensAvg. prompt tokens
完整上下文Full context0.71512,646
稠密检索Dense retrieval0.7531,338
BM250.7711,330
驻留缓存 + 归档回退Resident cache + archive fallback0.813–0.8301,368

展示表 8 的扩展上下文设置,含 300 道留出问题。F1 区间覆盖三个选择器种子;输入长度为完整提示词的平均 token 数。此处比较问答检索组件,归档、索引、检索延迟和能耗另行计量;完整设置见原表。The extended-context setting from Table 8, with 300 held-out questions. F1 ranges cover three selector seeds; token counts are averages for the complete prompt. This comparison covers the QA retrieval component. Archive storage, indexes, retrieval latency, and energy are separate costs; see the source table for full settings.

04 / 论文脉络The research program

从反馈到可复用记忆From feedback to reusable memory

五篇理论论文分别处理学习发生的时机、记忆的保存形式、写入的选择,以及两种状态如何协作。v4 把这些规则放进同一套架构。The five theory papers address when learning happens, how memory is represented, which writes to admit, and how the two states work together. v4 brings their rules into one architecture.

  • RTT 定义同一道题中的更新与再使用;片段整合规定何时可以回看已经发生的信息。RTT defines update and reuse within one problem; completed-segment consolidation defines when observed information can be reviewed.
  • 原子记忆行规定存储、版本和提交;预测式准入比较写入与保留原状的后续影响。Atomic Memory Rows defines storage, versions, and commits; Predictive Memory Admission compares the later consequences of writing and retaining the current state.
  • 双状态理论让策略与记忆可以分别读取、更新和干预,并分析两条路径的交互。Dual-state theory supports separate reads, updates, and interventions for policy and memory, and analyzes how the two paths interact.
阅读六章机制详解Read the six chapters

05 / 论文Publications

技术报告与配套论文Technical report and companion papers

v4 汇总整体架构、理论分析和实验结果。五篇配套论文分别讨论其中的具体问题。v4 brings together the architecture, theoretical analysis, and experimental results. Five companion papers develop individual parts of the theory.

配套理论论文Companion theory papers

五篇 R02 理论稿分别展开推理时训练、记忆表示、准入和整合。修订后的定义与结论以 v4 为准。Five R02 theory manuscripts develop reasoning-time training, memory representation, admission, and consolidation. Revised definitions and conclusions in v4 take precedence.

论文仓库Paper repository
01
POLICY / Φ

推理时训练Reasoning-Time Training

Reasoning-Time Training: Learning Before a Single Problem Ends

30 页30 pp.
02
MEMORY / M

原子记忆行Atomic Memory Rows

Atomic Memory Rows: A Bounded, Verifiable Substrate for Editable Reasoning

41 页41 pp.
03
ADMISSION

预测式记忆准入Predictive Memory Admission

Learning What to Remember: Predictive Admission for Bounded Reasoning-Time Memory

33 页33 pp.
04
DUAL STATE

推理时的双状态学习Dual-State Learning

MMLA-RTT: Dual-State Learning at Reasoning Time

31 页31 pp.
05
TIMING

完成片段整合Completed-Segment Consolidation

Causal Generation, Retrospective Consolidation: Completed-Segment Bidirectional Memory Without Temporal Leakage

31 页31 pp.

06 / 引用Citation

引用 MMLAv4Cite MMLAv4

BIBTEX / ARXIV:2606.28876V4
@article{zou2026mmla,
  title   = {{MMLA}: Memory-Mediated Learning Architecture for Predictive Dual-State Adaptation},
  author  = {Zou, Junyi and Donz, Avrova},
  year    = {2026},
  eprint  = {2606.28876},
  archivePrefix = {arXiv},
  primaryClass  = {cs.cl},
  note    = {Version 4},
  url     = {https://arxiv.org/abs/2606.28876v4}
}
MMLAv4 / ARCHITECTURE

双状态系统架构Dual-state system architecture

100%
MMLAv4 双状态架构图