Skip to content
Open
49 changes: 48 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -108,6 +108,53 @@ Anthropic direct (Claude) — drop `baseURL`, switch `provider`:

`embedding` is optional. When present, `dimensions` must match the Neo4j vector index dimension. For a fresh database, the plugin creates matching indexes during startup. If you change dimensions later, recreate the vector indexes or the Neo4j database.

### Memory decay (forgetting curve)

Each maintenance cycle scores every active node with a three-factor weighted model (recency + frequency + intrinsic) and bidirectionally transitions nodes across three tiers: `core` / `working` / `peripheral`. Nodes never get `status=deprecated` from decay — only manual deprecate / merge does that. Decay only adjusts `tier`, so all active nodes remain searchable.

The full formula, field mapping from the reference implementation, default-value rationale, and tuning guide live in **[`docs/decay.md`](docs/decay.md)**.

Minimal config (all fields optional, defaults shown):

```json
"decay": { "enabled": true }
```

Common overrides — for fuller control see `docs/decay.md` §4:

```json
"decay": {
"enabled": true,
"recencyHalfLifeDays": 30,
"peripheralCompositeThreshold": 0.15,
"workingAccessThreshold": 3
}
```

### Cron sessions

Sessions created by OpenClaw scheduled tasks can be configured independently of normal sessions. The host places the cron marker on the **sessionKey** (`sessionId` is a random UUID); real shapes are `cron:<jobId>`, `agent:<agentId>:cron:<jobId>`, or `agent:<agentId>:cron:<jobId>:run:<runId>`:

```json
"cron": {
"enabled": true,
"extract": true,
"finalizeAndMaintain": true
}
```

| Option | Default | Description |
| --- | --- | --- |
| `enabled` | `true` | Enable graph functionality inside cron sessions (recall injection + message buffering). When `false`, cron sessions skip automatic recall and message persistence; the `gm_*` tools remain available for explicit calls (manual escape hatch). |
| `extract` | `true` | Trigger knowledge extraction (LLM triples) in cron sessions via `afterTurn` / `compact`. When `false`, messages are still buffered and can be backfilled later with `openclaw graph-memory extract`. |
| `finalizeAndMaintain` | `true` | Run finalize (EVENT→SKILL promotion) and graph maintenance (decay / PageRank / communities) when a cron session ends. Disable when frequent cron runs make end-of-session global maintenance too costly. |

All three options default to **`true`**: cron sessions behave like normal sessions (recall, buffering, extraction, and end-of-session maintenance all enabled) unless explicitly disabled. `enabled: false` is the master switch — even with `extract` / `finalizeAndMaintain` set to `true`, nothing runs. Non-cron sessions are never affected by these options.

All three sub-options are optional; omitted fields keep the default `true` (e.g. with `"cron": { "extract": false }` only extraction is disabled — recall, buffering, and end-of-session maintenance stay on).

Caveat: when a cron job sets an explicit custom `sessionKey`, the host does not append the `cron` segment — such sessions cannot be detected and are treated as normal sessions.

### OAuth login (experimental)

```bash
Expand All @@ -127,7 +174,7 @@ conversation messages -> GmMessage nodes -> LLM triple extraction
-> embeddings -> vector recall + community expansion + GDS PPR
-> XML context injection

session end -> dedup -> global PageRank -> communities -> summaries
session end -> decay (forgetting curve) -> dedup -> global PageRank -> communities -> summaries
```

## Verify
Expand Down
22 changes: 22 additions & 0 deletions README_CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -81,6 +81,28 @@ bash setup-graph-memory-pro.sh --uninstall

`embedding` 可选。设置时,`dimensions` 必须与 Neo4j 向量索引维度一致。新数据库会在插件启动时按配置创建索引;更换维度后需要重建向量索引或 Neo4j 数据库。

### cron 会话行为控制

OpenClaw 定时任务创建的会话可以独立配置图谱行为。host 把 cron 标记放在 **sessionKey** 上(`sessionId` 是随机 UUID),实际形状为 `cron:<jobId>`、`agent:<agentId>:cron:<jobId>` 或 `agent:<agentId>:cron:<jobId>:run:<runId>`:

```json
"cron": {
"enabled": true,
"extract": true,
"finalizeAndMaintain": true
}
```

| 选项 | 默认 | 说明 |
| --- | --- | --- |
| `enabled` | `true` | 是否在 cron 会话内启用图谱功能(召回注入 + 消息入库)。关闭后 cron 会话不自动召回、不自动入库;`gm_*` 工具仍可手动调用(作为显式逃生通道)。 |
| `extract` | `true` | 是否在 cron 会话内触发知识提取(afterTurn / compact 的 LLM 三元组提取)。关闭后消息仍入库缓冲,之后可用 `openclaw graph-memory extract` 手动回填。 |
| `finalizeAndMaintain` | `true` | cron 会话结束时是否执行 finalize(EVENT→SKILL 晋升)和图维护(decay / PageRank / 社区检测)。定时任务频繁时可关闭,避免每次会话结束都跑全局维护。 |

三个选项**默认全部开启**:cron 会话默认使用图谱,需按需显式关闭。`enabled=false` 是总开关:即使 `extract`/`finalizeAndMaintain` 设为 `true` 也不生效。非 cron 会话不受这些选项影响。三个子项均可省略,未写的字段取默认值 `true`。

注意:若 cron 任务显式设置了自定义 `sessionKey`,host 不再附加 `cron` 段,此类会话无法被识别,将按普通会话处理。

### OAuth 登录(实验性)

```bash
Expand Down
175 changes: 175 additions & 0 deletions docs/decay.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,175 @@
# Memory Decay — 柔性评分模型

graph-memory-pro 的衰减机制采用**三因子加权评分 + tier 双向转换**,参考 [memory-lancedb-pro](https://github.com/CortexReach/memory-lancedb-pro) 的设计并映射到本仓库的图模型信号。

- **decay 不动 `status`**——只调整 `tier`(`core` / `working` / `peripheral`)。`status=deprecated` 仅由手动弃用(`gm_update mode=deprecate` / merge)触发。
- 每次 `gm_maintain` 或 `session_end` 维护的第 0 步执行:扫描所有 active 节点 → 评分 → tier 转换 → 写回 `decayScore` / `tier` / `decayComputedAt`。
- 评分结果可通过 `gm_stats` / CRUD API 查看;外层搜索目前**不读 decayScore 排序**(已由 PageRank + tier 隐含分层)。

---

## 1. 评分公式

```
composite = wR · recency + wF · frequency + wI · intrinsic
```

三个权重默认 `0.4 / 0.3 / 0.3`,**推荐**和为 1。运行时若和≠1 会自动按比例归一化(`wR' = wR / (wR+wF+wI)`),保证 `composite ∈ [0,1]`,避免用户覆盖单个权重导致评分越界。归一化在 `scoreNode()` 内进行,原始 `cfg.*Weight` 值不被修改。

### 1.1 Recency(时间衰减,权重 0.4)

Weibull 拉伸指数:

```
recency = exp( −λ · daysSinceLastAccess^β )

λ = ln(2) / effectiveHL
effectiveHL = recencyHalfLifeDays · exp( importanceModulation · importance )
```

- **半衰期调制**:重要记忆(高 `importance`)的 `effectiveHL` 更大 → 衰减更慢。对应艾宾浩斯曲线"重要事件保留更久"。
- **tier-β**:曲线形状随 tier 变化,反馈式调整衰减速度:

| tier | β | 效果 |
|---|---|---|
| `core` | 0.8 | 尾部衰减缓(核心知识保得久) |
| `working` | 1.0 | 标准指数衰减 |
| `peripheral` | 1.3 | 加速衰减(边缘知识更快被遗忘) |

### 1.2 Frequency(访问频率,权重 0.3)

```
frequency = base · ( 0.5 + 0.5 · recentnessBonus )

base = 1 − exp( −validatedCount / 5 )
recentnessBonus = exp( −avgAccessGapDays / 30 ) # 仅当 validatedCount > 1
avgAccessGapDays = ( lastAccessedAt − createdAt ) / ( validatedCount − 1 )
```

- 用 `validatedCount`(LLM 重新提取的次数)替代 lancedb-pro 的 `accessCount`(manual recall 触发的次数)。前者是更强的"重新确认"信号。
- `validatedCount ≤ 1` 时跳过 `recentnessBonus`,只返回 `base`(无法算平均间隔)。

### 1.3 Intrinsic(内在价值,权重 0.3)

```
intrinsic = importance · confidence

importance = pagerank / maxPagerank # 每次扫描时按当前批次归一化到 [0,1]
confidence = 1 − 1 / ( 1 + validatedCount ) # 饱和函数,收敛到 1
```

---

## 2. 字段映射(lancedb-pro → graph-memory-pro)

| lancedb-pro 字段 | 本仓库替代 | 说明 |
|---|---|---|
| `accessCount` | `validatedCount` | LLM 重新提取次数(强信号,原为 manual recall 触发) |
| `lastAccessedAt` | `lastAccessedAt` | 由 `upsertNode` 在任意写入路径刷新(重新提取、`gm_record`、`gm_update`、CRUD POST)。`mergeNodes` 故意不刷新(合并 ≠ 用户重新激活) |
| `importance` | `pagerank / maxPagerank` | 图结构重要性,每次扫描归一化 |
| `confidence` | `1 − 1/(1+validatedCount)` | 饱和置信度 |
| `tier` | `tier`(新增字段) | 与 `status` 正交 |

---

## 3. Tier 双向转换

| 转换 | 条件 |
|---|---|
| **core → working** | `composite < peripheralCompositeThreshold` **AND** `count < workingAccessThreshold` |
| **working → peripheral** | `composite < peripheralCompositeThreshold` **OR**(`ageDays > peripheralAgeDays` **AND** `count < workingAccessThreshold`) |
| **peripheral → working** | `count >= workingAccessThreshold` **AND** `composite >= workingCompositeThreshold` |
| **working → core** | `count >= coreAccessThreshold` **AND** `composite >= coreCompositeThreshold` **AND** `importance >= coreImportanceThreshold` |

- 新节点默认 `tier = "working"`。
- 节点保持 `status = active` 不变;tier 变化时仅更新 `updatedAt`,不改变搜索过滤行为。
- 不存在的"core→peripheral"和"peripheral→core"由两次相邻转换实现(经过 working)。

---

## 4. 默认值与调参指南

### 4.1 默认配置

```json
{
"decay": {
"enabled": true,
"recencyHalfLifeDays": 30,
"recencyWeight": 0.4,
"importanceModulation": 1.5,
"frequencyWeight": 0.3,
"intrinsicWeight": 0.3,
"betaCore": 0.8,
"betaWorking": 1.0,
"betaPeripheral": 1.3,
"coreAccessThreshold": 10,
"coreCompositeThreshold": 0.7,
"coreImportanceThreshold": 0.8,
"peripheralCompositeThreshold": 0.15,
"peripheralAgeDays": 60,
"workingAccessThreshold": 3,
"workingCompositeThreshold": 0.4
}
}
```

### 4.2 数值来源

| 参数 | 默认值 | 来源 |
|---|---|---|
| `recencyHalfLifeDays` | 30 | 艾宾浩斯曲线 ~25% 保留率拐点;同时与 lancedb-pro 的 `recencyHalfLifeDays` + `ACCESS_DECAY_HALF_LIFE_DAYS` 一致 |
| `importanceModulation` | 1.5 | lancedb-pro:`effectiveHL = 30 · exp(1.5 · importance)`,importance=1 时半衰期延长到 ~134 天 |
| `betaCore/Working/Peripheral` | 0.8 / 1.0 / 1.3 | lancedb-pro Weibull 形状参数 |
| 7 个 tier 转换阈值 | — | lancedb-pro `tier-manager` 默认值 |
| `recencyWeight / frequencyWeight / intrinsicWeight` | 0.4 / 0.3 / 0.3 | lancedb-pro 三因子权重,和为 1 |
| `validatedCount` 分母 | 5 | lancedb-pro 的 `1 − exp(−count/5)` 基础频率项(未改) |

### 4.3 常见调参场景

| 想要的效果 | 调整方向 |
|---|---|
| 记忆整体保留更久 | 调高 `recencyHalfLifeDays`(如 60)或调低 `peripheralCompositeThreshold`(更难降级) |
| 更激进遗忘 | 调低 `recencyHalfLifeDays`(如 14)或调高 `peripheralCompositeThreshold` |
| 重要知识显著保得久 | 调高 `importanceModulation`(半衰期调制更强) |
| 核心知识不易降级 | 调低 `betaCore`(更缓的尾部)或调高 `coreCompositeThreshold`(更难升 core,留在 working 也保得久) |
| 单次曝光更易遗忘 | 调高 `workingAccessThreshold`(promote 到 working 需要更多确认) |
| 永久禁用衰减 | `"enabled": false` |

### 4.4 与原布尔阈值方案的对照(向后兼容)

旧版本(`maxAgeDays` + `minCalls`)的布尔规则已被这套柔性评分取代。原默认值 `maxAgeDays=30, minCalls=2` 在新模型下大致对应于:

- 一个 `validatedCount=1`、`tier=working`、低 pagerank 的节点,约 30 天后 `recency` 跌破 0.15 → `composite` 跌破 `peripheralCompositeThreshold` → demote 到 `peripheral`。
- 关键差别:新模型**不会 deprecate**,只是降到 `peripheral` tier,搜索过滤仍包含它(只是 decayScore 较低)。

---

## 5. 数据库字段

| 字段 | 类型 | 写入者 | 说明 |
|---|---|---|---|
| `tier` | string | `applyDecay` / `upsertNode`(创建时初始化为 `working`) | `core` / `working` / `peripheral` |
| `lastAccessedAt` | int (epoch ms) | `upsertNode`(重新提取时) | decay 评分的时间基准 |
| `decayScore` | float (0~1) | `applyDecay` | 最近一次评分结果 |
| `decayComputedAt` | int (epoch ms) | `applyDecay` | 评分时间戳 |

旧节点缺这些字段时:
- `tier` 缺失 → 评分按 `working` 处理;首次 `applyDecay` 时自动写入 `working`
- `lastAccessedAt` 缺失 → 回退到 `updatedAt` / `createdAt`
- `decayScore` / `decayComputedAt` 缺失 → 在首次 `applyDecay` 前为 undefined,不影响评分

**Backfill 时机**:新字段在第一次 `applyDecay` 运行时为每个 active 节点批量写入。如果部署初始用 `decay.enabled=false`,字段会一直缺失直到切换为 `true` 后的第一次维护周期。在切换前的窗口期,对 raw DB 直接做 `tier` 过滤查询会返回 null/missing 而非 `"working"`——目前搜索路径不读 `tier`,但自定义查询需要留意。

---

## 6. 实现位置

| 文件 | 内容 |
|---|---|
| `src/graph/decay.ts` | 评分函数 + tier 决策 + `applyDecay()` 批处理 |
| `src/types.ts` | `DecayConfig` 接口、`NodeTier` 类型、`GmNode` 新字段、`DEFAULT_CONFIG.decay` |
| `src/store/store.ts` | `toNode` 字段映射、`upsertNode` 初始化 `tier` / `lastAccessedAt` |
| `src/graph/maintenance.ts` | 调用入口(step 0) |
| `test/decay.test.ts` | 评分函数 + tier 决策纯函数单元测试 |
| `openclaw.plugin.json` | 用户可见的配置 schema |
Loading