diff --git a/.gitignore b/.gitignore
index dd3ca40..141eb19 100644
--- a/.gitignore
+++ b/.gitignore
@@ -71,9 +71,14 @@ env/
# local storage
api/storage/
+apps/api/storage/
# agent memory (runtime data, user-specific — not for version control)
api/data/memory/
+apps/api/data/memory/
+
+# local runtime caches
+apps/api/tmp/
# daily rotating log files
api/logs/
diff --git a/apps/api/storage/notebooks/9a2371fa-8a99-4436-b330-f35bf43c893c/3acf8dfe-1723-4896-a9e6-23a547d5737a.md b/apps/api/storage/notebooks/9a2371fa-8a99-4436-b330-f35bf43c893c/3acf8dfe-1723-4896-a9e6-23a547d5737a.md
deleted file mode 100644
index 901fdac..0000000
--- a/apps/api/storage/notebooks/9a2371fa-8a99-4436-b330-f35bf43c893c/3acf8dfe-1723-4896-a9e6-23a547d5737a.md
+++ /dev/null
@@ -1,3 +0,0 @@
-# test
-
-this is a tiny ingestion regression check
\ No newline at end of file
diff --git a/apps/api/storage/notebooks/9a2371fa-8a99-4436-b330-f35bf43c893c/41b24de4-9d20-48d8-bc3c-9fe8fdeb4934.md b/apps/api/storage/notebooks/9a2371fa-8a99-4436-b330-f35bf43c893c/41b24de4-9d20-48d8-bc3c-9fe8fdeb4934.md
deleted file mode 100644
index 1822a81..0000000
--- a/apps/api/storage/notebooks/9a2371fa-8a99-4436-b330-f35bf43c893c/41b24de4-9d20-48d8-bc3c-9fe8fdeb4934.md
+++ /dev/null
@@ -1,7 +0,0 @@
-# Codex Ingestion Smoke Test
-
-This source is used to verify the pending -> processing -> indexed pipeline.
-
-- queue routing
-- after commit dispatch
-- worker consumption
diff --git a/apps/api/storage/notebooks/9a2371fa-8a99-4436-b330-f35bf43c893c/72708d3a-a42e-4a96-9fe2-28081462ed31.md b/apps/api/storage/notebooks/9a2371fa-8a99-4436-b330-f35bf43c893c/72708d3a-a42e-4a96-9fe2-28081462ed31.md
deleted file mode 100644
index 1822a81..0000000
--- a/apps/api/storage/notebooks/9a2371fa-8a99-4436-b330-f35bf43c893c/72708d3a-a42e-4a96-9fe2-28081462ed31.md
+++ /dev/null
@@ -1,7 +0,0 @@
-# Codex Ingestion Smoke Test
-
-This source is used to verify the pending -> processing -> indexed pipeline.
-
-- queue routing
-- after commit dispatch
-- worker consumption
diff --git a/apps/api/storage/notebooks/9a2371fa-8a99-4436-b330-f35bf43c893c/85239377-83ad-478b-b819-f36d894bb9e3.pdf b/apps/api/storage/notebooks/9a2371fa-8a99-4436-b330-f35bf43c893c/85239377-83ad-478b-b819-f36d894bb9e3.pdf
deleted file mode 100644
index e1571d1..0000000
Binary files a/apps/api/storage/notebooks/9a2371fa-8a99-4436-b330-f35bf43c893c/85239377-83ad-478b-b819-f36d894bb9e3.pdf and /dev/null differ
diff --git a/apps/api/storage/notebooks/9a2371fa-8a99-4436-b330-f35bf43c893c/a9b726fd-1b6f-4444-a297-0dbdb73fdd3c.md b/apps/api/storage/notebooks/9a2371fa-8a99-4436-b330-f35bf43c893c/a9b726fd-1b6f-4444-a297-0dbdb73fdd3c.md
deleted file mode 100644
index 1822a81..0000000
--- a/apps/api/storage/notebooks/9a2371fa-8a99-4436-b330-f35bf43c893c/a9b726fd-1b6f-4444-a297-0dbdb73fdd3c.md
+++ /dev/null
@@ -1,7 +0,0 @@
-# Codex Ingestion Smoke Test
-
-This source is used to verify the pending -> processing -> indexed pipeline.
-
-- queue routing
-- after commit dispatch
-- worker consumption
diff --git a/apps/api/storage/notebooks/9a2371fa-8a99-4436-b330-f35bf43c893c/a9e24316-a846-4e3c-8b20-52234859609e.md b/apps/api/storage/notebooks/9a2371fa-8a99-4436-b330-f35bf43c893c/a9e24316-a846-4e3c-8b20-52234859609e.md
deleted file mode 100644
index d3b8c27..0000000
--- a/apps/api/storage/notebooks/9a2371fa-8a99-4436-b330-f35bf43c893c/a9e24316-a846-4e3c-8b20-52234859609e.md
+++ /dev/null
@@ -1,240 +0,0 @@
-# LyraNote 本科毕设全局设计说明(论文版)
-
-> 文档用途:用于本科毕业论文中的“研究背景、系统设计、技术特点、关键难点与创新点”章节初稿。
->
-> 版本日期:2026-03-26
->
-> 适用对象:论文作者、答辩评审、项目协作者
-
----
-
-## 摘要
-
-LyraNote 是一个面向个人研究与知识管理场景的 AI 驱动系统,目标是将传统“知识存储工具”升级为“可持续协作的研究伙伴”。系统围绕“私有知识优先”的设计原则,整合了 RAG 检索增强、多 Agent 编排、长期记忆、主动感知、深度研究与可视化生成界面(GenUI)等能力,形成从信息采集、理解、生成到反馈演化的闭环。
-
-与常见 AI 笔记产品相比,LyraNote 不仅提供被动问答,还重点探索 AI 的主动性与持续性:在用户未显式提问时,系统仍可基于上下文状态进行低干扰主动辅助;在长周期任务中,系统支持定时调度与自动投递;在多轮交互中,系统通过多层记忆与画像机制持续提升个性化质量。该系统具有较强工程实现价值与研究讨论价值,适合作为“智能知识系统 + Agent 工程化”方向的本科毕设课题。
-
----
-
-## 1. 研究背景与意义
-
-### 1.1 研究背景
-
-当前主流知识管理工具(笔记类应用、文档类应用)普遍存在两类不足:
-
-1. 知识层面:多来源信息可存储但难以深度理解,跨文档关联弱,检索结果与用户真实任务脱节。
-2. AI 层面:大多停留在“你问我答”的被动模式,缺乏持续状态感知、长期个性化与任务自动执行能力。
-
-与此同时,通用大模型虽然具备强生成能力,但在真实研究流程中仍面临幻觉、上下文窗口限制、长任务中断、可解释性不足、部署复杂等问题。基于此,LyraNote 以“面向研究场景的 AI 原生第二大脑”为目标,尝试把 LLM 能力转化为可持续、可控、可落地的系统能力。
-
-### 1.2 研究意义
-
-1. 应用意义:提高个人研究与学习效率,减少“搜集-整理-输出”的重复劳动。
-2. 工程意义:构建可扩展的 AI 系统架构,验证多 Agent、记忆系统、事件驱动链路在真实产品中的可行性。
-3. 学术意义:为“从被动对话到主动协作”的 AI Agent 产品化路径提供可复用方法与实验设计。
-
----
-
-## 2. 研究目标与问题定义
-
-### 2.1 总体目标
-
-构建一个支持“知识摄取—语义检索—深度研究—结构化表达—长期演化”的 AI 研究笔记系统,实现 AI 从工具属性向伙伴属性的演进。
-
-### 2.2 核心研究问题
-
-1. 如何让 AI 回答建立在用户私有知识之上,而不是泛化互联网知识?
-2. 如何在复杂任务中提升质量与稳定性,避免单 Agent 架构的能力拥挤?
-3. 如何让 AI 在不打扰用户的前提下实现主动感知与主动辅助?
-4. 如何在工程上实现可扩展、可维护、可自托管的系统架构?
-
----
-
-## 3. 系统总体架构设计
-
-### 3.1 分层总体架构
-
-LyraNote 采用“前端交互层 + 后端服务层 + AI 能力层 + 基础设施层”的分层设计:
-
-1. 前端交互层(Next.js + Tiptap + SSE)
-用于笔记编辑、对话交互、研究进度可视化、结构化内容展示。
-2. 后端服务层(FastAPI)
-负责路由编排、鉴权、业务规则与 API 暴露。
-3. AI 能力层(Agents + Skills + Memory)
-负责检索、写作、深度研究、工具调用、记忆更新与主动洞察。
-4. 基础设施层(PostgreSQL/pgvector + Redis + Celery + S3 兼容存储)
-负责结构化数据、向量索引、异步任务、对象存储与调度。
-
-### 3.2 关键数据闭环
-
-系统形成“源→知识→对话→笔记→再入库”的闭环:
-
-1. 用户导入 PDF/网页/Markdown。
-2. 后端异步完成解析、分块、向量化并入库(RAG 索引)。
-3. 对话阶段基于检索证据生成回答并附引用。
-4. 用户将结果沉淀为笔记、摘要或报告(Artifact)。
-5. 笔记可再次作为来源进入知识库(Note-as-Source),形成持续演化。
-
-### 3.3 AI 核心架构
-
-#### 3.3.1 多 Agent 协作架构(System A)
-
-采用“主 Agent(Orchestrator)+ 专家 Agent(RAG/Research/Writing/Memory/Web)”的两层决策:
-
-1. 主 Agent 负责任务级路由(把问题交给最合适的专家)。
-2. 专家 Agent 负责步骤级执行(检索、搜索、写作、记忆更新等)。
-
-该设计缓解了单 ReAct 循环能力拥挤、复杂任务步骤不足的问题。
-
-#### 3.3.2 Soul 持续思维与表达机制(System B)
-
-通过活动感知、后台思维循环和表达门控机制,让 AI 在用户静默期仍可进行低频高价值思考,并在合适时机以可忽略的方式推送洞察。
-
-#### 3.3.3 用户画像与记忆体系(System C)
-
-系统从“碎片记忆”升级到“多层记忆 + 用户画像”:
-
-1. 记忆层存储偏好、事实、场景与反思信息。
-2. 画像层按周期进行综合,形成对用户研究方向、能力水平、表达偏好的立体表征。
-3. 画像结果回注到路由决策与回答策略,实现长期个性化。
-
----
-
-## 4. 核心特点(面向论文阐述)
-
-### 4.1 私有知识优先的 RAG 管道
-
-系统在 Query 改写、多路检索、混合排序、重排与去重环节进行优化,重点解决多轮指代、召回不足和内容冗余问题。根据优化设计预期:
-
-1. 召回率提升约 25%~35%。
-2. 精确率提升约 20%。
-3. 首次响应延迟下降约 30%。
-
-### 4.2 深度研究(Deep Research)能力
-
-区别于单轮问答,深度研究支持“规划→检索→递归扩展→综合成文”的多阶段流程,支持 Quick/Deep 两种预算模式,并通过 SSE 反馈研究过程。为提升鲁棒性,研究任务与连接流解耦,支持刷新恢复与断点续传。
-
-### 4.3 GenUI 结构化表达能力
-
-通过统一 `genui` 协议将模型输出映射为可交互组件(图表、表格、时间轴、矩阵、看板等),提升研究结果可读性与可操作性,使 AI 输出从“纯文本”升级为“可视化知识单元”。
-
-### 4.4 主动感知与自动化执行
-
-系统提出“微交互渗透→上下文智能→自主行动”的三层主动架构:
-
-1. 在索引完成、编辑停顿、页面切换等时机提供低打扰提示。
-2. 通过定时任务实现持续任务自动执行(采集、生成、投递)。
-3. 将 AI 从一次性交互扩展为长期运行助手。
-
-### 4.5 可自托管与工程可扩展性
-
-1. 认证从外部 Clerk 迁移到本地单用户方案(bcrypt + JWT + 初始化向导),降低部署门槛。
-2. 存储层引入 `StorageProvider` 抽象,统一本地与 S3 兼容后端,支持横向扩展与多环境部署。
-3. 后端引入分层重构思路(Router/Service/Modules/Infrastructure),提升可维护性与测试性。
-
----
-
-## 5. 关键难点与解决思路
-
-### 难点 1:检索质量与多轮语义理解
-
-问题:用户问题口语化、上下文省略、语义跨度大,导致检索召回不稳定。
-解决:引入对话感知 Query 改写、多 Query 并行、向量+全文混合检索、MMR 去重与 Cross-Encoder 重排。
-
-### 难点 2:深度研究任务长链路稳定性
-
-问题:若研究流程与 SSE 连接强耦合,用户刷新页面会导致任务中断。
-解决:任务后台化(独立运行)+ 事件缓冲 + 状态查询,支持断点续传与结果恢复。
-
-### 难点 3:上下文窗口与成本控制
-
-问题:深度研究在多轮检索后易超出上下文窗口,且联网搜索有调用成本。
-解决:学习结果摘要压缩、结果数量上限、检索阈值触发 Web Search、关键词去重与模式化预算(Quick/Deep)。
-
-### 难点 4:主动性与用户体验平衡
-
-问题:主动提示过多会造成干扰。
-解决:采用“适时、适量、可忽略”原则,限制触发频率与展示密度,允许用户关闭或忽略。
-
-### 难点 5:系统复杂度持续上升
-
-问题:功能扩展导致路由臃肿、职责混乱、可维护性下降。
-解决:推进 Service 分层、模块化 Agents、统一 Provider 接口、Schema 规范化,降低耦合度。
-
-### 难点 6:自托管部署与多环境兼容
-
-问题:文件存储、认证依赖、异步任务在不同部署环境中易出兼容性问题。
-解决:认证本地化、存储抽象化、任务队列标准化(Celery + Redis)、配置集中化管理。
-
----
-
-## 6. 论文可主张的创新点
-
-1. 提出并实现“多 Agent 协作 + Soul 持续感知 + 用户画像演化”的三系统耦合架构。
-2. 将 AI 交互从“被动问答”扩展到“主动感知与任务执行”的产品形态。
-3. 在研究报告场景中实现 Deep Research 的工程化闭环(多阶段推理、可恢复任务、可复用产物)。
-4. 提出并落地面向生成式应用的统一 GenUI 协议,增强 LLM 输出可视化表达能力。
-5. 结合记忆与技能系统,实现“能力可插拔 + 个性可演化”的长期 Agent 机制。
-
----
-
-## 7. 论文实验与评估建议(可直接写入实验章节)
-
-### 7.1 建议评估维度
-
-1. 检索效果:Recall@K、Precision@K、nDCG。
-2. 回答质量:事实一致性、引用覆盖率、结构完整度。
-3. 任务能力:复杂任务完成率、平均完成时长、失败重试率。
-4. 主动交互:建议点击率、忽略率、用户主观满意度。
-5. 个性化效果:开启/关闭记忆后的质量差异、长期 `quality_score` 变化。
-6. 工程性能:首 token 延迟、平均响应时延、任务吞吐量、资源消耗。
-
-### 7.2 建议对比实验
-
-1. 单 Agent vs 多 Agent(复杂任务完成率对比)。
-2. 无记忆 vs V1 记忆 vs V2 记忆(个性化质量对比)。
-3. 被动模式 vs 主动模式(交互效率与满意度对比)。
-4. 纯文本输出 vs GenUI 输出(信息理解效率与可用性对比)。
-
----
-
-## 8. 局限性与后续工作
-
-1. 深度研究仍依赖外部搜索服务,成本与稳定性受第三方接口影响。
-2. 多 Agent 路由效果对 Prompt 与结构化输出稳定性敏感。
-3. 主动机制目前主要面向单用户场景,多用户协作机制有待拓展。
-4. 评估体系仍需更大规模用户实验来验证泛化性。
-
-后续可重点推进:
-
-1. 更细粒度的路由评估与自动纠偏机制。
-2. 更强的多模态研究能力(图像/表格/代码执行)。
-3. 研究任务模板化与可复现实验流水线。
-4. 面向团队协作的多租户与权限机制。
-
----
-
-## 9. 文档依据(可作为论文内部参考来源)
-
-1. `architecture.md`(系统总架构、AI 流水线)
-2. `lyra-soul-system.md`(多 Agent + Soul + 画像)
-3. `memory-system-v2.md`(五层记忆体系)
-4. `skills-system.md`(可插拔技能架构)
-5. `deep-research.md`(多阶段研究流程与任务解耦)
-6. `rag-optimization.md`(检索优化链路)
-7. `genui-integration.md`(GenUI 协议)
-8. `proactive-ai-system.md`(主动感知体系)
-9. `scheduled-tasks.md`(定时任务自动化)
-10. `backend-architecture-refactor.md`(后端分层重构)
-11. `storage-system.md`(存储抽象)
-12. `single-user-auth.md`(单用户自托管认证)
-
----
-
-## 附:论文写作使用建议
-
-1. 可将第 1-3 节作为“绪论 + 总体设计”。
-2. 可将第 4-5 节拆为“关键技术实现”。
-3. 可将第 6-7 节作为“创新点与实验设计”。
-4. 可将第 8 节作为“总结与展望”。
-
diff --git a/patch_home.js b/patch_home.js
deleted file mode 100644
index 96c1888..0000000
--- a/patch_home.js
+++ /dev/null
@@ -1,140 +0,0 @@
-const fs = require("fs");
-const path = require("path");
-
-const filePath = path.join(process.cwd(), "apps/desktop/src/pages/home/home-page.tsx");
-let content = fs.readFileSync(filePath, "utf-8");
-
-// replace SUGGESTIONS
-content = content.replace(
- /const SUGGESTIONS = \[\s*"分析知识库核心主题",\s*"生成结构化研究摘要",\s*"对比不同来源观点",\s*"根据笔记制定学习计划",\s*\]/g,
- `const SUGGESTIONS = [
- { icon: , text: "帮我分析知识库中的核心主题" },
- { icon: , text: "为我的研究生成一份结构化摘要" },
- { icon: , text: "对比不同来源中的相似观点" },
- { icon: , text: "根据笔记内容生成学习计划" },
-]`
-);
-
-// replace render
-content = content.replace(
- /return \(\s*
[\s\S]*?\)\n\}/,
- \`return (
-
- {/* Title Area */}
-
-
-

-
-
-
- 有什么我可以帮你的?
-
-
- 基于你的知识库,我可以帮你分析、总结和探索任何内容
-
-
-
-
- {/* Suggestions Grid */}
-
- {SUGGESTIONS.map((s, idx) => (
- handleSuggestion(s.text)}
- className="flex items-center gap-2.5 px-4 py-3.5 rounded-xl text-left text-[12.5px] transition-colors"
- style={{
- background: "rgba(255,255,255,0.015)",
- border: "1px solid rgba(255,255,255,0.05)",
- color: "var(--color-text-secondary)",
- }}
- >
- {s.icon}
- {s.text}
-
- ))}
-
-
- {/* Input card */}
-
- textareaRef.current?.focus()}
- >
-
-
-
- )
-}\`
-);
-
-fs.writeFileSync(filePath, content, "utf-8");
diff --git a/scripts/create_db_table_design_docx.py b/scripts/create_db_table_design_docx.py
deleted file mode 100644
index 95607a8..0000000
--- a/scripts/create_db_table_design_docx.py
+++ /dev/null
@@ -1,470 +0,0 @@
-from pathlib import Path
-
-from docx import Document
-from docx.enum.table import WD_CELL_VERTICAL_ALIGNMENT, WD_ROW_HEIGHT_RULE, WD_TABLE_ALIGNMENT
-from docx.enum.text import WD_ALIGN_PARAGRAPH
-from docx.oxml import OxmlElement
-from docx.oxml.ns import qn
-from docx.shared import Cm, Mm, Pt
-
-
-HEADERS = ["字段名称", "类型", "长度", "字段说明"]
-COLUMN_WIDTHS_CM = (4.2, 2.1, 1.6, 7.1)
-PRIMARY_OUTPUT_DOCX = Path("output/doc/LyraNote数据库表设计.docx")
-VERSIONED_OUTPUT_DOCX = Path("output/doc/LyraNote数据库表设计-三线表-最新模型.docx")
-
-TABLES = [
- (
- "表3-1 用户信息表(users)",
- [
- ("id", "UUID", "36", "用户主键"),
- ("username", "varchar", "255", "用户名"),
- ("password_hash", "varchar", "255", "密码哈希"),
- ("email", "varchar", "255", "邮箱"),
- ("name", "varchar", "255", "用户姓名/昵称"),
- ("avatar_url", "text", "-", "用户头像地址"),
- ("google_id", "varchar", "255", "Google 登录标识"),
- ("github_id", "varchar", "255", "GitHub 登录标识"),
- ("oauth_unbound", "varchar", "64", "已解绑第三方登录标记"),
- ("created_at", "datetime", "-", "创建时间"),
- ],
- ),
- (
- "表3-2 应用配置表(app_config)",
- [
- ("key", "varchar", "255", "配置项主键"),
- ("value", "text", "-", "配置项值"),
- ],
- ),
- (
- "表3-3 笔记本信息表(notebooks)",
- [
- ("id", "UUID", "36", "笔记本主键"),
- ("user_id", "UUID", "36", "所属用户 ID"),
- ("title", "varchar", "500", "笔记本标题"),
- ("description", "text", "-", "笔记本描述"),
- ("status", "varchar", "50", "笔记本状态"),
- ("is_global", "boolean", "-", "是否全局知识库"),
- ("is_system", "boolean", "-", "是否系统笔记本"),
- ("system_type", "varchar", "50", "系统笔记本类型"),
- ("source_count", "integer", "-", "来源数量"),
- ("is_public", "boolean", "-", "是否公开"),
- ("published_at", "datetime", "-", "发布时间"),
- ("cover_emoji", "varchar", "10", "封面表情"),
- ("cover_gradient", "varchar", "50", "封面渐变样式"),
- ("created_at", "datetime", "-", "创建时间"),
- ("updated_at", "datetime", "-", "更新时间"),
- ],
- ),
- (
- "表3-4 笔记本摘要表(notebook_summaries)",
- [
- ("notebook_id", "UUID", "36", "所属笔记本 ID"),
- ("summary_md", "text", "-", "Markdown 摘要内容"),
- ("key_themes", "json", "-", "关键主题列表"),
- ("last_synced_at", "datetime", "-", "最近同步时间"),
- ],
- ),
- (
- "表3-5 知识来源表(sources)",
- [
- ("id", "UUID", "36", "来源主键"),
- ("notebook_id", "UUID", "36", "所属笔记本 ID"),
- ("title", "varchar", "500", "来源标题"),
- ("type", "varchar", "50", "来源类型"),
- ("status", "varchar", "50", "索引状态"),
- ("file_path", "text", "-", "本地文件路径"),
- ("url", "text", "-", "网页地址"),
- ("raw_text", "text", "-", "抽取后的原始文本"),
- ("summary", "text", "-", "来源摘要"),
- ("storage_key", "varchar", "500", "对象存储键"),
- ("storage_backend", "varchar", "20", "存储后端类型"),
- ("created_at", "datetime", "-", "创建时间"),
- ("updated_at", "datetime", "-", "更新时间"),
- ],
- ),
- (
- "表3-6 知识分块表(chunks)",
- [
- ("id", "UUID", "36", "分块主键"),
- ("source_id", "UUID", "36", "所属来源 ID"),
- ("notebook_id", "UUID", "36", "所属笔记本 ID"),
- ("content", "text", "-", "分块正文"),
- ("chunk_index", "integer", "-", "分块顺序号"),
- ("embedding", "vector", "-", "向量表示"),
- ("token_count", "integer", "-", "Token 数量"),
- ("metadata", "json", "-", "分块元数据"),
- ("created_at", "datetime", "-", "创建时间"),
- ("source_type", "varchar", "20", "分块来源类型"),
- ("note_id", "UUID", "36", "关联笔记 ID"),
- ],
- ),
- (
- "表3-7 对话信息表(conversations)",
- [
- ("id", "UUID", "36", "对话主键"),
- ("notebook_id", "UUID", "36", "关联笔记本 ID"),
- ("user_id", "UUID", "36", "所属用户 ID"),
- ("title", "varchar", "500", "对话标题"),
- ("source", "varchar", "20", "对话来源"),
- ("created_at", "datetime", "-", "创建时间"),
- ],
- ),
- (
- "表3-8 消息信息表(messages)",
- [
- ("id", "UUID", "36", "消息主键"),
- ("conversation_id", "UUID", "36", "所属对话 ID"),
- ("generation_id", "UUID", "36", "关联生成任务 ID"),
- ("role", "varchar", "50", "消息角色"),
- ("status", "varchar", "20", "消息状态"),
- ("content", "text", "-", "消息正文"),
- ("reasoning", "text", "-", "推理内容"),
- ("citations", "json", "-", "引用信息"),
- ("agent_steps", "json", "-", "Agent 执行步骤"),
- ("attachments", "json", "-", "附件信息"),
- ("speed", "json", "-", "流式速度指标"),
- ("mind_map", "json", "-", "思维导图结果"),
- ("diagram", "json", "-", "图形结果"),
- ("mcp_result", "json", "-", "MCP 工具结果"),
- ("ui_elements", "json", "-", "生成式界面元素"),
- ("parent_message_id", "UUID", "36", "父消息 ID"),
- ("created_at", "datetime", "-", "创建时间"),
- ],
- ),
- (
- "表3-9 对话摘要表(conversation_summaries)",
- [
- ("conversation_id", "UUID", "36", "所属对话 ID"),
- ("summary_text", "text", "-", "压缩后的对话摘要"),
- ("compressed_message_count", "integer", "-", "已压缩消息数量"),
- ("compressed_through", "datetime", "-", "摘要覆盖截止时间"),
- ("updated_at", "datetime", "-", "更新时间"),
- ],
- ),
- (
- "表3-10 深度研究任务表(research_tasks)",
- [
- ("id", "UUID", "36", "任务主键"),
- ("user_id", "UUID", "36", "所属用户 ID"),
- ("notebook_id", "varchar", "100", "关联笔记本 ID"),
- ("conversation_id", "UUID", "36", "关联对话 ID"),
- ("query", "text", "-", "研究问题"),
- ("mode", "varchar", "20", "研究模式"),
- ("status", "varchar", "20", "任务状态"),
- ("report", "text", "-", "研究报告"),
- ("deliverable_json", "json", "-", "结构化交付物"),
- ("timeline_json", "json", "-", "过程时间线"),
- ("web_sources_json", "json", "-", "外部资料列表"),
- ("error_message", "text", "-", "错误信息"),
- ("created_at", "datetime", "-", "创建时间"),
- ("completed_at", "datetime", "-", "完成时间"),
- ],
- ),
- (
- "表3-11 定时任务表(scheduled_tasks)",
- [
- ("id", "UUID", "36", "定时任务主键"),
- ("user_id", "UUID", "36", "所属用户 ID"),
- ("name", "varchar", "255", "任务名称"),
- ("description", "text", "-", "任务描述"),
- ("task_type", "varchar", "50", "任务类型"),
- ("schedule_cron", "varchar", "100", "调度表达式"),
- ("timezone", "varchar", "50", "时区"),
- ("parameters", "json", "-", "任务参数"),
- ("delivery_config", "json", "-", "投递配置"),
- ("enabled", "boolean", "-", "是否启用"),
- ("last_run_at", "datetime", "-", "上次执行时间"),
- ("next_run_at", "datetime", "-", "下次执行时间"),
- ("run_count", "integer", "-", "执行次数"),
- ("last_result", "text", "-", "最近执行结果"),
- ("last_error", "text", "-", "最近错误信息"),
- ("consecutive_failures", "integer", "-", "连续失败次数"),
- ("created_at", "datetime", "-", "创建时间"),
- ("updated_at", "datetime", "-", "更新时间"),
- ],
- ),
- (
- "表3-12 定时任务执行记录表(scheduled_task_runs)",
- [
- ("id", "UUID", "36", "执行记录主键"),
- ("task_id", "UUID", "36", "所属定时任务 ID"),
- ("status", "varchar", "20", "执行状态"),
- ("started_at", "datetime", "-", "开始时间"),
- ("finished_at", "datetime", "-", "结束时间"),
- ("duration_ms", "integer", "-", "执行耗时(毫秒)"),
- ("result_summary", "text", "-", "结果摘要"),
- ("error_message", "text", "-", "错误信息"),
- ("generated_content", "text", "-", "生成内容"),
- ("sources_count", "integer", "-", "涉及来源数量"),
- ("delivery_status", "json", "-", "投递状态"),
- ],
- ),
- (
- "表3-13 用户记忆表(user_memories)",
- [
- ("id", "UUID", "36", "用户记忆主键"),
- ("user_id", "UUID", "36", "所属用户 ID"),
- ("key", "varchar", "100", "记忆键"),
- ("value", "text", "-", "记忆值"),
- ("confidence", "float", "-", "置信度"),
- ("memory_type", "varchar", "20", "记忆类型"),
- ("memory_kind", "varchar", "32", "记忆类别"),
- ("access_count", "integer", "-", "访问次数"),
- ("last_accessed_at", "datetime", "-", "最后访问时间"),
- ("expires_at", "datetime", "-", "过期时间"),
- ("reinforced_by", "varchar", "36", "强化来源反思 ID"),
- ("updated_at", "datetime", "-", "更新时间"),
- ("embedding", "vector", "-", "记忆向量"),
- ("source", "varchar", "20", "记忆来源类型"),
- ("evidence", "text", "-", "证据链或来源记录"),
- ("conflict_flag", "boolean", "-", "冲突标记"),
- ],
- ),
- (
- "表3-14 用户画像表(user_portraits)",
- [
- ("user_id", "UUID", "36", "所属用户 ID"),
- ("portrait_json", "json", "-", "六维画像 JSON"),
- ("avatar_url", "text", "-", "画像头像地址"),
- ("synthesis_summary", "text", "-", "画像合成摘要"),
- ("version", "integer", "-", "画像版本号"),
- ("synthesized_at", "datetime", "-", "合成时间"),
- ("created_at", "datetime", "-", "创建时间"),
- ("updated_at", "datetime", "-", "更新时间"),
- ],
- ),
- (
- "表3-15 Agent 思想记录表(agent_thoughts)",
- [
- ("id", "UUID", "36", "思想记录主键"),
- ("user_id", "UUID", "36", "所属用户 ID"),
- ("visibility", "varchar", "20", "可见性状态"),
- ("content", "text", "-", "思想内容"),
- ("activity_context", "json", "-", "活动上下文快照"),
- ("notebook_id", "UUID", "36", "关联笔记本 ID"),
- ("created_at", "datetime", "-", "创建时间"),
- ],
- ),
- (
- "表3-16 主动洞察表(proactive_insights)",
- [
- ("id", "UUID", "36", "洞察主键"),
- ("user_id", "UUID", "36", "所属用户 ID"),
- ("notebook_id", "UUID", "36", "关联笔记本 ID"),
- ("insight_type", "varchar", "50", "洞察类型"),
- ("title", "varchar", "500", "洞察标题"),
- ("content", "text", "-", "洞察内容"),
- ("metadata", "json", "-", "洞察附加元数据"),
- ("is_read", "boolean", "-", "是否已读"),
- ("created_at", "datetime", "-", "创建时间"),
- ],
- ),
-]
-
-
-def set_east_asia_font(run, font_name: str) -> None:
- run.font.name = font_name
- run._element.rPr.rFonts.set(qn("w:eastAsia"), font_name)
-
-
-def set_doc_fonts(doc: Document) -> None:
- normal = doc.styles["Normal"]
- normal.font.name = "宋体"
- normal.font.size = Pt(12)
- normal._element.rPr.rFonts.set(qn("w:eastAsia"), "宋体")
-
-
-def set_cell_margins(cell, top: int = 40, start: int = 60, bottom: int = 40, end: int = 60) -> None:
- tc_pr = cell._tc.get_or_add_tcPr()
- tc_mar = tc_pr.first_child_found_in("w:tcMar")
- if tc_mar is None:
- tc_mar = OxmlElement("w:tcMar")
- tc_pr.append(tc_mar)
-
- for edge, value in {"top": top, "start": start, "bottom": bottom, "end": end}.items():
- element = tc_mar.find(qn(f"w:{edge}"))
- if element is None:
- element = OxmlElement(f"w:{edge}")
- tc_mar.append(element)
- element.set(qn("w:w"), str(value))
- element.set(qn("w:type"), "dxa")
-
-
-def set_cell_border(cell, **borders) -> None:
- tc_pr = cell._tc.get_or_add_tcPr()
- tc_borders = tc_pr.first_child_found_in("w:tcBorders")
- if tc_borders is None:
- tc_borders = OxmlElement("w:tcBorders")
- tc_pr.append(tc_borders)
-
- for edge, data in borders.items():
- element = tc_borders.find(qn(f"w:{edge}"))
- if element is None:
- element = OxmlElement(f"w:{edge}")
- tc_borders.append(element)
- for key, value in data.items():
- element.set(qn(f"w:{key}"), str(value))
-
-
-def clear_cell_borders(cell) -> None:
- nil = {"val": "nil"}
- set_cell_border(cell, top=nil, bottom=nil, left=nil, right=nil)
-
-
-def set_table_borders(table) -> None:
- tbl_pr = table._tbl.tblPr
- tbl_borders = tbl_pr.first_child_found_in("w:tblBorders")
- if tbl_borders is None:
- tbl_borders = OxmlElement("w:tblBorders")
- tbl_pr.append(tbl_borders)
-
- styles = {
- "top": {"val": "single", "sz": "12", "space": "0", "color": "000000"},
- "bottom": {"val": "single", "sz": "12", "space": "0", "color": "000000"},
- "left": {"val": "nil"},
- "right": {"val": "nil"},
- "insideH": {"val": "nil"},
- "insideV": {"val": "nil"},
- }
- for edge, data in styles.items():
- element = tbl_borders.find(qn(f"w:{edge}"))
- if element is None:
- element = OxmlElement(f"w:{edge}")
- tbl_borders.append(element)
- for key, value in data.items():
- element.set(qn(f"w:{key}"), str(value))
-
-
-def set_row_cant_split(row) -> None:
- tr_pr = row._tr.get_or_add_trPr()
- if tr_pr.find(qn("w:cantSplit")) is None:
- tr_pr.append(OxmlElement("w:cantSplit"))
-
-
-def set_cell_width(cell, width_cm: float) -> None:
- cell.width = Cm(width_cm)
- tc_pr = cell._tc.get_or_add_tcPr()
- tc_w = tc_pr.first_child_found_in("w:tcW")
- if tc_w is None:
- tc_w = OxmlElement("w:tcW")
- tc_pr.append(tc_w)
- tc_w.set(qn("w:w"), str(int(Cm(width_cm).twips)))
- tc_w.set(qn("w:type"), "dxa")
-
-
-def set_cell_text(cell, text: str, *, bold: bool = False) -> None:
- cell.text = ""
- cell.vertical_alignment = WD_CELL_VERTICAL_ALIGNMENT.CENTER
- paragraph = cell.paragraphs[0]
- paragraph.alignment = WD_ALIGN_PARAGRAPH.CENTER
- paragraph.paragraph_format.space_before = Pt(0)
- paragraph.paragraph_format.space_after = Pt(0)
- paragraph.paragraph_format.line_spacing = 1.0
- run = paragraph.add_run(text)
- run.bold = bold
- run.font.size = Pt(10.5)
- set_east_asia_font(run, "宋体")
-
-
-def format_three_line_table(table) -> None:
- table.alignment = WD_TABLE_ALIGNMENT.CENTER
- table.autofit = False
- set_table_borders(table)
-
- for row in table.rows:
- set_row_cant_split(row)
- row.height_rule = WD_ROW_HEIGHT_RULE.AT_LEAST
- row.height = Cm(0.8)
- for index, cell in enumerate(row.cells):
- set_cell_width(cell, COLUMN_WIDTHS_CM[index])
- set_cell_margins(cell)
- clear_cell_borders(cell)
-
- header_border = {"val": "single", "sz": "8", "space": "0", "color": "000000"}
- for cell in table.rows[0].cells:
- set_cell_border(cell, bottom=header_border)
-
-
-def add_caption(doc: Document, title: str) -> None:
- paragraph = doc.add_paragraph()
- paragraph.alignment = WD_ALIGN_PARAGRAPH.CENTER
- paragraph.paragraph_format.space_before = Pt(8)
- paragraph.paragraph_format.space_after = Pt(4)
- run = paragraph.add_run(title)
- run.bold = True
- run.font.size = Pt(12)
- set_east_asia_font(run, "宋体")
-
-
-def add_intro_paragraph(doc: Document, text: str) -> None:
- paragraph = doc.add_paragraph()
- paragraph.paragraph_format.first_line_indent = Pt(24)
- paragraph.paragraph_format.line_spacing = 1.5
- paragraph.paragraph_format.space_after = Pt(4)
- run = paragraph.add_run(text)
- run.font.size = Pt(12)
- set_east_asia_font(run, "宋体")
-
-
-def add_table(doc: Document, title: str, rows: list[tuple[str, str, str, str]]) -> None:
- add_caption(doc, title)
-
- table = doc.add_table(rows=len(rows) + 1, cols=len(HEADERS))
- for index, header in enumerate(HEADERS):
- set_cell_text(table.cell(0, index), header, bold=True)
-
- for row_index, row_data in enumerate(rows, start=1):
- for col_index, value in enumerate(row_data):
- set_cell_text(table.cell(row_index, col_index), value)
-
- format_three_line_table(table)
- doc.add_paragraph("")
-
-
-def build_doc() -> Document:
- doc = Document()
- set_doc_fonts(doc)
-
- section = doc.sections[0]
- section.page_width = Mm(210)
- section.page_height = Mm(297)
- section.top_margin = Cm(2.54)
- section.bottom_margin = Cm(2.54)
- section.left_margin = Cm(3.0)
- section.right_margin = Cm(3.0)
-
- title = doc.add_paragraph()
- title.alignment = WD_ALIGN_PARAGRAPH.LEFT
- title.paragraph_format.space_after = Pt(10)
- title_run = title.add_run("3.3.2 数据库表设计")
- title_run.bold = True
- title_run.font.size = Pt(14)
- set_east_asia_font(title_run, "黑体")
-
- add_intro_paragraph(
- doc,
- "根据 LyraNote 系统当前代码实现,数据库表主要围绕用户管理、知识组织、对话交互、研究任务、长期记忆和主动服务等几个方面展开设计。系统底层采用 PostgreSQL 作为核心关系型数据库,并结合 pgvector 扩展保存知识分块与长期记忆的向量表示。",
- )
- add_intro_paragraph(
- doc,
- "为体现当前版本的数据结构设计,本文选取用户、应用配置、笔记本、笔记本摘要、知识来源、知识分块、对话、消息、对话摘要、深度研究任务、定时任务、定时任务执行记录、用户记忆、用户画像、Agent 思想记录和主动洞察等 16 张核心数据表进行说明。各主要数据表设计如下所示。",
- )
-
- for title_text, rows in TABLES:
- add_table(doc, title_text, rows)
-
- return doc
-
-
-def main() -> None:
- PRIMARY_OUTPUT_DOCX.parent.mkdir(parents=True, exist_ok=True)
- doc = build_doc()
- doc.save(PRIMARY_OUTPUT_DOCX)
- doc.save(VERSIONED_OUTPUT_DOCX)
- print(PRIMARY_OUTPUT_DOCX)
- print(VERSIONED_OUTPUT_DOCX)
-
-
-if __name__ == "__main__":
- main()
diff --git a/scripts/fix_thesis_er_diagrams.py b/scripts/fix_thesis_er_diagrams.py
deleted file mode 100644
index 13d1e14..0000000
--- a/scripts/fix_thesis_er_diagrams.py
+++ /dev/null
@@ -1,639 +0,0 @@
-from __future__ import annotations
-
-from dataclasses import dataclass
-from pathlib import Path
-from typing import Iterable
-
-from docx import Document
-from docx.enum.text import WD_ALIGN_PARAGRAPH
-from docx.oxml import OxmlElement
-from docx.shared import Inches
-from docx.text.paragraph import Paragraph
-from PIL import Image, ImageDraw, ImageFont
-
-
-REPO_ROOT = Path(__file__).resolve().parents[1]
-SOURCE_DOC = Path("/Users/kaihuang/Documents/毕业论文/220501020064-黄凯-毕业论文-引用修订版_4.docx")
-TMP_DIR = REPO_ROOT / "tmp" / "docs" / "thesis-er-diagram-cleanup"
-OUTPUT_DIR = REPO_ROOT / "output" / "doc"
-OUTPUT_DOC = OUTPUT_DIR / "220501020064-黄凯-毕业论文-引用修订版_4-ER图修订版.docx"
-
-
-@dataclass(frozen=True)
-class EntitySpec:
- index: int
- figure_no: str
- entity_name: str
- caption: str
- description: str
- attrs: tuple[str, ...]
-
- @property
- def image_name(self) -> str:
- return f"{self.figure_no}-{self.entity_name}.png"
-
-
-ENTITY_SPECS: tuple[EntitySpec, ...] = (
- EntitySpec(
- index=1,
- figure_no="3-6",
- entity_name="用户",
- caption="图3-6 用户实体 E-R 图",
- description=(
- "(1)用户实体 E-R 图如图3-6所示。图 3-6 展示了用户实体的 E-R 结构。"
- "用户实体用于标识系统中的真实使用者,图中仅保留用户自身属性,"
- "包括账号信息、认证标识、头像信息以及创建时间等内容;"
- "与笔记本、会话、任务和长期记忆等实体之间的归属关系在后续关系约束中体现。"
- ),
- attrs=(
- "用户ID(PK)",
- "用户名",
- "密码哈希",
- "邮箱",
- "姓名",
- "头像地址",
- "Google标识",
- "GitHub标识",
- "OAuth解绑标记",
- "创建时间",
- ),
- ),
- EntitySpec(
- index=2,
- figure_no="3-7",
- entity_name="笔记本",
- caption="图3-7 笔记本实体 E-R 图",
- description=(
- "(2)笔记本实体 E-R 图如图3-7所示。图 3-7 展示了笔记本实体的 E-R 结构。"
- "笔记本是 LyraNote 组织私有知识的核心工作空间,"
- "图中保留了标题、描述、状态、公开性与封面等实体自有属性,"
- "而与用户、来源、对话和笔记等对象的关联不再作为属性列重复展示。"
- ),
- attrs=(
- "笔记本ID(PK)",
- "标题",
- "描述",
- "状态",
- "是否全局知识库",
- "是否系统笔记本",
- "系统类型",
- "来源数量",
- "是否公开",
- "发布时间",
- "封面表情",
- "封面渐变",
- "创建时间",
- "更新时间",
- ),
- ),
- EntitySpec(
- index=3,
- figure_no="3-8",
- entity_name="知识来源",
- caption="图3-8 知识来源实体 E-R 图",
- description=(
- "(3)知识来源实体 E-R 图如图3-8所示。图 3-8 展示了知识来源实体的 E-R 结构。"
- "知识来源记录用户导入的 PDF、网页和 Markdown 等资料,"
- "实体自身属性主要描述来源类型、索引状态、文本内容、摘要和存储信息,"
- "其与笔记本及知识分块之间的联系通过关系设计表达。"
- ),
- attrs=(
- "来源ID(PK)",
- "标题",
- "来源类型",
- "索引状态",
- "文件路径",
- "链接地址",
- "原始文本",
- "摘要",
- "存储键",
- "存储后端",
- "创建时间",
- "更新时间",
- ),
- ),
- EntitySpec(
- index=4,
- figure_no="3-9",
- entity_name="知识分块",
- caption="图3-9 知识分块实体 E-R 图",
- description=(
- "(4)知识分块实体 E-R 图如图3-9所示。图 3-9 展示了知识分块实体的 E-R 结构。"
- "知识分块是系统执行检索增强生成时使用的最小知识单元,"
- "图中仅保留分块正文、顺序编号、向量嵌入、元数据和来源类型等自身属性,"
- "不再将所属来源、笔记本或笔记的外键字段混入实体属性。"
- ),
- attrs=(
- "分块ID(PK)",
- "正文内容",
- "分块序号",
- "向量嵌入",
- "Token数量",
- "元数据",
- "来源类型",
- "创建时间",
- ),
- ),
- EntitySpec(
- index=5,
- figure_no="3-10",
- entity_name="对话",
- caption="图3-10 对话实体 E-R 图",
- description=(
- "(5)对话实体 E-R 图如图3-10所示。图 3-10 展示了对话实体的 E-R 结构。"
- "对话实体用于承载一次连续的交互会话,"
- "其自身属性主要包括对话标题、来源场景和创建时间;"
- "与用户、笔记本以及消息记录的关联关系通过实体联系表达,而不直接写入属性列表。"
- ),
- attrs=(
- "对话ID(PK)",
- "标题",
- "对话来源",
- "创建时间",
- ),
- ),
- EntitySpec(
- index=6,
- figure_no="3-11",
- entity_name="消息",
- caption="图3-11 消息实体 E-R 图",
- description=(
- "(6)消息实体 E-R 图如图3-11所示。图 3-11 展示了消息实体的 E-R 结构。"
- "消息实体是对话过程中的最小交互单元,"
- "保存角色、状态、正文、引用、推理轨迹以及富媒体结果等内容;"
- "对话归属、生成归属和父消息分支等外键信息不再作为本实体属性展示。"
- ),
- attrs=(
- "消息ID(PK)",
- "角色",
- "状态",
- "正文内容",
- "推理轨迹",
- "引用信息",
- "Agent步骤",
- "附件信息",
- "速度指标",
- "思维导图",
- "图表数据",
- "MCP结果",
- "UI元素",
- "创建时间",
- ),
- ),
- EntitySpec(
- index=7,
- figure_no="3-12",
- entity_name="深度研究任务",
- caption="图3-12 深度研究任务实体 E-R 图",
- description=(
- "(7)深度研究任务实体 E-R 图如图3-12所示。图 3-12 展示了深度研究任务实体的 E-R 结构。"
- "深度研究任务用于表示一次长生命周期的研究流程,"
- "实体自身属性包括研究问题、执行模式、状态、报告正文、结构化交付物和时间线等,"
- "与用户、会话和笔记本的关联通过关系约束体现。"
- ),
- attrs=(
- "任务ID(PK)",
- "研究问题",
- "研究模式",
- "状态",
- "研究报告",
- "交付物JSON",
- "时间线JSON",
- "外部资料列表",
- "错误信息",
- "创建时间",
- "完成时间",
- ),
- ),
- EntitySpec(
- index=8,
- figure_no="3-13",
- entity_name="定时任务",
- caption="图3-13 定时任务实体 E-R 图",
- description=(
- "(8)定时任务实体 E-R 图如图3-13所示。图 3-13 展示了定时任务实体的 E-R 结构。"
- "定时任务实体描述用户配置的持续服务模板,"
- "图中保留任务名称、调度表达式、参数配置、投递配置和运行状态等自身属性,"
- "而用户归属关系仅通过实体联系表示,不再以外键字段形式放入属性图。"
- ),
- attrs=(
- "任务ID(PK)",
- "任务名称",
- "描述",
- "任务类型",
- "调度表达式",
- "时区",
- "参数配置",
- "投递配置",
- "是否启用",
- "最近执行时间",
- "下次执行时间",
- "执行次数",
- "最近结果",
- "最近错误",
- "连续失败次数",
- "创建时间",
- "更新时间",
- ),
- ),
- EntitySpec(
- index=9,
- figure_no="3-14",
- entity_name="用户记忆",
- caption="图3-14 用户记忆实体 E-R 图",
- description=(
- "(9)用户记忆实体 E-R 图如图3-14所示。图 3-14 展示了用户记忆实体的 E-R 结构。"
- "用户记忆用于保存系统长期积累的偏好、事实与技能画像,"
- "实体自身属性包括键值内容、置信度、访问统计、失效时间、向量嵌入和证据来源等,"
- "用户归属关系在关系层处理,不再混入属性集合。"
- ),
- attrs=(
- "记忆ID(PK)",
- "记忆键",
- "记忆值",
- "置信度",
- "记忆类型",
- "记忆类别",
- "访问次数",
- "最近访问时间",
- "失效时间",
- "强化来源",
- "更新时间",
- "向量嵌入",
- "来源类型",
- "证据",
- "冲突标记",
- ),
- ),
- EntitySpec(
- index=10,
- figure_no="3-15",
- entity_name="用户画像",
- caption="图3-15 用户画像实体 E-R 图",
- description=(
- "(10)用户画像实体 E-R 图如图3-15所示。图 3-15 展示了用户画像实体的 E-R 结构。"
- "用户画像实体是对长期记忆与交互行为进行阶段性聚合后的结果,"
- "图中仅保留画像内容、头像地址、合成摘要、版本号以及时间相关属性,"
- "一对一归属关系由实体联系表达而不作为属性列出现。"
- ),
- attrs=(
- "画像JSON",
- "头像地址",
- "合成摘要",
- "版本号",
- "合成时间",
- "创建时间",
- "更新时间",
- ),
- ),
- EntitySpec(
- index=11,
- figure_no="3-16",
- entity_name="主动洞察",
- caption="图3-16 主动洞察实体 E-R 图",
- description=(
- "(11)主动洞察实体 E-R 图如图3-16所示。图 3-16 展示了主动洞察实体的 E-R 结构。"
- "主动洞察实体用于保存可以直接展示给用户的系统提示内容,"
- "其自身属性包括洞察类型、标题、正文、元数据、已读状态与创建时间,"
- "与用户和笔记本的关联关系则由实体之间的联系体现。"
- ),
- attrs=(
- "洞察ID(PK)",
- "洞察类型",
- "标题",
- "内容",
- "元数据",
- "已读状态",
- "创建时间",
- ),
- ),
- EntitySpec(
- index=12,
- figure_no="3-17",
- entity_name="系统配置",
- caption="图3-17 系统配置实体 E-R 图",
- description=(
- "(12)系统配置实体 E-R 图如图3-17所示。图 3-17 展示了系统配置实体的 E-R 结构。"
- "系统配置实体采用键值形式保存安装向导和运行期写入的配置项,"
- "由于该实体本身不承担业务归属关系,因此图中仅包含配置键与配置值两个自身属性。"
- ),
- attrs=(
- "配置键(PK)",
- "配置值",
- ),
- ),
- EntitySpec(
- index=13,
- figure_no="3-18",
- entity_name="笔记本摘要",
- caption="图3-18 笔记本摘要实体 E-R 图",
- description=(
- "(13)笔记本摘要实体 E-R 图如图3-18所示。图 3-18 展示了笔记本摘要实体的 E-R 结构。"
- "笔记本摘要实体用于保存系统自动生成的笔记本级聚合摘要,"
- "图中仅展示摘要内容、主题词以及同步时间等实体自身属性,"
- "其与笔记本的一对一绑定关系不再作为属性列重复绘制。"
- ),
- attrs=(
- "摘要内容",
- "核心主题",
- "最后同步时间",
- ),
- ),
- EntitySpec(
- index=14,
- figure_no="3-19",
- entity_name="会话摘要",
- caption="图3-19 会话摘要实体 E-R 图",
- description=(
- "(14)会话摘要实体 E-R 图如图3-19所示。图 3-19 展示了会话摘要实体的 E-R 结构。"
- "会话摘要实体用于压缩长对话中的历史消息,"
- "实体自身属性包含摘要正文、已压缩消息数量、压缩边界时间和更新时间等内容,"
- "与对话实体的一对一归属关系通过关系设计表示。"
- ),
- attrs=(
- "摘要正文",
- "压缩消息数",
- "压缩边界时间",
- "更新时间",
- ),
- ),
- EntitySpec(
- index=15,
- figure_no="3-20",
- entity_name="定时任务运行记录",
- caption="图3-20 定时任务运行记录实体 E-R 图",
- description=(
- "(15)定时任务运行记录实体 E-R 图如图3-20所示。图 3-20 展示了定时任务运行记录实体的 E-R 结构。"
- "该实体用于记录定时任务每一次实际执行的过程和结果,"
- "图中保留运行状态、开始与结束时间、执行耗时、结果摘要、错误信息以及投递状态等自身属性,"
- "所属任务关系不再写入属性图。"
- ),
- attrs=(
- "运行ID(PK)",
- "状态",
- "开始时间",
- "结束时间",
- "耗时毫秒",
- "结果摘要",
- "错误信息",
- "生成内容",
- "来源数量",
- "投递状态",
- ),
- ),
- EntitySpec(
- index=16,
- figure_no="3-21",
- entity_name="后台思考",
- caption="图3-21 后台思考实体 E-R 图",
- description=(
- "(16)后台思考实体 E-R 图如图3-21所示。图 3-21 展示了后台思考实体的 E-R 结构。"
- "后台思考实体用于保存 Lyra Soul 在后台循环中形成的思考记录,"
- "其自身属性包括可见性、思考内容、活动上下文和创建时间;"
- "与用户和笔记本的关联关系通过实体联系表达,不再作为属性字段直接绘制。"
- ),
- attrs=(
- "思考ID(PK)",
- "可见性",
- "思考内容",
- "活动上下文",
- "创建时间",
- ),
- ),
-)
-
-
-def ensure_dirs() -> None:
- TMP_DIR.mkdir(parents=True, exist_ok=True)
- OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
-
-
-def pick_font(size: int, *, bold: bool = False) -> ImageFont.FreeTypeFont | ImageFont.ImageFont:
- font_candidates = [
- "/System/Library/Fonts/PingFang.ttc",
- "/System/Library/Fonts/STHeiti Light.ttc",
- "/System/Library/Fonts/STHeiti Medium.ttc",
- "/System/Library/Fonts/Supplemental/Songti.ttc",
- "/Library/Fonts/Arial Unicode.ttf",
- ]
- for candidate in font_candidates:
- path = Path(candidate)
- if path.exists():
- try:
- index = 1 if bold and "PingFang" in path.name else 0
- return ImageFont.truetype(str(path), size=size, index=index)
- except OSError:
- continue
- return ImageFont.load_default()
-
-
-def wrap_text(draw: ImageDraw.ImageDraw, text: str, font: ImageFont.ImageFont, max_width: int) -> str:
- if draw.textbbox((0, 0), text, font=font)[2] <= max_width:
- return text
-
- lines: list[str] = []
- current = ""
- for ch in text:
- candidate = current + ch
- width = draw.textbbox((0, 0), candidate, font=font)[2]
- if width <= max_width or not current:
- current = candidate
- else:
- lines.append(current)
- current = ch
- if current:
- lines.append(current)
- return "\n".join(lines[:2])
-
-
-def chunk_counts(n: int) -> tuple[int, ...]:
- if n <= 6:
- return (n,)
- if n <= 12:
- first = (n + 1) // 2
- return (first, n - first)
- first = min(6, (n + 2) // 3)
- second = min(6, (n - first + 1) // 2)
- third = n - first - second
- return tuple(count for count in (first, second, third) if count)
-
-
-def generate_diagram(spec: EntitySpec) -> Path:
- canvas_w = 2200
- canvas_h = 1300
- image = Image.new("RGB", (canvas_w, canvas_h), "#ffffff")
- draw = ImageDraw.Draw(image)
-
- title_font = pick_font(54, bold=True)
- box_font = pick_font(56, bold=True)
- attr_font = pick_font(34)
-
- box_w = 460
- box_h = 120
- counts = chunk_counts(len(spec.attrs))
- row_ys = [110, 300, 490][: len(counts)]
- box_top = 770 if len(counts) >= 3 else 710 if len(counts) == 2 else 650
- box_left = (canvas_w - box_w) // 2
- box_right = box_left + box_w
- box_bottom = box_top + box_h
-
- ellipse_w = 260
- ellipse_h = 92
- gap_x = 36
-
- draw.text((90, 40), f"图 {spec.figure_no} {spec.entity_name} 实体", fill="#333333", font=title_font)
-
- draw.rounded_rectangle(
- (box_left, box_top, box_right, box_bottom),
- radius=18,
- outline="#8a7ae6",
- width=5,
- fill="#f7f2ff",
- )
- box_text = spec.entity_name
- text_bbox = draw.multiline_textbbox((0, 0), box_text, font=box_font, spacing=8, align="center")
- box_text_w = text_bbox[2] - text_bbox[0]
- box_text_h = text_bbox[3] - text_bbox[1]
- draw.multiline_text(
- ((canvas_w - box_text_w) / 2, box_top + (box_h - box_text_h) / 2 - 4),
- box_text,
- fill="#473a8b",
- font=box_font,
- align="center",
- spacing=8,
- )
-
- attr_iter = iter(spec.attrs)
- for row_count, row_y in zip(counts, row_ys):
- total_w = row_count * ellipse_w + (row_count - 1) * gap_x
- start_x = (canvas_w - total_w) // 2
- for i in range(row_count):
- attr = next(attr_iter)
- left = start_x + i * (ellipse_w + gap_x)
- top = row_y
- right = left + ellipse_w
- bottom = top + ellipse_h
- center_x = (left + right) // 2
- center_y = (top + bottom) // 2
-
- draw.line(
- [(center_x, bottom), (canvas_w // 2, box_top)],
- fill="#c6b9f8",
- width=3,
- )
- draw.ellipse(
- (left, top, right, bottom),
- outline="#9d8ff0",
- width=4,
- fill="#f3efff",
- )
-
- label = wrap_text(draw, attr, attr_font, ellipse_w - 24)
- bbox = draw.multiline_textbbox((0, 0), label, font=attr_font, spacing=4, align="center")
- text_w = bbox[2] - bbox[0]
- text_h = bbox[3] - bbox[1]
- draw.multiline_text(
- (center_x - text_w / 2, center_y - text_h / 2 - 2),
- label,
- fill="#2f2f2f",
- font=attr_font,
- align="center",
- spacing=4,
- )
-
- output = TMP_DIR / spec.image_name
- image.save(output)
- return output
-
-
-def insert_paragraph_after(paragraph: Paragraph, text: str = "", style: str | None = None) -> Paragraph:
- new_p = OxmlElement("w:p")
- paragraph._p.addnext(new_p)
- new_para = Paragraph(new_p, paragraph._parent)
- if style:
- new_para.style = style
- if text:
- new_para.add_run(text)
- return new_para
-
-
-def delete_paragraph(paragraph: Paragraph) -> None:
- element = paragraph._element
- parent = element.getparent()
- if parent is not None:
- parent.remove(element)
-
-
-def resolve_style(document: Document, preferred: str, fallback: str = "Normal") -> str:
- styles = {style.name for style in document.styles}
- if preferred in styles:
- return preferred
- return fallback if fallback in styles else "Normal"
-
-
-def rebuild_er_section(document: Document, image_paths: dict[str, Path]) -> None:
- body_style = resolve_style(document, "毕业设计(论文)正文")
- caption_style = resolve_style(document, "图标", "Normal")
- image_style = resolve_style(document, "No Spacing", "Normal")
-
- document.paragraphs[183].text = (
- "经过对系统分析并结合当前代码中的数据模型定义,"
- "本文将用户、笔记本、知识来源、知识分块、对话、消息、深度研究任务、定时任务、用户记忆、"
- "用户画像、主动洞察、系统配置、笔记本摘要、会话摘要、定时任务运行记录和后台思考共 16 个核心实体"
- "整理为属性型 E-R 图。为突出实体自身语义,图中只保留实体固有属性,"
- "不再将外键字段纳入属性列表,实体之间的归属和关联关系通过后续数据库约束说明体现。"
- )
- document.paragraphs[183].style = body_style
-
- document.paragraphs[189].text = (
- "在实体-关系(E-R)模型设计中,本文主要围绕关系型数据库 PostgreSQL 的核心业务实体进行建模。"
- "各实体的自有属性直接决定数据库表中的主要字段设计,而实体之间的关系则由主键、外键和关联约束进行表达。"
- "因此,以下实体图仅展示主键与非外键业务属性,"
- "不再把归属关系字段重复放入属性图中,以避免实体属性与关系属性混淆。"
- )
- document.paragraphs[189].style = body_style
-
- for paragraph in list(document.paragraphs[190:222]):
- delete_paragraph(paragraph)
-
- anchor = document.paragraphs[189]
- for spec in ENTITY_SPECS:
- desc_para = insert_paragraph_after(anchor, spec.description, body_style)
- anchor = desc_para
-
- image_para = insert_paragraph_after(anchor, style=image_style)
- image_para.alignment = WD_ALIGN_PARAGRAPH.CENTER
- image_para.add_run().add_picture(str(image_paths[spec.figure_no]), width=Inches(6.8))
- anchor = image_para
-
- caption_para = insert_paragraph_after(anchor, spec.caption, caption_style)
- caption_para.alignment = WD_ALIGN_PARAGRAPH.CENTER
- anchor = caption_para
-
-
-def validate_output(document_path: Path) -> None:
- document = Document(document_path)
- expected_captions = {spec.caption for spec in ENTITY_SPECS}
- actual_captions = {p.text.strip() for p in document.paragraphs if p.text.strip().startswith("图3-")}
- missing = expected_captions - actual_captions
- if missing:
- raise RuntimeError(f"输出文档缺少图题: {sorted(missing)}")
-
-
-def main() -> None:
- ensure_dirs()
- if not SOURCE_DOC.exists():
- raise FileNotFoundError(f"未找到输入文档: {SOURCE_DOC}")
-
- image_paths = {spec.figure_no: generate_diagram(spec) for spec in ENTITY_SPECS}
-
- document = Document(SOURCE_DOC)
- rebuild_er_section(document, image_paths)
- document.save(OUTPUT_DOC)
- validate_output(OUTPUT_DOC)
-
- print(OUTPUT_DOC)
-
-
-if __name__ == "__main__":
- main()
diff --git a/scripts/fix_thesis_er_diagrams_mermaid.py b/scripts/fix_thesis_er_diagrams_mermaid.py
deleted file mode 100644
index 617ae79..0000000
--- a/scripts/fix_thesis_er_diagrams_mermaid.py
+++ /dev/null
@@ -1,303 +0,0 @@
-from __future__ import annotations
-
-from pathlib import Path
-
-from docx import Document
-from docx.enum.text import WD_ALIGN_PARAGRAPH
-from docx.shared import Inches
-from PIL import Image, ImageDraw, ImageFont
-
-from fix_thesis_er_diagrams import (
- ENTITY_SPECS,
- OUTPUT_DIR,
- REPO_ROOT,
- delete_paragraph,
- ensure_dirs,
- insert_paragraph_after,
- resolve_style,
-)
-
-
-SOURCE_DOC = OUTPUT_DIR / "220501020064-黄凯-毕业论文-引用修订版_4-ER图修订版.docx"
-TMP_DIR = REPO_ROOT / "tmp" / "docs" / "thesis-er-diagram-mermaid"
-OUTPUT_DOC = OUTPUT_DIR / "220501020064-黄凯-毕业论文-引用修订版_4-ER图Mermaid版.docx"
-
-BG_COLOR = "#ffffff"
-BOX_FILL = "#fff4dd"
-BOX_BORDER = "#8b6f47"
-HEADER_TEXT = "#3b342c"
-ROW_ODD = "#ffffff"
-ROW_EVEN = "#f2f2f2"
-GRID = "#b9a88a"
-TEXT = "#333333"
-TYPE_TEXT = "#59534e"
-KEY_FILL = "#ece8ff"
-KEY_TEXT = "#51459b"
-
-
-def pick_font(size: int, *, bold: bool = False) -> ImageFont.FreeTypeFont | ImageFont.ImageFont:
- font_candidates = [
- ("/System/Library/Fonts/PingFang.ttc", 1 if bold else 0),
- ("/System/Library/Fonts/STHeiti Medium.ttc", 0),
- ("/System/Library/Fonts/Supplemental/Songti.ttc", 0),
- ]
- for candidate, index in font_candidates:
- path = Path(candidate)
- if path.exists():
- try:
- return ImageFont.truetype(str(path), size=size, index=index)
- except OSError:
- continue
- return ImageFont.load_default()
-
-
-def measure(draw: ImageDraw.ImageDraw, text: str, font: ImageFont.ImageFont) -> tuple[int, int]:
- bbox = draw.textbbox((0, 0), text, font=font)
- return bbox[2] - bbox[0], bbox[3] - bbox[1]
-
-
-def parse_attr(attr: str) -> tuple[str, str, str]:
- key = ""
- name = attr
- if attr.endswith("(PK)"):
- name = attr[:-4]
- key = "PK"
- elif attr.endswith("(UK)"):
- name = attr[:-4]
- key = "UK"
-
- attr_type = infer_type(name)
- return attr_type, name, key
-
-
-def infer_type(name: str) -> str:
- if name.endswith("ID"):
- return "uuid"
- if name.startswith("是否") or name in {"已读状态", "冲突标记"}:
- return "bool"
- if "时间" in name or "日期" in name:
- return "datetime"
- if "数量" in name or "次数" in name or "序号" in name or "版本号" in name or "耗时" in name:
- return "int"
- if "置信度" in name:
- return "float"
- if "嵌入" in name:
- return "vector"
- if (
- "JSON" in name
- or "元数据" in name
- or "列表" in name
- or "上下文" in name
- or name in {
- "参数配置",
- "投递配置",
- "核心主题",
- "活动上下文",
- "引用信息",
- "Agent步骤",
- "附件信息",
- "速度指标",
- "思维导图",
- "图表数据",
- "MCP结果",
- "UI元素",
- "画像JSON",
- "交付物JSON",
- "时间线JSON",
- "投递状态",
- }
- ):
- return "json"
- if (
- "内容" in name
- or "正文" in name
- or "摘要" in name
- or "描述" in name
- or "报告" in name
- or "错误" in name
- or "证据" in name
- or "轨迹" in name
- or "结果" in name
- or "文本" in name
- ):
- return "text"
- return "string"
-
-
-def make_mermaid_source(entity_name: str, attrs: tuple[str, ...]) -> str:
- rows = []
- for attr in attrs:
- attr_type, name, key = parse_attr(attr)
- if key:
- rows.append(f" {attr_type} {name} {key}")
- else:
- rows.append(f" {attr_type} {name}")
- return "erDiagram\n" + f" {entity_name} {{\n" + "\n".join(rows) + "\n }\n"
-
-
-def render_mermaid_style_entity(entity_name: str, attrs: tuple[str, ...]) -> tuple[Path, tuple[int, int]]:
- TMP_DIR.mkdir(parents=True, exist_ok=True)
-
- rows = [parse_attr(attr) for attr in attrs]
- mmd_path = TMP_DIR / f"{entity_name}.mmd"
- mmd_path.write_text(make_mermaid_source(entity_name, attrs), encoding="utf-8")
-
- title_font = pick_font(28, bold=True)
- row_font = pick_font(24)
- key_font = pick_font(18, bold=True)
-
- probe = Image.new("RGB", (10, 10), BG_COLOR)
- draw = ImageDraw.Draw(probe)
-
- type_width = max(measure(draw, r[0], row_font)[0] for r in rows)
- name_width = max(measure(draw, r[1], row_font)[0] for r in rows)
- header_width = measure(draw, entity_name, title_font)[0]
- has_key = any(r[2] for r in rows)
-
- cell_pad_x = 18
- header_h = 52
- row_h = 40
- border = 3
- key_w = 52 if has_key else 0
- type_col_w = type_width + cell_pad_x * 2
- name_col_w = max(name_width + cell_pad_x * 2, header_width + 48)
- table_w = type_col_w + name_col_w + key_w
- table_h = header_h + len(rows) * row_h
- image_w = table_w + border * 2
- image_h = table_h + border * 2
-
- image = Image.new("RGB", (image_w, image_h), BG_COLOR)
- draw = ImageDraw.Draw(image)
-
- left = border
- top = border
- right = left + table_w
- bottom = top + table_h
- draw.rounded_rectangle(
- (left, top, right, bottom),
- radius=10,
- fill=BOX_FILL,
- outline=BOX_BORDER,
- width=border,
- )
- draw.rectangle((left, top, right, top + header_h), fill=BOX_FILL, outline=None)
- draw.line((left, top + header_h, right, top + header_h), fill=BOX_BORDER, width=2)
-
- header_text_w, header_text_h = measure(draw, entity_name, title_font)
- draw.text(
- (left + (table_w - header_text_w) / 2, top + (header_h - header_text_h) / 2 - 2),
- entity_name,
- fill=HEADER_TEXT,
- font=title_font,
- )
-
- for idx, (attr_type, name, key) in enumerate(rows):
- row_top = top + header_h + idx * row_h
- row_bottom = row_top + row_h
- draw.rectangle(
- (left, row_top, right, row_bottom),
- fill=ROW_EVEN if idx % 2 else ROW_ODD,
- outline=None,
- )
- if idx < len(rows) - 1:
- draw.line((left, row_bottom, right, row_bottom), fill=GRID, width=1)
-
- draw.text((left + cell_pad_x, row_top + 7), attr_type, fill=TYPE_TEXT, font=row_font)
- draw.text((left + type_col_w + cell_pad_x, row_top + 7), name, fill=TEXT, font=row_font)
-
- draw.line((left + type_col_w, row_top, left + type_col_w, row_bottom), fill=GRID, width=1)
- if has_key:
- draw.line((right - key_w, row_top, right - key_w, row_bottom), fill=GRID, width=1)
- if key:
- pill_w = 34
- pill_h = 22
- pill_x = right - key_w + (key_w - pill_w) / 2
- pill_y = row_top + (row_h - pill_h) / 2
- draw.rounded_rectangle(
- (pill_x, pill_y, pill_x + pill_w, pill_y + pill_h),
- radius=10,
- fill=KEY_FILL,
- outline=None,
- )
- key_text_w, key_text_h = measure(draw, key, key_font)
- draw.text(
- (pill_x + (pill_w - key_text_w) / 2, pill_y + (pill_h - key_text_h) / 2 - 1),
- key,
- fill=KEY_TEXT,
- font=key_font,
- )
-
- png_path = TMP_DIR / f"{entity_name}.png"
- image.save(png_path)
- return png_path, (image_w, image_h)
-
-
-def rebuild_er_section(document: Document, image_meta: dict[str, tuple[Path, tuple[int, int]]]) -> None:
- body_style = resolve_style(document, "毕业设计(论文)正文")
- caption_style = resolve_style(document, "图标", "Normal")
- image_style = resolve_style(document, "No Spacing", "Normal")
-
- document.paragraphs[183].text = (
- "经过对系统分析并结合当前代码中的数据模型定义,本文将用户、笔记本、知识来源、知识分块、对话、消息、"
- "深度研究任务、定时任务、用户记忆、用户画像、主动洞察、系统配置、笔记本摘要、会话摘要、定时任务运行记录和后台思考"
- "共 16 个核心实体整理为属性型 E-R 图。为突出实体自身语义,图中只保留实体固有属性,不再将外键字段纳入属性列表;"
- "本次图形表现采用 Mermaid 风格的实体卡片形式,以便在版面上更紧凑、更统一地呈现各实体结构。"
- )
- document.paragraphs[183].style = body_style
-
- document.paragraphs[189].text = (
- "在实体-关系(E-R)模型设计中,本文主要围绕关系型数据库 PostgreSQL 的核心业务实体进行建模。"
- "各实体的自有属性直接决定数据库表中的主要字段设计,而实体之间的关系则由主键、外键和关联约束进行表达。"
- "因此,以下实体图仅展示主键与非外键业务属性,不再把归属关系字段重复放入属性图中;"
- "同时采用 Mermaid 风格的表格式实体表示,以减少图片冗余留白并提升可读性。"
- )
- document.paragraphs[189].style = body_style
-
- for paragraph in list(document.paragraphs[190:238]):
- delete_paragraph(paragraph)
-
- anchor = document.paragraphs[189]
- for spec in ENTITY_SPECS:
- desc_para = insert_paragraph_after(anchor, spec.description, body_style)
- anchor = desc_para
-
- image_para = insert_paragraph_after(anchor, style=image_style)
- image_para.alignment = WD_ALIGN_PARAGRAPH.CENTER
- image_path, (px_w, _px_h) = image_meta[spec.figure_no]
- width_inches = min(5.8, max(2.6, px_w / 230))
- image_para.add_run().add_picture(str(image_path), width=Inches(width_inches))
- anchor = image_para
-
- caption_para = insert_paragraph_after(anchor, spec.caption, caption_style)
- caption_para.alignment = WD_ALIGN_PARAGRAPH.CENTER
- anchor = caption_para
-
-
-def validate_output(document_path: Path) -> None:
- document = Document(document_path)
- expected = {spec.caption for spec in ENTITY_SPECS}
- actual = {p.text.strip() for p in document.paragraphs if p.text.strip().startswith("图3-")}
- missing = expected - actual
- if missing:
- raise RuntimeError(f"输出文档缺少图题: {sorted(missing)}")
-
-
-def main() -> None:
- ensure_dirs()
- TMP_DIR.mkdir(parents=True, exist_ok=True)
- if not SOURCE_DOC.exists():
- raise FileNotFoundError(f"未找到输入文档: {SOURCE_DOC}")
-
- image_meta = {}
- for spec in ENTITY_SPECS:
- image_meta[spec.figure_no] = render_mermaid_style_entity(spec.entity_name, spec.attrs)
-
- document = Document(SOURCE_DOC)
- rebuild_er_section(document, image_meta)
- document.save(OUTPUT_DOC)
- validate_output(OUTPUT_DOC)
- print(OUTPUT_DOC)
-
-
-if __name__ == "__main__":
- main()
diff --git a/scripts/generate_thesis_pdfs.py b/scripts/generate_thesis_pdfs.py
deleted file mode 100644
index 9caa641..0000000
--- a/scripts/generate_thesis_pdfs.py
+++ /dev/null
@@ -1,297 +0,0 @@
-from __future__ import annotations
-
-import html
-import re
-from pathlib import Path
-
-from reportlab.lib import colors
-from reportlab.lib.enums import TA_CENTER
-from reportlab.lib.pagesizes import A4
-from reportlab.lib.styles import ParagraphStyle, getSampleStyleSheet
-from reportlab.lib.units import mm
-from reportlab.pdfbase.cidfonts import UnicodeCIDFont
-from reportlab.pdfbase.pdfmetrics import registerFont
-from reportlab.platypus import (
- HRFlowable,
- ListFlowable,
- ListItem,
- PageBreak,
- Paragraph,
- SimpleDocTemplate,
- Spacer,
-)
-
-
-ROOT = Path(__file__).resolve().parent.parent
-SOURCE_DIR = ROOT / "docs" / "graduation-thesis"
-OUTPUT_DIR = ROOT / "output" / "pdf"
-
-DOCUMENTS = [
- {
- "source": SOURCE_DIR / "lyranote-undergrad-thesis-outline.md",
- "output": OUTPUT_DIR / "lyranote-undergrad-thesis-outline.pdf",
- "title": "LyraNote 本科毕业论文大纲",
- },
- {
- "source": SOURCE_DIR / "lyranote-undergrad-thesis-demo.md",
- "output": OUTPUT_DIR / "lyranote-undergrad-thesis-demo.pdf",
- "title": "LyraNote 本科毕业论文 Demo 稿",
- },
-]
-
-
-def register_fonts() -> None:
- registerFont(UnicodeCIDFont("STSong-Light"))
-
-
-def build_styles():
- styles = getSampleStyleSheet()
- body = ParagraphStyle(
- "BodyCN",
- parent=styles["BodyText"],
- fontName="STSong-Light",
- fontSize=11,
- leading=18,
- textColor=colors.HexColor("#1f2937"),
- spaceAfter=6,
- wordWrap="CJK",
- )
- title = ParagraphStyle(
- "TitleCN",
- parent=styles["Title"],
- fontName="STSong-Light",
- fontSize=22,
- leading=30,
- alignment=TA_CENTER,
- textColor=colors.HexColor("#111827"),
- spaceAfter=14,
- )
- h1 = ParagraphStyle(
- "Heading1CN",
- parent=styles["Heading1"],
- fontName="STSong-Light",
- fontSize=18,
- leading=26,
- textColor=colors.HexColor("#111827"),
- spaceBefore=10,
- spaceAfter=10,
- wordWrap="CJK",
- )
- h2 = ParagraphStyle(
- "Heading2CN",
- parent=styles["Heading2"],
- fontName="STSong-Light",
- fontSize=14,
- leading=22,
- textColor=colors.HexColor("#111827"),
- spaceBefore=10,
- spaceAfter=8,
- wordWrap="CJK",
- )
- h3 = ParagraphStyle(
- "Heading3CN",
- parent=styles["Heading3"],
- fontName="STSong-Light",
- fontSize=12,
- leading=18,
- textColor=colors.HexColor("#111827"),
- spaceBefore=8,
- spaceAfter=6,
- wordWrap="CJK",
- )
- meta = ParagraphStyle(
- "MetaCN",
- parent=body,
- fontSize=10.5,
- leading=17,
- alignment=TA_CENTER,
- textColor=colors.HexColor("#4b5563"),
- spaceAfter=4,
- )
- list_style = ParagraphStyle(
- "ListCN",
- parent=body,
- leftIndent=0,
- firstLineIndent=0,
- spaceAfter=2,
- )
- return {
- "body": body,
- "title": title,
- "h1": h1,
- "h2": h2,
- "h3": h3,
- "meta": meta,
- "list": list_style,
- }
-
-
-def escape_text(text: str) -> str:
- return html.escape(text).replace("\n", "
")
-
-
-def add_page_number(canvas, doc) -> None:
- canvas.setFont("STSong-Light", 9)
- canvas.setFillColor(colors.HexColor("#6b7280"))
- canvas.drawCentredString(A4[0] / 2, 10 * mm, f"{doc.page}")
-
-
-def flush_paragraph(story, buffer: list[str], style) -> None:
- if not buffer:
- return
- text = " ".join(line.strip() for line in buffer).strip()
- if text:
- story.append(Paragraph(escape_text(text), style))
- buffer.clear()
-
-
-def flush_list(story, items: list[str], list_type: str, styles) -> None:
- if not items:
- return
- bullet_type = "bullet" if list_type == "bullet" else "1"
- list_items = [
- ListItem(Paragraph(escape_text(item), styles["list"]))
- for item in items
- ]
- story.append(
- ListFlowable(
- list_items,
- bulletType=bullet_type,
- start="1",
- leftIndent=18,
- bulletFontName="STSong-Light",
- bulletFontSize=10.5,
- spaceAfter=6,
- )
- )
- items.clear()
-
-
-def markdown_to_story(text: str, doc_title: str, styles) -> list:
- story: list = []
- story.append(Spacer(1, 22 * mm))
- story.append(Paragraph(escape_text(doc_title), styles["title"]))
- story.append(Spacer(1, 4 * mm))
-
- paragraph_buffer: list[str] = []
- list_buffer: list[str] = []
- list_type: str | None = None
-
- for raw_line in text.splitlines():
- line = raw_line.rstrip()
- stripped = line.strip()
-
- if stripped == "<<
>>":
- flush_paragraph(story, paragraph_buffer, styles["body"])
- if list_type:
- flush_list(story, list_buffer, list_type, styles)
- list_type = None
- story.append(PageBreak())
- continue
-
- if stripped == "":
- flush_paragraph(story, paragraph_buffer, styles["body"])
- if list_type:
- flush_list(story, list_buffer, list_type, styles)
- list_type = None
- continue
-
- if stripped == "---":
- flush_paragraph(story, paragraph_buffer, styles["body"])
- if list_type:
- flush_list(story, list_buffer, list_type, styles)
- list_type = None
- story.append(Spacer(1, 2 * mm))
- story.append(HRFlowable(width="100%", thickness=0.5, color=colors.HexColor("#d1d5db")))
- story.append(Spacer(1, 3 * mm))
- continue
-
- if stripped.startswith("# "):
- flush_paragraph(story, paragraph_buffer, styles["body"])
- if list_type:
- flush_list(story, list_buffer, list_type, styles)
- list_type = None
- story.append(Paragraph(escape_text(stripped[2:].strip()), styles["h1"]))
- continue
-
- if stripped.startswith("## "):
- flush_paragraph(story, paragraph_buffer, styles["body"])
- if list_type:
- flush_list(story, list_buffer, list_type, styles)
- list_type = None
- story.append(Paragraph(escape_text(stripped[3:].strip()), styles["h2"]))
- continue
-
- if stripped.startswith("### "):
- flush_paragraph(story, paragraph_buffer, styles["body"])
- if list_type:
- flush_list(story, list_buffer, list_type, styles)
- list_type = None
- story.append(Paragraph(escape_text(stripped[4:].strip()), styles["h3"]))
- continue
-
- bullet_match = re.match(r"^[-*]\s+(.*)$", stripped)
- if bullet_match:
- flush_paragraph(story, paragraph_buffer, styles["body"])
- if list_type and list_type != "bullet":
- flush_list(story, list_buffer, list_type, styles)
- list_type = None
- list_type = "bullet"
- list_buffer.append(bullet_match.group(1).strip())
- continue
-
- ordered_match = re.match(r"^\d+\.\s+(.*)$", stripped)
- if ordered_match:
- flush_paragraph(story, paragraph_buffer, styles["body"])
- if list_type and list_type != "ordered":
- flush_list(story, list_buffer, list_type, styles)
- list_type = None
- list_type = "ordered"
- list_buffer.append(ordered_match.group(1).strip())
- continue
-
- if not story or len(story) <= 4:
- flush_paragraph(story, paragraph_buffer, styles["body"])
- if list_type:
- flush_list(story, list_buffer, list_type, styles)
- list_type = None
- story.append(Paragraph(escape_text(stripped), styles["meta"]))
- continue
-
- paragraph_buffer.append(stripped)
-
- flush_paragraph(story, paragraph_buffer, styles["body"])
- if list_type:
- flush_list(story, list_buffer, list_type, styles)
-
- return story
-
-
-def build_pdf(source_path: Path, output_path: Path, doc_title: str, styles) -> None:
- output_path.parent.mkdir(parents=True, exist_ok=True)
- text = source_path.read_text(encoding="utf-8")
- story = markdown_to_story(text, doc_title, styles)
-
- doc = SimpleDocTemplate(
- str(output_path),
- pagesize=A4,
- leftMargin=20 * mm,
- rightMargin=20 * mm,
- topMargin=18 * mm,
- bottomMargin=18 * mm,
- title=doc_title,
- author="Codex for LyraNote",
- )
- doc.build(story, onFirstPage=add_page_number, onLaterPages=add_page_number)
-
-
-def main() -> None:
- register_fonts()
- styles = build_styles()
- for item in DOCUMENTS:
- build_pdf(item["source"], item["output"], item["title"], styles)
- print(f"generated: {item['output'].relative_to(ROOT)}")
-
-
-if __name__ == "__main__":
- main()