Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
215 commits
Select commit Hold shift + click to select a range
5a8c47c
Add: document A5 FDWIC Submit atomic analysis
Jul 17, 2026
3a3de54
Update: 重命名 PA atomic 情况分析文档
Jul 17, 2026
a3f5ecc
test(a5): isolate AICore submit inline boundary
Jul 17, 2026
290dbda
test(a5): 增加独立 PA 调度性能复现用例
Jul 17, 2026
76df85c
test(atomic_probe): 验证 atomic 与 I-cache 等待周期归类
Jul 18, 2026
04ec9b9
perf(a5): 跳过HeapGuard首圈冗余原子读取
Jul 18, 2026
3d174a0
文档(a5): 记录fanin顺序实验与回退结论
Jul 18, 2026
3d0aaea
工具(a5): 建立独立PA每核标量PMU观察链路
Jul 18, 2026
67407cc
测试(a5): 细分PA依赖与frontier原子计数
Jul 18, 2026
8deefde
修复PMU owner活跃子核配置与恢复闭环
Jul 18, 2026
13431a2
工具(a5): 补齐 PA 调度器逐 atomic 泳道观测
Jul 18, 2026
c93bd65
工具(a5): 接通独立PA标量PMU所有权闭环
Jul 18, 2026
640efe5
测试(a5): 建立单次I-cache miss性能标尺
Jul 18, 2026
99971ac
工具(a5): 补齐独立PA Submit全窗PMU取证
Jul 18, 2026
c6daaeb
文档(a5): 固化PA原子与Submit PMU观察口径
Jul 18, 2026
5274945
合并(a5): 吸纳单次I-cache miss性能标尺
Jul 18, 2026
e66001f
新增 standalone CCEC 真实 Cube/Vector 负载
Jul 18, 2026
9aeda0d
测试(a5): 为standalone AscendC接入真实Cube与Vector负载
Jul 18, 2026
1d3a374
测试(a5): 补齐standalone CPU对等算术负载
Jul 18, 2026
0c9cebc
测试(a5): 增加真实引擎非均匀布局诊断
Jul 18, 2026
7bb118a
测试(a5): 默认启用standalone真实引擎负载
Jul 18, 2026
cbaf7c6
工具(a5): 接通真实PA atomic泳道与精确轮询聚合
Jul 18, 2026
44199a5
工具(a5): 固化Submit全窗I-cache逐核分析口径
Jul 18, 2026
187e54b
工具(a5): 完善standalone atomic合并泳道
Jul 18, 2026
5c2d39b
工具(a5): 固化standalone Submit PMU观测口径
Jul 18, 2026
8430f41
工具(a5): 增加EfDrain Submit PMU独立归因
Jul 18, 2026
7466e6f
工具(a5): 完善Submit PMU局部归因与I-cache可视报告
Jul 19, 2026
6caa269
工具(a5): 收敛Submit观测边界与排他泳道分析
Jul 19, 2026
d2d8ce2
优化(a5): 将低频winner调整为冷分支
Jul 19, 2026
cafa9ca
优化(a5): 恢复真实PA的loser热路布局
Jul 19, 2026
4436797
优化(a5): 外提atomic泳道冷路径消除取指回退
Jul 19, 2026
14c2429
文档(a5): 分离atomic与I-cache分析记录
Jul 19, 2026
dbb95bb
Support: 将 Submit 排他泳道分析移植到 FDWIC
Jul 19, 2026
911ecf9
文档(a5): 完善standalone排他泳道与性能分布分析
Jul 19, 2026
ba4334d
验证(pa): 交付 compete-first/lazy 三版独立对照
Jul 20, 2026
0d08c43
优化(pa): standalone采用compete-first eager提交流程
Jul 20, 2026
2899cc3
重构(pa): 真实路径接入compete-first eager提交流程
Jul 20, 2026
8ef337c
Docs: 完善 PA 调度器泳道主流程说明
Jul 21, 2026
2e92da1
文档(pa): 记录cursor分片跨路径性能对照
Jul 21, 2026
9ae512f
实验(pa): 分层final汇合选择二级G8
Jul 21, 2026
9b871d8
Update: select two-level G16 for standalone final barrier
Jul 21, 2026
c9ea57c
Update: use two-level G16 final barrier in FDWIC
Jul 21, 2026
caa7007
test: 新增 st_dev/ld_dev 多核同步探针
Jul 22, 2026
74e6d07
文档(a5): 建立PA Submit性能优化全过程记录
Jul 20, 2026
8b2702b
性能(a5): 建立真实PA低扰动perf-clock基线
Jul 20, 2026
fa0ebf0
验证(a5): 收口真实PA业务与atomic合并泳道
Jul 20, 2026
073caba
观察(a5): 建立真实PA Submit整窗PMU证据链
Jul 21, 2026
35a51bf
观察(a5): 建立真实PA构参区间单阶段PMU
Jul 21, 2026
a1d895d
校准(a5): 量化真实PA局部PMU观察器开销
Jul 21, 2026
e0ac51c
观察(a5): 建立真实PA Materialize单阶段PMU
Jul 21, 2026
fac560d
观察(a5): 建立真实PA Claim单阶段PMU
Jul 21, 2026
7f6f40b
观察(a5): 建立真实PA Register单阶段PMU
Jul 21, 2026
262c830
观察(a5): 建立真实PA Submit间隙单阶段PMU
Jul 21, 2026
e1c77de
校准(a5): 交错量化三类观察构建整体影响
Jul 21, 2026
ed36fb9
分析(a5): 收口真实PA独占设备波动来源
Jul 21, 2026
539e4c1
分析(a5): 收口Submit整窗PMU波动载体
Jul 21, 2026
71cf0b1
观察(a5): 建立Submit窗内Kernel低容量聚合
Jul 21, 2026
968b371
分析(a5): 收口Kernel与residual波动载体
Jul 21, 2026
8ecbbee
分析(a5): 记录设备状态与低功耗取证边界
Jul 21, 2026
90299d1
分析(a5): 排除Materialize为波动主载体
Jul 21, 2026
1af240c
分析(a5): 排除Claim为波动主载体
Jul 21, 2026
5818037
分析(a5): 排除SubmitTransition为波动主载体
Jul 21, 2026
92816a1
分析(a5): 排除ArgBuild为波动主载体
Jul 21, 2026
2d2d932
分析(a5): 排除Register为波动主载体
Jul 21, 2026
636abe7
观察(a5): 完善I-cache逐核时间报告
Jul 21, 2026
40096ee
观察(a5): 为Submit PMU绑定实际构建身份
Jul 21, 2026
019fb80
验证(a5): 闭合Submit PMU构建身份三件套
Jul 21, 2026
fe7e266
观察(a5): 建立EfDrain控制段I-cache归因
Jul 21, 2026
d9117a8
验证(a5): 闭合EfDrain控制段上板数据
Jul 21, 2026
474e198
分析(a5): 形成EfDrain控制段首轮稳态证据
Jul 21, 2026
a270b5f
分析(a5): 撤回BlockWon慢路冷外提
Jul 21, 2026
6eace28
测试(a5): 补齐I-cache观察链组合闭环
Jul 21, 2026
7b6cd34
测试(a5): 直接闭合FDWIC TensorMap清退语义
Jul 21, 2026
6b85d4f
优化候选(a5): 跳过PrepareMap空任务头写回
Jul 21, 2026
5bbce5a
文档(a5): 记录PrepareMap候选离线证据
Jul 21, 2026
b04f592
文档(a5): 清理性能记录中的多余换行
Jul 22, 2026
7bd59a8
文档(a5): 记录分角色I-cache miss精确标尺
Jul 22, 2026
ac0dd63
测试(a5): 验证Vector计算的scalar busy归类
Jul 22, 2026
75d2538
测试记录(a5): 归档7月23日Submit观测报告
Jul 23, 2026
1273492
记录(a5): 归档Submit-PMU v3阶段报告
Jul 23, 2026
287e0d8
测试记录(a5): 收敛7月23日Submit-PMU最新报告
Jul 23, 2026
499fac1
补齐 PrepareMap 的 Submit-PMU 归因
Jul 22, 2026
1b15737
补齐 Fanin 的动态 Submit-PMU 归因
Jul 22, 2026
ad4d96e
观测(a5): 重建纯Scalar Submit-PMU v2并补齐WinnerBuild归因
Jul 22, 2026
11359ce
记录(a5): 登记Submit-PMU v2全量重采产物
Jul 22, 2026
7a6b5ad
观测(a5): 补齐AllocComplete纯Scalar归因
Jul 22, 2026
7cd5542
观测(a5): 补齐LoserReplay纯Scalar归因
Jul 22, 2026
a9a2c71
观测(a5): 汇总Submit全span双证据链
Jul 22, 2026
38af799
观测(a5): 完善Submit-PMU v3并量化分段观察扰动
Jul 23, 2026
5c8583b
观测(a5): 拆清Submit分段记录开销与业务参考值
Jul 23, 2026
6fbea59
报告(a5): 收紧Submit观测表格布局
Jul 23, 2026
1726a77
文档: 同步核内全分布设计说明
Jul 23, 2026
150040a
文档(a5): 固化shared TensorMap审查与standalone优先路线
Jul 24, 2026
ab176e3
构建(a5): 为standalone建立TensorMap模式身份与ABI门禁
Jul 24, 2026
6944bfd
调度(a5): 将standalone private TensorMap同构为有界桶环
Jul 24, 2026
9fcb27f
调度(a5): 在standalone接入有序shared TensorMap基线
Jul 24, 2026
fd10916
调度(a5): 用ordered winner推进shared回收前沿
Jul 25, 2026
3a96fa0
调度(a5): 在standalone接入fresh-output符号链
Jul 25, 2026
95b4ea4
调度(a5): 用shared heap收敛winner输出物化
Jul 25, 2026
1291748
调度(a5): 将shared重构参收敛到winner
Jul 25, 2026
bb48200
观测(a5): 为standalone建立低扰动perf-clock基线
Jul 25, 2026
b5f4edc
文档(a5): 记录shared全局turn的b256活性反例
Jul 25, 2026
8d60f93
调度(a5): 将shared符号发布从全局turn中拆出
Jul 25, 2026
aa6e314
调度(a5): 让shared分片heap具备并发分配语义
Jul 25, 2026
3543854
调度(a5): 将shared符号提交收敛到winner构建后
Jul 25, 2026
dc22d07
调度(a5): 从shared PA Case1移除全局提交前沿
Jul 25, 2026
eed1c38
观测(a5): 冻结shared S4.6并完成配对基线
Jul 25, 2026
2514ef1
调度(a5): 让shared ready descriptor直写私有slot
Jul 25, 2026
9914a47
观测(a5): 记录ready直写与S4.6配对结果
Jul 25, 2026
d004269
调度(a5): 融合shared描述符直写与slot填充扫描
Jul 25, 2026
8d2268c
观测(a5): 对照standalone与参考shared调度时间
Jul 25, 2026
e832028
调度(a5): shared no-wrap完成路径移除frontier推进
Jul 25, 2026
a6ff810
观测(a5): 冻结验证shared无frontier性能收益
Jul 25, 2026
e83283f
调度(a5): 固定shared Alloc候选并消减无效Claim
Jul 25, 2026
64be531
观测(a5): 记录shared Alloc候选过滤配对回退
Jul 25, 2026
f41e283
调度(a5): 让shared Alloc非候选完整早退
Jul 25, 2026
d3dadce
回退(a5): 撤销shared Alloc固定候选与完整早退
Jul 25, 2026
bccf56f
分析(a5): 收敛shared下一候选为纯INPUT延迟解析
Jul 25, 2026
b516409
调度(a5): 延迟解析shared纯INPUT描述符
Jul 25, 2026
9d43c06
分析(a5): 记录pure INPUT首轮冻结配对回退
Jul 25, 2026
6275e32
优化(a5): 收敛shared延迟解析代码体积
Jul 25, 2026
91d2352
回退(a5): 撤销shared纯INPUT延迟解析
Jul 25, 2026
629ae13
分析(a5): 收敛shared loser轻返回候选
Jul 25, 2026
b2fe435
优化(a5): shared TensorMap loser 在 Claim 后直接返回
Jul 25, 2026
3a816ff
回退(a5): 撤销无稳定收益的 shared loser 快返
Jul 25, 2026
327de85
优化(a5): 分散shared Alloc Claim到24个owner
Jul 25, 2026
4964447
回退(a5): 撤销性能中性的shared 24-owner Claim
Jul 25, 2026
e24e579
实验(a5): 建立shared Vector cursor四分片迁址对照
Jul 25, 2026
0b7762b
记录(a5): 确认shared Vector四分片迁址收益
Jul 25, 2026
ee42b8c
实验(a5): 启用shared Vector同址八分片
Jul 25, 2026
7f57ee2
记录(a5): 确认shared Vector同址八分片收益
Jul 25, 2026
bab00e3
实验(a5): 建立shared Cube cursor四分片迁址对照
Jul 25, 2026
319077a
回退(a5): 撤销性能回退的shared Cube迁址
Jul 25, 2026
e719cd1
实验(a5): 建立shared Vector十六线容量控制
Jul 25, 2026
48ce9fc
记录(a5): 固化Vector十六线容量控制配对
Jul 25, 2026
2e7a0c7
优化(a5): 启用shared Vector十六分片
Jul 25, 2026
bf7a707
回退(a5): 撤销未过门槛的Vector十六分片
Jul 25, 2026
82c0828
优化(a5): 前置shared worker热控制字段
Jul 25, 2026
1572d7d
回退(a5): 撤销中性收益的shared热字段前置
Jul 25, 2026
5a89038
构建(a5): 建立FDWIC TensorMap双模式身份与三镜像隔离
Jul 26, 2026
54e0d42
分析(a5): 辨析shared迁移review并收敛standalone观察缺口
Jul 26, 2026
7053e60
观测(a5): 将PollBatch启用位改为紧凑索引
Jul 26, 2026
9ad5c98
观测(a5): 补齐shared heap原子泳道
Jul 26, 2026
cc53d22
正确性(a5): 建立shared多级writer发布门原语并隔离private
Jul 26, 2026
e19b519
正确性(a5): 用真实PA双组参数验证shared writer intent
Jul 26, 2026
24eb97e
测试(a5): 集中整理PA scheduler门槛用例
Jul 26, 2026
93f0c85
正确性(a5): 接通shared双组Finish协议与异常收敛
Jul 26, 2026
ee0fe8c
正确性(a5): 建立shared动态task plan门槛
Jul 26, 2026
2c163fa
正确性(a5): 让shared设备回放消费动态task plan
Jul 26, 2026
10cfbe9
正确性(host): 建立shared权威动态task plan
Jul 26, 2026
2ddb710
正确性(泳道): 按Alloc边界恢复shared动态task身份
Jul 26, 2026
40b8c06
正确性(PMU): 对齐shared动态Submit计划
Jul 26, 2026
12defd4
正确性(shared): 接受G0合法零计算输出
Jul 26, 2026
d4d917a
验证(shared): 闭合动态task的A5 PMU矩阵
Jul 26, 2026
ab1da1b
验证(shared): 固化真实B256泳道与性能基线
Jul 26, 2026
ce1ae7c
正确性(shared): 在worker启动前完成heap容量准入
Jul 26, 2026
08d40dd
正确性(shared): 闭合G2放门后Build失败
Jul 26, 2026
47d22e3
文档(shared): 冻结standalone PA迁移基线
Jul 26, 2026
9fc3681
性能(shared): 收敛loser轻路径与稀疏观测
Jul 26, 2026
c4c4e4c
重构(a5): 抽取TensorMap双后端门面并冻结private行为
Jul 26, 2026
4053973
测试(standalone): 参数化16K连续分桶环并补齐容量边界
Jul 26, 2026
ce0cdd2
构建(standalone): 将默认TensorMap容量纳入CCEC产物身份
Jul 26, 2026
9343a34
正确性(fdwic): 闭合TensorMap容量失败传播
Jul 26, 2026
a765a3c
正确性(fdwic): 严格校验TensorMap历史窗口
Jul 26, 2026
c8c81ac
构建(fdwic): 将TensorMap容量纳入三镜像身份
Jul 26, 2026
ae3ef37
测试(fdwic): 冻结private TensorMap逻辑语义
Jul 26, 2026
68f5145
重构(fdwic): 将private TensorMap迁移为连续分桶环
Jul 26, 2026
9577be4
架构(fdwic): 冻结shared TensorMap尾部状态
Jul 26, 2026
0c3b362
重构(fdwic): 抽取TensorMap共用逻辑原语
Jul 26, 2026
3943a82
功能(fdwic): 实现有序shared TensorMap环原语
Jul 26, 2026
177a09c
正确性(fdwic): 用CAS加固shared TensorMap发布
Jul 26, 2026
ebe3ff2
正确性(fdwic): 区分shared TensorMap部分发布
Jul 26, 2026
6e7e8af
接口(fdwic): 冻结shared TensorMap错误合同
Jul 26, 2026
f2c715f
功能(fdwic): 接入shared TensorMap基础Submit事务
Jul 26, 2026
9d5cd45
测试(fdwic): 固定shared future-turn多核恢复
Jul 26, 2026
3428186
正确性(fdwic): 让shared等待在远端fatal后收敛
Jul 26, 2026
6e9294a
测试(fdwic): 固定shared三核FinalDrain正向闭环
Jul 26, 2026
57cb097
测试(fdwic): 固定shared joint最后一核完成闭环
Jul 26, 2026
ebe1e47
正确性(fdwic): 补齐PA G2与manual_dep依赖合同
Jul 27, 2026
a19d7e2
测试(fdwic): 闭合shared PA G2跨核writer链
Jul 27, 2026
35c4e59
工程(atomic-probe): 固化中文standalone的预提交边界
Jul 27, 2026
adb87c6
正确性(pa-scheduler): 用CAS加固writer-ready发布门
Jul 27, 2026
b072ce0
工程(atomic-probe): 扩大中文注释检查豁免范围
Jul 27, 2026
8a0b832
正确性(pa-scheduler): 建立通用WriterIntentSet基础协议
Jul 27, 2026
c1a9617
正确性(pa-scheduler): 为shared symbol补齐不可变writer历史
Jul 27, 2026
cddc98b
正确性(pa-scheduler): 闭合shared symbol历史的A5跨核门槛
Jul 27, 2026
b94a22b
正确性(pa-scheduler): 建立shared ordinary reader完成前沿
Jul 27, 2026
9190ae4
正确性(pa-scheduler): 闭合shared满环慢reader回收门槛
Jul 27, 2026
172ab93
构建(pa-scheduler): 闭合shared reader协议的CCEC实例化
Jul 27, 2026
05bcc05
重构(pa-scheduler): 泛化shared协议A5门槛载体
Jul 27, 2026
7a29796
正确性(pa-scheduler): 闭合shared reader回收A5门槛
Jul 27, 2026
41ae20c
正确性(pa-scheduler): 容忍shared合法前缀并发回收
Jul 27, 2026
a66ebff
正确性(pa-scheduler): 建立reader前沿整批追加
Jul 27, 2026
92eb1bf
正确性(pa-scheduler): 让TensorMap前沿发布失败不改状态
Jul 27, 2026
b50c53c
正确性(pa-scheduler): 仅串行化shared TensorMap插入
Jul 27, 2026
c587635
正确性(pa-scheduler): 闭合shared有序插入分组前沿
Jul 27, 2026
7070603
观测(pa-scheduler): 拆分shared Register内部耗时
Jul 28, 2026
37b3368
测试(atomic): 补齐TaskCell共线与独占行DCCI验证
Jul 28, 2026
b0f846f
优化(pa-scheduler): 收敛shared有序插入与输出发布
Jul 28, 2026
b2e03e5
优化(pa-scheduler): 消减shared有序Register重复工作
Jul 28, 2026
c3aaf99
优化(pa-scheduler): 将shared静态提交计划移出有序区
Jul 28, 2026
bbd1877
文档(pa-scheduler): 分析SharedOutputRef依赖查询选型
Jul 28, 2026
915fc24
文档(pa-scheduler): 重写SharedOutputRef依赖方案对比
Jul 28, 2026
e42aba5
完善(pa-scheduler): 增加winner与loser闭合路径分析
Jul 29, 2026
9fe630b
优化(pa-scheduler): 延后shared winner上下文初始化
Jul 28, 2026
9a813bc
优化(pa-scheduler): 消除shared EfDrain重复fanin读取
Jul 28, 2026
9e97dbb
文档(pa-scheduler): 按1%门槛复核shared三项优化
Jul 29, 2026
ad52c01
测试(cache-preload): 补齐 CCEC 与 AscendC 预取探针
Jul 29, 2026
f0903f6
测试(cache-preload): 区分 store-only 与 GM 发布预取收益
Jul 29, 2026
0f51a06
测试(cache-preload): 验证持续泳道写预取收益
Jul 29, 2026
9a529e3
测试(cache-preload): 补充 shared 路径预取模型与 A5 结论
Jul 29, 2026
95a293f
优化(pa-scheduler): 压缩shared完整泳道并闭合Atomic与DCCI
Jul 29, 2026
aa1698a
测试(shared-tensormap): 增加A5跨核可见性探针
Jul 29, 2026
325b48d
文档(pa-scheduler): 记录Register差异与16B泳道收益
Jul 29, 2026
c838f51
文档(shared-tensormap): 补齐有序单写与INOUT协议分析
Jul 29, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
14 changes: 12 additions & 2 deletions .pre-commit-config.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,10 @@
# See LICENSE in the root of the software repository for the full text of the License.
# -----------------------------------------------------------------------------------------------------------

# This self-contained probe intentionally carries generated binaries, Chinese
# source annotations, and performance outputs for offline comparison.
exclude: ^tests/atomic_probe/lazy_lamda_sample/

repos:

# Common
Expand All @@ -23,7 +27,10 @@ repos:
entry: python tests/lint/check_english_only.py
language: python
language_version: python3
exclude: ^(3rdparty/|docs/zh-cn/|README\.zh-CN\.md)
# The atomic-probe tree contains standalone probes, usage guides and
# investigation records that are intentionally maintained in Chinese
# for the A5 device-performance workflow.
exclude: ^(3rdparty/|docs/zh-cn/|README\.zh-CN\.md|tests/atomic_probe/)

- repo: https://github.com/pre-commit/pre-commit-hooks
rev: v4.6.0
Expand All @@ -39,7 +46,10 @@ repos:
hooks:
- id: clang-format
types_or: [c++, c]
exclude: ^3rdparty/
# This self-contained probe preserves its hand-reviewed CCEC layout;
# formatting a touched legacy header rewrites thousands of unrelated
# lines and obscures the protocol delta under test.
exclude: ^(3rdparty/|tests/atomic_probe/pa_scheduler/)

- repo: local
hooks:
Expand Down
225 changes: 213 additions & 12 deletions conftest.py
Original file line number Diff line number Diff line change
Expand Up @@ -149,6 +149,55 @@ def pytest_addoption(parser):
help="Enable L2 swimlane. Bare flag=level 4 (full). "
"1=AICore timing, 2=+dispatch/fanout, 3=+sched phases, 4=+orch phases",
)
parser.addoption(
"--fdwic-tensormap",
action="store",
choices=["private", "shared"],
default="private",
help="Select the compile-time TensorMap artifact family for the a5/a5sim "
"fully_distributed_within_core runtime. The default is private.",
)
parser.addoption(
"--fdwic-profile",
action="store",
choices=[
"none",
"perf-clock",
"perf-clock-kernel",
"submit-pmu-none",
"submit-pmu-arg-build",
"submit-pmu-empty-bracket",
"submit-pmu-materialize",
"submit-pmu-claim",
"submit-pmu-register",
"submit-pmu-submit-transition",
"submit-pmu-efdrain-control",
"submit-pmu-prepare-map",
"submit-pmu-fanin",
"submit-pmu-winner-build-control",
"submit-pmu-alloc-complete-control",
"submit-pmu-loser-replay",
],
default="none",
help="Select a private fully_distributed_within_core evidence build. "
"perf-clock keeps only the first/last Submit device clock per core; "
"perf-clock-kernel additionally aggregates linked-kernel time/calls inside that per-core window; "
"submit-pmu-none keeps one full Submit-sequence scalar/I-cache PMU window per core; "
"submit-pmu-arg-build attributes the Claim-to-Materialize eager-build interval; "
"submit-pmu-empty-bracket calibrates the adjacent begin/end observer cost at Claim.end; "
"submit-pmu-materialize attributes the current Materialize business span; "
"submit-pmu-claim attributes the current Claim business span; "
"submit-pmu-register attributes the RegisterOutputs call body; "
"submit-pmu-submit-transition attributes adjacent Submit gaps; "
"submit-pmu-efdrain-control attributes EfDrain scalar control while excluding linked Kernel calls; "
"submit-pmu-prepare-map attributes the dist_submit_prepare_map call body; "
"submit-pmu-fanin attributes the dynamic Kernel-winner Fanin span; "
"submit-pmu-winner-build-control attributes scalar control inside the complete WinnerBuild time boundary "
"while excluding linked Kernel calls; "
"submit-pmu-alloc-complete-control attributes scalar control inside the complete AllocComplete time boundary "
"while excluding linked Kernel calls; "
"submit-pmu-loser-replay attributes the real Kernel-loser drain_block_won call body.",
)
parser.addoption(
"--use-example-exec-time",
action="store_true",
Expand Down Expand Up @@ -433,6 +482,85 @@ def _configure_sanitizer(config):
)


def _configure_fdwic_profile(config):
"""Validate and publish the private real-A5 FDWIC evidence profile."""
fdwic_profile = config.getoption("--fdwic-profile", default="none")
if fdwic_profile == "none":
os.environ.pop("PTO_FDWIC_PROFILE", None)
return
if fdwic_profile not in {
"perf-clock",
"perf-clock-kernel",
"submit-pmu-none",
"submit-pmu-arg-build",
"submit-pmu-empty-bracket",
"submit-pmu-materialize",
"submit-pmu-claim",
"submit-pmu-register",
"submit-pmu-submit-transition",
"submit-pmu-efdrain-control",
"submit-pmu-prepare-map",
"submit-pmu-fanin",
"submit-pmu-winner-build-control",
"submit-pmu-alloc-complete-control",
"submit-pmu-loser-replay",
}:
raise pytest.UsageError(f"unsupported --fdwic-profile {fdwic_profile!r}")

platform = config.getoption("--platform", default=None)
runtime = config.getoption("--runtime", default=None)
level = config.getoption("--level", default=None)
if platform != "a5":
raise pytest.UsageError(f"--fdwic-profile {fdwic_profile} requires --platform a5")
if runtime not in {None, "fully_distributed_within_core"}:
raise pytest.UsageError(f"--fdwic-profile {fdwic_profile} only supports runtime fully_distributed_within_core")
if level not in {None, 2}:
raise pytest.UsageError(f"--fdwic-profile {fdwic_profile} only supports SceneTest level 2")
if config.getoption("--rounds", default=1) != 1:
raise pytest.UsageError(
f"--fdwic-profile {fdwic_profile} requires --rounds 1 because its per-case artifact is single-run"
)
conflicting = []
for option, label in (
("--enable-l2-swimlane", "--enable-l2-swimlane"),
("--dump-args", "--dump-args"),
("--enable-pmu", "--enable-pmu"),
("--enable-dep-gen", "--enable-dep-gen"),
("--enable-scope-stats", "--enable-scope-stats"),
("--enable-device-log-timing", "--enable-device-log-timing"),
("--enable-swimlane-overhead", "--enable-swimlane-overhead"),
("--use-example-exec-time", "--use-example-exec-time"),
):
if config.getoption(option, default=0):
conflicting.append(label)
if conflicting:
raise pytest.UsageError(
f"--fdwic-profile {fdwic_profile} must run without other diagnostics: " + ", ".join(conflicting)
)
os.environ["PTO_FDWIC_PROFILE"] = fdwic_profile


def _configure_fdwic_tensormap(config):
"""Validate and publish the explicit FDWIC TensorMap artifact family."""
mode = config.getoption("--fdwic-tensormap", default="private")
if mode == "private":
os.environ.pop("PTO_FDWIC_TENSORMAP_MODE", None)
return
if mode != "shared":
raise pytest.UsageError(f"unsupported --fdwic-tensormap {mode!r}")

platform = config.getoption("--platform", default=None)
runtime = config.getoption("--runtime", default=None)
level = config.getoption("--level", default=None)
if platform not in {"a5", "a5sim"}:
raise pytest.UsageError(f"--fdwic-tensormap {mode} requires --platform a5 or a5sim")
if runtime not in {None, "fully_distributed_within_core"}:
raise pytest.UsageError(f"--fdwic-tensormap {mode} only supports runtime fully_distributed_within_core")
if level not in {None, 2}:
raise pytest.UsageError(f"--fdwic-tensormap {mode} only supports SceneTest level 2")
os.environ["PTO_FDWIC_TENSORMAP_MODE"] = mode


def pytest_configure(config):
"""Register custom markers and apply global config."""
config.addinivalue_line("markers", "platforms(list): supported platforms for standalone ST functions")
Expand All @@ -445,6 +573,8 @@ def pytest_configure(config):
)

_configure_sanitizer(config)
_configure_fdwic_tensormap(config)
_configure_fdwic_profile(config)

# Configure logging unconditionally (not only when --log-level is passed) so
# simpler's own WARNINGs — e.g. the device-log-timing "no device log written"
Expand Down Expand Up @@ -613,6 +743,65 @@ def sort_key(item):

items.sort(key=sort_key)

fdwic_profile = config.getoption("--fdwic-profile", default="none")
if fdwic_profile in {
"perf-clock",
"perf-clock-kernel",
"submit-pmu-none",
"submit-pmu-arg-build",
"submit-pmu-empty-bracket",
"submit-pmu-materialize",
"submit-pmu-claim",
"submit-pmu-register",
"submit-pmu-submit-transition",
"submit-pmu-efdrain-control",
"submit-pmu-prepare-map",
"submit-pmu-fanin",
"submit-pmu-winner-build-control",
"submit-pmu-alloc-complete-control",
"submit-pmu-loser-replay",
}:
incompatible = []
for item in items:
if any(m.name == "skip" for m in item.iter_markers()):
continue
cls = getattr(item, "cls", None)
if cls is None:
incompatible.append(item.nodeid)
continue
if getattr(cls, "_st_level", None) != 2 or getattr(cls, "_st_runtime", None) != (
"fully_distributed_within_core"
):
incompatible.append(item.nodeid)
if incompatible:
sample = ", ".join(incompatible[:3])
more = "" if len(incompatible) <= 3 else f" (+{len(incompatible) - 3} more)"
raise pytest.UsageError(
f"--fdwic-profile {fdwic_profile} only accepts level-2 fully_distributed_within_core tests; "
f"incompatible item(s): {sample}{more}"
)

fdwic_tensormap = config.getoption("--fdwic-tensormap", default="private")
if fdwic_tensormap == "shared":
incompatible = []
for item in items:
if any(m.name == "skip" for m in item.iter_markers()):
continue
cls = getattr(item, "cls", None)
if (
cls is None
or getattr(cls, "_st_level", None) != 2
or getattr(cls, "_st_runtime", None) != "fully_distributed_within_core"
):
incompatible.append(item.nodeid)
if incompatible:
sample = ", ".join(incompatible[:3])
more = "" if len(incompatible) <= 3 else f" (+{len(incompatible) - 3} more)"
raise pytest.UsageError(
"--fdwic-tensormap shared only accepts level-2 fully_distributed_within_core tests; "
f"incompatible item(s): {sample}{more}"
)

# L3 perf collection is not supported yet: a single L3 case forks N chip-processes
# that all write l2_swimlane_records_<ts>.json to the same directory with
# second-precision timestamps, so they trample each other. Block the
Expand Down Expand Up @@ -1267,6 +1456,29 @@ def _l2_poisoned():
return set()


def _fdwic_worker_build_config(cls, platform, runtime):
"""Prepare mode-aware FDWIC worker arguments and pool identity."""
if runtime != "fully_distributed_within_core" or platform not in {"a5", "a5sim"}:
return {}, ""

from simpler_setup.scene_test import ( # noqa: PLC0415
_fdwic_tensormap_mode,
get_aicore_path_override,
)

cache_key = (cls.__qualname__, platform, runtime)
cls.compile_chip_callable(platform)
tensormap_mode = _fdwic_tensormap_mode()
kwargs = {"fdwic_tensormap_mode": tensormap_mode}
pool_token = f"{tensormap_mode}:"
aicore_override = get_aicore_path_override(cache_key)
if aicore_override is not None:
aicore_override = aicore_override.resolve()
kwargs["aicore_path_override"] = aicore_override
pool_token += str(aicore_override)
return kwargs, pool_token


@pytest.fixture()
def st_worker(request, st_platform, device_pool, _l2_worker_pool, _l2_poisoned):
"""Per-test Worker.
Expand Down Expand Up @@ -1297,18 +1509,7 @@ def st_worker(request, st_platform, device_pool, _l2_worker_pool, _l2_poisoned):

from simpler.worker import Worker # noqa: PLC0415

kwargs = {}
aicore_pool_token = ""
if runtime == "fully_distributed_within_core" and st_platform in {"a5", "a5sim"}:
from simpler_setup.scene_test import get_aicore_path_override # noqa: PLC0415

cache_key = (cls.__qualname__, st_platform, runtime)
cls.compile_chip_callable(st_platform)
aicore_override = get_aicore_path_override(cache_key)
if aicore_override is not None:
aicore_override = aicore_override.resolve()
kwargs["aicore_path_override"] = aicore_override
aicore_pool_token = str(aicore_override)
kwargs, aicore_pool_token = _fdwic_worker_build_config(cls, st_platform, runtime)

# L2 share: reuse any Worker already created for this runtime image in
# the current process. Under xdist, each worker process is sliced to a
Expand Down
71 changes: 68 additions & 3 deletions docs/dfx/l2-swimlane-profiling.md
Original file line number Diff line number Diff line change
Expand Up @@ -127,9 +127,10 @@ runs):

```text
<output_prefix>/
├── l2_swimlane_records.json # raw runtime output
├── name_map_<case>.json # optional func_id → name mapping
└── merged_swimlane.json # Perfetto trace (added by converter)
├── l2_swimlane_records.json # raw runtime output
├── name_map_<case>.json # optional func_id → name mapping
├── merged_swimlane.json # Perfetto trace (added by converter)
└── swimlane_exclusive_analysis.json # FDWIC schema-v4 only
```

Filenames are fixed (no per-file timestamp) — the directory is the
Expand Down Expand Up @@ -190,6 +191,70 @@ join key between `aicore_tasks` and `aicpu_tasks` is
canonical producer of `task_token_raw`; AICPU only stamps the
dispatch / finish timestamps and the per-core join token.

#### Fully-distributed-within-core schema-v4

The A5 fully-distributed-within-core (FDWIC) executor adds a strict
per-scalar-lane hierarchy at every enabled collection level. It does
not infer task kind from the task ID: `Submit.aux` is the source of
truth (`0` = kernel, `1` = allocation), and `Submit.flags & 1` records
the winner state.

The exact exclusive child order is:

| Submit path | Exclusive children in timestamp order |
| ----------- | ------------------------------------- |
| Kernel winner | `EfDrain`, `Materialize`, `PrepareMap`, `Claim`, `Fanin`, `Register`, `WinnerBuild` |
| Kernel loser | `EfDrain`, `Materialize`, `PrepareMap`, `Claim`, `Register`, `LoserReplay` |
| Alloc winner | `EfDrain`, `Materialize`, `PrepareMap`, `Register`, `Claim`, `AllocComplete` |
| Alloc loser | `EfDrain`, `Materialize`, `PrepareMap`, `Register`, `Claim` |

The `Submit` timestamp starts after `dist_submit_begin()` and ends
before publishing the Submit record and returning from the API. Its
residual therefore includes unmarked control and intermediate trace
record writes. Exact closure describes the instrumented observation
window; it is not a claim that instrumentation has zero performance
impact.

`LoserReplay` measures the production kernel-loser
`drain_block_won()` call. An allocation loser has no corresponding
action, so its Claim-to-Submit-end suffix remains a measured residual;
the tooling does not create a synthetic phase. `DrainWon`, `Atomic`,
`ClockBaseline`, `Commit`, and `RingBp` are nested observations and are
therefore reported as non-additive overlays.

Each core emits exactly one adjacent pair of top-level parents:
`OrchestrationReplay` followed by `FinalDrain`. This worker-completion
window starts after the startup barrier and ends
before clock baselines, trace flush, and finish publication. Those
observation/lifecycle operations are deliberately outside the additive
business partition.

The analyzer validates the following identities with raw integer cycles
before any cycle-to-µs conversion:

```text
Submit = exclusive Submit children + SubmitResidual
SubmitEnvelope = SubmitUnion + BetweenSubmitResidual
OrchestrationReplay = Setup + SubmitUnion + BetweenSubmitResidual + Tail
FinalDrain = KernelUnion + FinalDrainResidual
WorkerCompletion = OrchestrationReplay + FinalDrain
```

Production orchestration can execute ready kernels while tensor-data
access waits between Submit calls. Such kernels are valid inside the
orchestration residual. Inside a Submit they are valid only in
`EfDrain`, `WinnerBuild`, or `AllocComplete`; a Kernel in Submit
residual or crossing a partition boundary is invalid. The report keeps
the cross-core wall-clock makespan separate from aggregate per-core
work, because summing core cycles is not elapsed wall time.

For schema-v4 input, `swimlane_converter` also writes
`swimlane_exclusive_analysis.json`. The merged trace contains explicit
`submit_residual`, `submit_tail_gap`, and
`between_submit_residual` spans derived from the validated raw records.
This schema change adds no I-cache/PMU metric; existing level-4 atomic
records remain overlays.

#### Reader output (µs domain)

After `read_perf_data()` joins the streams and converts to
Expand Down
Loading