TMA

图片来源:Intel VTune Profiler Cookbook:Top-down Microarchitecture Analysis Method。 1.CPU Architecture CPU流水线架构 前端:取指令、解码、把uOP发送到后端 后端:allocate uOP、执行uOP,执行完成(retirement) 流水线架构图 ROB:ReOrder Buffer,乱序缓冲区,指令是乱序执行的,但要按照原顺序提交运行结果(in-order retirement),用来暂存乱序的指令结果,一般实现是一个环形队列 后端很多个执行端口有不同的执行单元 L1 I/D cache,L2 cache一般每核私有,LLC一般共享 pipeline slot:一个周期内处理一个uop所需的硬件资源/位置,intel CPU FE/BE:4 slots/cycle。 slot的状态在allocate处获取 一个slot在一个周期内是空的,被归于stall,使用TMA可以将stall归因到FE/BE,前端无法用uOP填充槽:FE-bound;前端已经准备好uop,但是后端无法处理:BE-bound 2. TMA TMA L1:每个pipeline slot被分为四类 Bad speculation:预测指令取错了,结果不能写回(flush pipeline) retire:可以从ROB把结果写回CPU state TMA L2 Backend Bound(mostly) Core Bound port utilization divider xxx Memory Bound(mostly) L1/L2/L3/DRAM Bound:从cache到主存 store bound xxx Frontend Bound(并不常见) Frontend Latency: Frontend Bandwidth:前端处理小于4uop的周期 Bad Speculation branch mispredict machine clears Retiring 3. CPU Metrics Dictionary Cycles-based:所占cycles比例,以时间度量通常没有slots-based按照资源度量准确 ...

September 20, 2026 · 2 min · sudo