AllReduce 训练一直锯齿?手把手教你用 hccn_tool + 交换机命令,30 分钟定位 RoCEv2 拥塞根因
如果你的 AllReduce 训练出现 Step Time 锯齿、吞吐掉 30%、Wireshark 看到 PSN 重传——大概率不是网卡坏了,而是拥塞控制链路出了三件套:ECN 门限太晚、PFC 被迫兜底、DSCP→TC 映射不一致。
这篇文章记录了一次真实的端到端排障过程,从 NPU 侧的 hccn_tool 命令开始,一路追到交换机侧的 PFC 帧计数、ECN 门限、队列丢包和 DiffServ 域配置。每一步都给出华为官方命令原文 + 真实回显格式 + 排障逻辑,不编字段、不脑补输出。
读完你会带走:
- 一套完整的 RoCEv2 拥塞排障思路(先判“被暂停”还是“真丢包”)
- 7 条 NPU 侧命令 + 6 条交换机侧命令的官方用法
- 一个可以直接贴在运维手册里的排障检查单
0. 故障现象
智算集群跑 AllReduce 大消息:
- Step Time 从 1.2s 跳到 1.8s,呈锯齿状
- 集合通信带宽掉 ~30%
- 抓包能看到 RoCE PSN 重传
- 不是网卡 down、不是光模块误码、不是 QP Error 大面积崩
第一直觉:拥塞控制链路出问题,不是“链路断了”。
环境:
Node A0~A7(Leaf1,Atlas 800T A2 ×8)
Node B0~B7(Leaf2,Atlas 800T A2 ×8)
\ /
Spine1
- RoCEv2 / UDP 4791
- DSCP 24 → TC3 → PFC pri3(数据)
- DSCP 25 → TC6 → PFC pri6(CNP)
- Leaf/Spine:PFC + ECN
- NPU:DCQCN 开启
1. 第一轮:先判断“是真丢包,还是被暂停”
1.1 NPU 侧看 DCQCN 有没有开
hccn_tool -i 0 -dcqcn -g status
dcqcn enable status: enable
dcqcn enable alg mode: 0
0=DCQCN。✅ 源端拥塞控制是开的。
1.2 看 RoCE/MAC/PFC 统计
hccn_tool -i 0 -stat -g
packet statistics:
mac_rx_pfc_pkt_num:182033
mac_tx_pfc_pkt_num:0
mac_rx_pfc_pri3_pkt_num:182033
roce_rx_cnp_pkt_num:312
roce_tx_cnp_pkt_num:0
roce_rx_all_pkt_num:981231
roce_new_pkt_rty_num:124
roce_out_of_order_num:55
roce_unexpected_ack_num:38
roce_rx_err_pkt_num:7
读图:
- mac_rx_pfc_pkt_num 18 万 → NPU 被对端/上行口 PFC 暂停过很多次
- roce_rx_cnp_pkt_num 312 → 交换机确实在标 ECN,CNP 回来了
- roce_new_pkt_rty_num 124 → RoCE 层已经发生重传
- mac_tx_pfc_pkt_num=0 → 本端没主动反压对端,问题在上游/交换机
结论:不是“源端不收敛”,是“ECN 标了但 PFC 还是被触发”,反压从网络侧回来。
2. 第二轮:NPU 被 PFC 暂停了多久
hccn_tool -i 0 -pfc_stat -g
空闲态:
tx_pfc0_duration_time=0.000ms
...
tx_pfc3_duration_time=0.000ms
...
rx_pfc3_duration_time=0.000ms
tx_pfc3_duration_warn_cnt: 0
tx_pfc3_duration_err_cnt: 0
本机异常态(排障示例):
tx_pfc3_duration_time=325.418ms
rx_pfc3_duration_time=19.206ms
tx_pfc3_duration_warn_cnt: 2
tx_pfc3_duration_err_cnt: 0
含义(Atlas A2 官方):
- tx_pfcN_duration_time:发送端因该优先级被 PFC 暂停的累计时间
- rx_pfcN_duration_time:接收端向对端发 PFC 的累计时间
- tx_pfcN_duration_warn_cnt:watchdog 检测周期内反压占比超告警门限次数
- tx_pfcN_duration_err_cnt:超错误门限次数
判断:
- 只有 pfc3(RoCE 数据)反压大
- TX 被暂停 325ms 在训练同步里非常致命,AllReduce 必然等
- warn_cnt=2 说明 PFC storm watchdog 已经观察到“反压时间占比异常”,但还没到 err/自动关 PFC
到这里已经能下结论:NPU 侧症状是“被网络反压”,根因在交换机/队列/映射,不在 NPU 本身。
3. 第三轮:交换机上看“谁在发 PFC、谁在丢包”
3.1 PFC 反压帧
display dcb pfc interface 100GE 1/0/32
Interface Queue Received(Frames) ReceivedRate(pps) DeadlockNum Transmitted(Frames) TransmittedRate(pps) RecoveryNum
100GE1/0/32 3 0 - 0 190622 - 0
100GE1/0/32 6 0 - 0 0 - 0
(数值为排障示例)
Leaf1 上行口向外发了 19 万个 PFC 帧 → Leaf1 上行队列拥塞,反过来暂停 Spine。
再看 Leaf2 下行口:
display dcb pfc interface 100GE 1/0/1
Interface Queue Received(Frames) ReceivedRate(pps) DeadlockNum Transmitted(Frames) TransmittedRate(pps) RecoveryNum
100GE1/0/1 3 160244 - 0 0 - 0
100GE1/0/1 6 45 - 0 28 - 0
Node B0 的 NPU 收到 16 万 PFC 帧 → 被 Leaf2 暂停。
反压链:
Spine 上行拥塞
→ Leaf 上行队列超 XOFF
→ Leaf 向 Spine 发 PFC(Transmitted(Frames) 大)
→ Spine 向下游 Leaf 扩散反压
→ Leaf 向 NPU 发 PFC(NPU Received(Frames) 大)
→ NPU 停发 → AllReduce 等 → Step 锯齿
3.2 ECN 到底标了多少
display qos ecn statistics interface 100GE 1/0/32
Interface ECN-marked packets
100GE1/0/32 96627
(排障示例)
display qos ecn statistics interface 100GE 1/0/32 verbose
Interface Queue ECN-marked packets
100GE1/0/32 3 96627
100GE1/0/32 6 0
⚠️ 这条命令只有 ECN-marked packets,没有 ECN mark pps、没有 WRED drops、没有 Tail drops。
3.3 ECN 门限是不是太晚
display qos ecn threshold interface 100GE 1/0/32
Interface Queue Min(KB) Max(KB) Probability(%) Queue Size(KB)
100GE1/0/32 3 287 293 40 305
(排障示例)
ECN High=293KB,Queue Size 已经 305KB → 标完 ECN 队列还在涨,说明:
- ECN 渐进区太窄(287→293 只有 6KB)
- Max 离 PFC XOFF 太近
- DCQCN 收到 CNP 再降速,RTT + 反应时间过去后队列已经撞 PFC
3.4 队列有没有丢包:普通模式
display qos queue statistics interface 100GE 1/0/32
Queue statistics info: interface 100GE1/0/32
Queue CIR/PIR Passed Pass Rate Dropped Drop Rate Drop Time
(% or kbps) (Packets/Bytes) (pps/bps) (Packets/Bytes) (pps/bps)
3 0 99233123/... 123k/... 262013/... 520/... 2026-09-13 11:02:14
6 0 195356/... 152/... 0/0 -/0 -
官方普通模式字段就是 Passed / Dropped / Drop Time 这种表。
Queue3 Dropped 非 0 → 出向队列确实丢过包。
3.5 丢在哪个芯片/哪个 VOQ:verbose 模式
display qos queue statistics interface 100GE 1/0/32 verbose
Queue Slot Chip VOQ Dropped(Packets)
0 1 0 4111 0
1 1 0 4110 0
2 1 0 4109 0
3 1 0 4108 262013
4 1 0 4107 0
5 1 0 4106 0
6 1 0 4105 0
7 1 0 4104 0
→ 丢包集中在Queue3 / Slot1 / Chip0 / VOQ 4108,正好是无损 RoCE 数据队列。
4. 第四轮:流有没有进错队列(AI 集群高频坑)
for i in $(seq 0 7); do hccn_tool -i $i -dscp_to_tc -g; done
device 0: dscp 24 -> tc 3
device 1: dscp 24 -> tc 3
device 2: dscp 24 -> tc 2
device 3: dscp 24 -> tc 3
device 4: dscp 24 -> tc 3
device 5: dscp 24 -> tc 2
device 6: dscp 24 -> tc 3
device 7: dscp 24 -> tc 3
交换机侧:
display diffserv domain ds1
ip-dscp-inbound 24 phb af3 green
ip-dscp-inbound 25 phb af4 green
AF3→TC3,AF4→TC6。
问题:
- device2 / device5 把 DSCP24 映射到 TC2
- 交换机把 RoCE 数据当 TC3 做无损
- 这两张卡的流量进了非无损 TC2 → 不受 PFC/ECN 保护,拥塞时直接丢
这能解释为什么:
- 整体 ECN 在跑
- 但还是有 Dropped(Packets)
- 而且重传分布不均匀
5. 根因(排障收口)
层 | 现象 | 证据 |
NPU | 被 PFC 暂停 325ms | tx_pfc3_duration_time=325.418ms |
NPU | ECN/CNP 有,但重传发生 | roce_rx_cnp_pkt_num=312、roce_new_pkt_rty_num=124 |
交换机 PFC | Leaf 上行狂发 PFC | Transmitted(Frames)=190622 |
交换机 ECN | 门限太晚 | Min 287 / Max 293 / Queue Size 305 |
交换机队列 | Queue3 丢包 | 普通模式 Dropped 非 0;verbose VOQ 4108 Dropped 262013 |
NPU 配置 | DSCP→TC 不一致 | device2/5 是 tc2,交换机期望 tc3 |
一句话根因:
ECN 门限离 PFC XOFF 太近,DCQCN 降速来不及,PFC 被迫兜底;同时部分 NPU 的 DSCP→TC 映射错误,使一部分 RoCE 流量脱离无损队列,最终表现为 AllReduce 锯齿、PSN 重传、Queue3 VOQ 丢包。
6. 修复动作
- 把 ECN Max 拉到明显低于 PFC XOFF(留几个 MTU + 足够 DCQCN 反应时间)
- ECN 低/高门限拉开(例如 30%~70% 缓冲,而不是 287/293KB 这种 6KB 窄区)
- 统一所有 NPU:dscp 24 -> tc 3、dscp 25 -> tc 6
- 检查 headroom / shared buffer,抑制 incast 微突发
- PFC storm watchdog 先观察,确认无抖动后再评估 mode 2 自动关 PFC
- 用 display qos queue statistics ... verbose 持续看 VOQ 丢包是否归零
7. 排障检查单
【NPU】
hccn_tool -i 0 -link -g
hccn_tool -i 0 -stat -g
hccn_tool -i 0 -dcqcn -g status
hccn_tool -i 0 -pfc_stat -g
hccn_tool -i 0 -pfc_storm -g
hccn_tool -i 0 -dscp_to_tc -g
hccn_tool -i 0 -qp_info -g
【Switch】
display dcb pfc interface 100GE 1/0/32
display dcb pfc interface 100GE 1/0/1
display qos ecn statistics interface 100GE 1/0/32
display qos ecn statistics interface 100GE 1/0/32 verbose
display qos ecn threshold interface 100GE 1/0/32
display qos queue statistics interface 100GE 1/0/32
display qos queue statistics interface 100GE 1/0/32 verbose
display diffserv domain ds1