In terms of “narrative-pushing with suspect data/sources” this...


What's behind the delay?
DeepSeek's R1 model thrust it into the national spotlight. The company was tasked with accelerating Huawei's AI ecosystem by training its successor model on the Ascend chip. Despite having a dedicated Huawei team on site, it didn't work.
DeepSeek abandoned efforts to train on Ascend, a decision that has cost them precious time in this fast-moving race. It has instead doubled down on Nvidia for training, underscoring the challenges facing China’s drive to be technologically self-sufficient. It is still working with Huawei to make its upcoming model run on Ascend for inference.
It's unclear when R2 will come. Liang Wenfeng has said internally that he is dissatisfied with R2’s progress and has been pushing to spend more time building an advanced model that can sustain the company’s lead in the AI field.
More here ft.com/content/eb9846… with @zijing_wu
The claim that reliance on Huawei’s Ascend AI chips had caused significant delays in DeepSeek’s release date struck me as odd.
DeepSeek is no longer resource-constrained, so why would it place so many of its eggs in Huawei’s basket? They are not a foolish company.
Most likely they setup a parallel team to start working experimentally on an Ascend-optimized model and continued to work with nVidia chips for the main training runs.

I'd view it more as parallel paths of two branches of a tech tree that run semi-independently. And as you know China's M.O. pm tech development is typically "why not both?".
Specifically on DS: Any alleged delays (and as @zephyr_z9 points out, there are some real question marks on whether that May date was a real vs. speculated target) are unlikely to be related to trying to ramp up the learning curve on Ascend chips unless they were foolish to rely solely on Ascend chips for training. And they don't strike me as foolish company.
Yes, there is a potential opportunity cost / resource constraint issue, but DeepSeek's emergence onto the scene late last year ensured that internal resource constraints (i.e. human capital and chips/compute) should no longer be the bottlenecks they had been for its existence up to that point.
reuters.com/world/china/hu…

Deepseek's training runs are on fp8
Ascend doesn't even support that format
As @zephyr_z9 points out, and as was reported earlier in the year in the wake of the “DeepSeek moment”, DeepSeek is likely consulting closely with Huawei on its upcoming next chip design, the Ascend 920.
I’m going to go out on a limb and predict that it will support FP8, for example.

920 mass production will begin in Q1 2026
Whale certainly tested the hardware and gave feedback

@zephyr_z9

🧠 Hybrid inference: Think & Non-Think — one model, two modes
⚡️ Faster thinking: DeepSeek-V3.1-Think reaches answers in less time vs. DeepSeek-R1-0528
🛠️ Stronger agent skills: Post-training boosts tool use and multi-step agent tasks
Try it now — toggle Think/Non-Think via the "DeepThink" button: chat.deepseek.com
1/5

在人工智能训练和推理加速的竞赛中,浮点数(Floating Point)的表示方式正成为关键突破口。作为计算机中用于表示小数的核心手段,浮点数由三部分构成:符号位(Sign)、指数(Exponent)和尾数(Mantissa)。符号位决定正负,指数决定小数点的位置,尾数则影响精度。内存位数越多,浮点数的表示精度越高,但同时带来的计算和存储开销也更大。
如下图为浮点数据类型的结构。所有显示的值(在 FP16,BF16,FP8 E4M3和FP8 E5M2中)都是最接近数值0.3952的表示形式。
在主流硬件生态中,NVIDIA GPU目前支持E4M3与E5M2两种FP8格式,并通过绑定硬件和软件的深度优化提升了其适用性。例如,NVIDIA引入了per-tensor scaling,per-block scaling等动态缩放策略,以解决FP8动态范围不足,容易溢出的难题。同时,在Tensor Core中也专门增加了FP8指令集,使得FP8在H100 GPU上能够充分释放算力。
在新一代Blackwell架构中,NVIDIA更进一步提出了“微缩浮点格式”(Microscaling formats),涵盖MXFP8(8位)、MXFP6(6位)、MXFP4(4位)等多种新型表示方式。研究显示,在高质量数据集上,一个8亿参数的模型若采用MXFP8-E4M3格式,并配合优化的数值转换策略,训练结果几乎可与传统的BF16持平。这意味着在Blackwell平台上,MXFP8正在成为兼顾性能与精度的最佳选择。
与之相比,中国团队DeepSeek在V3.1模型中提出的UE8M0 FP8格式,走了一条完全不同的道路。UE8M0采取极简设计:8位全部用于指数(Exponent),尾数(Mantissa)为零。换言之,它牺牲了精度,以换取更大的动态范围。在这种格式下,最接近刚才图片内提到的数值0.3952的表示形式为0.5。可以很明显地看出来,精度差异较大,但是这种“极端化”的方案不仅减少了硬件实现复杂度,也为未来中国技术栈在模型训练、部署和推理中的数值优化提供了新的可能性。
1,FP8/UE8M0的优势与权衡
🔹 显存与带宽显著节省:相较于FP16和BF16,8-bit表示可将内存占用与传输成本大幅降低,有利于支持更大规模模型、更高并行度或更多批处理。
🔹 吞吐与能效提升:更窄的数据通路意味着在相同内核与内存带宽下,系统可处理更多算子,整体吞吐率和能效显著提升。
🔹 成本与部署门槛下降:低精度带来更高的性价比,对于数据中心及国产算力环境尤为重要,使大模型在受限带宽或成本条件下的部署成为可能。
🔹 软硬件协同优化:当模型与硬件围绕低精度格式协同设计时(如DeepSeek专门针对“国产芯片优化”),能够释放软硬件一体化的性能潜力。
但需要注意的是:更低位宽必然带来精度与鲁棒性下降,尤其是UE8M0这类极端“无尾数”设计,更依赖于训练、量化、校准等算法补偿,以及硬件支持机制。FP8在训练与推理端的应用边界,仍是学术界和工程界研究的活跃话题。
2,UE8M0的战略思维:软件先行推动硬件适配
UE8M0的“发起”方式具有鲜明的战略意义。不同于传统由硬件厂商先定义数据格式,DeepSeek选择在模型端率先采用并公开声明使用UE8M0格式,将其训练与scale策略与该精度绑定。
这等于由大模型端主动提出标准,迫使硬件和工具链进行适配。媒体普遍认为,这一举措是“模型先行推动硬件协同”的里程碑事件,加速了国产软硬件一体化的生态建设。
3,战略协同:AI与半导体的一体化生态
诚如笔者浅见:DeepSeek的高明之处在于其战略协同。公开资料显示,已有超过15家国内企业正在调整硬件以适配DeepSeek模型,覆盖电信、汽车、移动科技等多个领域,其中包括华为、中国移动等行业巨头。
这种合作并非单向:
🔹 对半导体厂商而言,DeepSeek模型成为性能与效率的标杆,推动其改进设计。
🔹 对DeepSeek而言,合作确保了其AI工具的落地基础,开发者与企业正在加速采用。
结果是形成一个自我强化的正反馈生态:软件与硬件同步演进,速度甚至可能超过美国碎片化的“AI公司依赖外部芯片”模式。
至此,看我推文比较久的小伙伴们或许还记得,我曾写过一篇解读DeepSeek论文的文章:《洞见 —— 硬件与模型协同设计,突破Scaling挑战》(x.com/Compute_King/s…)。如今,看到国内AI企业在这条道路上迈出关键一步,实在令人欣喜。
4,国产芯片代表:寒武纪与华为的FP8路径
🔹 寒武纪(Cambricon)690系列
据多家媒体报道,寒武纪MLU370、思元590及最新的思元690均已支持FP8或“Block FP8”。其NeuWare软件栈在低精度支持上提供了完整的工具链,包括量化、混合精度调度以及对主流框架的适配。
在硬件层面,寒武纪的MLU架构通过算子引擎、片上缓存和张量内核优化,实现了高吞吐的低精度计算。媒体称思元690在低精度算力与能效上提升明显,已能够兼容DeepSeek模型。
需要强调的是,公司公开资料并未披露是否支持UE8M0这类极端格式,实际效果依赖SDK与模型方的适配验证。
🔹 华为(Ascend/昇腾)
华为提出了HiFloat8(HiF8)方案(arxiv.org/pdf/2409.16626),不同于E4M3/E5M2,而是一种“渐进式(tapered precision)”设计,根据数值区间动态分配指数与尾数,以在范围与精度之间取得平衡。
华为的Ascend系列已在OptiQuant、Atlas等平台上支持量化和混合精度,并将HiF8作为未来关键方向。与寒武纪偏重推理优化不同,华为强调同时支持训练的前向与反向传播,力图构建更通用的FP8训练方案。
5,大局观:AI已是国家战略
中国的AI发展早已超越实验室阶段,成为国家战略的重要组成部分。通过将AI软件与国产半导体深度结合,北京正在减少对外部技术的依赖,并为未来创新绘制蓝图。
DeepSeek的UE8M0 FP8优化,不仅是数值表示的一次尝试,更是中国在AI软硬件协同上的战略突破。
对投资者而言,启示清晰:
🔹 AI的未来不仅仅是算法,而是完整的生态系统。
🔹 DeepSeek与国产半导体生态的绑定,正在塑造这一趋势。
最终,问题不是中国能否实现AI自主,而是多块能够实现。而凭借UE8M0 FP8优化与深度产业整合,DeepSeek无疑是目前最值得关注的AI企业之一。




