纯 .NET 手写 CUDA kernel,GLM-5.3-Flash decode 跑出 llama.cpp 的 2 倍

8 月 26 日,智谱 GLM-5.3-Flash(MIT 许可,首日开源)和阿里 Qwen3.8-Flash-Next 同日发布。三天后,纯 .NET 推理引擎 TensorSharp 把两者都接进了主干——两个全新的 GGUF 架构 id(
glm5next
qwen4exp),README 与架构卡片同步入库。

赞(0)
未经允许不得转载:小狮博客 » 纯 .NET 手写 CUDA kernel,GLM-5.3-Flash decode 跑出 llama.cpp 的 2 倍
分享到: 更多 (0)

联系我们