From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple Silicon
From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple Silicon
图1:CUDA到MLX优化转换图。CUDA优化知识可以转化为原生架构的MLX策略,而不是逐条指令复制。我们正进入计算的新纪元。硬件正在迅速变化——不仅仅是更快的GPU,还有来自不同供应商的越来越多的芯片,每种芯片都有自己的架构,通常针对特定的AI工作负载进行优化。软件也在以同样快的速度变化,如今AI编码工具在几分钟内生成的内容,几年前需要数月的努力才能完成。随着计算的重心如今集中在AI上,GPU内核是其成功的关键组成部分。这些是在GPU内部运行的底层程序,编写高效的内核远非显而易见——需要数年的专业知识才能掌握。将内核从一个供应商的硬件转移到另一个供应商的硬件更加困难,通常意味着要从头开始重新发现相同的优化方法。例如,CUDA生态系统已经积累了数十年来通过艰苦努力获得的内核专业知识:针对注意力机制、状态空间模型和其他关键操作的手动调优实现,这些代表了数千小时的工程努力。较新的硬件生态系统(Apple Silicon,定制AI加速器)
Figure 1: CUDA-to-MLX optimization translation map. CUDA optimization knowledge can be translated into architecture-native MLX strategies rather than copied instruction-for-instruction. We face a new epoch in computing. Hardware is changing rapidly — not just faster GPUs, but a growing range of chips from different vendors, each with its own architecture and often tailored to specific AI workloads. Software is changing just as fast, and AI coding tools now generate in minutes what took months of effort a few years ago. With so much of computing now centered on AI, GPU kernels are a crucial component of its success. These are the low-level programs that run inside the GPU, and writing efficient ones is far from obvious — it takes years of expertise to get right. Transferring a kernel from one vendor’s hardware to another is harder still, and often means rediscovering the same optimizations from scratch. The CUDA ecosystem, for example, has accumulated decades of hard-won kernel expertise: hand-tuned implementations of attention, state space models, and other critical operations representing thousands of engineering hours. Newer hardware ecosystems (Apple Silicon, custom AI accelerat
应来源方要求,这里只提供摘要与原文入口。完整内容请阅读原文。
来源:Berkeley AI Research · bair.berkeley.edu