metal-autotune

metal-autotune

Make an MLX model faster on Apple Silicon without changing its answers.

An AI writes custom Metal kernels for your model. A separate harness keeps one only if the outputs still match the original and the whole workload gets measurably faster.

Baseline · mx.compile · 1.00×Group norm, MetalBench, M4

Measured against MLX's own compiler

Whole-model speedups, confirmed by timing. Every kernel had to preserve correctness before it shipped.

    1.16×geomean vs compiled MLX
    1.46×geomean vs eager MLX

    Apple M4 · vs compiled MLX

    MetalBench

    A public, KernelBench-style suite for Apple GPUs. We ran its 35 standard workloads, each a fused pair of operations.

    • Group normalization 2.34×
    • Instance normalization 2.23×
    • Cross-entropy loss 1.54×
    • Scaled dot-product 1.47×
    • SwiGLU 1.39×
    1.16×geomean vs compiled MLX
    1.46×geomean vs eager MLX

    A kernel ships only if the whole model is faster and still rightOnly verified wins ship

    1. 01

      You bring a model

      model.py returns your MLX model from build(): mlx-lm, mflux, your own code. manifest.yaml says which inputs to speed up and how many attempts to allow.

    2. 02

      An AI writes kernels

      Claude Code, Codex or Gemini CLI proposes Metal kernels for one region of the model at a time, on your own subscription.

    3. 03

      The harness decides

      Every candidate must preserve correctness against the original model. Then the whole workload must get faster by a margin it can confirm. A kernel that is only fast on its own doesn't count.

    4. 04

      You apply the bundle

      Copy artifact/ into your app and call model = apply(model). Untested input sizes quietly run the original code.

    1. 1

      Bring a model

      Any MLX model, plus the inputs to speed up.

      model.py · manifest.yaml
    2. 2

      AI writes kernels

      Claude Code, Codex or Gemini, on your own plan.

      kernel void fused(…)
    3. 3

      Harness verifies

      Outputs must match and the whole model must get faster.

      ✓ match · ✓ faster
    4. 4

      Apply the bundle

      One line in your app.

      model = apply(model)

    Runs on your Mac

    metal-autotune runs locally on Apple Silicon, with the AI coding CLI you already use. Want it on your models? Get in touch.

    Apple Silicon MacClaude Code · Codex · Gemini CLI