metal-autotune

metal-autotune

Make an MLX model faster on Apple Silicon without changing its answers.

An AI writes custom Metal kernels for your model. A separate harness keeps one only if the outputs still match the original and the whole workload gets measurably faster.

Baseline · mx.compile · 1.00×Group norm, MetalBench, M4

Results

Measured against MLX's own compiler

    How it works

    A kernel ships only if the whole model is faster and still right

    1. 01

      You bring a model

      model.py returns your MLX model from build(): mlx-lm, mflux, your own code. manifest.yaml says which inputs to speed up and how many attempts to allow.

    2. 02

      An AI writes kernels

      Claude Code, Codex or Gemini CLI proposes Metal kernels for one region of the model at a time, on your own subscription.

    3. 03

      The harness decides

      Outputs must match the original, bit for bit or within a tolerance you set. Then the whole workload must get faster by a margin it can confirm. A kernel that is only fast on its own doesn't count.

    4. 04

      You apply the bundle

      Copy artifact/ into your app and call model = apply(model). Untested input sizes quietly run the original code.

    Sometimes nothing survives.That's a real answer, not a crash: no candidate beat the baseline by a margin the tool could confirm.

    Get started

    Runs on your Mac, with your AI subscription

    Free, with no account. Everything runs locally: your model, your GPU, and the AI you already sign in to.

    Terminal · the tiny example, no downloads
    # once: Apple's C++ compiler
    $ xcode-select --install
    $ git clone https://github.com/shivamg05/metal-autotune.git
    $ cd metal-autotune
    $ uv sync --locked
    $ uv run autotune run examples/tiny_mlp.yaml --judge claude-cli
    Or paste into Claude Code, from the repo
    Follow RUNNING.md to run metal-autotune on manifest.yaml with --judge claude-cli using a fresh work directory under runs/. Give concise updates at milestones: initial measurements, region changes, accepted improvements, errors, and final validation. Report the final result and artifact path, or explain why nothing shipped.

    You need

    • An Apple Silicon Mac
    • uv, which installs Python 3.12 and MLX 0.32.2 for you
    • Claude Code, Codex or Gemini CLI, signed in. It writes the kernels.
    • Patience: a real model can take hours. Keep other GPU work quiet so timings stay honest.

    To run your own model, write a build() file and a manifest: usage guide · write a manifest · model catalog