Make an MLX model faster on Apple Silicon without changing its answers.
An AI writes custom Metal kernels for your model. A separate harness keeps one only if the outputs still match the original and the whole workload gets measurably faster.
Whole-model speedups, confirmed by timing. Every kernel had to preserve correctness before it shipped.
Apple M4 · vs compiled MLX
A public, KernelBench-style suite for Apple GPUs. We ran its 35 standard workloads, each a fused pair of operations.
model.py returns your MLX model from build(): mlx-lm, mflux, your own code. manifest.yaml says which inputs to speed up and how many attempts to allow.
Claude Code, Codex or Gemini CLI proposes Metal kernels for one region of the model at a time, on your own subscription.
Every candidate must preserve correctness against the original model. Then the whole workload must get faster by a margin it can confirm. A kernel that is only fast on its own doesn't count.
Copy artifact/ into your app and call model = apply(model). Untested input sizes quietly run the original code.
Any MLX model, plus the inputs to speed up.
model.py · manifest.yamlClaude Code, Codex or Gemini, on your own plan.
kernel void fused(…)Outputs must match and the whole model must get faster.
✓ match · ✓ fasterOne line in your app.
model = apply(model)metal-autotune runs locally on Apple Silicon, with the AI coding CLI you already use. Want it on your models? Get in touch.