Retrieved article excerpt
Open article · Retrieved 2026-09-15T21:23:18.085602+00:00
# I Came, I Prompted, I Left Part 2: Building a GPU Driver From Scratch in One Month
Previous blog post: <https://codyho.dev/blog/hypervisor-macbook-neo/>
## What We Did
**TL;DR:** Niklas and I built a fully OpenGL ES 3.0 compliant GPU driver for the
M4 Mac Mini and MacBook Neo in about a month, a process which normally takes
years. Here is Chrome and Firefox running WebGL on the M4 Mac Mini with working
compositing:
Chrome and Firefox running three.js WebGL demos on our driver.
Most importantly, the driver is fast enough to run Minecraft at 200fps:
Minecraft running at 212fps on the M4 Mac Mini.
Building this driver involved reverse engineering the AGX’s (Apple’s name for
the GPU) incredibly complicated firmware ABI and user-space components. This
was all done in a transparent, verifiably clean room manner using well
established techniques. The code is not yet ready for end users, but we are
looking to get it to end users as soon as possible.
## How We Did It
Previously, I built a hypervisor to reverse engineer macOS. Now the goal became
to actually do something useful with it, and what better target than writing a
GPU driver. The GPU is effectively a requirement for any modern system,
otherwise everything needs to be CPU rendered which is orders of magnitude
slower and less power efficient. Our goal was to implement conformant OpenGL
(and soon, Vulkan) drivers for the M4 Mac Mini and MacBook Neo.
Normally, building a GPU driver is an endeavor that takes years; our goal was
to do it in days. It turns out that days was overly optimistic, but weeks is
still a massive improvement. In those weeks we have:
- Reverse engineered the M4, A18 Pro, and (mostly) M5 user space using only
live probing, discovering hardware-supported features and instructions not
emitted by Apple’s driver
- Built a fully working user-space driver, including a new custom IR/shader
compiler, command stream builder, and many more components
- Reverse engineered, from scratch, the full AGX firmware ABI using traces from
the hypervisor I previously built
- Implemented a full Linux kernel driver for said firmware ABI
Throughout this process, we have not looked at any Apple binaries, only
hardware traces (from our hypervisor) and shaders we built ourselves. For
user-space graphics RE, we were careful to treat any required Apple blobs as
opaque objects. We had a friend write documentation on these blobs [1](https://codyho.dev/blog/gpu-driver/#fn:1) so we
could write a clean room implementation ourselves (which was mostly built by
just blindly trying stuff until it worked). We have published all of our
experiments so that anyone can verify the provenance of our work (see the twin
`agx-re` repos under [Deliverables](https://codyho.dev/blog/gpu-driver/#deliverables)).
This blog post is divided into two parts, user and kernel space. This mirrors
the split in all modern GPU drivers: the kernel is responsible for interfacing
with the firmware, allocating buffers, and managing scheduling, while the
actual contents of those buffers and what is being scheduled are opaque.
User space is responsible for actually understanding how the GPU works and
filling those buffers with stuff.
## Kernel Space
On Apple Silicon, the kernel driver does not interface directly with the
hardware. Instead, it talks to the GPU firmware running a custom RTOS called
RTKit. That means that the first step to a kernel driver is not talking to
hardware, it’s figuring out the firmware ABI.
The firmware ABI was by far the most annoying part of this project, because
rather than doing the sane thing of coming up with a reasonable ABI with nice
interfaces, Apple essentially took a regular kernel driver, cut it in half, and
then put half of it in the AGX and called it firmware, with the other half of
the kernel driver communicating using shared structs in memory. Many of these
structs have firmware owned fields (which we must never modify and which we
must learn from reverse engineering) interleaved with host controlled fields.
For an idea of how complicated the ABI is, this is what the shared memory
tree looks like on the M1/M2:
The M1/M2 firmware ABI
Asahi Lina famously figured all of this out over grueling 12-hour days to build
the M1/M2 kernel driver, an amazing technical accomplishment. Unfortunately,
the A18 Pro firmware ABI (I started my RE work on the MacBook Neo and later
pivoted to the M4 Mac Mini) is *significantly* more complicated than the
already very complicated M1 firmware ABI:
The A18 Pro firmware ABI
**What the F@!#, Apple.** Note how the A18 has:
- 1.5x as many structs
- twice as many pointers
- a significantly more complicated process for submitting work
There are many other issues that add friction to the RE process [2](https://codyho.dev/blog/gpu-driver/#fn:2). I did
have some documentation on the firmware ABI, but it was highly incomplete and
honestly was not very useful [3](https://codyho.dev/blog/gpu-driver/#fn:3).
My approach was simple and based on the approach used to successfully reverse
engineer the M1/M2 machines: watch what macOS did, replay it, then try to do it
ourselves, which is made possible by the hypervisor.
When I described this approach to the LLM, it took replay extremely literally:
the first thing it did was wait for the first firmware visible event (these are
called “kicks”), then *saved a copy of the entire GPU memory state*. After a
reboot, it copied the saved memory state straight back into host memory,
performed the kick, and saw the output pages change. It would then try to
reconstruct these objects in code, following all the pointers and making sense
of the contents. Over successive experiments, Codex would reduce the number of
pages it copied until there was no more replayed state and everything was built
from source. [4](https://codyho.dev/blog/gpu-driver/#fn:4) Amazingly, I noticed Codex had good taste regarding when it
should poke the hardware some more and when it should just run the hypervisor
and capture the state itself.
There were three major issues, and all were caused by our inability to
get a clean capture of host work:
The first issue was render work submitted *after* the GPU firmware started. We
could prestage work before the firmware started, start the GPU, and that work
would be completed as expected, but once the firmware started any work
submitted would just be ACKed and retired without actually doing anything. Once
the firmware has started, capturing state is much harder because everything
becomes dynamic and the firmware becomes a stateful object with state you can’t
easily replay.
I had to step in at this point and examine Codex’s process. It turns out it was
trying to replay a capture very late in the AGX’s lifecycle, where there had
already been many previous events. When I told it to choose a capture far
earlier in the AGX’s lifecycle, the very first capture after firmware start,
Codex was able to almost immediately discover the issue (it was missing
a single byte descriptor). This took a few days.
The second, and only major blocking, issue was compute. The AGX, broadly,
supports two kinds of work: compute and render. In the regular GUI path,
compute work is only scheduled after a significant amount of render work was
already executed. Thus, it took a long time to get a clean capture of a compute
workload, and when Codex finally did it was 336 MB and impossible to replay (it
tried, for a long time). It also tried to construct the objects itself by
looking at the capture, and spent over a week doing this, but was ultimately
unsuccessful. There was just too much nonsense to sift through. This was
exacerbated by issues on my side– after getting render working, I expected
that submitting compute work would be simpler (the firmware ABI for compute is
indeed simpler, so I was correct here), but lost my humility and thought it
would be a cakewalk that would only take a few hours. Thus, I didn’t scaffold
out the task properly for the LLM.
The fix actually was given to me in another Codex session. In essence:
1. Disable the GUI by booting into single-user mode; this means no render work
would be done.
2. Install a LaunchDaemon to run at the earliest possible point, the moment
Metal (Apple’s proprietary graphics framework) became available.
3. Run a tiny Metal program that we supplied
4. Capture and replay this tiny, pure compute trace.
The trace was captured successfully. Within a few hours, Codex had
deconstructed it, and within a few days, Codex had compute working. As for why
the original compute codebase didn’t work… Codex has no idea. The working one
and the broken one look very similar.
In hindsight, this should have been the strategy from the start– smallest
possible capture, run in single-user mode so as not to perturb results. I
learned from my mistakes here for the final issue:
Partial renders ended up being one of the hardest things to figure out. They
occur when the Tiled Vertex Buffer (TVB) isn’t large enough to store the
current geometry (ie, there’s just too many triangles to draw). In these cases,
there are two options, and the driver needs to support both: either increase
the size of the TVB, or perform a partial render, ie, render part of
the geometry, then reload the buffer with the rest of the triangles, and finish
the partial render. These partial renders turned out to be very, very finicky,
even more so than the rest of the work because they essentially mean adding
save and resume to the GPU driver.
The workflow I discovered earlier came in very handy here. Codex was able to
replay one partial render transaction, and then modified our Metal shader to
perform multiple partial renders (this is pretty easy by just hammering a
single tile with thousands of triangles until a partial render is triggered)
and then learned how to replay these. Once Codex had a successful replay, it
was only a matter of time until it learned how to build it ourselves.
### Building the Kernel Driver
Moving from a Python prototype driver to a fully featured Linux driver took
three days, and one of those days was almost totally wasted because Codex, for
some reason I still do not understand, chose to tackle partial renders first
(by far the hardest task) instead of doing compute first (the easiest task).
Once I told it to do compute first, everything went smoothly.
At all high level, the entire process was, basically:
1. Rewrite the existing `drm-shim` in Rust following the exact same pattern;
this gives us a synchronous Rust driver.
2. Rewrite the frontend to be asynchronous; the actual GPU submission remains
synchronous.
3. Refactor the GPU submissions to be asynchronous and, instead of polling,
listen for firmware events and associate work with a fence
4. Implement some low-hanging optimizations, such as batched work submission.
This is all pretty routine engineering work that LLMs are definitely capable
of.
The only notable thing I found is that Codex aggressively used the hypervisor
to debug why its code didn’t work, including capturing the full address space
and comparing it to known good samples. This sort of systematic debugging is
why Codex is by far my favorite coding agent.
## User Space
The A18 Pro user space is very different from the M1/M2; it has new descriptor
formats, a new ISA, and a bunch of other new things. Aside from being a tile
based deferred renderer designed to run Metal, it’s just a different GPU.
The good news is user-space RE has a very well defined process. Simply write a
small Metal program, compile it, run it, see what changed, then take it apart
and start fiddling with the bits until we understand what all of them do. If
you’re thinking this sounds like the sort of boring, repetitive, rote work that
LLMs are very good at, you would be correct.
The RE work occurred in two phases. For the first phase, I had Claude look at
every possible Metal program it could find and trying to build a disassemble