Retrospectively Reverse-Engineering Apple’s Neural Engine
I stopped working on the reverse-engineered Apple Neural Engine (ANE) driver three years ago, upon a sad mini realization that the ANE block is just not that useful, and I could be doing more useful things, and moved onto upstreaming other, more useful, blocks. The ANE’s architecture was too opinionated to build a general-purpose accelerator platform around it, and a linux driver effectively opening ANE hardware API access could not broaden the class of workloads it could do. Even macOS only regularly uses their own ANE to generate upsampled preview images in Finder.
https://github.com/eiln/ane/tree/main
The M5 (2025)‘s headline feature was “LLM performance”, and they also conveniently folded the ANE cores inside the GPU cores — I knew it was coming, but it officially feels like the beginning of the end for the standalone NPU. So, in honor of the ANE’s apparent demise, we will do something even more useless: go back and reverse-engineer the ANE on the M1, finish what we started. It’s been three years (fuck), and I should know more than I did when I first worked on this.
If the goal three years ago was to make the ANE useful by running ops on it; this time, it’s more about mapping the full internal architecture — compute, datapath, scheduler, memory, and execution model — because those internal design decisions reveal the assumptions about ML workloads that Apple was willing to commit to silicon first in the A11 Bionic (2017), and what that says about the shift from CNN-era NPUs to today’s GPUs running transformer workloads.
The 16 compute cores are probably the least interesting part of the ANE. Apple originally targeted dense image-processing CNN workloads, which consists of dense tensor reductions with predictable reuse. The M1 ANE compute core is a large parallel array of multiply-accumulate (MAC) units, but that alone says almost nothing about what workloads it was designed for and accels at.
A convolutional layer does a dot product between an activation window and learned kernel weights, and attention does a dot product between a query and key vector. A dot product is a dot product, and a MAC does just that. What specialized ANE to the 2017 CNN models is not the MAC, but dataflow surrounding the MACs: when and where MAC inputs and outputs enter, stay, move. The assumption that transformers broke, especially with autoregressive decode, was predictable reuse patterns, which the ANE exploited to architect a dataflow efficient enough to run on phones. The M5 decision confirms that ANE’s compute core remained still useful for transformers, but inside a different dataflow.
Still, here’s the datapath inside each of the 16 compute cores:
ANE has 16 parallel compute cores. Each compute core has 128 FP16 (or 256 INT8) parallel multiply-accumulate (MAC) lanes. Each MAC lane performs the recurrence:
Multiply two operands (a) and (b), and then add the product to the running sum (accumulator).
Repeating the MAC operation over T cycles computes a T-term dot product:
A MAC lane thus performs a scalar reduction over time. A 16-core ANE has 2048 parallel MAC lanes,
So each cycle performs 2048 parallel reductions spatially, with time being the only reduction axis:
An individual MAC lane does not know what dimension of the matrix or tensor it is reducing over. It’s important to note that a dot product vs matrix multiplication vs convolution arises from how the operands are mapped and scheduled onto the core. The ANE core (with the exception of kernel memory, discussed later) does not encode a 4-channel CNN layer into the hardware.
Internally, the MAC datapath consists of a multiplier, adder, and a 32-bit accumulator register. Each cycle, the adder adds the fresh multiplier output with the previous sum, which then becomes the new running sum.
This feedback path keeps the partial sum in memory local to the MAC lane, so it does not need fetched from an external memory far away, between MAC cycles.
Regarding resolution, it does fixed-point reduction with FP16 at readout. The multiplier is 16-bit, accumulated in a 32-bit register as Q16.16, then read out as FP16 via sign-extend and etc. Working in integer (hex) FP16 representation, to probe the accumulator range, build a CoreML ANE program that computes a dot product with a vector of all (1)s, so each multiplier results in a bounded v, but the running sum in the accumulator keeps growing:
Since 32768 is itself a valid FP16 word (0x7800), the ANE’s 0x7c00 can’t be FP16 output overflow, the clamp happens inside the accumulator, at (2^{15}). Thus the accumulator saturates at (2^{15}), exactly the range of a signed 32-bit fixed-point value with 16 fractional bits.
For a fused layer, the ANE computes:
Importantly, completed MAC sums feed directly into the post-MAC activation block, avoiding an intermediate memory round-trip. This is possible because the activation is pointwise: once a scalar reduction is complete, its activation depends only on that scalar and can be applied immediately.
To determine how the ANE implements tanh(), compile a CoreML model containing a single TANH activation layer and inspect the resulting compiled hardware register file (hwx). The coefficient region contains 33 consecutive FP16 words beginning at 0x4288:
Those 33 FP16 words match 33 IEEE LE FP16 quantized samples of (tanh(x)):
Now switch to RELU activation layer:
Thus, mode 2 selects a custom 33-entry lookup table. 33 points defines 32 intervals. With (R=3), the knots are
covering ([0,4]) with spacing (1/8). The input maps into the table as (u=2^R|x|), so (R) sets the knot spacing. The resolution is smoother than its 33 bin; I suspect that adjacent entries are linearly interpolated. To test, build an impulse LUT with a single spike:
Then sweep the input across the two cells around (T_8). The measured output forms a triangle: magnitude rises linearly from (0) at (|x|=7/8) to (1) at (|x|=1), then falls linearly to (0) at (|x|=9/8).
Thus, we know that mode 2 implements a 33-entry piecewise-linear LUT. (R) scales the input into LUT coordinates,
so the knot spacing is (Delta x=2^{-R}). (lfloor urfloor) and (lceil urceil) select the adjacent entries, and (alpha=u-lfloor urfloor) gives the interpolation weight between them.
CoreML also supports a linear scaling and bias (ax + b) transform. I then suspected (ax + b) could share the linear interpolation hardware of mode 2. To confirm, construct a CoreML model with a ReLU with a constant scale and offset:
If the compiler folds the constant scale and offset into the convolution:
Decoding model.espresso.weights confirms exactly this folded transformation on ReLU:
And the register file hexdiff shows how bias and activation are fused into the same post-MAC path at compile time:
Extremely cursed idea: use nonlinear interpolation to compute an additional kernel pass, or quantize int8 into int4 weights.
The ane driver source code is disappointingly boring. The driver never gives the ANE a CONV, MATMUL, or RELU opcode to run. All the neural operations have all already been compiled into a command stream of task descriptors (TDs), and the driver software simply loads the task to memory, sets the pointer to the opaque task blob via (TM_ADDR, TM_SIZE), and submits the staged task by ringing the doorbell (TM_PUSH).
https://github.com/eiln/ane/blob/main/ane/src/ane_tm.c#L87
The hardware then owns the submission until completion, and raises an interrupt to the ARM64 core when it’s done.
This (boring) command submission frontend resembles that of a GPU’s, think NVIDIA’s pushbuffer/PBDMA. The software submits a command stream resident in memory, and the GPU’s command processor walks over command stream and dispatches the commands, without knowing what that command executes.
TM_ADDR and TM_INFO are global staging registers, and TM_PUSH atomically commits that staged launch state, given that nothing happens until TM_PUSH is written (“magic”). TM_INFO in particular stores the total number of descriptors in the supplied stream:
TM_INFO register naturally maps onto a hardware counter:
Why the “minus 1”? Encoding length - 1 is an RTL-friendly way to terminate a zero-based counter out of the critical path. But note how, compared to GPU commands which parse a variable-length stream of descriptors in a ringbuffer, ANE only receives the total count, indicating that descriptors are fixed-size.
Going one layer deeper, what’s in a task queue (TQ) that the task manager selects from?
There’s 8 copies of the same register block (indexed by qid (0…7)), structured as:
Notice how TM_PUSH executes a task referenced in TM_ADDR/TM_SIZE by attaching a qid:
The natural interpretation is that the descriptor stream specifies what task to run, while the qid selects the launch context the descriptor runs under. The resident TQ context (BAR, NID) is much like a GPU hardware channel. Here’s my driver populating a single TQ to launch it:
https://github.com/eiln/ane/blob/main/ane/src/ane_tm.c#L70
The only thing important here is the 32-entry BAR table (base address register). We’ll get into task descriptors next, but the compiled ANE command stream only references virtual addresses by relative offsets, and BAR provides the base IOVA (IOMMU peripheral virtual address) relocation address. An ANE virtual address access needs a hard-coded BAR base offset supplied at compile time, meaning it lacks GPU-style load/store instructions that dynamically issue load/stores from virtual address.
The task manager walks over and executes chain of fixed-size task descriptors:
What’s in each TD? Here’s a hexdump of the TD for the simplest 1x1 convolution:
Important is that a TD is not an executable instruction stream. ANE has no ISA. TD is a sequence of “ControlDMA” (I made this name up) burst-write packets writes to the ANE’s hardware configuration registers, such as input dimension, input/output address, activation function. Each ControlDMA packet consists of a 32-bit transfer word followed by N consecutive 32-bit register values:
Notice the “minus 1” termination count again. ControlDMA is a flexible unidirectional DMA engine that copies N 32-bit words from IOMMU virtual DRAM into the ANE’s physical register space. For example, KernelDMASrc’s packet header in TD is 0xf401f800:
This is not a LOAD_WEIGHTS instruction. It’s copying 0xf4 or 62 consecutive words into the KernelDMA register offset starting at 0x1f800. And those KernelDMA configuration values can tell KernelDMA where to load the weights from.
Since each section writes to one MMIO register block, TD divides cleanly into ANE’s datapath sections:
A TD is effectively a serialized register-file dump of the ANE’s datapath registers. Each “ANE program” is simply the configuration for one pass through the datapath. We can configure how the fixed datapath operates (subject to the knobs it exposes), but not what operations the datapath is capable of performing, or how those operations are sequenced.
When the “magic” atomic word is written to task manager to execute a TD, roughly, the sequence of what happens:
ANE is a fixed-function dataflow engine, not a GPU executing arbitrary instructions. The TD configures a domain-specific datapath. Constraining the hardware interface usually means smaller area, deterministic movement, lower latency, and less power drawn. ANE’s compiler can explicitly schedule what the tensors do, but that also means the compiler must explicitly schedule what the tensors do. This is a tradeoff, but a justified one: we usually know what the model looks like at compile time. Dynamic execution is not what limits ANE. ANE’s processor interface is relatively generic, and it simply launches tasks, and the tasks can describe transformers.
For example making tensor sizes fixed at compile time does not mean it can’t handle variable-length tensors: for example, a growing KV cache can be traversed by looping over the size, and dispatch overhead is negligible relative to the elephant in the room here, that is, memory-streaming bandwidth. What actually shaped ANE for CNNs over transformers is memory movement.
It’s always good to identify our current slowest link, so we can optimize what actually matters.
Apple’s unified memory lets the ANE access buffers from the system DRAM pool accessible by the CPU and GPU. It does not mean the ANE zero-copy streams directly out of that DRAM pool. ANE must first copy any memory into its local “ANE memory” or SRAM. Any bandwidth-limited task will thus be limited by ANE’s local memory streaming throughput.
M1 ANE reports (11text{ TOP/s}) at (68text{ GB/s}) at system DRAM bandwidth. A MAC performs two operations but consumes two FP16 operands, or 4 bytes:
If every MAC operand streamed from DRAM, sustaining (11text{ TOP/s}) would require streaming
which is over 300x times the reported (68text{ GB/s}) system DRAM capacity. Thus peak ANE MAC throughpu