A graphical desktop for the ZX Spectrum
A graphical desktop for the ZX Spectrum 48K, written in Z80 assembly.
Overlapping windows with a z order and focus, pull down menus, a heap, an event queue, a storage layer with swappable backends, dialogues, controls, a notepad, a clock, a calendar, a two pane file manager, and a settings panel that actually changes things. It all fits in 48K on a machine from 1982, and it can drag a window inside a single 69,888 T state frame.
It runs on the real thing, not just an emulator.
Back in the eighties I wanted an Atari ST and couldn’t afford one. What I really wanted was GEM: the desktop, the windows, the menu bar that was always there, the feeling that the machine was a place rather than a prompt. I had a Spectrum instead, and I spent a long time wondering how much of that you could do on it. I started writing bits of it, and never finished.
So this is that, finished. It isn’t a port of GEM and doesn’t pretend to be. It’s what the idea turns into when you push it up against a 3.5 MHz Z80, 48K of RAM, a one bit display with attribute clash, and a video chip that steals cycles from the CPU while it paints. A lot of the answers turned out to be more interesting than the question, and nearly all of them came from measuring the machine rather than reasoning about it.
It’s a fun project and a labour of love, and the reason it’s written up at this length is that the measurements are the useful bit. If you’re building something on this hardware, the numbers below cost me a lot of evenings. They’re yours.
The toolchain is local and small: pasmo 0.5.5, built from source into tools/.
esxDOS runs here too, on an emulated DivMMC with a 64MB card image. tools/ isn’t in the repository because pasmo, the emulators and the esxDOS ROM aren’t mine to redistribute, so you’ll need to build the image yourself from an esxDOS release and a DivMMC card image. Once it exists, launch tools/esxdos/esxdos.szx and esxDOS is already resident; Machine > NMI gets you its file browser.
Build flags, all passed through —equ:
build.sh must pass —name explicitly, because pasmo takes the tape header name from the output path exactly as written and would otherwise put build/zxde in the header. It also refuses to build a tape if the code has grown into the buffer region, because that overrun is silent otherwise: the first window grab writes over the program and a few seconds later the machine drops into BASIC with an unrelated error.
Dragging a window under Fuse needs the space bar, not the mouse button. Fuse for macOS, 1.9.2, stops delivering Kempston mouse movement while a button is held, so the pointer freezes at the moment a drag begins and the window never follows. Point at the title bar, hold SPACE, move, release. The mouse is fine for everything else and the buttons themselves register correctly; it is only movement that stops.
This is the emulator, not the desktop, and MOUSETEST=1 is how I proved it: it reads the mouse ports once each with nothing between the port and the screen, and the counters still stand still while a button is down. RiBtn in ReadInput is what makes SPACE work, and it’s there so the desktop is usable on a machine with no mouse at all. A real Kempston mouse drags normally.
SCRIPT=1 is the one worth knowing about. It drives the desktop from a list of synthetic input events instead of the mouse, so an interaction (open a menu, pick an item, drag the window over another one, type into the field, save) runs the same way every time and can be compared against the last run rather than watched.
Device layer. DevFillRect, DevFillDesk, DeskFillCol, AddrAt, BlitRect, RectGrab. Everything above works in byte columns and pixel rows and never touches the screen’s third and interleave layout directly. This is the boundary a port swaps out, and I drew it on day one for exactly that reason.
Frame discipline. The main loop halts on the interrupt, does all pointer work in the top border, waits for the beam if the window moved, redraws, then reads input and dispatches at the end of the frame. Input goes last so that the cost before the beam wait is constant, which is what makes the scheduler exact.
Event queue. A sixteen slot ring of four byte events. EvPoll turns raw input into pointer moves, button presses and keys; EvDispatch drains it through a handler table.
Hit testing. A five byte row per control, front to back in z order, FF terminated, refilled from the model before every search so there is no second copy of the window position to go stale. A closed menu gets a height of zero, which can never match, so HitTest has no special case for it.
Transient surfaces. A four deep arena, each surface up to 16 by 96, pushed and popped. Two stacks rather than one, because the pixels under a menu are pushed by something that has no panel record at all.
Windows. A record swapped into a live copy, the same trick as the panel record, because the window position is referenced ninety three times across six files. The z order is also the paint order reversed and the hit test order.
Storage. A registry of backends, each a fourteen byte row of id, capability bits and six entry points. The six operations are hand laid JP instructions whose operands are patched on selection, so dispatch costs ten T states and clobbers no registers.
Controls. A panel is a stack of rows and every control is a row, so a row’s index is three shifts rather than a search. Four types, of which the useful one is a cycle: a checkbox is a cycle whose limit is two, and a radio group is a cycle whose limit is N. The settings panel, the file list and the save box are all the same code with different tables.
Two things in that map are worth explaining.
The slow region exists because the ULA steals cycles below 8000 while the display is being painted, so code there runs perhaps a third slower. The rule for it is one line: nothing in it may run inside a frame. Nothing there is on the drag path, the pointer path or in the interrupt, and the timings were unchanged to the T state across the move. It starts at 6000 rather than at the top of the system variables because the tape loader indexes the system variable area through IY and the BASIC loader itself lives just above it, and a CODE block that overwrote the program doing the loading would be a novel way to fail.
The heap stops a page short of the top of memory rather than at FFFF. Every walk computes the next block as address plus header plus size, and a block ending at 10000 would wrap to nought and compare as below the base of the heap. Stopping at FF00 costs 256 bytes of 8,272 and removes the entire class of failure.
This is the part I’d want if I were reading someone else’s repo. Every figure below was measured by running the code and timing it against the machine’s own clock. None of it comes from counting instructions, and the few claims that are derived say so.
The 48K interrupt period is exactly 69,888 T states and nothing a program does can move it, so it is the only usable clock on the machine. So a timing run syncs on HALT, optionally delays a known number of T states to place the routine at a chosen point in the frame, calls it, then counts turns of a sixteen T loop until the next interrupt. Comparing against an empty calibration run cancels every fixed overhead:
118 T is the exact cost of the returning interrupt path, counted instruction by instruction. It is exact rather than estimated because the handler, its variables, the counting loop and the stack all live above 8000, where the ULA never steals a cycle. Only the routine being timed touches contended memory.
Measured blind against three delays of known length:
Worst error is 7 T in 104,000, or 0.007%. The third crosses a frame boundary, which is what confirms the 118 T figure.
Every budget in the project started from a pessimistic 50% penalty on screen writes, taken from the folklore. Sweeping a 1,024 byte fill across the frame:
The two border figures agree to 1 T, which is the uncontended cost. The worst case inside the display is 13,473 T. That is 14.7%, or about 1.7 T per contended byte written. Budgets built on the 50% figure are roughly a third too conservative, and mine were.
WinDraw split into its four phases and each timed separately. The phases sum to within 330 T of the whole, which is the four call and return pairs plus the sixteen T counting granularity.
Twenty seven characters of title and body text, at roughly 1,330 T each. The fills, which I had assumed were the problem, are a fifth of it. Optimising the fill would have bought a few per cent of a drag frame and I’d have spent a week on it.
An earlier note recorded the cell aware push fill at 7.4 T per byte. The real routine costs 11.5 T per byte, 54% more. The benchmark had measured the technique; the routine carries per row address arithmetic the benchmark never paid. Of the 183 T a row costs, 88 T is the push chain actually writing pixels and 95 T is the register exchange, the address step, the cell boundary test and the loop.
Over half the cost of the fastest fill on the machine is not writing pixels. That generalises: on this processor, per row overhead is the thing to attack, not per byte throughput.
This is the one I’d most like other people on this hardware to know about, because it produces a fault that looks like anything except its cause.
The fills held DI for their whole run, because SP walks through screen memory and stops being a stack. The received wisdom is that this delays the interrupt. It does not. The Spectrum asserts INT for only 32 T states and then withdraws it, so a DI window that covers those 32 T destroys the interrupt rather than postponing it.
It was observed before it was understood. A fill placed at 62,399 T into the frame reported crossing no frame boundary when it plainly crossed one. Rescoring it as a lost interrupt gives 11,745 T against 11,744 T for the same fill in the top border, and the two agreeing to 1 T is what confirmed the diagnosis.
Then it was measured properly. With an interrupt injected after every single instruction of an 8 by 24 fill, the interrupt was refused at 458 of 500 instruction boundaries in one fill routine and 710 of 752 in the other.
The obvious fix does not work. Re-enabling interrupts between rows sounds right and fails, because the interrupt is not pending, it is gone. An EI window a few T states wide, once every 236 T, catches it about one row in thirty. Making the DI region short is not the same as making it absent, and only absent is a fix.
What works is owing the last push. The fills now run with interrupts enabled throughout. What DI was protecting was SP, so the fix is to guarantee that the two bytes of return address always land somewhere that is about to be overwritten anyway. SP takes two kinds of value: inside the rectangle, where a push writes exactly what the chain’s next push will write, and the low point after the last push of a row, which the chain never returns to. So the chain is made one push shorter than the row, and the leftmost two bytes are owed — paid one iteration later, once SP has moved into the next row and can no longer reach them.
Afterwards: 0 of 607 and 0 of 904 instruction boundaries destroy the interrupt, the rectangle is byte for byte identical every time, nothing outside it is touched, and the screen checksums are unchanged either side of the change.
That’s a real cost, and it’s worth it. The drag path is mostly the column fill, which writes through HL, never touched SP and was never at risk.
A frame watchdog counts interrupts in the handler and frames in the main loop; the difference is frames dropped. The handler is the awkward half, because it fires inside a fill where the stack is not a stack.
The first version borrowed IX and counted with INC (IX+0), and the interrupt sweep above failed on its first run. That instruction sets the flags, and a fill holds a live carry across the address step that finds the end of a row and the branch that decides whether the row crossed a boundary. An interrupt in that gap stole the carry and the row stepped to the wrong address.
INC IX is a sixteen bit increment and sixteen bit increments leave the flags alone. So the counter is a word, the handler is transparent, and it costs 104 T fifty times a second with not one flag or register altered. The roadmap had predicted that a watchdog would have caught the interrupt bug earlier. What actually happened is that the interrupt test caught the watchdog.
Letting the handler push is also a dead end. A handler that pushes AF uses four bytes below SP rather than two, so the owed region has to double and every row pays another 22 T.
The starting position was hopeless: a full erase and redraw of a window cost 96,010 T against a frame of 69,888 T. That is 137% of a frame, and it is why dragging ran at 25 Hz. Four changes, in order, each measured:
A drag frame now ends at 59,858 T, inside the frame, and dragging runs at 50 Hz.