17 Epilogue: long mode, UEFI, SMP and where to go next

The kernel you have at the end of chapter 15 boots from a hard disk, runs in protected mode with paging, handles interrupts, schedules processes, runs user programs in ring 3 behind system calls, reads them from an ext2 filesystem, and lets a process fork itself and exec another program in a fresh address space. Chapter 16 taught you what to do on the days it does not boot. It is a real operating system in the sense that matters for this book: every layer between the CPU reset vector and a user-space printf is yours, and you can single-step through all of it with gdb. It is also, deliberately, a 32-bit operating system that boots the way PCs booted in 1995. This closing chapter explains what separates it from the kernels running on the machine in front of you, which manual chapters to read next, and ends with three projects that are the examination of this book. None of these topics needs more than the knowledge you already have; they need the same method: read the official document, write a small experiment, inspect it.

17.1 64-bit: long mode

Every x86 processor sold since 2006 implements the 64-bit extension that AMD, who designed it, calls long mode, and that Intel calls IA-32e mode (chapter 3, “Two vendors, one architecture”, tells the story). From here on we cite both manuals side by side: the Intel SDM, as in the rest of the book, and the AMD64 Architecture Programmer’s Manual (APM) Volume 2, System Programming (AMD 2026), whose chapter 14 is the original description of the mechanism. Appendix D, Reading the Intel and AMD manuals, maps every chapter of this book to both.

Intel SDM Volume 3A, section 2.2 “Modes of Operation” (AMD APM Volume 2, section 1.3 “Operating Modes”, figure 1-6) shows the state diagram: a processor starts in real-address mode, enters protected mode by setting CR0.PE as we did in chapter 9, and enters long mode from protected mode by enabling paging with a specific page-table format and a flag in a model-specific register. The sequence is in Volume 3A, section 12.8.5 “Initializing IA-32e Mode” (AMD APM Volume 2, section 14.6 “Enabling and Activating Long Mode”, in particular 14.6.1 “Activating Long Mode”), and it is only about a dozen instructions longer than our switch to protected mode:

  1. Build page tables in the 4-level format of Volume 3A, section 5.5 “4-Level Paging and 5-Level Paging” (AMD APM Volume 2, section 5.3 “Long-Mode Page Translation”): PML4, page-directory-pointer tables, page directories, page tables, each entry 8 bytes instead of 4, and load CR3 with the address of the PML4. The bit layout of an entry is the same idea as the 32-bit entries of chapter 12, with a wider physical address field and a new no-execute bit at the top (Intel Table 5-15 to Table 5-20; AMD figures 5-20 to 5-23).
  2. Set CR4.PAE (bit 5), which selects that page-table format (Volume 3A, section 2.5 “Control Registers”; AMD APM Volume 2, section 3.1.3 “CR4 Register”).
  3. Set the LME bit (bit 8) of the IA32_EFER model-specific register, which AMD simply calls EFER, with the wrmsr instruction. Its address, C000_0080h, and its bits are the same on both vendors: Volume 3A, section 2.2.1 “Extended Feature Enable Register” and Volume 4 for the MSR list; AMD APM Volume 2, section 3.1.7 “Extended Feature Enable Register (EFER)” and Appendix A “MSR Cross-Reference”. The register was defined by AMD, which is why its address is not in the Intel range.
  4. Set CR0.PG. The processor is now in long mode and sets EFER.LMA (bit 10) itself; it still executes 32-bit code, which both manuals call compatibility mode.
  5. Load a GDT containing a code descriptor with the L bit (bit 53 of the descriptor, “64-bit code segment”, Volume 3A, section 3.4.5 and section 6.2.1 “Code-Segment Descriptor in 64-bit Mode”; AMD APM Volume 2, section 4.8.1 “Code-Segment Descriptors”, figure 4-20) set, and far-jump to it. The code after the jump is 64-bit.

Both manuals list the same consistency checks (Intel 12.8.5; AMD 14.6.2 “Consistency Checks”): setting LME while paging is on, or enabling paging with LME set and PAE clear, raises a general-protection fault. Appendix C, A minimal long-mode bootstrap, performs the five steps in one boot sector, prints a line from 64-bit code, and shows the switch in gdb; read it when you want to see the real thing in 110 lines of assembly.

What changes for a kernel written in C: the general-purpose registers become 64 bits wide (rax, rbx, …) and eight more appear (r8 to r15); the calling convention passes the first six integer arguments in registers (rdi, rsi, rdx, rcx, r8, r9) instead of on the stack, which is the System V AMD64 ABI document rather than the i386 one we used in chapter 4; segmentation is mostly switched off (base and limit are ignored for CS, DS, ES and SS, which is why a 64-bit GDT is so short), so the TSS is used only for stack switching and the interrupt-stack table (Volume 3A, section 7.14 “Exception and Interrupt Handling in 64-bit Mode”; AMD APM Volume 2, section 8.9 “Long-Mode Interrupt Control Transfers” and 12.2.5 “64-Bit Task State Segment”); the interrupt frame and the gate descriptors are 16 bytes wide (Volume 3A, sections 7.14.1 and 7.14.2; AMD 8.9.1 and 8.9.3); and syscall/sysret replace int 0x80 as the fast way to enter the kernel. Those two instructions are AMD’s: they were introduced with the AMD-K6 (AMD 1998) and are the only fast system call of long mode on AMD processors (APM Volume 2, section 6.1.1 “SYSCALL and SYSRET”; section 6.1.2 is titled “SYSENTER and SYSEXIT (Legacy Mode Only)”); Intel adopted them for 64-bit mode (Volume 3A, section 6.8.8 “Fast System Calls in 64-Bit Mode”, and the instruction pages of Volume 2B). Everything else, from the PIC to ext2, is unchanged.

If you want to redo Part III in 64-bit, the smallest useful change is to let the bootloader perform the five steps above and then jump to a kernel compiled with -m64 -mcmodel=kernel -mno-red-zone (the red zone is a 128-byte area below the stack pointer that the AMD64 ABI lets functions use without adjusting rsp; an interrupt arriving at that moment would overwrite it, so kernels disable it). gcc and ld in the container already support this; the elf_x86_64 emulation of ld replaces elf_i386, and qemu-system-x86_64 replaces qemu-system-i386. Appendix C ends with the list of what else the C kernel has to change.

17.1.1 What differs between Intel and AMD, and what does not

A kernel that follows this book never has to ask which vendor it runs on, and that is worth stating precisely. The following are specified identically by both manuals, down to the bit: real-address mode and the switch to protected mode, segment descriptors, the GDT, IDT and TSS, the exception vectors and their error codes, the 32-bit and the 4-level page-table formats including the no-execute bit, EFER and the long-mode activation sequence, the 64-bit interrupt frame, syscall and sysret in 64-bit mode, and 5-level paging, which both enable with the same bit, CR4.LA57 (Volume 3A, section 5.5; AMD APM Volume 2, section 3.1.3 and figures 5-18, 5-25 and 5-31). The devices of chapters 10, 11 and 14 (serial port, PIC, PIT, keyboard controller, VGA, ATA) belong to the chipset, not to the processor, and do not know which CPU is plugged in.

What differs, and where to read about it:

17.2 UEFI and the end of the BIOS

The BIOS services we used in chapters 7 and 9 (INT 13h for the disk, INT 10h for the screen, INT 15h for the memory map) are the oldest software interface still in use on PCs; they date from 1981 and run in real mode. Since about 2012 PCs ship with UEFI firmware instead, which starts the processor directly in 64-bit mode, reads a FAT filesystem on a dedicated partition, loads a PE executable (the Windows file format, not ELF) from it, and hands that executable a table of C function pointers for disk, console, memory and graphics access. There is no 512-byte boot sector and no INT instruction. Most firmware still offers a compatibility support module that emulates the BIOS, which is what QEMU’s SeaBIOS is, but it is being removed from new machines.

The specification is at https://uefi.org/specifications; the chapters to read are 2 (boot manager), 4 (the system table) and 7 (boot services, in particular AllocatePages and GetMemoryMap, which replace INT 15h, E820h), and 13 for the file protocol. QEMU can run UEFI firmware with the OVMF image shipped by most distributions (-bios OVMF.fd), and the gnu-efi project provides the linker script and crt0 needed to produce a PE executable with gcc. A UEFI loader for the kernel of chapter 15 is a good project: it reads the ELF file with the file protocol, allocates memory with the boot services, calls ExitBootServices, and jumps to the entry point with a memory map in hand, replacing the entire bootloader of chapter 9 with C. The FAT filesystem it boots from is the first capstone project below.

17.3 Multiple processors

QEMU has been emulating a single CPU for us; -smp 4 gives four. On x86 all processors but one start halted; the running one, the bootstrap processor, wakes the others with inter-processor interrupts sent through the local APIC, a per-processor interrupt controller that supersedes the 8259A of chapter 11. The procedure is in Volume 3A, chapter 13 “Advanced Programmable Interrupt Controller (APIC)” and section 11.4 “Multiple-Processor (MP) Initialization” (AMD APM Volume 2, chapter 16, section 16.5 “Interprocessor Interrupts (IPI)”, and section 14.1.4 “Multiple Processor Initialization”): find the processors in the ACPI tables (the MADT), send an INIT IPI then two STARTUP IPIs carrying the page number of a 16-bit trampoline, and let each processor run the real-mode-to-protected-mode sequence of chapter 9 on its own. Every structure in our kernel that is written from an interrupt handler or from two tasks (the run queue, the heap’s free list, the tick counter) then needs a lock; the lock-prefixed instructions and xchg of Volume 3A, chapter 11 “Multiple-Processor Management” are the building blocks, and Volume 3A, section 11.2 “Memory Ordering” (AMD APM Volume 2, section 7.2 “Multiprocessor Memory Access Ordering”) explains the memory ordering rules that make them necessary.

The I/O APIC (Volume 3A, section 13.1, and the chipset datasheet) replaces the PIC for routing device interrupts to any processor, and the high precision event timer or the local APIC timer replace the PIT. Our drivers from chapters 10 and 11 keep working on a multiprocessor machine as long as only one processor services interrupts, which is why the PIC is still a reasonable first step.

17.4 Other things a real kernel has

17.5 Capstone projects

Every part of this book closed with a milestone project, and every chapter with exercises whose answers are in Appendix E. The three projects below are different: they are the examination. There are no answers for them anywhere in the book or its repository, and there will not be, because a solution you can read is not a solution you can write. Each project comes with a rubric instead: a checklist of what a complete solution does, and a way to know that it works. Treat the checklist the way a reviewer would treat your pull request. Each project needs between a few hundred and a thousand lines, builds on the kernel of chapter 15 without changing its structure, and should keep make test green at every step; chapter 16 is your companion when it does not.

17.5.1 Project 1: a FAT12 reader

The kernel of chapter 14 reads one filesystem, ext2, from one disk. Give it a second disk with a second filesystem: FAT12, the 1980 design that UEFI still boots from and every SD card still carries. The format is specified in Microsoft’s FAT32 File System Specification (the document also covers FAT12 and FAT16; look for the BPB, the FAT and the directory-entry layouts), and the OSDev wiki page “FAT” is a good second reading. The image is built on the host, because the container does not ship dosfstools: mkfs.fat -F 12 -C fat.img 1440 creates a 1.44 MB image, and mcopy -i fat.img file ::/ from mtools (or a loop mount) copies files into it; commit the Makefile rule, not the image. Attach it as the second drive of the IDE controller that chapter 14 added: a second -drive with if=none and a second ide-hd device on bus=ide.0 with unit=1, which is the ATA slave of the primary channel.

A complete solution:

How you know it works: build the image with three files, one of them larger than one cluster and one in a subdirectory, and compare the bytes the kernel prints on the serial port with mtype -i fat.img ::/FILE.TXT on the host; make test greps for both the ext2 Hello from user space and a line from the FAT file. Then run fsck.fat -v fat.img and compare its numbers with the ones your superblock parser printed.

Stretch goals: write support (allocate clusters, update both FATs, create a directory entry, then verify with fsck.fat), long file names (the VFAT entries you skipped), and a tiny virtual filesystem layer in which / is ext2 and /fat is the second disk, so that exec can run a program copied onto the FAT image.

17.5.2 Project 2: a shell

Chapter 15 gave you fork, exec and wait; the keyboard driver of chapter 11 turns scan codes into characters. The missing piece between them is a shell: a program that reads a line, splits it into words, runs the program named by the first word with the others as arguments, and waits for it. The design document is the POSIX sh description, of which you implement a tenth.

A complete solution:

How you know it works: make test boots, feeds the shell a script through the serial port (QEMU’s -serial is bidirectional; make read accept the serial port as a second console, which is also how you will debug it), runs ls, echo one two, hello, and a program that faults, and greps the serial output for each expected line and for the prompt after the fault.

Stretch goals: | with a pipe system call and a dup that lets read and write go through file descriptors instead of fixed devices; > redirection into a file created with the ext2 write support of chapter 14; & to run a process without waiting; and Ctrl-C, which needs the beginnings of signals.

17.5.3 Project 3: a priority scheduler with sleeping and accounting

The scheduler of chapter 13 is round-robin: every ready task runs for one time slice, forever, all tasks are equal, and task_wake_sleepers looks at every task on every tick to find the sleeping ones. Replace it with a scheduler that has priorities, sleeps efficiently, and keeps accounts, then write the program that shows the accounts.

A complete solution:

How you know it works: start three CPU-bound tasks with priorities 0, 5 and 10 and one that sleeps for a second between messages; after ten seconds ps must show the CPU time divided as your rule says it should be (write the expected ratio down before running), the sleeper with almost no CPU time, and the idle task with the rest. make test greps for the ps output of a run with two tasks of known priorities and a tolerance.

Stretch goals: a time slice that depends on priority; getrusage-style accounting of user and kernel time separately (the system call entry is where the line is); a top that re-reads the table every second using sleep; and, with -smp 2, one run queue per processor, which is the beginning of the chapter on multiple processors you will write yourself.

17.6 Where to go next

Writing an operating system is fun. Keep going.