17 Epilogue: long mode, UEFI, SMP and where to go next
The kernel you have at the end of chapter 15 boots from a hard disk,
runs in protected mode with paging, handles interrupts, schedules
processes, runs user programs in ring 3 behind system calls, reads them
from an ext2 filesystem, and lets a process fork itself and
exec another program in a fresh address space. Chapter 16
taught you what to do on the days it does not boot. It is a real
operating system in the sense that matters for this book: every layer
between the CPU reset vector and a user-space printf is
yours, and you can single-step through all of it with gdb. It is also,
deliberately, a 32-bit operating system that boots the way PCs booted in
1995. This closing chapter explains what separates it from the kernels
running on the machine in front of you, which manual chapters to read
next, and ends with three projects that are the examination of this
book. None of these topics needs more than the knowledge you already
have; they need the same method: read the official document, write a
small experiment, inspect it.
17.1 64-bit: long mode
Every x86 processor sold since 2006 implements the 64-bit extension that AMD, who designed it, calls long mode, and that Intel calls IA-32e mode (chapter 3, “Two vendors, one architecture”, tells the story). From here on we cite both manuals side by side: the Intel SDM, as in the rest of the book, and the AMD64 Architecture Programmer’s Manual (APM) Volume 2, System Programming (AMD 2026), whose chapter 14 is the original description of the mechanism. Appendix D, Reading the Intel and AMD manuals, maps every chapter of this book to both.
Intel SDM Volume 3A, section 2.2 “Modes of Operation” (AMD APM Volume 2, section 1.3 “Operating Modes”, figure 1-6) shows the state diagram: a processor starts in real-address mode, enters protected mode by setting CR0.PE as we did in chapter 9, and enters long mode from protected mode by enabling paging with a specific page-table format and a flag in a model-specific register. The sequence is in Volume 3A, section 12.8.5 “Initializing IA-32e Mode” (AMD APM Volume 2, section 14.6 “Enabling and Activating Long Mode”, in particular 14.6.1 “Activating Long Mode”), and it is only about a dozen instructions longer than our switch to protected mode:
- Build page tables in the 4-level format of Volume 3A, section 5.5 “4-Level Paging and 5-Level Paging” (AMD APM Volume 2, section 5.3 “Long-Mode Page Translation”): PML4, page-directory-pointer tables, page directories, page tables, each entry 8 bytes instead of 4, and load CR3 with the address of the PML4. The bit layout of an entry is the same idea as the 32-bit entries of chapter 12, with a wider physical address field and a new no-execute bit at the top (Intel Table 5-15 to Table 5-20; AMD figures 5-20 to 5-23).
- Set CR4.PAE (bit 5), which selects that page-table format (Volume 3A, section 2.5 “Control Registers”; AMD APM Volume 2, section 3.1.3 “CR4 Register”).
- Set the LME bit (bit 8) of the
IA32_EFERmodel-specific register, which AMD simply callsEFER, with thewrmsrinstruction. Its address,C000_0080h, and its bits are the same on both vendors: Volume 3A, section 2.2.1 “Extended Feature Enable Register” and Volume 4 for the MSR list; AMD APM Volume 2, section 3.1.7 “Extended Feature Enable Register (EFER)” and Appendix A “MSR Cross-Reference”. The register was defined by AMD, which is why its address is not in the Intel range. - Set CR0.PG. The processor is now in long mode and sets EFER.LMA (bit 10) itself; it still executes 32-bit code, which both manuals call compatibility mode.
- Load a GDT containing a code descriptor with the L bit (bit 53 of the descriptor, “64-bit code segment”, Volume 3A, section 3.4.5 and section 6.2.1 “Code-Segment Descriptor in 64-bit Mode”; AMD APM Volume 2, section 4.8.1 “Code-Segment Descriptors”, figure 4-20) set, and far-jump to it. The code after the jump is 64-bit.
Both manuals list the same consistency checks (Intel 12.8.5; AMD 14.6.2 “Consistency Checks”): setting LME while paging is on, or enabling paging with LME set and PAE clear, raises a general-protection fault. Appendix C, A minimal long-mode bootstrap, performs the five steps in one boot sector, prints a line from 64-bit code, and shows the switch in gdb; read it when you want to see the real thing in 110 lines of assembly.
What changes for a kernel written in C: the general-purpose registers
become 64 bits wide (rax, rbx, …) and eight
more appear (r8 to r15); the calling
convention passes the first six integer arguments in registers
(rdi, rsi, rdx, rcx,
r8, r9) instead of on the stack, which is the
System V AMD64 ABI document rather than the i386 one we used in
chapter 4; segmentation is mostly switched off (base and limit are
ignored for CS, DS, ES and SS, which is why a 64-bit GDT is so short),
so the TSS is used only for stack switching and the interrupt-stack
table (Volume 3A, section 7.14 “Exception and Interrupt Handling in
64-bit Mode”; AMD APM Volume 2, section 8.9 “Long-Mode Interrupt Control
Transfers” and 12.2.5 “64-Bit Task State Segment”); the interrupt frame
and the gate descriptors are 16 bytes wide (Volume 3A, sections 7.14.1
and 7.14.2; AMD 8.9.1 and 8.9.3); and
syscall/sysret replace int 0x80
as the fast way to enter the kernel. Those two instructions are AMD’s:
they were introduced with the AMD-K6 (AMD
1998) and are the only fast system call of long mode on AMD
processors (APM Volume 2, section 6.1.1 “SYSCALL and SYSRET”; section
6.1.2 is titled “SYSENTER and SYSEXIT (Legacy Mode Only)”); Intel
adopted them for 64-bit mode (Volume 3A, section 6.8.8 “Fast System
Calls in 64-Bit Mode”, and the instruction pages of Volume 2B).
Everything else, from the PIC to ext2, is unchanged.
If you want to redo Part III in 64-bit, the smallest useful change is
to let the bootloader perform the five steps above and then jump to a
kernel compiled with -m64 -mcmodel=kernel -mno-red-zone
(the red zone is a 128-byte area below the stack pointer that the AMD64
ABI lets functions use without adjusting rsp; an interrupt
arriving at that moment would overwrite it, so kernels disable it). gcc
and ld in the container already support this; the
elf_x86_64 emulation of ld replaces elf_i386,
and qemu-system-x86_64 replaces
qemu-system-i386. Appendix C ends with the list of what
else the C kernel has to change.
17.1.1 What differs between Intel and AMD, and what does not
A kernel that follows this book never has to ask which vendor it runs
on, and that is worth stating precisely. The following are specified
identically by both manuals, down to the bit: real-address mode and the
switch to protected mode, segment descriptors, the GDT, IDT and TSS, the
exception vectors and their error codes, the 32-bit and the 4-level
page-table formats including the no-execute bit, EFER and
the long-mode activation sequence, the 64-bit interrupt frame,
syscall and sysret in 64-bit mode, and 5-level
paging, which both enable with the same bit, CR4.LA57 (Volume 3A,
section 5.5; AMD APM Volume 2, section 3.1.3 and figures 5-18, 5-25 and
5-31). The devices of chapters 10, 11 and 14 (serial port, PIC, PIT,
keyboard controller, VGA, ATA) belong to the chipset, not to the
processor, and do not know which CPU is plugged in.
What differs, and where to read about it:
- Virtualization. The extensions are two different designs: Intel’s VT-x, in SDM Volume 3C, chapters 26 to 33, from “Introduction to Virtual Machine Extensions” to “VMX Instruction Reference”, and AMD’s AMD-V, which the APM calls Secure Virtual Machine (Volume 2, chapter 15). A hypervisor has two back-ends; a kernel has none.
- The other fast system call. AMD supports
sysenter/sysexitin legacy mode only (APM Volume 2, section 6.1.2), whereas Intel supports them in IA-32e mode too (Volume 3A, section 6.8.7.1) butsyscall/sysretin 64-bit mode only (section 6.8.8: “not supported in compatibility mode”), while AMD’ssyscallalso works from 32-bit code, with its own target registers (theSTARandCSTARMSRs of APM figure 6-1). A 64-bit kernel that usessyscallfor 64-bit programs is portable; one that wants 32-bit programs to make fast system calls must checkcpuid. - Model-specific registers.
EFERis shared, but AMD defines bits in it that Intel does not (SVME, bit 12, enables AMD-V; compare AMD figure 3-9 with Intel figure 2-4, which has onlySCE,LME,LMAandNXE), and each vendor has its own list of MSRs: Intel’s is a whole volume, Volume 4; AMD’s is Appendix A of APM Volume 2, withSYSCFGand the like in section 3.2. Writing an MSR the processor does not have raises#GP, so a kernel checkscpuidfirst. cpuidleaves. Leaf 0 (the vendor string), leaf 1 and the extended leaf8000_0001hare common, and the feature bits that matter here are in the latter on both vendors: long mode isEDX[29],syscallisEDX[11], no-execute isEDX[20](Volume 3A, sections 5.1.4 and 6.8.8; AMD APM Volume 3, section E.4.2). Beyond that the leaves diverge: cache and processor topology, for instance, are Intel’s leaves04Hand0BH(Volume 3A, section 11.9.2) and AMD’s8000_001Dhand8000_001Eh(APM Volume 3, sections E.4.15 and E.4.16). A kernel that walks topology needs two code paths.- The local APIC. The xAPIC and x2APIC programming models are the same on both (Volume 3A, chapter 13; AMD APM Volume 2, chapter 16), but AMD adds extended registers behind the “Extended APIC Feature Register” (APM Volume 2, section 16.3.5), which a portable kernel ignores.
17.2 UEFI and the end of the BIOS
The BIOS services we used in chapters 7 and 9 (INT 13h
for the disk, INT 10h for the screen, INT 15h
for the memory map) are the oldest software interface still in use on
PCs; they date from 1981 and run in real mode. Since about 2012 PCs ship
with UEFI firmware instead, which starts the processor directly
in 64-bit mode, reads a FAT filesystem on a dedicated partition, loads a
PE executable (the Windows file format, not ELF) from it, and hands that
executable a table of C function pointers for disk, console, memory and
graphics access. There is no 512-byte boot sector and no
INT instruction. Most firmware still offers a
compatibility support module that emulates the BIOS, which is
what QEMU’s SeaBIOS is, but it is being removed from new machines.
The specification is at https://uefi.org/specifications; the chapters to read
are 2 (boot manager), 4 (the system table) and 7 (boot services, in
particular AllocatePages and GetMemoryMap,
which replace INT 15h, E820h), and 13 for the file
protocol. QEMU can run UEFI firmware with the OVMF image
shipped by most distributions (-bios OVMF.fd), and the
gnu-efi project provides the linker script and crt0 needed
to produce a PE executable with gcc. A UEFI loader for the kernel of
chapter 15 is a good project: it reads the ELF file with the file
protocol, allocates memory with the boot services, calls
ExitBootServices, and jumps to the entry point with a
memory map in hand, replacing the entire bootloader of chapter 9 with C.
The FAT filesystem it boots from is the first capstone project
below.
17.3 Multiple processors
QEMU has been emulating a single CPU for us; -smp 4
gives four. On x86 all processors but one start halted; the running one,
the bootstrap processor, wakes the others with
inter-processor interrupts sent through the local
APIC, a per-processor interrupt controller that supersedes the
8259A of chapter 11. The procedure is in Volume 3A, chapter 13 “Advanced
Programmable Interrupt Controller (APIC)” and section 11.4
“Multiple-Processor (MP) Initialization” (AMD APM Volume 2, chapter 16,
section 16.5 “Interprocessor Interrupts (IPI)”, and section 14.1.4
“Multiple Processor Initialization”): find the processors in the ACPI
tables (the MADT), send an INIT IPI then two STARTUP IPIs carrying the
page number of a 16-bit trampoline, and let each processor run the
real-mode-to-protected-mode sequence of chapter 9 on its own. Every
structure in our kernel that is written from an interrupt handler or
from two tasks (the run queue, the heap’s free list, the tick counter)
then needs a lock; the lock-prefixed instructions and
xchg of Volume 3A, chapter 11 “Multiple-Processor
Management” are the building blocks, and Volume 3A, section 11.2 “Memory
Ordering” (AMD APM Volume 2, section 7.2 “Multiprocessor Memory Access
Ordering”) explains the memory ordering rules that make them
necessary.
The I/O APIC (Volume 3A, section 13.1, and the chipset datasheet) replaces the PIC for routing device interrupts to any processor, and the high precision event timer or the local APIC timer replace the PIT. Our drivers from chapters 10 and 11 keep working on a multiprocessor machine as long as only one processor services interrupts, which is why the PIC is still a reasonable first step.
17.4 Other things a real kernel has
- A proper ELF loader and dynamic linking. Our
chapter 14 loader handles static executables, and chapter 15 gives each
one its own address space. Shared libraries add the
.dynamicsection,PT_INTERP(the program interpreter,ld-linux.soon Linux) and the relocations of chapter 8 applied at run time; the ELF specification you already read covers it, andreadelf -don any Linux program shows the data. - A virtual filesystem layer and a page cache, so that ext2 is one of several filesystems and disk blocks are cached in memory. The ext2 specification was chosen because it is small; ext4, FAT32 and ISO 9660 follow the same pattern of a superblock, allocation structures and directory entries.
- PCI enumeration and real device drivers. The serial
port, keyboard, PIT and ATA PIO devices live at fixed legacy I/O ports,
which let us skip discovering them. Everything else (network cards, AHCI
disk controllers, USB, graphics) is found by walking PCI configuration
space (ports
0xCF8/0xCFC, or the memory-mapped region ACPI describes), which is specified in the PCI Local Bus Specification and summarized well on the OSDev wiki. An AHCI driver is the natural next step after chapter 14 on the q35 machine. - Power management and ACPI. Shutting the machine down, reading the real time clock, enumerating processors: all of it goes through the ACPI tables the firmware leaves in memory.
- Signals, pipes, a shell, a C library. With
fork,exec,wait,read,writeandopenas system calls you can run a port of a small shell and of a C library such as musl; at that point your operating system can compile and run programs written by other people. The shell is the second capstone project.
17.5 Capstone projects
Every part of this book closed with a milestone project, and every
chapter with exercises whose answers are in Appendix E. The three
projects below are different: they are the examination. There are no
answers for them anywhere in the book or its repository, and there will
not be, because a solution you can read is not a solution you can write.
Each project comes with a rubric instead: a checklist of what a
complete solution does, and a way to know that it works. Treat the
checklist the way a reviewer would treat your pull request. Each project
needs between a few hundred and a thousand lines, builds on the kernel
of chapter 15 without changing its structure, and should keep
make test green at every step; chapter 16 is your companion
when it does not.
17.5.1 Project 1: a FAT12 reader
The kernel of chapter 14 reads one filesystem, ext2, from one disk.
Give it a second disk with a second filesystem: FAT12, the 1980 design
that UEFI still boots from and every SD card still carries. The format
is specified in Microsoft’s FAT32 File System Specification
(the document also covers FAT12 and FAT16; look for the BPB, the FAT and
the directory-entry layouts), and the OSDev wiki page “FAT” is a good
second reading. The image is built on the host, because the container
does not ship dosfstools:
mkfs.fat -F 12 -C fat.img 1440 creates a 1.44 MB image, and
mcopy -i fat.img file ::/ from mtools (or a
loop mount) copies files into it; commit the Makefile rule, not the
image. Attach it as the second drive of the IDE controller that chapter
14 added: a second -drive with if=none and a
second ide-hd device on bus=ide.0 with
unit=1, which is the ATA slave of the primary
channel.
A complete solution:
How you know it works: build the image with three files, one of them
larger than one cluster and one in a subdirectory, and compare the bytes
the kernel prints on the serial port with
mtype -i fat.img ::/FILE.TXT on the host;
make test greps for both the ext2
Hello from user space and a line from the FAT file. Then
run fsck.fat -v fat.img and compare its numbers with the
ones your superblock parser printed.
Stretch goals: write support (allocate clusters, update both FATs,
create a directory entry, then verify with fsck.fat), long
file names (the VFAT entries you skipped), and a tiny virtual
filesystem layer in which / is ext2 and
/fat is the second disk, so that exec can run
a program copied onto the FAT image.
17.5.2 Project 2: a shell
Chapter 15 gave you fork, exec and
wait; the keyboard driver of chapter 11 turns scan codes
into characters. The missing piece between them is a shell: a program
that reads a line, splits it into words, runs the program named by the
first word with the others as arguments, and waits for it. The design
document is the POSIX sh description, of which you
implement a tenth.
A complete solution:
How you know it works: make test boots, feeds the shell
a script through the serial port (QEMU’s -serial is
bidirectional; make read accept the serial port as a second
console, which is also how you will debug it), runs ls,
echo one two, hello, and a program that
faults, and greps the serial output for each expected line and for the
prompt after the fault.
Stretch goals: | with a pipe system call
and a dup that lets read and
write go through file descriptors instead of fixed devices;
> redirection into a file created with the ext2 write
support of chapter 14; & to run a process without
waiting; and Ctrl-C, which needs the beginnings of
signals.
17.5.3 Project 3: a priority scheduler with sleeping and accounting
The scheduler of chapter 13 is round-robin: every ready task runs for
one time slice, forever, all tasks are equal, and
task_wake_sleepers looks at every task on every tick to
find the sleeping ones. Replace it with a scheduler that has priorities,
sleeps efficiently, and keeps accounts, then write the program that
shows the accounts.
A complete solution:
How you know it works: start three CPU-bound tasks with priorities 0,
5 and 10 and one that sleeps for a second between messages; after ten
seconds ps must show the CPU time divided as your rule says
it should be (write the expected ratio down before running),
the sleeper with almost no CPU time, and the idle task with the rest.
make test greps for the ps output of a run
with two tasks of known priorities and a tolerance.
Stretch goals: a time slice that depends on priority;
getrusage-style accounting of user and kernel time
separately (the system call entry is where the line is); a
top that re-reads the table every second using
sleep; and, with -smp 2, one run queue per
processor, which is the beginning of the chapter on multiple processors
you will write yourself.
17.6 Where to go next
- The OSDev wiki (https://wiki.osdev.org) is organized exactly like Part III: one page per hardware mechanism, with references into the manuals. You now meet its “required knowledge” page with room to spare.
- xv6 (MIT) is a 64-bit teaching kernel of about ten
thousand lines with a book-length commentary; it implements everything
in the list above in the simplest possible way, and reading it after
this book will feel familiar, because its structure (bootloader,
kernel/,user/, a Makefile that builds a filesystem image) is the structure you built. - Linux Insides and the Linux source itself, with
arch/x86/bootandarch/x86/kernel/head_64.Sas the starting points: the boot sequence of chapter 9 and the 64-bit sequence above are both there, in production form. - The Intel and AMD manuals remain the final reference. You have read parts of Volumes 1, 2 and 3 of the Intel SDM, and Appendix D tells you where the same material is in the AMD APM; Volume 3 of the SDM has chapters on virtualization (VMX), performance monitoring and system management mode that open other fields entirely, and APM Volume 2 has the AMD side of each.
Writing an operating system is fun. Keep going.