Appendix E Answers to the “Check your understanding” questions

Each chapter ends with a handful of questions. They are not an exam: they are the questions we would ask you after you have read the chapter, to find out together whether the important points went through. Try to answer in writing before looking here; the explanations are short on purpose, and the chapter has the long version. The sections below follow the chapters in order, one per chapter from chapter 0 to chapter 16, and the answers are numbered as the questions are.

E.1 Chapter 0: Setting up the development environment

  1. A position-independent executable embeds no absolute address, because the operating system moves it to a random address at every run. The compiler must therefore generate code that computes “where am I” at run time, which is the call __x86.get_pc_thunk.ax followed by add and lea in the first listing. On bare metal nothing relocates our code, and the whole point of Part III is to choose and control fixed addresses, so this machinery only hides what we want to see.

  2. The thunk returns the address of the instruction after the call, which is the only way 32-bit x86 code can read its own eip. The second program knows its address at link time, 0x8049166, so the address of the string is a constant, 0x804a008, pushed directly. Same program, but the second one is readable by a beginner.

  3. The kernel of Part III runs in 32-bit protected mode, and the registers, calling convention and instruction encodings you learn in Part I are exactly those you will meet in gdb while debugging that kernel. Teaching 64-bit in Part I would mean learning two of everything. The 64-bit differences are covered once, in chapter 17 and Appendix C, when you know the 32-bit version.

  4. QEMU is the computer: it executes the BIOS, our bootloader and our kernel as a real machine would. gdb is the window into that computer: it stops the emulated CPU and shows registers and memory. Without QEMU you would need real hardware and a way to stop it; without gdb you would only see a black screen and would have to guess what went wrong, which is the situation chapter 16 teaches you to escape.

  5. gdb disables address-space randomization for the program it debugs, so that addresses are the same from one run to the next; it does this with a system call that Docker’s default filter forbids. Without the option, gdb warns at every run and the addresses of hosted programs in chapter 6 move each time, which makes the printed listings impossible to follow.

  6. First the flags (BOOKFLAGS forgotten or mistyped), then the tool versions against the table in this chapter, then your own typing. The book’s listings were produced by running the commands in the repository under the stated versions and the build checks them, so a difference is almost always on the reader’s side; and the first two checks take a minute while suspecting the book leads nowhere.

  7. The hosted programs of Part I are ELF files, and chapters 5 and 8 dissect that format because the kernel uses it too. gcc on Windows produces PE files and on macOS Mach-O files, so readelf and objdump output, section names, linker scripts and the whole of chapter 5 would differ. The container gives the same Linux tools on every host, which is why it is the recommended option.

E.2 Chapter 1: Domain documents

  1. The language is only the means of expressing a solution; what is missing is knowledge of the problem domain, which for an operating system is the hardware: how a CPU switches modes, how a device is programmed through its registers, what the firmware leaves in memory. A programmer who knows C perfectly but not the domain cannot even read the manuals that describe what the code must do, let alone write correct code.

  2. A requirement document describes the problem domain and the effects the software must produce in it; a specification states the rules between inputs and outputs, that is, the interface. You are reading a requirement when the text explains how the world behaves (how a serial line carries bits), and a specification when it tells you which value to put where to obtain a behavior (write this bit of this register to set the baud rate).

  3. A drop-down menu and a linked list are implementation choices, not facts about the problem domain. Implementation changes often, so the document must be updated each time or it starts lying; soon the effort is given up, the document is abandoned, and everybody concludes that documentation is useless. A document that sticks to the domain stays true while the implementation evolves.

  4. Hardware cannot be patched after it is manufactured, so a defect found by a customer is a recall, financially and reputationally devastating. The interface must therefore be fixed and described down to the last bit before the chip ships. Software is cheap to change, so its specification is often incomplete and written after the fact.

  5. “Functional Description” is the requirement document: it describes the domain, what the device does and why. “Register Description” is the specification: the list of registers, bits and the values they accept. When you need to program a device you go to the register description for the interface, but you understand what the bits mean only after reading the functional description; chapter 10 does exactly this with the serial port.

  6. What matters is that a complete specification exists at the end, because every program that uses the interface depends on it. It can be written before, during or after the implementation; writing it as parts of the interface become stable is often the most practical. What must not happen is for it to be left unwritten, or for implementation details to leak into it.

  7. Software-only domains (compilers, cryptography) produce software for other software, so the requirement and the specification tend to merge into one document. An operating system is both: it serves programs, but it also has to produce effects in hardware, which is an external domain with its own documents written by hardware engineers. To write one you must learn enough of that external domain to read those documents, which is what the three manuals listed at the end of the chapter are for.

E.3 Chapter 2: From hardware to software: Layers of abstraction

  1. Any Boolean function can be written using only NAND (or only NOR), a result Peirce proved; a set of functions with this property is called functionally complete. Since every digital device is a combination of Boolean functions, a device that can be built from gates can be built from NAND gates alone, which is why the 74HC00 and the Apollo Guidance Computer are enough to make computers.

  2. A transistor is an electrical component, a resistor whose value depends on a voltage; a logic gate is a function on bits, implemented with a few transistors wired in a fixed CMOS pattern. The designer thinks in gates, because the pattern that turns a gate into transistors is always the same and can be applied mechanically afterwards; thinking in transistors would mean translating every idea by hand and verifying it at the wrong level.

  3. It makes the translation from gates to circuits a mechanical rule that never needs to be reconsidered. Because it holds for every gate, a design expressed in gates can be “compiled” into a circuit automatically, and the designer can ignore the circuit entirely. This is the recurring pattern that lets the gate layer exist as a language of its own.

  4. The user needs the truth table because it is the interface, the complete description of what the device does for every input; the internals are an implementation that could change (a different transistor technology, a different layout) without changing the table. The user loses nothing by not knowing them, exactly as a C programmer loses nothing by not knowing how printf is written, as long as the table is correct and complete.

  5. The decoder reads the opcode bits of the instruction and selects the circuit that implements that instruction, routing the operand bits to it. Without it, a human would have to pick the right chip by hand and feed it signals, which is what the chapter does before introducing the opcode; the decoder is what turns a pile of devices into a machine that executes a program by itself.

  6. Each instruction has a small circuit that performs it; the bit pattern of the opcode is what selects that circuit, as a name selects a function. The “body” is the circuit itself, the gates that compute the result when the decoder activates them. Assembly then gives the name a mnemonic (hlt for 11110100), but the mapping from name to circuit is unchanged.

  7. Each layer takes the patterns that recur in the layer below, gives them a simpler notation, and drops the details that do not recur: logic gates name a fixed transistor pattern, machine instructions name fixed gate circuits, assembly names bit patterns, C names repeated blocks of assembly. Without a recurring pattern there is nothing to name, so no language and no abstraction could be built on top of that layer.

  8. All languages can express the same computations, but a high-level language hides the hardware to free the programmer from it, and the hiding carries extra code (memory management, run-time checks) that the programmer cannot control. An operating system has to control the hardware directly and predictably, so it needs a language that is a thin layer over the machine, where one can see what the hardware does for a given line; C is that language, and chapter 4 shows the mapping line by line.

E.4 Chapter 3: Computer Architecture

  1. The CPU is the only device that executes the programmer’s code; every other device only reacts to values written to its registers and ports. So the program controls a device by having the CPU read and write those registers, either through I/O instructions (in, out) or through memory addresses that the hardware maps onto the device. A device driver is nothing but the knowledge of which register to write, with what value, in what order.

  2. Both are hardware cells that software reads and writes, but a register stores a value, while a port delegates the value to another circuit and does not keep it. Writing to a register changes a state that can be read back; writing to a port triggers an operation, and reading the same port may return something unrelated, for instance the next byte received on a serial line.

  3. You can rely on everything the ISA defines: the instructions, the registers, the operating modes, the descriptor and page-table formats, the exceptions, because both vendors implement them identically and the manuals of both describe them. You cannot rely on anything that belongs to the organization or the hardware: how many cycles an instruction takes, the cache sizes, vendor-specific model-specific registers and CPUID leaves, and the virtualization extensions, which differ between VT-x and AMD-V.

  4. A microcontroller is a small computer built to control other hardware; a system-on-chip is a complete general-purpose computer on one chip, able to run an operating system and many applications. Both are programmable, so the difference is in resources and intent. An SoC in the place of a microcontroller would work, but it costs far more and most of its hardware would sit idle running a program that needs almost nothing.

  5. Memory is an array of bytes with addresses; the CPU fetches whatever bytes eip points to and executes them if they form valid instructions. Only the software decides which bytes are code and which are data. For a kernel this means that a wrong jump executes data as instructions, and a wrong write turns code into garbage; protection mechanisms, from segment types in chapter 9 to the no-execute bit, exist to let the kernel tell the CPU what it meant.

  6. Q35 is the most recent chipset QEMU emulates, and QEMU is our computer. What moved since 2007 is the location of functions: the memory controller and then the graphics went into the CPU package, the Northbridge disappeared and the Southbridge became the PCH. What stayed is the software view: PCI configuration space, the legacy devices at their traditional I/O ports, ACPI tables. The kernel of this book talks to that software view, so it is not tied to 2007.

  7. In real-address mode the register holds a plain number, multiplied by 16 and added to the offset, with no checks; any code reaches any byte of the 1 MB space. In protected mode it holds a selector, an index into a descriptor table, and the base, limit, type and privilege level come from the descriptor. The CPU checks every access against those fields and raises an exception when they are violated, which is where memory protection comes from, and why chapter 9 must build the GDT before anything else.

  8. Hardware interrupts arrive only when IF is set, so clearing it with cli guarantees that the code until the next sti runs without being interrupted on this CPU; nothing else is needed to make the section atomic. On a second processor IF is a separate bit in a separate eflags, so that processor keeps running and can enter the same critical section; a multiprocessor kernel needs a lock in memory as well, which chapter 13 discusses.

E.5 Chapter 4: x86 Assembly and C

  1. 01 c8. With bits 32, 32-bit operands are the default and the operand-size override prefix 66 is not needed; the opcode and the ModR/M byte do not change, since the register numbers are the same. In 32-bit mode the prefix selects the non-default size, 16 bits, so 66 01 c8 is add ax, cx: the same bytes are two different instructions depending on the mode of the processor, which is why a disassembler must be told which mode the code is meant for before it can decode a flat binary.

  2. A displacement takes part in an address computation, added to the base and index registers or used alone as an absolute address; an immediate is an operand value used as is. In mov DWORD PTR [ebp-0x8],0xabcdef, -0x8 is an 8-bit displacement and 0xabcdef a 32-bit immediate. The processor does not tell them apart by their values but by the opcode and the ModR/M byte, which fix how many displacement and immediate bytes follow and in which order.

  3. eb fe is a relative jump: the displacement, -2, is added to eip after the instruction, so the processor computes the target from where it is. Code made of relative jumps runs unchanged wherever it is loaded, which is what the linker wants for loops and local branches, and what a bootloader will like in chapter 7. An absolute address is needed when the target is not at a fixed distance: a jump through a register or memory (FF /4, as for a function pointer or a switch table) and a far jump (EA, FF /5) that also loads cs, which we will use to enter protected mode.

  4. word at +0, byte at +2, one padding byte, dword at +4: each variable is placed at the next address that is a multiple of its size, so where the padding lands depends on the declaration order. The reason is in section 4.1.1 of volume 1: a word or doubleword that crosses a 4-byte boundary costs two memory accesses instead of one, and some instructions on 16-byte operands raise #GP on unaligned data. Padding is cheaper than slow accesses; reordering the declarations is what a programmer can do to avoid it.

  5. idiv divides the 64-bit value edx:eax by its operand; cdq copies the sign bit of eax into every bit of edx, so that edx:eax is i sign-extended to 64 bits. Without it, edx holds whatever the previous instruction left: for a negative i, the dividend would be a huge positive number, the quotient wrong, and if the quotient did not fit in 32 bits the processor would raise a divide error, #DE, the same exception as a division by zero. One idiv computes both results: i / j keeps eax, i % j keeps edx.

  6. [ebp] holds the saved ebp of the caller and [ebp+4] the return address pushed by call; the arguments start at [ebp+8] and a third one would be at [ebp+0x10]. Pushing in reverse order puts the first argument at the lowest address, right above the return address, so that it is at [ebp+8] whatever the number of arguments: a function with a variable number of arguments, such as printf, finds its first argument without knowing how many follow.

  7. In this convention, cdecl, the caller removes the arguments because only the caller knows how many it pushed: printf("%d %d", a, b) and printf("hi") call the same function with different numbers of arguments. A callee that knows its argument count can clean up itself with ret imm16, the form of RET with an immediate in volume 2, which pops the return address and then adds the immediate to esp; this is the stdcall convention of the Windows API, and the add esp in the caller disappears.

  8. jle becomes jbe, “below or equal”, the unsigned condition, and nothing else changes. cmp is a subtraction that sets all the flags at once, ZF, SF, OF and CF; whether the comparison is signed or unsigned is decided by the conditional jump, which chooses which flags to test: jle tests ZF=1 or SF≠OF, jbe tests CF=1 or ZF=1 (appendix B of volume 1). The same cmp serves both; the compiler expresses the type only through the Jcc it emits.

E.6 Chapter 5: The Anatomy of a Program

  1. The two tables serve two readers. The operating system loads segments: a few contiguous ranges with permissions, which is all it needs to map the file into memory and jump to the entry point. The linker and the inspection tools work on sections: named pieces with a type, a symbol table and relocations. Without its section header table, hello still runs, since the loader never reads it; readelf -S, objdump -d by section, gdb symbols and linking against the file stop working.

  2. The variable would move to .data, which holds initialized data and is stored in the file: the file would grow by 4 bytes, and .comment and everything after it would shift by 4. In memory nothing changes: the .bss bytes were allocated and zeroed by the loader anyway; they are now read from the file instead.

  3. .dynsym and .dynstr are needed at run time, by the dynamic linker, to find puts and __libc_start_main in the shared C library, so they are allocated (A) and have an address in the memory image. .symtab, .strtab and .shstrtab are needed only by tools working on the file, the linker, readelf, gdb; they are not loaded, have no address, and strip can remove the first two without changing what the program does.

  4. A section header is a fixed-size entry so that entry i is at e_shoff + i * e_shentsize without parsing the previous ones; a string of variable length does not fit in such an entry. sh_name is an offset into .shstrtab, the string table whose index the ELF header gives in e_shstrndx; readelf reads that section and prints the null-terminated string at the offset, exactly as example 5.12 showed by hand.

  5. .rel.text listed the places in the code of hello.o whose address was unknown at compile time, such as the string passed to printf and the call itself. The linker resolved them, patched the bytes and dropped the section. .rel.dyn and .rel.plt are the relocations the linker could not resolve, because their targets live in the shared C library, which is loaded at an address known only at run time; the dynamic linker applies them when the program starts.

  6. The separation costs a few pages: each segment starts on its own 4 KB page, so the file has padding and the process a few more mappings. In return, no page is both writable and executable: a bug that writes through a wild pointer into the code, or that jumps into data, is reported immediately as a segmentation fault, where a single read-write-execute segment would let the program run on corrupted code and fail later, somewhere else. Our kernel, on the bare metal, will have no such protection until we set it up ourselves.

  7. The C runtime would be skipped: the stack would hold the raw argc, argv and environment that the kernel placed there, not the arguments main expects, stdio would not be initialized, constructors would not run. The dynamic linker still runs before the entry point, so printf can execute, but when main returns, its ret pops whatever is at the top of the stack, the value of argc, as a return address, and the program crashes with a segmentation fault; nobody calls exit, so even the text that printf buffered may never appear. _start exists to do all that before and after main.

  8. The loader does not copy sections one by one; it maps each LOAD segment as one block, a range of file bytes onto a range of addresses, page by page. Within the block, a byte keeps its distance from the start, so a section’s address is the segment address plus its offset within the segment. Since mapping works in whole pages, the file offset and the virtual address of a segment must be congruent modulo the page size, 4 KB; the GNU linker goes further and starts each segment on a fresh page in the file, hence the offsets 0x1000, 0x2000, so that no page holds bytes of two segments with different permissions.

E.7 Chapter 6: Runtime inspection and debug

  1. Line 3 produces no code, so gdb moves to the first line that does. break main and break 3 both land on 0x8049177 because gdb skips the prologue of the function, using the line table: the first row of main (line 4, the opening brace) is the prologue, and the second row is where ebp is set up and the arguments and locals can be read correctly. break *0x8049166 stops at the very first instruction, which is what you want to watch the prologue itself, or when debugging assembly without debug information.

  2. gdb keeps the list of addresses where it has written 0xCC, with the original byte of each. On a SIGTRAP, it checks whether eip - 1 is one of them: if so, it reports the breakpoint number, restores the byte and moves eip back by one so that the original instruction runs. Our own int 3 is not in the list, so gdb passes the signal on as it is and leaves eip where the processor put it, after the int 3, at 0x8049178, which is why it stops on the printf line without rewinding anything.

  3. From eax: the calling convention returns an int there, and finish runs until the return address and then reads the register. The debug information tells gdb the return type of add1000, from the DW_AT_type of its subprogram DIE: without it, gdb would not know whether to read al, eax, edx:eax for a 64-bit integer, or nothing for a void function, nor how to print the value. On a function without debug information, finish stops but prints no value.

  4. The program breaks first. Code compiled with -no-pie embeds absolute addresses, such as push 0x804a008 for the string of example 6.10 or mov eax,ds:0x804c00c in chapter 4, which would point to nothing at the new location. Even position-independent code would be undebuggable: gdb would plant breakpoints at the linked addresses, where there is no code, and map eip to no line, because the DWARF entries and the line table are expressed in linked addresses (section 6.4.3). Both must agree, which is why the load address of our kernel will be fixed in the linker script.

  5. It needs DW_AT_name and DW_AT_comp_dir of the compilation unit to build the path /tmp/hello.c, and the line table to map eip to line 5; it then opens the file and prints its fifth line. Without the file, everything that depends on the executable alone still works: breakpoints by line or function, next and step, backtraces, variables and disassemble. Only the source text is missing: gdb warns hello.c: No such file or directory, shows 7 in hello.c instead of the line, and x/i $eip becomes the way to see what runs.

  6. Memory is bytes without a type; the format letter decides how gdb reads them. x/i decodes whatever it is given as instructions, exactly as objdump -D did in chapter 4 on .data: the 78 byte of bf becomes a js that swallows the next byte, and the listing has no relation to the data. On data, use x/8xb or x/4xw to see the bytes, or print with the variable’s type when there is debug information, so that gdb decodes the bytes the way the program does.

  7. With int 3, the debugger would have to decode the current instruction to know its length and, for a jump, a call or a return, compute where execution goes next, and it would have to patch memory before every step. With the trap flag set, the processor raises interrupt 1 after each instruction, whatever it is, with no decoding and no patching. ni needs to let a call run to completion: when the current instruction is a call, it plants a temporary breakpoint at the return address, the instruction after the call, and continues; otherwise it would single-step through the whole of puts.

E.8 Chapter 7: Bootloader

  1. The directive times 510 - ($-$$) db 0 pads the file with zeros up to offset 510, whatever the size of the code, so dw 0xAA55 always lands on the last two bytes of the sector (as 55 aa, little-endian). With 511 bytes of code, the expression 510 - ($-$$) is negative, and nasm refuses to assemble the file: the error is caught at build time rather than as a “not a bootable disk” message at boot time.

  2. nasm -f bin computes every label from offset 0 of the file, so mov si, msg is assembled as mov si, 2; at run time the string is at 0x7c02, and 0000:0002 is the interrupt vector table, which starts with a zero byte where a print loop stops. org 0x7c00 makes the labels agree with where the BIOS puts the sector. The chapter’s bootloaders never read msg, and jmp boot is a relative jump, so no absolute address is ever used, and they get away with it.

  3. The far jump loaded CS = 0x50 and IP = 0; the CPU is executing at physical 0x50 × 16 + 0 = 0x500. info registers shows the segment and offset, and x/10xb $cs*16+$eip shows the bytes of sample. gdb reports addresses from EIP alone, so it says 0x0 and does not recognize the breakpoint it set at 0x500; the hook-stop prints [cs:eip] at every stop so that the real location is always visible.

  4. b8 50 00 is mov ax, 50h and 8e c0 is mov es, ax. QEMU describes the CPU to gdb as a 32-bit target, and gdb decodes 32-bit instructions, where b8 takes a 4-byte immediate and swallows the next instruction. The bytes in memory are right, so the check is to read them with x/8xb and compare with hd or the nasm -l listing; breakpoints, stepping and the registers are unaffected.

  5. bootloader and os are directories, and a directory is a file that exists and is always up to date, so without .PHONY make would say “nothing to be done” and never run the child Makefiles. bootdisk is also phony, and declared so, but no file of that name exists, so the problem would not show up until someone created one.

  6. It is impossible to boot or test a stale image: whatever changed in the sources is rebuilt before QEMU starts. The cost is a run of the child Makefiles at every make qemu, which is negligible here, since an unchanged object file is not rebuilt.

  7. The stack is wherever the BIOS left SS:SP when it jumped to the bootloader, 0000:6f08 under SeaBIOS; the BIOS chose it for its own use, and the interrupt frame of int 0x13 is pushed there. Another BIOS may leave the stack anywhere, including on top of the buffer at 0x500 or of the bootloader, so a bootloader sets SS:SP itself, to a region it knows to be free, before the first call, push or int.

E.9 Chapter 8: Linking and loading on bare metal

  1. call takes a displacement relative to the end of the instruction, and the assembler leaves -4 as a placeholder because the relocation formula, S+A−PS + A - P, is applied to the start of the 4-byte field, not to the end of the instruction: the -4 is the addend that corrects for the field’s own size. After linking, S=0x8049172S = 0x8049172, A=−4A = -4 and P=0x8049167P = 0x8049167 give 7, which is foo minus the address of the next instruction, 0x804916b.

  2. A PT_PHDR entry tells the program where its own program header table is in memory; the dynamic linker uses it to find PT_DYNAMIC. The table is in memory only if a PT_LOAD segment covers it, so the ELF specification allows PT_PHDR only in that case, and ld now enforces it. We keep it because it documents the layout and because putting FILEHDR PHDRS on the first LOAD segment is what makes the bootloader’s scheme work: the header is loaded at the start of the file, so e_entry is at 0x18 from the load address.

  3. They are the ELF magic number, 0x7f followed by ELF: the bootloader copies the whole file to 0x500, so the header is at 0x500 and .text, at file offset 0x500, is at 0xa00. The linker computed e_entry and every address in the debug information assuming the segment starts at address 0, where its file offset is, so both were 0x500 short of reality.

  4. -nmagic turns off page alignment: without it, ld gives every LOAD segment an alignment of 0x1000 and places sections at file offsets congruent to their addresses modulo 0x1000, so .text at 0x600 goes to file offset 0x600. With it, the alignment comes from the sections, and ALIGN(0x100) puts .text at file offset 0x100 and address 0x600, so the segment starts at address 0x500, file offset 0, exactly where the bootloader puts it.

  5. Debug information, the symbol table and the section headers: everything gdb needs and the CPU does not. Only LOAD segments must be in memory, and ours is 0x106 bytes, so 17 sectors cover it many times over. gdb reads the symbols and the DWARF data from build/os/os on the host, through symbol-file, never from the memory of the virtual machine.

  6. Everything except the bytes of the .text section: symbols, line numbers, section headers. The flat binary is what the BIOS can load and the ELF is what gdb can read, so both are kept. symbol-file replaces the current symbol table; add-symbol-file adds a second one, which is what you need to see the bootloader and the kernel in one session.

  7. Only by luck: 55 89 e5 90 5d c3 decode in 16-bit mode as push bp, mov bp, sp, nop, pop bp, ret, the same operations on 16-bit registers. An instruction with a 32-bit immediate or displacement, such as mov DWORD PTR [ebp-4], 5, decodes differently in 16-bit mode, takes the wrong number of bytes and derails everything after it. Chapter 9 switches to protected mode before any C code runs.

  8. -ffreestanding tells gcc that there is no C library: it stops treating main as special and stops assuming that functions named memcpy or printf have their standard meaning, so it will not replace a loop by a call to a library that does not exist. -nostdlib only changes what is linked. Without -fno-stack-protector, a function with a local array reads a canary through gs, which points to nothing on bare metal, and calls __stack_chk_fail, which is not there: a fault or an undefined reference at link time.

E.10 Chapter 9: Protected mode and x86 descriptors

  1. The bootloader’s table lives inside the boot sector at 0x7C00, memory that nothing protects and that the kernel is free to reuse (the bootloader itself writes the boot information block of chapter 12 at 0x500, a few hundred bytes below it). The processor reads descriptors from the table whenever a segment register is loaded, so a GDT that gets overwritten is a crash waiting to happen. Building the table in the kernel’s own .bss also lets later chapters add entries (user segments, the TSS) without touching the 512-byte bootloader.
  2. The processor is in protected mode from the moment PE is set, but CS still holds the real-mode value 0 and its hidden part still describes a 16-bit segment. A near jump does not reload CS, so the code at protected_mode is decoded as 16-bit code, and the bits 32 instructions there are misread (mov ax, 0x10 becomes a different instruction followed by garbage). Only a far jump, far call or iret loads CS with a selector and brings in the 32-bit descriptor.
  3. The limit field is 20 bits wide, so it cannot hold 0xFFFFFFFF. With G = 1 the limit is counted in 4 KiB units: 0xFFFFF means 0x100000 pages, which is 0x100000000 bytes, all 4 GiB. G = 0 would make the same value mean 1 MiB. The granularity bit is what makes a 20-bit field cover a 32-bit space, and QEMU’s monitor info registers shows the scaled result, ffffffff.
  4. Each segment register caches its descriptor in a hidden part at the moment it is loaded, and every address computation uses that cache (Intel SDM Volume 3A, section 3.4.3). lgdt only changes the GDTR; the hidden parts still hold the descriptors loaded from the bootloader’s table, which are valid. Without the ljmp, CS would keep the old descriptor, which happens to be identical, so nothing would fail now; but the kernel would depend on the boot sector’s table for CS until the first interrupt or far transfer, and any later change to the kernel GDT’s code entry would not take effect.
  5. 0x13 is 0x10 | 3: index 2 of the GDT with RPL = 3, the data segment requested at privilege level 3. The processor compares max(CPL, RPL) with the descriptor’s DPL; with RPL = 3 and DPL = 0 the check fails and the mov raises #GP. Chapter 13 adds descriptors with DPL = 3 precisely so that such selectors can be loaded.
  6. The gate is outside the processor, in the chipset, and the bootloader is the natural place to open it, once, before anyone depends on memory above 1 MiB. With the gate closed, bit 20 of every address is forced to zero, so a write to 0x100000 lands at 0x000000. Chapter 12 maps and uses all of the 128 MiB; a kernel with A20 closed would see its page tables above 1 MiB silently alias the first megabyte, corrupt the IVT or the boot information block, and fail in ways that have nothing to do with paging.
  7. A hosted program is loaded by the operating system’s ELF loader, which reads the program headers and, for each LOAD segment, copies FileSiz bytes and then zero-fills MemSiz - FileSiz bytes; .bss is the difference. Our bootloader copies the file verbatim and never reads a program header, so the memory of .bss receives whatever followed .data in the file, here .debug_aranges. The loader’s missing job is exactly that zero-fill, which _start performs from the two linker symbols.
  8. Without packed, gcc may insert padding so that each member is naturally aligned: base_middle and access are bytes and need none, but a compiler is free to pad after them, and on some targets the structure grows. With this exact member list on i386 the layout happens to stay 8 bytes, which is why the bug is dangerous: the code works until a member is reordered or the target changes. The processor reads 8 raw bytes per entry, so any padding shifts every later field and the selectors describe garbage, usually a #GP at the first segment load.

E.11 Chapter 10: Talking to devices: the serial port and the VGA text console

  1. The UART shifts bits out at 38400 baud, about 260 microseconds per character, while the CPU can hand it a byte every few nanoseconds. Writing without waiting overwrites the transmit holding register before the previous byte has moved to the shift register, so most characters are lost and the output is a few scattered letters. QEMU hides the bug because its emulated UART delivers a byte to the host file or terminal the instant it is written; the register is always empty, so the broken driver works perfectly until it meets a real chip.

  2. The hang happens before any serial byte is written, so the log would be empty: no greeting, nothing. The script would time out, print an empty serial output and exit 1, and the CI would report a failure without any clue about where the kernel stopped. With the serial port first, the greeting is already in the log when the VGA driver hangs, and the log shows exactly how far the kernel got. That is why the order in putc is a design decision and not a detail.

  3. The keyword states a fact about the hardware, not about the compiler flags: the memory and the ports have side effects the compiler cannot see, and a driver that only works because optimization happens to be off is wrong, not lucky. At -O2, without volatile on inb, the compiler would see a loop whose body has no visible effect and whose condition reads the same “pure” function with the same argument; it would read the port once and spin forever on the result, or drop the loop entirely, and serial_putc would either hang or write without waiting.

  4. The bytes of EBX would be stored high byte first, so the array would hold uneG Ieni letn: the same twelve characters in groups of four, each group reversed. The code is correct only because the register-to-memory byte order of the machine matches the order in which the processor packed the characters, which is itself little-endian by design. It does not matter in practice, since cpuid exists only on x86, which is little-endian, but it shows that “copy the register to memory” is a statement about the architecture, not about C.

  5. vga_putc advances the column after storing a character and wraps to the next row as soon as the column reaches 80, so a line of exactly 80 characters moves the cursor to the next row by itself, and the \n that follows moves it down once more. A terminal defers the wrap: after the 80th character the cursor stays in a pending state at the end of the row, and a newline simply ends the line, while a further printable character triggers the wrap. The fix is to wrap lazily too: do not advance the row when the column reaches 80, but keep a flag, and wrap at the start of the next printable character, clearing the flag on \n.

  6. The processor raises exception 6, invalid opcode, and looks up vector 6 in the IDT. The kernel has not loaded an IDT, so the IDTR still holds whatever the BIOS left, a real-mode interrupt vector table that makes no sense as protected-mode gates; the descriptor the CPU fetches is invalid, which raises a general protection fault, whose gate is equally invalid, which raises a double fault, and the third failure in the sequence is a triple fault: the processor resets. The machine reboots a few microseconds after the cpuid, and the serial log shows the greeting and nothing else, over and over.

  7. The C standard says that arguments passed through ... undergo the default argument promotions, so the caller pushed an int, never a char; va_arg(args, char) asks for a type that cannot have been passed, which is undefined behavior, and gcc warns about it even though on this machine both occupy four bytes. On a 64-bit build the arguments no longer sit on the stack at all: the System V AMD64 convention passes the first six integer arguments in registers, and va_list becomes a structure that remembers how many register arguments have been consumed before switching to the stack. va_arg still works, because the compiler implements it, but the comment in printf.c about “the next double words on the stack” would no longer be true, and a %p would fetch eight bytes, so print_pointer would need sixteen digits.

E.12 Chapter 11: Interrupts

  1. The BIOS leaves the master 8259A delivering IRQ 0 to 7 on vectors 8 to 15, which in protected mode are the exceptions #DF, #TS, #NP, #SS, #GP and #PF. A timer tick would arrive on vector 8 and the kernel would print “Exception 8: Double Fault” and halt, a hundred times a second if it did not. pic_init moves the lines to vectors 32 to 47 with ICW2 before sti enables delivery.
  2. The two macros exist so that the stack has the same shape for every vector: the processor pushes an error code for #DF and the stub must not add a second one. If isr8 pushed a fake 0, struct registers would be shifted by one word for that vector (eip would read the real error code, cs the saved EIP), and add esp, 8 before iret would leave one word too many, so iret would pop the saved EIP as EFLAGS and jump to the saved CS value. The symptom is a triple fault from inside the handler of a double fault.
  3. #BP is a trap: the saved EIP points after the int3, so returning continues the program. #DE is a fault: the saved EIP points at the idiv itself, so returning re-executes it and raises #DE again. Without a fix (the kernel has none for a division by zero), isr_dispatch would print the register dump forever; halting is the honest reaction.
  4. In 32-bit code push ds pushes 4 bytes, the segment selector zero-extended, as the PUSH description in Volume 2B says. If the four fields were uint16_t the structure would be 8 bytes shorter than the stack frame, and every field after ds would be read 8 bytes too low in memory, two words early: regs->eip would hold the vector number, regs->err_code the saved EAX, and regs->int_no the saved ECX, so the dispatcher would index its handler table with whatever ECX happened to contain.
  5. For any IRQ: the end-of-interrupt command was not sent to the PIC, so the line stays “in service” and the chip never raises it again; isr_dispatch sends it after the handler for exactly this reason, so in our code the cause would be a handler that never returns, or an EOI sent to the master only for a slave line. For the keyboard: the handler did not read port 0x60, and the 8042 does not raise IRQ 1 again until its output buffer has been read.
  6. Both are 8-byte gates with the same fields; the only difference is that an interrupt gate clears IF when the processor enters the handler and a trap gate leaves it alone (Intel SDM Volume 3A, section 7.12.1.3). With interrupts off, a handler cannot be interrupted by another hardware interrupt, so it can touch shared data (the tick counter, the run queue of chapter 13) without a race, and nested timer ticks cannot pile up on the stack. Exercise 11.7 shows what the trap-gate version does under load.
  7. sti enables interrupts after the instruction that follows it completes (Volume 2B, “STI”), so the pending tick could only be delivered after the mov. The delay exists for sti; hlt: if an interrupt could be taken between the two, the handler would run and hlt would then stop the processor with nothing to wake it up; with the delay, hlt is reached with interrupts enabled and the pending interrupt wakes it.
  8. esp_dummy is the value of ESP before pusha, that is, the address of the vector number. pusha then pushed 8 words and the stub 4 segment registers, 12 words or 48 bytes, which is why regs, the stack pointer after those pushes, is 48 bytes lower. Above the vector number sit the error code, EIP, CS and EFLAGS, 5 words in all, so esp_dummy + 20 is where ESP was when the processor started pushing, the interrupted code’s stack pointer; the QEMU log confirmed it with SP=0010:0008ffe0.

E.13 Chapter 12: Memory management

  1. There is no instruction that reports the amount of RAM: the memory controller in the chipset knows, and only the firmware has talked to it, so the map is a BIOS service, INT 15h, EAX=E820h. It also describes holes and reserved ranges (VGA memory, the BIOS ROM, ACPI tables, PCI apertures) that a plain size could not express. BIOS services are 16-bit real-mode code reached through the real-mode interrupt vector table, which the kernel replaces with its IDT, so the bootloader must collect the map before CR0.PE is set and leave it at 0x500 for the kernel.
  2. The usable ranges are the only ones the BIOS guarantees; everything it does not mention (from 0xa0000 to 0xf0000 on QEMU, where the VGA buffer and option ROMs live) is simply absent from the map. Starting from “all free” and subtracting the reserved entries would leave those unlisted ranges free, and the allocator would eventually hand out the VGA frame buffer or a ROM as a page table. Starting from “all used” makes silence mean “not RAM”.
  3. A frame that is only partly covered by a reserved range must stay reserved, so a used range rounds outwards (start down, end up); a frame that is only partly covered by a usable range is not entirely RAM, so a free range rounds inwards (start up, end down). On QEMU the first usable entry ends at 0x9fc00, in the middle of frame 0x9f, which holds the Extended BIOS Data Area; the inward rounding keeps that frame out of the free pool even before the first-megabyte line marks it used.
  4. The page the next instruction lives in must be mapped at its own address, with the directory entry and the table entry both present, and CR3 must hold the physical address of a directory whose entries point to tables that are themselves correct; the identity map guarantees all of it. If the mapping is wrong, the fetch raises #PF, the processor tries to read the IDT and the handler through the same broken tables, and the result is a double then a triple fault. In the -d int log this is a v=0e entry whose CR2 equals the address of the instruction after mov cr0, followed by v=08 and a reset.
  5. The identity map reserves the virtual range from 0 to pmm_memory_top(), 128 MiB, for the frames themselves: every physical address must stay a valid pointer. Mapping the heap at 0x00400000 would replace the identity mapping of the frames at 4 MiB with heap pages, so the allocator could still hand out those frames but the kernel could no longer write to them through their physical address, and the page tables that live there would become unreachable. 0xC0000000 is above every address the machine has RAM at.
  6. The processor caches translations in the TLB and only walks the tables on a miss, so a store into a page table does not change what the TLB answers for a page it has already translated. invlpg drops the cached entry for that page. Without it, after paging_unmap the TLB could still translate 0x30000000 to frame 0x11f000 and a read through the window would succeed instead of faulting; worse, after the frame is reused, the stale entry would read someone else’s data.
  7. No. Bit 0 of the error code only says whether the fault was caused by a non-present entry or by a protection violation; a missing directory entry and a missing table entry both clear it, and CR2 only gives the address. The handler has to walk the tables itself: read kernel_directory[cr2 >> 22], and if that entry is present, read the page-table entry it points to, which is exercise 12.5.
  8. CR3 and a directory entry hold only bits 31:12 of the address they point to, so a directory at 0x14010 would be used as if it were at 0x14000; the arrays must start on a 4 KiB boundary. aligned(PAGE_SIZE) on the two arrays forces .bss to be 4 KiB aligned, and ld raises the alignment of the segment that contains it to the largest alignment of its sections, so the LOAD segment reports 0x1000 where chapter 11 had 0x100; readelf -S shows the .bss alignment of 4096 and nm the directory at 0x14000.

E.14 Chapter 13: Processes: context switching, scheduling, user mode and system calls

  1. It would not survive, and nobody is at fault, because the compiler never does that: EAX, ECX and EDX are caller-saved in the System V i386 ABI, so the code generated for the thread saves anything it still needs before the call to yield() and reloads it afterwards. context_switch is an ordinary function from the caller’s point of view, and it saves exactly the registers the ABI says a function must preserve. If a value in EAX were lost across yield(), the compiler would have broken the calling convention; the scheduler only honors it.

  2. ESP0 is consulted only on a transfer from ring 3 to ring 0, and a task is in ring 3 only when its kernel stack is empty: it was entered through a gate, the handler ran, and iret emptied it again on the way out. A task in the middle of a system call is in ring 0, and an interrupt in ring 0 stays on the current stack without reading the TSS. For a kernel thread the field is never used at all, since it is never in ring 3; the scenario cannot arise, which is why “the top of the stack” is always right. If the processor did switch to ESP0 on a ring 0 interrupt, it would overwrite the frames below the stack top, and that is precisely why the architecture does not.

  3. No. With a trap gate IF stays set on entry, and once the EOI is sent the PIC is free to raise IRQ 0 again. The next tick would interrupt the handler on top of itself, push a second frame on the same kernel stack and call schedule() from inside schedule(); nested often enough, the stack overflows, and the run queue is modified by two activations of the same code. The early EOI is safe for us only because the interrupt gate keeps IF clear until iret, so that a tick arriving during the handler is held by the PIC and delivered after.

  4. Because the processor defines the current privilege level as the RPL of the selector in CS, and it refuses to load CS with a selector whose descriptor has a lower DPL than the requested privilege; the only way to execute with CPL = 3 is to load a descriptor that has DPL = 3. The segments must exist to make the processor believe the code is unprivileged, not to confine it. What confines it is bit 2 of the page-table entry for the page of ticks, the U/S bit, together with the same bit in the directory entry above it; both must be set for a ring 3 access, and the kernel’s pages have neither.

  5. task_wake_sleepers would see a sleeping task whose wake_at is in the future and leave it; schedule() from the timer would skip it, switch to another task, and come back to it only when it is woken, whereupon it would resume inside task_sleep and call schedule() a second time, which simply gives the processor away for one more turn. Harmless here, but only because the data is consistent at that moment; the line that makes the question moot is irq_save() at the top of task_sleep, which keeps the tick out until the switch has happened.

  6. ticks - wake_at is -0x90000000 as an unsigned subtraction, and read as a signed 32-bit number that is +0x70000000, non-negative: the task is woken on the next tick, after one hundredth of a second instead of 280 days. The idiom treats a difference with the top bit set as “in the past”, so the longest future it can express is 0x7FFFFFFF ticks, about 248 days at 100 Hz, and a sleep longer than that is a sleep of one tick. A kernel that needs longer timeouts uses a 64-bit tick count or an absolute deadline with a different comparison.

  7. An exception, or a non-maskable interrupt: cli masks only the interrupts that come through the PIC, so a page fault inside a critical section still runs the fault handler on top of it. task_sleep may switch because the data it shares with the timer handler, the run queue and its own state, is consistent at the moment schedule() is called; the critical section is closed in the sense that matters, even though IF is still clear. kmalloc in the middle of splitting a block has a list in an inconsistent state, and a switch there would let another task walk it. The rule is not “do not switch with interrupts off” but “do not switch with shared data half-modified”.

  8. Take the user stack page, ptr = 0x7FFFF000, and len = 0xFFFFF010: the sum wraps to 0x7FFFE010, which is below the split, so test 2 passes, and since end is now below ptr the page loop does not execute at all, so test 3 passes too. The handler would then loop putc four billion times, reading the stack page, then the unmapped page at 0x80000000, where the kernel takes a page fault in ring 0 and panics. The wrap test exists because a single overflowed addition turns two correct tests into a lie, and WRITE_MAX exists so that even a correct check is not asked to walk a million pages.

E.15 Chapter 14: File system

  1. While BSY is set the drive is updating its registers, and the specification says the other status bits are not valid then; a stale DRQ left over from the previous sector could be read as “data ready” for the next one. A loop that tested DRQ before BSY would then pull 256 words of the previous sector, or of nothing, from the data register and desynchronize the transfer. Clearing of BSY is the only edge that means the drive has finished thinking.
  2. The drive acknowledges a write when the data is in its cache, so without FLUSH CACHE a power failure right after ata_write_sectors returns can lose the sector, and worse, the drive may commit cached sectors in a different order from the one the file system chose, which breaks the “mark used before pointing to it” reasoning. Nothing is lost while the power stays on. Linux flushes only at barriers its file systems request, typically when a journal transaction commits, because a flush after every sector costs more than the writes themselves.
  3. The block size is in the superblock, so it is unknown until the superblock has been read; the only thing known in advance is that the superblock starts at byte 1024, which is sectors 2 and 3 whatever the block size. With 4 KiB blocks the two raw sectors are the same, but the superblock now sits in the middle of block 0, so s_first_data_block is 0 and the group descriptor table is block 1; s_first_data_block + 1 expresses both cases.
  4. file_block returns 0, ext2_read_file fills that kilobyte with zeros, and the file reads as if it held zeros there: a hole in a sparse file. Block 0 holds the boot record, which ext2 never allocates to a file, so 0 can mean “no block” without ambiguity. The first user programs had a block of zero padding between the headers and the code, mke2fs -d stored it as a hole, and the loader, which did not check for 0, read block 0 of the file system and loaded the boot record’s bytes in place of the code.
  5. A crash after the bitmap write leaves a bit set with a count that is one too high, or, counts first, a count that is one too low; in both cases pass 5 of e2fsck recomputes the counts from the bitmaps and prints “Free inodes count wrong”. No data is involved, so neither order is worse. Writing the inode and the directory entry before the bitmap is different: the inode is in use, reachable by name, yet marked free, and the next alloc_inode would hand the same number to another file, so two names would share one inode.
  6. Every entry’s rest would be smaller than the 16 bytes needed, the loop would end and dir_add_entry would print “directory is full” and return -1, leaving an inode allocated and written but unreachable (which e2fsck would move to lost+found). Growing the directory needs alloc_block for a new block, a single entry in it whose rec_len is the whole block, and then the directory’s inode written back with the block number in i_block and i_size increased by one block.
  7. Pass 1 checks that i_size covers the blocks the inode has, and would report that the inode’s size is wrong, “should be” the size of the allocated blocks, and also that i_blocks is wrong if the kernel had not counted them. lost+found is fine because mke2fs set its i_size to 12288, the size of the 12 blocks it preallocated so that e2fsck can link orphans into it without allocating anything on a damaged disk.
  8. An inode number identifies the file but not the position in it, so two tasks reading the same file would share one offset, or the kernel would have nowhere to store one; the same task opening the file twice would be indistinguishable. Real kernels keep an open file description (inode, current offset, mode) per open, a table of pointers to them per process indexed by the descriptor, and the inode shared underneath; fork copies the table, which is why a parent and child share an offset after fork but not after two separate opens.

E.16 Chapter 15: Address spaces: fork and exec

  1. The parent would keep writing to the shared frame at full speed, and the child, which reads the same frame, would see every write the parent makes after the fork: a local variable the parent changes would change under the child’s feet. The child would notice, not the parent, since the parent’s view is always the “right” one; and the child’s first write would still copy the page, taking a snapshot of whatever the parent had done by then, so the bug would depend on timing. Copy-on-write only works if both sides fault on their first write, which is why the parent’s entries are modified in place.

  2. The parent’s TLB still holds the writable translation of the stack page it was using just before the system call, so its first stores after fork go straight into the shared frame without a fault, and the child, when it runs, reads them: it sees the parent’s stored return value, or a few bytes of whatever the parent wrote next, in its own stack. The entry stays in the TLB until something evicts it, a mov cr3 on a switch to another task or just other translations crowding it out, after which the next write faults and is copied correctly; so a busy machine with frequent switches would show the bug rarely and an idle one more often. This is exactly the class of bug section 5.10.4.2 of the SDM exists to prevent.

  3. Nothing observable: with 0x202 the few instructions between the popfd in context_switch and the iret in isr_return would run with interrupts enabled, and a tick landing there would run the timer handler on the child’s kernel stack below the frame, perhaps switch away and come back, and return to the same iret. In chapter 13 the prepared frame’s EFLAGS was the value the new task would run with, because the start function went straight into the thread or into enter_user_mode, which pushes its own 0x202; nothing else would ever have set IF. Here the iret loads the program’s EFLAGS from the copied frame, so the value on the prepared stack only matters for a dozen ring 0 instructions, and 0x2 makes them look like the end of any other handler.

  4. After paging_switch_directory the address in EBX is translated through the new directory: it points into hello’s freshly loaded page, or into its new stack, or into nothing, so elf_load would be asked for a garbage path (it was called earlier, so the real damage comes later) and task_set_name would copy bytes of the new program, or fault in ring 0 and panic. If the name were kept as a pointer, as in chapter 14, it would point into the old program’s page, whose frame was returned to the allocator and may already be somebody else’s page table. The kernel must own every byte of a user argument it uses after the address space changes.

  5. The store would succeed silently, in ring 0, into the frame both processes map: the child would find the parent’s status value in its own copy of that variable, at the same virtual address, without ever having written it. The parent would not notice anything, because its page-table entry stays read-only and COW, so its next user-mode write would still fault and copy the page, and the copy would contain the status it expects. The bug is visible only in the child, and only if the child reads that variable before taking its own copy of the page, which makes it the kind of bug that appears once a month.

  6. task_exit runs on the task’s own kernel stack and cannot free it while standing on it, which is chapter 13’s reason for the split between exit and reap; it could free the page directory and keep the rest, which is what Linux does (a zombie owns almost nothing). Copying the status into the parent at exit time would need a slot per child in the parent, or a list, since a parent may have several children that exit before it calls wait; the zombie’s own structure is that list, with no new allocation and no limit on the number of children, and the parent’s wait is the one place that reads it. Unix made the same choice, which is why zombies exist.

  7. The count stops at 255 while 256 entries point at the frame. As the processes exit one by one, each pmm_free_frame decrements: after the first exit the count says 254 while 255 users remain, and after 255 exits it reaches 0 and the frame is returned to the bitmap while one process still maps it. That process keeps reading and writing a frame the allocator will hand out again as somebody else’s page table or page, which is a use-after-free, worse than a leak. The cheapest correct fix is to refuse: make pmm_frame_ref report failure at 255 and have fork return -1, as Unix returns EAGAIN when it cannot create a process; the next cheapest is a 16-bit count.

  8. P = 0 outside every lazy region: a program dereferencing a null or wild pointer, say *(int *)0x10000, faults with error code 0x4 or 0x6; demand_fault finds no region and declines, and the handler prints and kills the task, as before. P = 1 inside a lazy region: a page of .bss that was demanded, then shared read-only by a fork, then written; the error code is 0x7, the page is present and marked COW, and cow_fault copies it, exactly as for a stack page. Trying demand_fault first would be wrong in that case: the address is in a lazy region, so it would allocate a fresh zeroed frame and map it over the shared one, and the process would lose the contents of its page. Bit 0 is what says whether the page exists, and only a page that does not exist may be made out of zeros.

  9. The frame goes into the child’s directory only: demand_fault maps it in current_task->page_directory, and the child is current. The parent’s entry stays zero, and when the parent reads the address later it takes its own demand fault and gets its own zeroed frame, so it reads zeros, which is exactly what it would have read had fork copied the page eagerly (a page nobody had touched held zeros). Two zeroed frames describe the same contents as one shared frame; nothing is observable, and nothing had to be copied or counted. Without the two fields, the child’s bss_start and bss_end would be zero, so a touch of any untouched .bss page would be reported as page not present and kill the child, while its stack would keep growing on demand, because user_stack_top is copied separately and the stack window is derived from it; the pages the parent had already demanded before the fork would work too, since they are present and shared copy-on-write.

E.17 Chapter 16: When it does not boot: debugging a kernel

  1. It tells you that the kernel was still running C code with a working serial driver when it stopped, and that the stop came between two calls of serial_putc, so the last print is a lower bound on how far the boot got. It does not tell you why: a triple fault, a hang in a polling loop and a cli; hlt all cut the output the same way. The next step is the QEMU log for a fault or an attached gdb for a hang.

  2. sti enables interrupts after the instruction that follows it completes, so that sti; hlt cannot be interrupted between the two and leave the processor halted with nothing to wake it (Intel SDM Volume 2B, “STI”). The EIP recorded for the first interrupt is therefore the second instruction after sti, not sti itself; if you expect the fault “at sti” you will look at the wrong line, and the real lesson of such an EIP is “the first interrupt that was ever allowed through is the one that failed”.

  3. Not necessarily. A #GP with error code 0 is not about a descriptor at all: it is what the processor raises for a privileged instruction in ring 3 (cli, hlt, in), for an access through a segment outside its limit, or for a malformed iret frame. If the gate’s DPL were the problem the error code would be 0x402, naming gate 0x80. A v=80 followed by v=0d e=0000 means the system call was taken and the handler, or the return from it, executed something ring 3 may not; look at the saved EIP.

  4. 0x5 is P = 1, W/R = 0, U/S = 1: a read, from ring 3, of a page that is present but reserved to the kernel. A user program read kernel memory, which is exactly what the page tables are for; the kernel should kill that task and continue, as chapter 13’s handler does, since the kernel’s own state is intact. Error code 0x0 at the same address would be a kernel-mode read of a page that is not mapped, which is impossible unless the kernel’s own tables or pointers are wrong, and that one is a kernel bug worth a panic.

  5. gdb unwinds by following saved frame pointers and return addresses, and the assembly stubs have no frame: isr_common pushes registers without push ebp; mov ebp, esp, so the unwinder takes the wrong words for the caller’s EBP and return address and prints garbage or stops. Across an iret or a context_switch the stack simply changes, which no unwinder can follow. A user program built with -fomit-frame-pointer has the same problem in every function: EBP is an ordinary register and the chain does not exist, which is why debug builds keep the frame pointer and why release kernels carry unwind tables instead.

  6. A reset requested through hardware rather than through a fault: a write to port 0x92 with bit 0 set (the fast-A20 port, which also resets the processor), the 0xFE command to the keyboard controller, the reset register at 0xCF9 on q35, or a jump to the reset vector at 0xFFFFFFF0. None of these is an exception, so -d int is silent. Confirm with -d cpu_reset, whose CPU Reset block records the registers at the moment of the reset, including the EIP of the out instruction, and with -no-reboot, which intercepts every reset request and stops the machine instead.

  7. The driver could check the error register and the final status after the transfer, and it does check ERR and DF; but a sector full of zeros is a perfectly valid sector, and the drive returned exactly the sector it was asked for. Validating the meaning of the data (a magic number, a mode field, a block number that must not be 0) belongs to the layer that knows the format, which is ext2.c, and that is where the fix goes. Each layer can only check that its own input made sense.

  8. Because compiling proves nothing about a kernel: there is no runtime, no type checker for selectors or page-table bits, and the bugs of this chapter all compile without a warning. make test boots the image and checks the behavior, so a version that passes it is a known-good point, and a regression is then a search between two such points, which git bisect does in log2(n) steps. Committing every compiling version gives a history in which “good” is undefined.