Appendix E Answers to the “Check your understanding” questions
Each chapter ends with a handful of questions. They are not an exam: they are the questions we would ask you after you have read the chapter, to find out together whether the important points went through. Try to answer in writing before looking here; the explanations are short on purpose, and the chapter has the long version. The sections below follow the chapters in order, one per chapter from chapter 0 to chapter 16, and the answers are numbered as the questions are.
E.1 Chapter 0: Setting up the development environment
A position-independent executable embeds no absolute address, because the operating system moves it to a random address at every run. The compiler must therefore generate code that computes “where am I” at run time, which is the
call __x86.get_pc_thunk.axfollowed byaddandleain the first listing. On bare metal nothing relocates our code, and the whole point of Part III is to choose and control fixed addresses, so this machinery only hides what we want to see.The thunk returns the address of the instruction after the
call, which is the only way 32-bit x86 code can read its owneip. The second program knows its address at link time,0x8049166, so the address of the string is a constant,0x804a008, pushed directly. Same program, but the second one is readable by a beginner.The kernel of Part III runs in 32-bit protected mode, and the registers, calling convention and instruction encodings you learn in Part I are exactly those you will meet in
gdbwhile debugging that kernel. Teaching 64-bit in Part I would mean learning two of everything. The 64-bit differences are covered once, in chapter 17 and Appendix C, when you know the 32-bit version.QEMU is the computer: it executes the BIOS, our bootloader and our kernel as a real machine would. gdb is the window into that computer: it stops the emulated CPU and shows registers and memory. Without QEMU you would need real hardware and a way to stop it; without gdb you would only see a black screen and would have to guess what went wrong, which is the situation chapter 16 teaches you to escape.
gdb disables address-space randomization for the program it debugs, so that addresses are the same from one run to the next; it does this with a system call that Docker’s default filter forbids. Without the option, gdb warns at every
runand the addresses of hosted programs in chapter 6 move each time, which makes the printed listings impossible to follow.First the flags (
BOOKFLAGSforgotten or mistyped), then the tool versions against the table in this chapter, then your own typing. The book’s listings were produced by running the commands in the repository under the stated versions and the build checks them, so a difference is almost always on the reader’s side; and the first two checks take a minute while suspecting the book leads nowhere.The hosted programs of Part I are ELF files, and chapters 5 and 8 dissect that format because the kernel uses it too. gcc on Windows produces PE files and on macOS Mach-O files, so
readelfandobjdumpoutput, section names, linker scripts and the whole of chapter 5 would differ. The container gives the same Linux tools on every host, which is why it is the recommended option.
E.2 Chapter 1: Domain documents
The language is only the means of expressing a solution; what is missing is knowledge of the problem domain, which for an operating system is the hardware: how a CPU switches modes, how a device is programmed through its registers, what the firmware leaves in memory. A programmer who knows C perfectly but not the domain cannot even read the manuals that describe what the code must do, let alone write correct code.
A requirement document describes the problem domain and the effects the software must produce in it; a specification states the rules between inputs and outputs, that is, the interface. You are reading a requirement when the text explains how the world behaves (how a serial line carries bits), and a specification when it tells you which value to put where to obtain a behavior (write this bit of this register to set the baud rate).
A drop-down menu and a linked list are implementation choices, not facts about the problem domain. Implementation changes often, so the document must be updated each time or it starts lying; soon the effort is given up, the document is abandoned, and everybody concludes that documentation is useless. A document that sticks to the domain stays true while the implementation evolves.
Hardware cannot be patched after it is manufactured, so a defect found by a customer is a recall, financially and reputationally devastating. The interface must therefore be fixed and described down to the last bit before the chip ships. Software is cheap to change, so its specification is often incomplete and written after the fact.
“Functional Description” is the requirement document: it describes the domain, what the device does and why. “Register Description” is the specification: the list of registers, bits and the values they accept. When you need to program a device you go to the register description for the interface, but you understand what the bits mean only after reading the functional description; chapter 10 does exactly this with the serial port.
What matters is that a complete specification exists at the end, because every program that uses the interface depends on it. It can be written before, during or after the implementation; writing it as parts of the interface become stable is often the most practical. What must not happen is for it to be left unwritten, or for implementation details to leak into it.
Software-only domains (compilers, cryptography) produce software for other software, so the requirement and the specification tend to merge into one document. An operating system is both: it serves programs, but it also has to produce effects in hardware, which is an external domain with its own documents written by hardware engineers. To write one you must learn enough of that external domain to read those documents, which is what the three manuals listed at the end of the chapter are for.
E.3 Chapter 2: From hardware to software: Layers of abstraction
Any Boolean function can be written using only NAND (or only NOR), a result Peirce proved; a set of functions with this property is called functionally complete. Since every digital device is a combination of Boolean functions, a device that can be built from gates can be built from NAND gates alone, which is why the 74HC00 and the Apollo Guidance Computer are enough to make computers.
A transistor is an electrical component, a resistor whose value depends on a voltage; a logic gate is a function on bits, implemented with a few transistors wired in a fixed CMOS pattern. The designer thinks in gates, because the pattern that turns a gate into transistors is always the same and can be applied mechanically afterwards; thinking in transistors would mean translating every idea by hand and verifying it at the wrong level.
It makes the translation from gates to circuits a mechanical rule that never needs to be reconsidered. Because it holds for every gate, a design expressed in gates can be “compiled” into a circuit automatically, and the designer can ignore the circuit entirely. This is the recurring pattern that lets the gate layer exist as a language of its own.
The user needs the truth table because it is the interface, the complete description of what the device does for every input; the internals are an implementation that could change (a different transistor technology, a different layout) without changing the table. The user loses nothing by not knowing them, exactly as a C programmer loses nothing by not knowing how
printfis written, as long as the table is correct and complete.The decoder reads the opcode bits of the instruction and selects the circuit that implements that instruction, routing the operand bits to it. Without it, a human would have to pick the right chip by hand and feed it signals, which is what the chapter does before introducing the opcode; the decoder is what turns a pile of devices into a machine that executes a program by itself.
Each instruction has a small circuit that performs it; the bit pattern of the opcode is what selects that circuit, as a name selects a function. The “body” is the circuit itself, the gates that compute the result when the decoder activates them. Assembly then gives the name a mnemonic (
hltfor11110100), but the mapping from name to circuit is unchanged.Each layer takes the patterns that recur in the layer below, gives them a simpler notation, and drops the details that do not recur: logic gates name a fixed transistor pattern, machine instructions name fixed gate circuits, assembly names bit patterns, C names repeated blocks of assembly. Without a recurring pattern there is nothing to name, so no language and no abstraction could be built on top of that layer.
All languages can express the same computations, but a high-level language hides the hardware to free the programmer from it, and the hiding carries extra code (memory management, run-time checks) that the programmer cannot control. An operating system has to control the hardware directly and predictably, so it needs a language that is a thin layer over the machine, where one can see what the hardware does for a given line; C is that language, and chapter 4 shows the mapping line by line.
E.4 Chapter 3: Computer Architecture
The CPU is the only device that executes the programmer’s code; every other device only reacts to values written to its registers and ports. So the program controls a device by having the CPU read and write those registers, either through I/O instructions (
in,out) or through memory addresses that the hardware maps onto the device. A device driver is nothing but the knowledge of which register to write, with what value, in what order.Both are hardware cells that software reads and writes, but a register stores a value, while a port delegates the value to another circuit and does not keep it. Writing to a register changes a state that can be read back; writing to a port triggers an operation, and reading the same port may return something unrelated, for instance the next byte received on a serial line.
You can rely on everything the ISA defines: the instructions, the registers, the operating modes, the descriptor and page-table formats, the exceptions, because both vendors implement them identically and the manuals of both describe them. You cannot rely on anything that belongs to the organization or the hardware: how many cycles an instruction takes, the cache sizes, vendor-specific model-specific registers and CPUID leaves, and the virtualization extensions, which differ between VT-x and AMD-V.
A microcontroller is a small computer built to control other hardware; a system-on-chip is a complete general-purpose computer on one chip, able to run an operating system and many applications. Both are programmable, so the difference is in resources and intent. An SoC in the place of a microcontroller would work, but it costs far more and most of its hardware would sit idle running a program that needs almost nothing.
Memory is an array of bytes with addresses; the CPU fetches whatever bytes
eippoints to and executes them if they form valid instructions. Only the software decides which bytes are code and which are data. For a kernel this means that a wrong jump executes data as instructions, and a wrong write turns code into garbage; protection mechanisms, from segment types in chapter 9 to the no-execute bit, exist to let the kernel tell the CPU what it meant.Q35 is the most recent chipset QEMU emulates, and QEMU is our computer. What moved since 2007 is the location of functions: the memory controller and then the graphics went into the CPU package, the Northbridge disappeared and the Southbridge became the PCH. What stayed is the software view: PCI configuration space, the legacy devices at their traditional I/O ports, ACPI tables. The kernel of this book talks to that software view, so it is not tied to 2007.
In real-address mode the register holds a plain number, multiplied by 16 and added to the offset, with no checks; any code reaches any byte of the 1 MB space. In protected mode it holds a selector, an index into a descriptor table, and the base, limit, type and privilege level come from the descriptor. The CPU checks every access against those fields and raises an exception when they are violated, which is where memory protection comes from, and why chapter 9 must build the GDT before anything else.
Hardware interrupts arrive only when IF is set, so clearing it with
cliguarantees that the code until the nextstiruns without being interrupted on this CPU; nothing else is needed to make the section atomic. On a second processor IF is a separate bit in a separateeflags, so that processor keeps running and can enter the same critical section; a multiprocessor kernel needs a lock in memory as well, which chapter 13 discusses.
E.5 Chapter 4: x86 Assembly and C
01 c8. Withbits 32, 32-bit operands are the default and the operand-size override prefix66is not needed; the opcode and the ModR/M byte do not change, since the register numbers are the same. In 32-bit mode the prefix selects the non-default size, 16 bits, so66 01 c8isadd ax, cx: the same bytes are two different instructions depending on the mode of the processor, which is why a disassembler must be told which mode the code is meant for before it can decode a flat binary.A displacement takes part in an address computation, added to the base and index registers or used alone as an absolute address; an immediate is an operand value used as is. In
mov DWORD PTR [ebp-0x8],0xabcdef,-0x8is an 8-bit displacement and0xabcdefa 32-bit immediate. The processor does not tell them apart by their values but by the opcode and the ModR/M byte, which fix how many displacement and immediate bytes follow and in which order.eb feis a relative jump: the displacement,-2, is added toeipafter the instruction, so the processor computes the target from where it is. Code made of relative jumps runs unchanged wherever it is loaded, which is what the linker wants for loops and local branches, and what a bootloader will like in chapter 7. An absolute address is needed when the target is not at a fixed distance: a jump through a register or memory (FF /4, as for a function pointer or aswitchtable) and a far jump (EA,FF /5) that also loadscs, which we will use to enter protected mode.wordat+0,byteat+2, one padding byte,dwordat+4: each variable is placed at the next address that is a multiple of its size, so where the padding lands depends on the declaration order. The reason is in section 4.1.1 of volume 1: a word or doubleword that crosses a 4-byte boundary costs two memory accesses instead of one, and some instructions on 16-byte operands raise#GPon unaligned data. Padding is cheaper than slow accesses; reordering the declarations is what a programmer can do to avoid it.idivdivides the 64-bit valueedx:eaxby its operand;cdqcopies the sign bit ofeaxinto every bit ofedx, so thatedx:eaxisisign-extended to 64 bits. Without it,edxholds whatever the previous instruction left: for a negativei, the dividend would be a huge positive number, the quotient wrong, and if the quotient did not fit in 32 bits the processor would raise a divide error,#DE, the same exception as a division by zero. Oneidivcomputes both results:i / jkeepseax,i % jkeepsedx.[ebp]holds the savedebpof the caller and[ebp+4]the return address pushed bycall; the arguments start at[ebp+8]and a third one would be at[ebp+0x10]. Pushing in reverse order puts the first argument at the lowest address, right above the return address, so that it is at[ebp+8]whatever the number of arguments: a function with a variable number of arguments, such asprintf, finds its first argument without knowing how many follow.In this convention, cdecl, the caller removes the arguments because only the caller knows how many it pushed:
printf("%d %d", a, b)andprintf("hi")call the same function with different numbers of arguments. A callee that knows its argument count can clean up itself withret imm16, the form ofRETwith an immediate in volume 2, which pops the return address and then adds the immediate toesp; this is the stdcall convention of the Windows API, and theadd espin the caller disappears.jlebecomesjbe, “below or equal”, the unsigned condition, and nothing else changes.cmpis a subtraction that sets all the flags at once,ZF,SF,OFandCF; whether the comparison is signed or unsigned is decided by the conditional jump, which chooses which flags to test:jletestsZF=1orSF≠OF,jbetestsCF=1orZF=1(appendix B of volume 1). The samecmpserves both; the compiler expresses the type only through theJccit emits.
E.6 Chapter 5: The Anatomy of a Program
The two tables serve two readers. The operating system loads segments: a few contiguous ranges with permissions, which is all it needs to map the file into memory and jump to the entry point. The linker and the inspection tools work on sections: named pieces with a type, a symbol table and relocations. Without its section header table,
hellostill runs, since the loader never reads it;readelf -S,objdump -dby section,gdbsymbols and linking against the file stop working.The variable would move to
.data, which holds initialized data and is stored in the file: the file would grow by 4 bytes, and.commentand everything after it would shift by 4. In memory nothing changes: the.bssbytes were allocated and zeroed by the loader anyway; they are now read from the file instead..dynsymand.dynstrare needed at run time, by the dynamic linker, to findputsand__libc_start_mainin the shared C library, so they are allocated (A) and have an address in the memory image..symtab,.strtaband.shstrtabare needed only by tools working on the file, the linker,readelf,gdb; they are not loaded, have no address, andstripcan remove the first two without changing what the program does.A section header is a fixed-size entry so that entry
iis ate_shoff + i * e_shentsizewithout parsing the previous ones; a string of variable length does not fit in such an entry.sh_nameis an offset into.shstrtab, the string table whose index the ELF header gives ine_shstrndx;readelfreads that section and prints the null-terminated string at the offset, exactly as example 5.12 showed by hand..rel.textlisted the places in the code ofhello.owhose address was unknown at compile time, such as the string passed toprintfand the call itself. The linker resolved them, patched the bytes and dropped the section..rel.dynand.rel.pltare the relocations the linker could not resolve, because their targets live in the shared C library, which is loaded at an address known only at run time; the dynamic linker applies them when the program starts.The separation costs a few pages: each segment starts on its own 4 KB page, so the file has padding and the process a few more mappings. In return, no page is both writable and executable: a bug that writes through a wild pointer into the code, or that jumps into data, is reported immediately as a segmentation fault, where a single read-write-execute segment would let the program run on corrupted code and fail later, somewhere else. Our kernel, on the bare metal, will have no such protection until we set it up ourselves.
The C runtime would be skipped: the stack would hold the raw
argc,argvand environment that the kernel placed there, not the argumentsmainexpects,stdiowould not be initialized, constructors would not run. The dynamic linker still runs before the entry point, soprintfcan execute, but whenmainreturns, itsretpops whatever is at the top of the stack, the value ofargc, as a return address, and the program crashes with a segmentation fault; nobody callsexit, so even the text thatprintfbuffered may never appear._startexists to do all that before and aftermain.The loader does not copy sections one by one; it maps each
LOADsegment as one block, a range of file bytes onto a range of addresses, page by page. Within the block, a byte keeps its distance from the start, so a section’s address is the segment address plus its offset within the segment. Since mapping works in whole pages, the file offset and the virtual address of a segment must be congruent modulo the page size, 4 KB; the GNU linker goes further and starts each segment on a fresh page in the file, hence the offsets0x1000,0x2000, so that no page holds bytes of two segments with different permissions.
E.7 Chapter 6: Runtime inspection and debug
Line 3 produces no code, so gdb moves to the first line that does.
break mainandbreak 3both land on0x8049177because gdb skips the prologue of the function, using the line table: the first row ofmain(line 4, the opening brace) is the prologue, and the second row is whereebpis set up and the arguments and locals can be read correctly.break *0x8049166stops at the very first instruction, which is what you want to watch the prologue itself, or when debugging assembly without debug information.gdb keeps the list of addresses where it has written
0xCC, with the original byte of each. On aSIGTRAP, it checks whethereip - 1is one of them: if so, it reports the breakpoint number, restores the byte and moveseipback by one so that the original instruction runs. Our ownint 3is not in the list, so gdb passes the signal on as it is and leaveseipwhere the processor put it, after theint 3, at0x8049178, which is why it stops on theprintfline without rewinding anything.From
eax: the calling convention returns anintthere, andfinishruns until the return address and then reads the register. The debug information tells gdb the return type ofadd1000, from theDW_AT_typeof its subprogram DIE: without it, gdb would not know whether to readal,eax,edx:eaxfor a 64-bit integer, or nothing for avoidfunction, nor how to print the value. On a function without debug information,finishstops but prints no value.The program breaks first. Code compiled with
-no-pieembeds absolute addresses, such aspush 0x804a008for the string of example 6.10 ormov eax,ds:0x804c00cin chapter 4, which would point to nothing at the new location. Even position-independent code would be undebuggable: gdb would plant breakpoints at the linked addresses, where there is no code, and mapeipto no line, because the DWARF entries and the line table are expressed in linked addresses (section 6.4.3). Both must agree, which is why the load address of our kernel will be fixed in the linker script.It needs
DW_AT_nameandDW_AT_comp_dirof the compilation unit to build the path/tmp/hello.c, and the line table to mapeipto line 5; it then opens the file and prints its fifth line. Without the file, everything that depends on the executable alone still works: breakpoints by line or function,nextandstep, backtraces, variables anddisassemble. Only the source text is missing: gdb warnshello.c: No such file or directory, shows7 in hello.cinstead of the line, andx/i $eipbecomes the way to see what runs.Memory is bytes without a type; the format letter decides how gdb reads them.
x/idecodes whatever it is given as instructions, exactly asobjdump -Ddid in chapter 4 on.data: the78byte ofbfbecomes ajsthat swallows the next byte, and the listing has no relation to the data. On data, usex/8xborx/4xwto see the bytes, orprintwith the variable’s type when there is debug information, so that gdb decodes the bytes the way the program does.With
int 3, the debugger would have to decode the current instruction to know its length and, for a jump, a call or a return, compute where execution goes next, and it would have to patch memory before every step. With the trap flag set, the processor raises interrupt 1 after each instruction, whatever it is, with no decoding and no patching.nineeds to let acallrun to completion: when the current instruction is acall, it plants a temporary breakpoint at the return address, the instruction after thecall, and continues; otherwise it would single-step through the whole ofputs.
E.8 Chapter 7: Bootloader
The directive
times 510 - ($-$$) db 0pads the file with zeros up to offset 510, whatever the size of the code, sodw 0xAA55always lands on the last two bytes of the sector (as55 aa, little-endian). With 511 bytes of code, the expression510 - ($-$$)is negative, andnasmrefuses to assemble the file: the error is caught at build time rather than as a “not a bootable disk” message at boot time.nasm -f bincomputes every label from offset 0 of the file, somov si, msgis assembled asmov si, 2; at run time the string is at0x7c02, and0000:0002is the interrupt vector table, which starts with a zero byte where a print loop stops.org 0x7c00makes the labels agree with where the BIOS puts the sector. The chapter’s bootloaders never readmsg, andjmp bootis a relative jump, so no absolute address is ever used, and they get away with it.The far jump loaded
CS = 0x50andIP = 0; the CPU is executing at physical0x50 × 16 + 0 = 0x500.info registersshows the segment and offset, andx/10xb $cs*16+$eipshows the bytes ofsample.gdbreports addresses fromEIPalone, so it says0x0and does not recognize the breakpoint it set at0x500; thehook-stopprints[cs:eip]at every stop so that the real location is always visible.b8 50 00ismov ax, 50hand8e c0ismov es, ax. QEMU describes the CPU togdbas a 32-bit target, andgdbdecodes 32-bit instructions, whereb8takes a 4-byte immediate and swallows the next instruction. The bytes in memory are right, so the check is to read them withx/8xband compare withhdor thenasm -llisting; breakpoints, stepping and the registers are unaffected.bootloaderandosare directories, and a directory is a file that exists and is always up to date, so without.PHONYmakewould say “nothing to be done” and never run the child Makefiles.bootdiskis also phony, and declared so, but no file of that name exists, so the problem would not show up until someone created one.It is impossible to boot or test a stale image: whatever changed in the sources is rebuilt before QEMU starts. The cost is a run of the child Makefiles at every
make qemu, which is negligible here, since an unchanged object file is not rebuilt.The stack is wherever the BIOS left
SS:SPwhen it jumped to the bootloader,0000:6f08under SeaBIOS; the BIOS chose it for its own use, and the interrupt frame ofint 0x13is pushed there. Another BIOS may leave the stack anywhere, including on top of the buffer at0x500or of the bootloader, so a bootloader setsSS:SPitself, to a region it knows to be free, before the firstcall,pushorint.
E.9 Chapter 8: Linking and loading on bare metal
calltakes a displacement relative to the end of the instruction, and the assembler leaves-4as a placeholder because the relocation formula, , is applied to the start of the 4-byte field, not to the end of the instruction: the-4is the addend that corrects for the field’s own size. After linking, , and give7, which isfoominus the address of the next instruction,0x804916b.A
PT_PHDRentry tells the program where its own program header table is in memory; the dynamic linker uses it to findPT_DYNAMIC. The table is in memory only if aPT_LOADsegment covers it, so the ELF specification allowsPT_PHDRonly in that case, andldnow enforces it. We keep it because it documents the layout and because puttingFILEHDR PHDRSon the firstLOADsegment is what makes the bootloader’s scheme work: the header is loaded at the start of the file, soe_entryis at0x18from the load address.They are the ELF magic number,
0x7ffollowed byELF: the bootloader copies the whole file to0x500, so the header is at0x500and.text, at file offset0x500, is at0xa00. The linker computede_entryand every address in the debug information assuming the segment starts at address0, where its file offset is, so both were0x500short of reality.-nmagicturns off page alignment: without it,ldgives everyLOADsegment an alignment of0x1000and places sections at file offsets congruent to their addresses modulo0x1000, so.textat0x600goes to file offset0x600. With it, the alignment comes from the sections, andALIGN(0x100)puts.textat file offset0x100and address0x600, so the segment starts at address0x500, file offset0, exactly where the bootloader puts it.Debug information, the symbol table and the section headers: everything
gdbneeds and the CPU does not. OnlyLOADsegments must be in memory, and ours is0x106bytes, so 17 sectors cover it many times over.gdbreads the symbols and the DWARF data frombuild/os/oson the host, throughsymbol-file, never from the memory of the virtual machine.Everything except the bytes of the
.textsection: symbols, line numbers, section headers. The flat binary is what the BIOS can load and the ELF is whatgdbcan read, so both are kept.symbol-filereplaces the current symbol table;add-symbol-fileadds a second one, which is what you need to see the bootloader and the kernel in one session.Only by luck:
55 89 e5 90 5d c3decode in 16-bit mode aspush bp,mov bp, sp,nop,pop bp,ret, the same operations on 16-bit registers. An instruction with a 32-bit immediate or displacement, such asmov DWORD PTR [ebp-4], 5, decodes differently in 16-bit mode, takes the wrong number of bytes and derails everything after it. Chapter 9 switches to protected mode before any C code runs.-ffreestandingtells gcc that there is no C library: it stops treatingmainas special and stops assuming that functions namedmemcpyorprintfhave their standard meaning, so it will not replace a loop by a call to a library that does not exist.-nostdlibonly changes what is linked. Without-fno-stack-protector, a function with a local array reads a canary throughgs, which points to nothing on bare metal, and calls__stack_chk_fail, which is not there: a fault or an undefined reference at link time.
E.10 Chapter 9: Protected mode and x86 descriptors
- The bootloader’s table lives inside the boot sector at
0x7C00, memory that nothing protects and that the kernel is free to reuse (the bootloader itself writes the boot information block of chapter 12 at0x500, a few hundred bytes below it). The processor reads descriptors from the table whenever a segment register is loaded, so a GDT that gets overwritten is a crash waiting to happen. Building the table in the kernel’s own.bssalso lets later chapters add entries (user segments, the TSS) without touching the 512-byte bootloader. - The processor is in protected mode from the moment
PEis set, butCSstill holds the real-mode value 0 and its hidden part still describes a 16-bit segment. A near jump does not reloadCS, so the code atprotected_modeis decoded as 16-bit code, and thebits 32instructions there are misread (mov ax, 0x10becomes a different instruction followed by garbage). Only a far jump, far call oriretloadsCSwith a selector and brings in the 32-bit descriptor. - The limit field is 20 bits wide, so it cannot hold
0xFFFFFFFF. WithG = 1the limit is counted in 4 KiB units:0xFFFFFmeans0x100000pages, which is0x100000000bytes, all 4 GiB.G = 0would make the same value mean 1 MiB. The granularity bit is what makes a 20-bit field cover a 32-bit space, and QEMU’smonitor info registersshows the scaled result,ffffffff. - Each segment register caches its descriptor in a hidden part at the
moment it is loaded, and every address computation uses that cache
(Intel SDM Volume 3A, section 3.4.3).
lgdtonly changes the GDTR; the hidden parts still hold the descriptors loaded from the bootloader’s table, which are valid. Without theljmp,CSwould keep the old descriptor, which happens to be identical, so nothing would fail now; but the kernel would depend on the boot sector’s table forCSuntil the first interrupt or far transfer, and any later change to the kernel GDT’s code entry would not take effect. 0x13is0x10 | 3: index 2 of the GDT withRPL = 3, the data segment requested at privilege level 3. The processor comparesmax(CPL, RPL)with the descriptor’sDPL; withRPL = 3andDPL = 0the check fails and themovraises#GP. Chapter 13 adds descriptors withDPL = 3precisely so that such selectors can be loaded.- The gate is outside the processor, in the chipset, and the
bootloader is the natural place to open it, once, before anyone depends
on memory above 1 MiB. With the gate closed, bit 20 of every address is
forced to zero, so a write to
0x100000lands at0x000000. Chapter 12 maps and uses all of the 128 MiB; a kernel with A20 closed would see its page tables above 1 MiB silently alias the first megabyte, corrupt the IVT or the boot information block, and fail in ways that have nothing to do with paging. - A hosted program is loaded by the operating system’s ELF loader,
which reads the program headers and, for each
LOADsegment, copiesFileSizbytes and then zero-fillsMemSiz - FileSizbytes;.bssis the difference. Our bootloader copies the file verbatim and never reads a program header, so the memory of.bssreceives whatever followed.datain the file, here.debug_aranges. The loader’s missing job is exactly that zero-fill, which_startperforms from the two linker symbols. - Without
packed,gccmay insert padding so that each member is naturally aligned:base_middleandaccessare bytes and need none, but a compiler is free to pad after them, and on some targets the structure grows. With this exact member list on i386 the layout happens to stay 8 bytes, which is why the bug is dangerous: the code works until a member is reordered or the target changes. The processor reads 8 raw bytes per entry, so any padding shifts every later field and the selectors describe garbage, usually a#GPat the first segment load.
E.11 Chapter 10: Talking to devices: the serial port and the VGA text console
The UART shifts bits out at 38400 baud, about 260 microseconds per character, while the CPU can hand it a byte every few nanoseconds. Writing without waiting overwrites the transmit holding register before the previous byte has moved to the shift register, so most characters are lost and the output is a few scattered letters. QEMU hides the bug because its emulated UART delivers a byte to the host file or terminal the instant it is written; the register is always empty, so the broken driver works perfectly until it meets a real chip.
The hang happens before any serial byte is written, so the log would be empty: no greeting, nothing. The script would time out, print an empty serial output and exit 1, and the CI would report a failure without any clue about where the kernel stopped. With the serial port first, the greeting is already in the log when the VGA driver hangs, and the log shows exactly how far the kernel got. That is why the order in
putcis a design decision and not a detail.The keyword states a fact about the hardware, not about the compiler flags: the memory and the ports have side effects the compiler cannot see, and a driver that only works because optimization happens to be off is wrong, not lucky. At
-O2, withoutvolatileoninb, the compiler would see a loop whose body has no visible effect and whose condition reads the same “pure” function with the same argument; it would read the port once and spin forever on the result, or drop the loop entirely, andserial_putcwould either hang or write without waiting.The bytes of
EBXwould be stored high byte first, so the array would holduneGIeniletn: the same twelve characters in groups of four, each group reversed. The code is correct only because the register-to-memory byte order of the machine matches the order in which the processor packed the characters, which is itself little-endian by design. It does not matter in practice, sincecpuidexists only on x86, which is little-endian, but it shows that “copy the register to memory” is a statement about the architecture, not about C.vga_putcadvances the column after storing a character and wraps to the next row as soon as the column reaches 80, so a line of exactly 80 characters moves the cursor to the next row by itself, and the\nthat follows moves it down once more. A terminal defers the wrap: after the 80th character the cursor stays in a pending state at the end of the row, and a newline simply ends the line, while a further printable character triggers the wrap. The fix is to wrap lazily too: do not advance the row when the column reaches 80, but keep a flag, and wrap at the start of the next printable character, clearing the flag on\n.The processor raises exception 6, invalid opcode, and looks up vector 6 in the IDT. The kernel has not loaded an IDT, so the IDTR still holds whatever the BIOS left, a real-mode interrupt vector table that makes no sense as protected-mode gates; the descriptor the CPU fetches is invalid, which raises a general protection fault, whose gate is equally invalid, which raises a double fault, and the third failure in the sequence is a triple fault: the processor resets. The machine reboots a few microseconds after the
cpuid, and the serial log shows the greeting and nothing else, over and over.The C standard says that arguments passed through
...undergo the default argument promotions, so the caller pushed anint, never achar;va_arg(args, char)asks for a type that cannot have been passed, which is undefined behavior, and gcc warns about it even though on this machine both occupy four bytes. On a 64-bit build the arguments no longer sit on the stack at all: the System V AMD64 convention passes the first six integer arguments in registers, andva_listbecomes a structure that remembers how many register arguments have been consumed before switching to the stack.va_argstill works, because the compiler implements it, but the comment inprintf.cabout “the next double words on the stack” would no longer be true, and a%pwould fetch eight bytes, soprint_pointerwould need sixteen digits.
E.12 Chapter 11: Interrupts
- The BIOS leaves the master 8259A delivering IRQ 0 to 7 on vectors 8
to 15, which in protected mode are the exceptions
#DF,#TS,#NP,#SS,#GPand#PF. A timer tick would arrive on vector 8 and the kernel would print “Exception 8: Double Fault” and halt, a hundred times a second if it did not.pic_initmoves the lines to vectors 32 to 47 with ICW2 beforestienables delivery. - The two macros exist so that the stack has the same shape for every
vector: the processor pushes an error code for
#DFand the stub must not add a second one. Ifisr8pushed a fake 0,struct registerswould be shifted by one word for that vector (eipwould read the real error code,csthe savedEIP), andadd esp, 8beforeiretwould leave one word too many, soiretwould pop the savedEIPasEFLAGSand jump to the savedCSvalue. The symptom is a triple fault from inside the handler of a double fault. #BPis a trap: the savedEIPpoints after theint3, so returning continues the program.#DEis a fault: the savedEIPpoints at theidivitself, so returning re-executes it and raises#DEagain. Without a fix (the kernel has none for a division by zero),isr_dispatchwould print the register dump forever; halting is the honest reaction.- In 32-bit code
push dspushes 4 bytes, the segment selector zero-extended, as the PUSH description in Volume 2B says. If the four fields wereuint16_tthe structure would be 8 bytes shorter than the stack frame, and every field afterdswould be read 8 bytes too low in memory, two words early:regs->eipwould hold the vector number,regs->err_codethe savedEAX, andregs->int_nothe savedECX, so the dispatcher would index its handler table with whateverECXhappened to contain. - For any IRQ: the end-of-interrupt command was not sent to the PIC,
so the line stays “in service” and the chip never raises it again;
isr_dispatchsends it after the handler for exactly this reason, so in our code the cause would be a handler that never returns, or an EOI sent to the master only for a slave line. For the keyboard: the handler did not read port0x60, and the 8042 does not raise IRQ 1 again until its output buffer has been read. - Both are 8-byte gates with the same fields; the only difference is
that an interrupt gate clears
IFwhen the processor enters the handler and a trap gate leaves it alone (Intel SDM Volume 3A, section 7.12.1.3). With interrupts off, a handler cannot be interrupted by another hardware interrupt, so it can touch shared data (the tick counter, the run queue of chapter 13) without a race, and nested timer ticks cannot pile up on the stack. Exercise 11.7 shows what the trap-gate version does under load. stienables interrupts after the instruction that follows it completes (Volume 2B, “STI”), so the pending tick could only be delivered after themov. The delay exists forsti; hlt: if an interrupt could be taken between the two, the handler would run andhltwould then stop the processor with nothing to wake it up; with the delay,hltis reached with interrupts enabled and the pending interrupt wakes it.esp_dummyis the value ofESPbeforepusha, that is, the address of the vector number.pushathen pushed 8 words and the stub 4 segment registers, 12 words or 48 bytes, which is whyregs, the stack pointer after those pushes, is 48 bytes lower. Above the vector number sit the error code,EIP,CSandEFLAGS, 5 words in all, soesp_dummy + 20is whereESPwas when the processor started pushing, the interrupted code’s stack pointer; the QEMU log confirmed it withSP=0010:0008ffe0.
E.13 Chapter 12: Memory management
- There is no instruction that reports the amount of RAM: the memory
controller in the chipset knows, and only the firmware has talked to it,
so the map is a BIOS service,
INT 15h, EAX=E820h. It also describes holes and reserved ranges (VGA memory, the BIOS ROM, ACPI tables, PCI apertures) that a plain size could not express. BIOS services are 16-bit real-mode code reached through the real-mode interrupt vector table, which the kernel replaces with its IDT, so the bootloader must collect the map beforeCR0.PEis set and leave it at0x500for the kernel. - The usable ranges are the only ones the BIOS guarantees; everything
it does not mention (from
0xa0000to0xf0000on QEMU, where the VGA buffer and option ROMs live) is simply absent from the map. Starting from “all free” and subtracting the reserved entries would leave those unlisted ranges free, and the allocator would eventually hand out the VGA frame buffer or a ROM as a page table. Starting from “all used” makes silence mean “not RAM”. - A frame that is only partly covered by a reserved range must stay
reserved, so a used range rounds outwards (start down, end up); a frame
that is only partly covered by a usable range is not entirely RAM, so a
free range rounds inwards (start up, end down). On QEMU the first usable
entry ends at
0x9fc00, in the middle of frame0x9f, which holds the Extended BIOS Data Area; the inward rounding keeps that frame out of the free pool even before the first-megabyte line marks it used. - The page the next instruction lives in must be mapped at its own
address, with the directory entry and the table entry both present, and
CR3must hold the physical address of a directory whose entries point to tables that are themselves correct; the identity map guarantees all of it. If the mapping is wrong, the fetch raises#PF, the processor tries to read the IDT and the handler through the same broken tables, and the result is a double then a triple fault. In the-d intlog this is av=0eentry whoseCR2equals the address of the instruction aftermov cr0, followed byv=08and a reset. - The identity map reserves the virtual range from 0 to
pmm_memory_top(), 128 MiB, for the frames themselves: every physical address must stay a valid pointer. Mapping the heap at0x00400000would replace the identity mapping of the frames at 4 MiB with heap pages, so the allocator could still hand out those frames but the kernel could no longer write to them through their physical address, and the page tables that live there would become unreachable.0xC0000000is above every address the machine has RAM at. - The processor caches translations in the TLB and only walks the
tables on a miss, so a store into a page table does not change what the
TLB answers for a page it has already translated.
invlpgdrops the cached entry for that page. Without it, afterpaging_unmapthe TLB could still translate0x30000000to frame0x11f000and a read through the window would succeed instead of faulting; worse, after the frame is reused, the stale entry would read someone else’s data. - No. Bit 0 of the error code only says whether the fault was caused
by a non-present entry or by a protection violation; a missing directory
entry and a missing table entry both clear it, and
CR2only gives the address. The handler has to walk the tables itself: readkernel_directory[cr2 >> 22], and if that entry is present, read the page-table entry it points to, which is exercise 12.5. CR3and a directory entry hold only bits 31:12 of the address they point to, so a directory at0x14010would be used as if it were at0x14000; the arrays must start on a 4 KiB boundary.aligned(PAGE_SIZE)on the two arrays forces.bssto be 4 KiB aligned, andldraises the alignment of the segment that contains it to the largest alignment of its sections, so theLOADsegment reports0x1000where chapter 11 had0x100;readelf -Sshows the.bssalignment of 4096 andnmthe directory at0x14000.
E.14 Chapter 13: Processes: context switching, scheduling, user mode and system calls
It would not survive, and nobody is at fault, because the compiler never does that:
EAX,ECXandEDXare caller-saved in the System V i386 ABI, so the code generated for the thread saves anything it still needs before thecalltoyield()and reloads it afterwards.context_switchis an ordinary function from the caller’s point of view, and it saves exactly the registers the ABI says a function must preserve. If a value inEAXwere lost acrossyield(), the compiler would have broken the calling convention; the scheduler only honors it.ESP0is consulted only on a transfer from ring 3 to ring 0, and a task is in ring 3 only when its kernel stack is empty: it was entered through a gate, the handler ran, andiretemptied it again on the way out. A task in the middle of a system call is in ring 0, and an interrupt in ring 0 stays on the current stack without reading the TSS. For a kernel thread the field is never used at all, since it is never in ring 3; the scenario cannot arise, which is why “the top of the stack” is always right. If the processor did switch toESP0on a ring 0 interrupt, it would overwrite the frames below the stack top, and that is precisely why the architecture does not.No. With a trap gate
IFstays set on entry, and once the EOI is sent the PIC is free to raise IRQ 0 again. The next tick would interrupt the handler on top of itself, push a second frame on the same kernel stack and callschedule()from insideschedule(); nested often enough, the stack overflows, and the run queue is modified by two activations of the same code. The early EOI is safe for us only because the interrupt gate keepsIFclear untiliret, so that a tick arriving during the handler is held by the PIC and delivered after.Because the processor defines the current privilege level as the RPL of the selector in
CS, and it refuses to loadCSwith a selector whose descriptor has a lower DPL than the requested privilege; the only way to execute withCPL = 3is to load a descriptor that hasDPL = 3. The segments must exist to make the processor believe the code is unprivileged, not to confine it. What confines it is bit 2 of the page-table entry for the page ofticks, the U/S bit, together with the same bit in the directory entry above it; both must be set for a ring 3 access, and the kernel’s pages have neither.task_wake_sleeperswould see a sleeping task whosewake_atis in the future and leave it;schedule()from the timer would skip it, switch to another task, and come back to it only when it is woken, whereupon it would resume insidetask_sleepand callschedule()a second time, which simply gives the processor away for one more turn. Harmless here, but only because the data is consistent at that moment; the line that makes the question moot isirq_save()at the top oftask_sleep, which keeps the tick out until the switch has happened.ticks - wake_atis-0x90000000as an unsigned subtraction, and read as a signed 32-bit number that is+0x70000000, non-negative: the task is woken on the next tick, after one hundredth of a second instead of 280 days. The idiom treats a difference with the top bit set as “in the past”, so the longest future it can express is0x7FFFFFFFticks, about 248 days at 100 Hz, and a sleep longer than that is a sleep of one tick. A kernel that needs longer timeouts uses a 64-bit tick count or an absolute deadline with a different comparison.An exception, or a non-maskable interrupt:
climasks only the interrupts that come through the PIC, so a page fault inside a critical section still runs the fault handler on top of it.task_sleepmay switch because the data it shares with the timer handler, the run queue and its own state, is consistent at the momentschedule()is called; the critical section is closed in the sense that matters, even thoughIFis still clear.kmallocin the middle of splitting a block has a list in an inconsistent state, and a switch there would let another task walk it. The rule is not “do not switch with interrupts off” but “do not switch with shared data half-modified”.Take the user stack page,
ptr = 0x7FFFF000, andlen = 0xFFFFF010: the sum wraps to0x7FFFE010, which is below the split, so test 2 passes, and sinceendis now belowptrthe page loop does not execute at all, so test 3 passes too. The handler would then loopputcfour billion times, reading the stack page, then the unmapped page at0x80000000, where the kernel takes a page fault in ring 0 and panics. The wrap test exists because a single overflowed addition turns two correct tests into a lie, andWRITE_MAXexists so that even a correct check is not asked to walk a million pages.
E.15 Chapter 14: File system
- While
BSYis set the drive is updating its registers, and the specification says the other status bits are not valid then; a staleDRQleft over from the previous sector could be read as “data ready” for the next one. A loop that testedDRQbeforeBSYwould then pull 256 words of the previous sector, or of nothing, from the data register and desynchronize the transfer. Clearing ofBSYis the only edge that means the drive has finished thinking. - The drive acknowledges a write when the data is in its cache, so
without
FLUSH CACHEa power failure right afterata_write_sectorsreturns can lose the sector, and worse, the drive may commit cached sectors in a different order from the one the file system chose, which breaks the “mark used before pointing to it” reasoning. Nothing is lost while the power stays on. Linux flushes only at barriers its file systems request, typically when a journal transaction commits, because a flush after every sector costs more than the writes themselves. - The block size is in the superblock, so it is unknown until the
superblock has been read; the only thing known in advance is that the
superblock starts at byte 1024, which is sectors 2 and 3 whatever the
block size. With 4 KiB blocks the two raw sectors are the same, but the
superblock now sits in the middle of block 0, so
s_first_data_blockis 0 and the group descriptor table is block 1;s_first_data_block + 1expresses both cases. file_blockreturns 0,ext2_read_filefills that kilobyte with zeros, and the file reads as if it held zeros there: a hole in a sparse file. Block 0 holds the boot record, which ext2 never allocates to a file, so 0 can mean “no block” without ambiguity. The first user programs had a block of zero padding between the headers and the code,mke2fs -dstored it as a hole, and the loader, which did not check for 0, read block 0 of the file system and loaded the boot record’s bytes in place of the code.- A crash after the bitmap write leaves a bit set with a count that is
one too high, or, counts first, a count that is one too low; in both
cases pass 5 of
e2fsckrecomputes the counts from the bitmaps and prints “Free inodes count wrong”. No data is involved, so neither order is worse. Writing the inode and the directory entry before the bitmap is different: the inode is in use, reachable by name, yet marked free, and the nextalloc_inodewould hand the same number to another file, so two names would share one inode. - Every entry’s
restwould be smaller than the 16 bytes needed, the loop would end anddir_add_entrywould print “directory is full” and return -1, leaving an inode allocated and written but unreachable (whiche2fsckwould move tolost+found). Growing the directory needsalloc_blockfor a new block, a single entry in it whoserec_lenis the whole block, and then the directory’s inode written back with the block number ini_blockandi_sizeincreased by one block. - Pass 1 checks that
i_sizecovers the blocks the inode has, and would report that the inode’s size is wrong, “should be” the size of the allocated blocks, and also thati_blocksis wrong if the kernel had not counted them.lost+foundis fine becausemke2fsset itsi_sizeto 12288, the size of the 12 blocks it preallocated so thate2fsckcan link orphans into it without allocating anything on a damaged disk. - An inode number identifies the file but not the position in it, so
two tasks reading the same file would share one offset, or the kernel
would have nowhere to store one; the same task opening the file twice
would be indistinguishable. Real kernels keep an open file description
(inode, current offset, mode) per
open, a table of pointers to them per process indexed by the descriptor, and the inode shared underneath;forkcopies the table, which is why a parent and child share an offset afterforkbut not after two separateopens.
E.16 Chapter 15: Address spaces: fork and exec
The parent would keep writing to the shared frame at full speed, and the child, which reads the same frame, would see every write the parent makes after the fork: a local variable the parent changes would change under the child’s feet. The child would notice, not the parent, since the parent’s view is always the “right” one; and the child’s first write would still copy the page, taking a snapshot of whatever the parent had done by then, so the bug would depend on timing. Copy-on-write only works if both sides fault on their first write, which is why the parent’s entries are modified in place.
The parent’s TLB still holds the writable translation of the stack page it was using just before the system call, so its first stores after
forkgo straight into the shared frame without a fault, and the child, when it runs, reads them: it sees the parent’s stored return value, or a few bytes of whatever the parent wrote next, in its own stack. The entry stays in the TLB until something evicts it, amov cr3on a switch to another task or just other translations crowding it out, after which the next write faults and is copied correctly; so a busy machine with frequent switches would show the bug rarely and an idle one more often. This is exactly the class of bug section 5.10.4.2 of the SDM exists to prevent.Nothing observable: with
0x202the few instructions between thepopfdincontext_switchand theiretinisr_returnwould run with interrupts enabled, and a tick landing there would run the timer handler on the child’s kernel stack below the frame, perhaps switch away and come back, and return to the sameiret. In chapter 13 the prepared frame’sEFLAGSwas the value the new task would run with, because the start function went straight into the thread or intoenter_user_mode, which pushes its own0x202; nothing else would ever have setIF. Here theiretloads the program’sEFLAGSfrom the copied frame, so the value on the prepared stack only matters for a dozen ring 0 instructions, and0x2makes them look like the end of any other handler.After
paging_switch_directorythe address inEBXis translated through the new directory: it points intohello’s freshly loaded page, or into its new stack, or into nothing, soelf_loadwould be asked for a garbage path (it was called earlier, so the real damage comes later) andtask_set_namewould copy bytes of the new program, or fault in ring 0 and panic. If the name were kept as a pointer, as in chapter 14, it would point into the old program’s page, whose frame was returned to the allocator and may already be somebody else’s page table. The kernel must own every byte of a user argument it uses after the address space changes.The store would succeed silently, in ring 0, into the frame both processes map: the child would find the parent’s
statusvalue in its own copy of that variable, at the same virtual address, without ever having written it. The parent would not notice anything, because its page-table entry stays read-only and COW, so its next user-mode write would still fault and copy the page, and the copy would contain the status it expects. The bug is visible only in the child, and only if the child reads that variable before taking its own copy of the page, which makes it the kind of bug that appears once a month.task_exitruns on the task’s own kernel stack and cannot free it while standing on it, which is chapter 13’s reason for the split between exit and reap; it could free the page directory and keep the rest, which is what Linux does (a zombie owns almost nothing). Copying the status into the parent at exit time would need a slot per child in the parent, or a list, since a parent may have several children that exit before it callswait; the zombie’s own structure is that list, with no new allocation and no limit on the number of children, and the parent’swaitis the one place that reads it. Unix made the same choice, which is why zombies exist.The count stops at 255 while 256 entries point at the frame. As the processes exit one by one, each
pmm_free_framedecrements: after the first exit the count says 254 while 255 users remain, and after 255 exits it reaches 0 and the frame is returned to the bitmap while one process still maps it. That process keeps reading and writing a frame the allocator will hand out again as somebody else’s page table or page, which is a use-after-free, worse than a leak. The cheapest correct fix is to refuse: makepmm_frame_refreport failure at 255 and haveforkreturn -1, as Unix returnsEAGAINwhen it cannot create a process; the next cheapest is a 16-bit count.P = 0 outside every lazy region: a program dereferencing a null or wild pointer, say
*(int *)0x10000, faults with error code0x4or0x6;demand_faultfinds no region and declines, and the handler prints and kills the task, as before. P = 1 inside a lazy region: a page of.bssthat was demanded, then shared read-only by afork, then written; the error code is0x7, the page is present and marked COW, andcow_faultcopies it, exactly as for a stack page. Tryingdemand_faultfirst would be wrong in that case: the address is in a lazy region, so it would allocate a fresh zeroed frame and map it over the shared one, and the process would lose the contents of its page. Bit 0 is what says whether the page exists, and only a page that does not exist may be made out of zeros.The frame goes into the child’s directory only:
demand_faultmaps it incurrent_task->page_directory, and the child is current. The parent’s entry stays zero, and when the parent reads the address later it takes its own demand fault and gets its own zeroed frame, so it reads zeros, which is exactly what it would have read hadforkcopied the page eagerly (a page nobody had touched held zeros). Two zeroed frames describe the same contents as one shared frame; nothing is observable, and nothing had to be copied or counted. Without the two fields, the child’sbss_startandbss_endwould be zero, so a touch of any untouched.bsspage would be reported aspage not presentand kill the child, while its stack would keep growing on demand, becauseuser_stack_topis copied separately and the stack window is derived from it; the pages the parent had already demanded before the fork would work too, since they are present and shared copy-on-write.
E.17 Chapter 16: When it does not boot: debugging a kernel
It tells you that the kernel was still running C code with a working serial driver when it stopped, and that the stop came between two calls of
serial_putc, so the last print is a lower bound on how far the boot got. It does not tell you why: a triple fault, a hang in a polling loop and acli; hltall cut the output the same way. The next step is the QEMU log for a fault or an attached gdb for a hang.stienables interrupts after the instruction that follows it completes, so thatsti; hltcannot be interrupted between the two and leave the processor halted with nothing to wake it (Intel SDM Volume 2B, “STI”). TheEIPrecorded for the first interrupt is therefore the second instruction aftersti, notstiitself; if you expect the fault “atsti” you will look at the wrong line, and the real lesson of such anEIPis “the first interrupt that was ever allowed through is the one that failed”.Not necessarily. A
#GPwith error code 0 is not about a descriptor at all: it is what the processor raises for a privileged instruction in ring 3 (cli,hlt,in), for an access through a segment outside its limit, or for a malformediretframe. If the gate’s DPL were the problem the error code would be0x402, naming gate0x80. Av=80followed byv=0d e=0000means the system call was taken and the handler, or the return from it, executed something ring 3 may not; look at the savedEIP.0x5is P = 1, W/R = 0, U/S = 1: a read, from ring 3, of a page that is present but reserved to the kernel. A user program read kernel memory, which is exactly what the page tables are for; the kernel should kill that task and continue, as chapter 13’s handler does, since the kernel’s own state is intact. Error code0x0at the same address would be a kernel-mode read of a page that is not mapped, which is impossible unless the kernel’s own tables or pointers are wrong, and that one is a kernel bug worth a panic.gdb unwinds by following saved frame pointers and return addresses, and the assembly stubs have no frame:
isr_commonpushes registers withoutpush ebp; mov ebp, esp, so the unwinder takes the wrong words for the caller’sEBPand return address and prints garbage or stops. Across aniretor acontext_switchthe stack simply changes, which no unwinder can follow. A user program built with-fomit-frame-pointerhas the same problem in every function:EBPis an ordinary register and the chain does not exist, which is why debug builds keep the frame pointer and why release kernels carry unwind tables instead.A reset requested through hardware rather than through a fault: a write to port
0x92with bit 0 set (the fast-A20 port, which also resets the processor), the0xFEcommand to the keyboard controller, the reset register at0xCF9on q35, or a jump to the reset vector at0xFFFFFFF0. None of these is an exception, so-d intis silent. Confirm with-d cpu_reset, whoseCPU Resetblock records the registers at the moment of the reset, including theEIPof theoutinstruction, and with-no-reboot, which intercepts every reset request and stops the machine instead.The driver could check the error register and the final status after the transfer, and it does check
ERRandDF; but a sector full of zeros is a perfectly valid sector, and the drive returned exactly the sector it was asked for. Validating the meaning of the data (a magic number, a mode field, a block number that must not be 0) belongs to the layer that knows the format, which isext2.c, and that is where the fix goes. Each layer can only check that its own input made sense.Because compiling proves nothing about a kernel: there is no runtime, no type checker for selectors or page-table bits, and the bugs of this chapter all compile without a warning.
make testboots the image and checks the behavior, so a version that passes it is a known-good point, and a regression is then a search between two such points, whichgit bisectdoes in log2(n) steps. Committing every compiling version gives a history in which “good” is undefined.