Preface
Greetings!
You’ve probably asked yourself at least once how an operating system is written from the ground up. You might even have years of programming experience under your belt, yet your understanding of operating systems may still be a collection of abstract concepts not grounded in actual implementation. To those who’ve never built one, an operating system may seem like magic: a mysterious thing that can control hardware while handling a programmer’s requests via the API of their favorite programming language. Learning how to build an operating system seems intimidating and difficult; no matter how much you learn, it never feels like you know enough. You’re probably reading this book right now to gain a better understanding of operating systems to be a better software engineer.
If that is the case, this book is for you. By going through this book, you will be able to find the missing pieces that are essential and enable you to implement your own operating system from scratch! Yes, from scratch, without going through any existing operating system layer to prove to yourself that you are an operating system developer. You may ask,“Isn’t it more practical to learn the internals of Linux?”.
Yes…
and no.
Learning Linux can help your workflow at your day job. However, if you follow that route, you still won’t achieve the ultimate goal of writing an actual operating system. By writing your own operating system, you will gain knowledge that you will not be able to glean just from learning Linux.
Here’s a list of some benefits of writing your own OS:
You will learn how a computer works at the hardware level, and you will learn to write software to manage that hardware directly.
You will learn the fundamentals of operating systems, allowing you to adapt to any operating system, not just Linux
To hack on Linux internals suitably, you’ll need to write at least one operating system on your own. This is just like applications programming: to write a large application, you’ll need to start with simple ones.
You will open pathways to various low-level programming domains such as reverse engineering, exploits, building virtual machines, game console emulation and more. Assembly language will become one of your most indispensable tools for low-level analysis. (But that does not mean you have to write your operating system in Assembly!)
Writing an operating system is fun!
Why another book on Operating Systems?
There are many books and courses on this topic made by famous professors and experts out there already. Who am I to write a book on such an advanced topic? While it’s true that many quality resources exist, I find them lacking. Do any of them show you how to compile your C code and the C runtime library independent of an existing operating system? Most books on operating system design and implementation only discuss the software side; how the operating system communicates with the hardware is skipped. Important hardware details are skipped, and it’s difficult for a self-learner to find relevant resources on the Internet. The aim of this book is to bridge that gap: not only will you learn how to program hardware directly, but also how to read official documents from hardware vendors to program it. You no longer have to seek out resources to help yourself interpret hardware manuals and documentation: you can do it yourself. Lastly, I wrote this book from an autodidact’s perspective. I made this book as self-contained as possible so you can spend more time learning and less time guessing or seeking out information on the Internet.
One of the core focuses of this book is to guide you through the process of reading official documentation from vendors to implement your software. Official documents from hardware vendors like Intel are critical for implementing an operating system or any other software that directly controls the hardware. At a minimum, an operating system developer needs to be able to comprehend these documents and implement software based on a set of hardware requirements. Thus, the first chapter is dedicated to discussing relevant documents and their importance.
Another distinct feature of this book is that it is “Hello World” centric. Most examples revolve around variants of a “Hello World” program, which will acquaint you with core concepts. These concepts must be learned before attempting to write an operating system. Anything beyond a simple “Hello World” example gets in the way of teaching the concepts, thus lengthening the time spent on getting started writing an operating system.
Let’s dive in. With this book, I hope to provide enough foundational knowledge that will open doors for you to make sense of other resources. This book will be beneficial to students who’ve just finished their first C/C++ course greatly. Imagine how cool it would be to show prospective employers that you’ve already built an operating system.
Prerequisites
Basic knowledge of circuits
Basic Concepts of Electricity: atoms, electrons, proton, neutron, current flow.
Ohm’s law
If you are unfamiliar with these concepts, you can quickly learn them here: http://www.allaboutcircuits.com/textbook/, by reading chapter 1 and chapter 2.
C programming. In particular:
Variable and function declarations/definitions
While and for loops
Pointers and function pointers
Fundamental algorithms and data structures in C
Linux basics:
Know how to navigate directory with the command line
Know how to invoke a command with options
Know how to pipe output to another program
Touch typing. Since we are going to use Linux, touch typing helps. I know typing speed does not relate to problem-solving, but at least your typing speed should be fast enough not to let it get in the way and degrade the learning experience.
In general, I assume that the reader has basic C programming knowledge, and can use an IDE to build and run a program.
What you will learn in this book
How to write an operating system from scratch by reading hardware datasheets. In the real world, you will not be able to consult Google for a quick answer.
Write code independently. It’s pointless to copy and paste code. Real learning happens when you solve problems on your own. Some examples are provided to help kick start your work, but most problems are yours to conquer. However, the solutions are available online for you after giving a good try.
A big picture of how each layer of a computer related to each other, from hardware to software.
How to use Linux as a development environment and common tools for low-level programming.
How a program is structured so that an operating system can run.
How to debug a program running directly on hardware with
gdband QEMU.Linking and loading on bare metal 32-bit x86, with pure C. No standard library. No runtime overhead. Chapter 17 explains what changes in 64-bit mode.
What this book is not about
Electrical Engineering: The book discusses some concepts from electronics and electrical engineering only to the extent of how software operates on bare metal.
How to use Linux or any OS types of books: Though Linux is used as a development environment and as a medium to demonstrate high-level operating system concepts, it is not the focus of this book.
Linux Kernel development: There are already many high-quality books out there on this subject.
Operating system books focused on algorithms: This book focuses more on an actual hardware platform - Intel x86 - and how to write an OS that utilizes the OS support from the hardware platform.
What is new in the second edition
The first edition was written against Ubuntu 16.04 and gcc 5.4. A few years later, readers who typed the commands in the book got different output: compilers started producing position-independent executables by default, the linker started rejecting the linker script of chapter 8, and nobody could tell whether the book or their machine was at fault. Part 3 was also never finished; it stopped at a list of chapter headings. This edition fixes both problems:
A new chapter 0 sets up the development environment. It lists the tools, explains every compiler flag the book uses, and provides a container image with pinned tool versions so that the output printed in the book is the output you get.
Every listing in chapters 4 to 8 was regenerated with those tools. The bootloader in chapter 7 now reads from a hard-disk image with the BIOS extended read service, and the linker script in chapter 8 links with a current
ld.Part 3 is written. Chapters 9 to 15 grow one kernel, chapter by chapter, from protected mode to processes that fork and exec programs loaded from an ext2 filesystem; chapter 16 is about debugging a kernel that does not boot, and chapter 17 explains what changes in 64-bit mode, under UEFI and with several processors. Every chapter ends with exercises and a few “Check your understanding” questions (answers in Appendix E), each part closes with a milestone project, and the epilogue proposes three capstone projects with a rubric. The source code of each chapter is in the repository of the book and boots in QEMU.
The book stays in 32-bit protected mode. The first edition announced x86_64 in this preface but never used it; the concepts are the same, the Intel manual chapters to read are the same, and 32-bit code is shorter to show.
The errata reported by readers over the years are applied, and the source of the book is now plain Markdown, so a correction is a one-line pull request.
The organization of the book
- Chapter 0
-
sets up the development environment: the tools, the compiler flags used throughout the book and a container image with the exact tool versions that produced the listings. It ends with a smoke test that boots the bootloader of chapter 7, so that any installation problem is found before chapter 1.
- Part 1
-
provides a foundation for learning operating system.
Chapter 1 briefly explains the importance of domain documents. Documents are crucial for the learning experience, so they deserve a chapter.
Chapter 2 explains the layers of abstractions from hardware to software. The idea is to provide insight into how code runs physically.
Chapter 3 provides the general architecture of a computer, then introduces a sample computer model that you will use to write an operating system.
Chapter 4 introduces the x86 assembly language through the use of the Intel manuals, along with commonly used instructions. This chapter gives detailed examples of how high-level syntax corresponds to low-level assembly, enabling you to read generated assembly code comfortably. It is necessary to read assembly code when debugging an operating system.
Chapter 5 dissects ELF in detail. Only by understanding how the structure of a program at the binary level, you can build one that runs on bare metal.
Chapter 6 introduces
gdbdebugger with extensive examples for commonly used commands. After acquainting the reader withgdb, it then provides insight on how a debugger works. This knowledge is essential for building a debuggable program on the bare metal.
- Part 2
-
presents how to write a bootloader to bootstrap a kernel. Hence the name “Groundwork”. After mastering this part, the reader can continue with the next part, which is a guide for writing an operating system. However, if the reader does not like the presentation, he or she can look elsewhere, such as OSDev Wiki: http://wiki.osdev.org/.
Chapter 7 introduces what the bootloader is, how to write one in assembly, and how to load it on QEMU, a hardware emulator. This process involves typing repetitive and long commands, so GNU Make is applied to improve productivity by automating the repetitive parts and simplifying the interaction with the project. This chapter also demonstrates the use of GNU Make in context.
Chapter 8 introduces linking by explaining the relocation process when combining object files. In addition to a bootloader and an operating system written in C, this is the last piece of the puzzle required for building debuggable programs on bare metal, including the bootloader written in Assembly and an operating system written in C.
- Part 3
-
provides guidance on how to write an operating system, as you should implement an operating system on your own and be proud of your creation. The guidance consists of simpler and coherent explanations of necessary concepts, from hardware to software, to implement the features of an operating system. Without such guidance, you will waste time gathering information spread through various documents and the Internet. Each chapter names the sections of the Intel manuals to read, adds one feature to the kernel started in Part 2, shows how to debug it with
gdband QEMU, and ends with exercises.
Chapter 9 enters protected mode for good: segment descriptors, the GDT and the segment registers as the CPU really uses them. This is where the kernel gets its own memory layout.
Chapter 10 talks to the first devices, the serial port and the VGA text buffer, through I/O ports, and writes a freestanding
printf. Everything after this chapter is debugged with the output of this one.Chapter 11 introduces interrupts: the IDT, CPU exceptions, the programmable interrupt controller, the timer and the keyboard.
Chapter 12 manages memory: the memory map reported by the BIOS, a physical frame allocator, paging, page faults and a kernel heap.
Chapter 13 creates processes: context switching, a scheduler, the transition to user mode with the TSS, and system calls.
Chapter 14 reads an ext2 file system from an ATA disk and loads a user program from it, which turns the kernel into something that deserves the name of operating system.
Chapter 15 gives each process its own address space and adds the four system calls of the Unix process model:
forkwith copy-on-write,exec,waitandexit. The kernel now starts a singleinitprogram, and pages of the.bssand the stack are allocated on demand by the page-fault handler.Chapter 16 is about the kernel that does not boot. It gives a method, a table of symptoms with the tool that confirms each one, and hunts three bugs planted on purpose, with the real gdb and QEMU transcripts.
Chapter 17 is an epilogue. It explains what changes in 64-bit mode, under UEFI firmware and with several processors, and where to go next.
- Appendices
-
Appendix A is a reference of the toolchain and the compiler flags, and of the editions of the manuals the book cites. Appendix B lists the errata of the first edition. Appendix C is a minimal, tested long-mode bootstrap. Appendix D is a reading order for the Intel and AMD manuals, chapter by chapter. Appendix E holds the answers to the “Check your understanding” questions. A glossary closes the book. (The GNU Make and
gdbscript summaries stay in chapter 7, where they are first needed.)
Acknowledgments
Thank you, my beloved family. Thank you, the contributors.