5 The Anatomy of a Program

Every program consists of code and data, and only those two components make up a program. However, if a program consisted purely of code and data of its own, then from the perspective of an operating system (as well as of a human), there would be no way to know which block of binary is a program and which is just raw data, where in the program to start execution, which region of memory should be protected and which is free to modify. For that reason, each program carries extra metadata to communicate with the operating system how to handle the program.

When a source file is compiled, the generated machine code is stored into an object file, which is just a block of binary. One or more object files can be combined to produce an executable binary, which is a complete program runnable in an operating system.

readelf is a program that recognizes and displays the ELF metadata of a binary file, be it an object file or an executable binary. ELF, or Executable and Linkable Format, is the content at the very beginning of an executable to provide an operating system the necessary information to load the executable into main memory and run it. ELF can be thought of as similar to the table of contents of a book. In a book, a table of contents lists the page numbers of the main sections, subsections, sometimes even figures and tables for easy lookup. Similarly, ELF lists various sections used for code and data, and the memory addresses of each symbol along with other information.

An ELF binary is composed of:

ELF: linking view vs. execution view (source: Wikipedia).

Later we will compile our kernel as an ELF executable with GCC, and explicitly specify how segments are created and where they are loaded in memory through the use of a linker script, a text file that instructs the linker how to generate a binary. For now, we will examine the anatomy of an ELF executable in detail.

Throughout the chapter, the executable under the microscope is the following Hello World:

hello.c

#include <stdio.h>

int main(int argc, char *argv[])
{
    printf("Hello World\n");
    return 0;
}

The listings of the first part of this chapter are of a 64-bit executable, so that you see the 64-bit variant of the format once: the structures are the same as in a 32-bit file, only the width of the addresses and the size of the headers differ. For this one program we therefore leave -m32 out of the flags of chapter 0:

$ gcc -no-pie -fno-pie -fno-asynchronous-unwind-tables -fcf-protection=none -O0 hello.c -o hello

From Example 5.7 onward, every program is compiled with gcc $BOOKFLAGS, that is, as a 32-bit executable like everywhere else in the book.

5.1 Reference documents

The ELF specification is bundled as a man page in Linux:

$ man elf

It is a useful resource to understand and implement ELF. However, it will be much easier to use after you finish this chapter, as the specification mixes implementation details in it.

The default specification is a generic one, which every ELF implementation follows. However, each platform provides extra features unique to it. The ELF specification for x86 is currently maintained on Github by H.J. Lu: https://github.com/hjl-tools/x86-psABI/wiki/X86-psABI.

Platform-dependent details are referred to as “processor specific” in the generic ELF specification. We will not explore these details, but study the generic details, which are enough for crafting an ELF binary image for our operating system.

5.2 ELF header

To see the information of an ELF header:

$ readelf -h hello

The output:

ELF Header:
  Magic:   7f 45 4c 46 02 01 01 00 00 00 00 00 00 00 00 00
  Class:                             ELF64
  Data:                              2's complement, little endian
  Version:                           1 (current)
  OS/ABI:                            UNIX - System V
  ABI Version:                       0
  Type:                              EXEC (Executable file)
  Machine:                           Advanced Micro Devices X86-64
  Version:                           0x1
  Entry point address:               0x401040
  Start of program headers:          64 (bytes into file)
  Start of section headers:          13856 (bytes into file)
  Flags:                             0x0
  Size of this header:               64 (bytes)
  Size of program headers:           56 (bytes)
  Number of program headers:         14
  Size of section headers:           64 (bytes)
  Number of section headers:         30
  Section header string table index: 29

Let’s go through each field:

Magic

Displays the raw bytes that uniquely identify a file as an ELF executable binary. Each byte gives a brief piece of information.

In the example, we have the following magic bytes:

Magic:   7f 45 4c 46 02 01 01 00 00 00 00 00 00 00 00 00

Examine byte by byte:

Byte Description
7f 45 4c 46 Predefined values. The first byte is always 7F, the remaining 3 bytes represent the string "ELF".
02 See Class field below.
01 See Data field below.
01 See Version field below.
00 See OS/ABI field below.
00 See ABI Version field below.
00 00 00 00 00 00 00 Padding bytes. These bytes are unused and are always set to 0. Padding bytes are added for proper alignment, and are reserved for future use when more information is needed.
Class

A byte in the Magic field. It specifies the class or capacity of a file.

Possible values:

Value Description
0 Invalid class
1 32-bit objects
2 64-bit objects
Data

A byte in the Magic field. It specifies the data encoding of the processor-specific data in the object file.

Possible values:

Value Description
0 Invalid data encoding
1 Little endian, 2’s complement
2 Big endian, 2’s complement
Version

A byte in Magic. It specifies the ELF header version number.

Possible values:

Value Description
0 Invalid version
1 Current version
OS/ABI

A byte in the Magic field. It specifies the target operating system ABI. Originally, it was a padding byte.

Possible values: Refer to the latest ABI document, as it is a long list of different operating systems.

ABI Version

A byte in the Magic field. It specifies the version of the ABI named by the OS/ABI byte. For UNIX - System V there is only one, so the value is 0.

Type

Identifies the object file type.

Value Description
0 No file type
1 Relocatable file
2 Executable file
3 Shared object file
4 Core file
0xff00 Processor specific, lower bound
0xffff Processor specific, upper bound

The values from 0xff00 to 0xffff are reserved for a processor to define additional file types meaningful to it.

Machine

Specifies the required architecture value for an ELF file e.g. x86_64, MIPS, SPARC, etc. In the example, the machine is of x86_64 architecture.

Possible values: Please refer to the latest ABI document, as it is a long list of different architectures.

Version

Specifies the version number of the current object file (not the version of the ELF header, as the above Version field specified).

Entry point address

Specifies the memory address where the very first code is executed. In a normal application program the entry point is not main but _start, a small routine supplied by the C library that prepares the environment and then calls main; in the example the entry point is 0x401040, and later in this chapter the symbol table shows _start at exactly that address. It can be any function by explicitly specifying the function name to the linker (-e <name>). For the operating system we are going to write, this is the single most important field that we need to retrieve to bootstrap our kernel, and everything else can be ignored.

Start of program headers

The offset of the program header table, in bytes. In the example, this number is 64 bytes, which means the 65th byte, or <start address> + 64, is the start address of the program header table. That is, if a program is loaded at address 0x10000 in memory, then the start address is 0x10000 (the very first byte of the Magic field, where the value 0x7f resides) and the start address of the program header table is 0x10000 + 0x40 = 0x10040.

Start of section headers

The offset of the section header table in bytes, similar to the start of program headers. In the example, it is 13856 bytes into the file.

Flags

Holds processor-specific flags associated with the file. No flag is defined for x86, so the value is always 0x0, as in the example. Other architectures use this field to record, for example, which variant of the instruction set or of the ABI the file was compiled for.

Size of this header

Specifies the total size of the ELF header in bytes. In the example, it is 64 bytes, which is equivalent to Start of program headers. Note that these two numbers are not necessarily equivalent, as the program header table might be placed far away from the ELF header. The only fixed component in the ELF executable binary is the ELF header, which appears at the very beginning of the file.

Size of program headers

Specifies the size of each program header in bytes. In the example, it is 56 bytes.

Number of program headers

Specifies the total number of program headers. In the example, the file has a total of 14 program headers.

Size of section headers

Specifies the size of each section header in bytes. In the example, it is 64 bytes.

Number of section headers

Specifies the total number of section headers. In the example, the file has a total of 30 section headers. In a section header table, the first entry in the table is always an empty section.

Section header string table index

Specifies the index of the header in the section header table that points to the section that holds all the null-terminated section names. In the example, the index is 29, which means it’s the entry at index 29 of the table (counting from 0).

5.3 Section header table

As we know already, code and data compose a program. However, not all types of code and data have the same purpose. For that reason, instead of a big chunk of code and data, they are divided into smaller chunks, and each chunk must satisfy these conditions (according to gABI):

To get all the headers from an executable binary e.g. hello, use the following command:

$ readelf -S hello

Here is a sample output (do not worry if you don’t understand the output. Just skim to get your eyes familiar with it. We will dissect it soon enough):

There are 30 section headers, starting at offset 0x3620:

Section Headers:
  [Nr] Name              Type             Address           Offset
       Size              EntSize          Flags  Link  Info  Align
  [ 0]                   NULL             0000000000000000  00000000
       0000000000000000  0000000000000000           0     0     0
  [ 1] .note.gnu.pr[...] NOTE             0000000000400350  00000350
       0000000000000020  0000000000000000   A       0     0     8
  [ 2] .note.gnu.bu[...] NOTE             0000000000400370  00000370
       0000000000000024  0000000000000000   A       0     0     4
  [ 3] .interp           PROGBITS         0000000000400394  00000394
       000000000000001c  0000000000000000   A       0     0     1
  [ 4] .gnu.hash         GNU_HASH         00000000004003b0  000003b0
       000000000000001c  0000000000000000   A       5     0     8
  [ 5] .dynsym           DYNSYM           00000000004003d0  000003d0
       0000000000000060  0000000000000018   A       6     1     8
  [ 6] .dynstr           STRTAB           0000000000400430  00000430
       0000000000000048  0000000000000000   A       0     0     1
  [ 7] .gnu.version      VERSYM           0000000000400478  00000478
       0000000000000008  0000000000000002   A       5     0     2
  [ 8] .gnu.version_r    VERNEED          0000000000400480  00000480
       0000000000000030  0000000000000000   A       6     1     8
  [ 9] .rela.dyn         RELA             00000000004004b0  000004b0
       0000000000000030  0000000000000018   A       5     0     8
  [10] .rela.plt         RELA             00000000004004e0  000004e0
       0000000000000018  0000000000000018  AI       5    23     8
  [11] .init             PROGBITS         0000000000401000  00001000
       0000000000000017  0000000000000000  AX       0     0     4
  [12] .plt              PROGBITS         0000000000401020  00001020
       0000000000000020  0000000000000010  AX       0     0     16
  [13] .text             PROGBITS         0000000000401040  00001040
       0000000000000106  0000000000000000  AX       0     0     16
  [14] .fini             PROGBITS         0000000000401148  00001148
       0000000000000009  0000000000000000  AX       0     0     4
  [15] .rodata           PROGBITS         0000000000402000  00002000
       0000000000000010  0000000000000000   A       0     0     4
  [16] .eh_frame_hdr     PROGBITS         0000000000402010  00002010
       0000000000000024  0000000000000000   A       0     0     4
  [17] .eh_frame         PROGBITS         0000000000402038  00002038
       0000000000000080  0000000000000000   A       0     0     8
  [18] .note.ABI-tag     NOTE             00000000004020b8  000020b8
       0000000000000020  0000000000000000   A       0     0     4
  [19] .init_array       INIT_ARRAY       0000000000403df8  00002df8
       0000000000000008  0000000000000008  WA       0     0     8
  [20] .fini_array       FINI_ARRAY       0000000000403e00  00002e00
       0000000000000008  0000000000000008  WA       0     0     8
  [21] .dynamic          DYNAMIC          0000000000403e08  00002e08
       00000000000001d0  0000000000000010  WA       6     0     8
  [22] .got              PROGBITS         0000000000403fd8  00002fd8
       0000000000000010  0000000000000008  WA       0     0     8
  [23] .got.plt          PROGBITS         0000000000403fe8  00002fe8
       0000000000000020  0000000000000008  WA       0     0     8
  [24] .data             PROGBITS         0000000000404008  00003008
       0000000000000010  0000000000000000  WA       0     0     8
  [25] .bss              NOBITS           0000000000404018  00003018
       0000000000000008  0000000000000000  WA       0     0     1
  [26] .comment          PROGBITS         0000000000000000  00003018
       000000000000001f  0000000000000001  MS       0     0     1
  [27] .symtab           SYMTAB           0000000000000000  00003038
       0000000000000330  0000000000000018          28    18     8
  [28] .strtab           STRTAB           0000000000000000  00003368
       00000000000001a1  0000000000000000           0     0     1
  [29] .shstrtab         STRTAB           0000000000000000  00003509
       0000000000000116  0000000000000000           0     0     1
Key to Flags:
  W (write), A (alloc), X (execute), M (merge), S (strings), I (info),
  L (link order), O (extra OS processing required), G (group), T (TLS),
  C (compressed), x (unknown), o (OS specific), E (exclude),
  D (mbind), l (large), p (processor specific)

readelf keeps this table 80 columns wide and abbreviates a section name that does not fit: .note.gnu.pr[...] is .note.gnu.property and .note.gnu.bu[...] is .note.gnu.build-id. The option -W (wide) prints every name in full, at the price of one long line per section; we will use it later for the symbol tables.

The first line:

There are 30 section headers, starting at offset 0x3620

summarizes the total number of sections in the file, and the offset in the file where the table starts. Then comes the listing section by section with the following header, which is also the format of each section output:

[Nr] Name              Type             Address           Offset
     Size              EntSize          Flags  Link  Info  Align

Each section has two lines with different fields:

Nr

The index of each section.

Name

The name of each section.

Type

This field (in a section header) identifies the type of each section. Types are used to classify sections.

Address

The starting virtual address of each section. Note that the addresses are virtual only when a program runs in an OS with support for virtual memory enabled. In our OS, we run on the bare metal, so the addresses will all be physical.

Offset

is a distance in bytes, from the first byte of a file to the start of an object, such as a section or a segment in the context of an ELF binary file.

Size

The size in bytes of each section.

EntSize

Some sections hold a table of fixed-size entries, such as a symbol table. For such a section, this member gives the size in bytes of each entry. The member contains 0 if the section does not hold a table of fixed-size entries.

Flags

describes attributes of a section. Flags together with a type define the purpose of a section. Two sections can be of the same type, but serve different purposes. For example, even though .data and .text share the same type, .data holds the initialized data of a program while .text holds the executable instructions of a program. For that reason, .data is given read and write permission, but not execute. Any attempt to execute code in .data is denied by the running OS: in Linux, such invalid section usage gives a segmentation fault.

ELF gives information to enable an OS with such a protection mechanism. However, running on bare metal, nothing can prevent us from doing anything. Our OS can execute code in the data section, and vice versa, write to the code section.

Section Flags
Flag Description
W Bytes in this section are writable during execution.
A Memory is allocated for this section during process execution. Some control sections do not reside in the memory image of an object file; this attribute is off for those sections.
X The section contains executable instructions.
M The data in the section may be merged to eliminate duplication. Each element in the section is compared against other elements in sections with the same name, type and flags. Elements that would have identical values at program run-time may be merged.
S The data elements in the section consist of null-terminated character strings. The size of each character is specified in the section header’s EntSize field.
I The Info field of this section header holds an index of a section header. Otherwise, the number is the index of something else.
L Preserve section ordering when linking. If this section is combined with other sections in the output file, it must appear in the same relative order with respect to those sections, as the linked-to section appears with respect to sections the linked-to section is combined with. Applies when the Link field of this section’s header references another section (the linked-to section).
O This section requires special OS-specific processing (beyond the standard linking rules) to avoid incorrect behavior. If a link editor encounters sections whose headers contain OS-specific values it does not recognize by Type or Flags values defined by the ELF standard, the link editor should combine those sections.
G This section is a member (perhaps the only one) of a section group.
T This section holds Thread-Local Storage, meaning that each thread has its own distinct instance of this data. A thread is a distinct execution flow of code. A program can have multiple threads that pack different pieces of code and execute separately, at the same time. We will learn more about threads when writing our kernel.
C The section data is compressed. Compilers use it for debugging sections, which can be large.
x Unknown flag to readelf. It happens because the linking process can be done manually with a linker like GNU ld (we will do so later). That is, section flags can be specified manually, and some flags are for a customized ELF that the open-source readelf doesn’t know of.
o All bits included in this flag are reserved for operating system-specific semantics.
E The link editor is to exclude this section from the executable and shared library that it builds when those objects are not to be further relocated.
D A GNU extension: the section is to be placed in memory with special attributes, set with the mbind system call.
l Specific large section for the x86_64 architecture. This flag is not specified in the Generic ABI but in the x86_64 ABI.
p All bits included in this flag are reserved for processor-specific semantics. If meanings are specified, the processor supplement explains them.
Link and Info

are numbers that reference the indexes of sections, symbol table entries, hash table entries. The Link field only holds the index of a section, while the Info field holds an index of a section, a symbol table entry or a hash table entry, depending on the type of a section.

Later when writing our OS, we will handcraft the kernel image by explicitly linking the object files (produced by gcc) through a linker script. We will specify the memory layout of sections by specifying at what addresses they will appear in the final image. But we will not assign any section flag and let the linker take care of it. Nevertheless, knowing which flag does what is useful.

Align

is a value that enforces that the offset of a section should be divisible by the value. Only 0 and positive integral powers of two are allowed. Values 0 and 1 mean the section has no alignment constraint.

Example 5.1. Output of the .interp section:

  [Nr] Name              Type             Address           Offset
       Size              EntSize          Flags  Link  Info  Align
  [ 3] .interp           PROGBITS         0000000000400394  00000394
       000000000000001c  0000000000000000   A       0     0     1

Nr is 3.

Type is PROGBITS, which means this section is part of the program.

Address is 0x0000000000400394, which means the section is loaded at this virtual memory address at runtime.

Offset is 0x00000394 bytes into the file.

Size is 0x000000000000001c in bytes.

EntSize is 0, which means this section does not have any fixed-size entry.

Flags are A (Allocatable), which means this section consumes memory at runtime.

Link and Info are 0 and 0, which means this section links to no section or entry in any table.

Align is 1, which means no alignment.

Example 5.2. Output of the .text section:

  [Nr] Name              Type             Address           Offset
       Size              EntSize          Flags  Link  Info  Align
  [13] .text             PROGBITS         0000000000401040  00001040
       0000000000000106  0000000000000000  AX       0     0     16

Nr is 13.

Type is PROGBITS, which means this section is part of the program.

Address is 0x0000000000401040, which means the section is loaded at this virtual memory address at runtime.

Offset is 0x00001040 bytes into the file.

Size is 0x0000000000000106 in bytes.

EntSize is 0, which means this section does not have any fixed-size entry.

Flags are A (Allocatable) and X (Executable), which means this section consumes memory and can be executed as code at runtime.

Link and Info are 0 and 0, which means this section links to no section or entry in any table.

Align is 16, which means the starting address of the section should be divisible by 16, or 0x10. Indeed, it is: 0x1040/0x10=0x104.

5.4 Understand Section in-depth

In this section, we will learn different details of section types and the purposes of special sections e.g. .bss, .text, .data, etc, by looking at each section one by one. We will also examine the content of each section as a hexdump with the commands:

$ readelf -x <section name|section number> <file>

For example, if you want to examine the content of the section with index 24 (the .data section in the sample output) in the file hello:

$ readelf -x 24 hello

Equivalently, using the name instead of the index works:

$ readelf -x .data hello

If a section contains strings e.g. a string symbol table, the flag -x can be replaced with -p.

NULL

marks a section header as inactive and does not have an associated section. The NULL section is always the first entry of the section header table. It means, any useful section starts from 1.

Example 5.3. The sample output of the NULL section:

  [Nr] Name              Type             Address           Offset
       Size              EntSize          Flags  Link  Info  Align
  [ 0]                   NULL             0000000000000000  00000000
       0000000000000000  0000000000000000           0     0     0

Examining the content, the section is empty:

$ readelf -x 0 hello
Section '' has no data to dump.
NOTE

marks a section with special information that other programs will check for conformance, compatibility, etc, by a vendor or a system builder.

Example 5.4. In the sample output, we have 3 NOTE sections:

  [Nr] Name              Type             Address           Offset
       Size              EntSize          Flags  Link  Info  Align
  [ 1] .note.gnu.pr[...] NOTE             0000000000400350  00000350
       0000000000000020  0000000000000000   A       0     0     8
  [ 2] .note.gnu.bu[...] NOTE             0000000000400370  00000370
       0000000000000024  0000000000000000   A       0     0     4
...
  [18] .note.ABI-tag     NOTE             00000000004020b8  000020b8
       0000000000000020  0000000000000000   A       0     0     4

Examine the third one, .note.ABI-tag, by index or by name:

$ readelf -x 18 hello

we have:

Hex dump of section '.note.ABI-tag':
  0x004020b8 04000000 10000000 01000000 474e5500 ............GNU.
  0x004020c8 00000000 03000000 02000000 00000000 ................

A note is a name, a type and a payload. Here the name is GNU (the bytes 47 4e 55 00), the type 01 00 00 00 means “ABI tag”, and the payload 00 00 00 00 03 00 00 00 02 00 00 00 00 00 00 00 reads, as four little-endian integers: operating system 0 (Linux), minimum kernel version 3.2.0.

PROGBITS

indicates a section holding the main content of a program, either code or data.

Example 5.5. There are many PROGBITS sections:

  [Nr] Name              Type             Address           Offset
       Size              EntSize          Flags  Link  Info  Align
  [ 3] .interp           PROGBITS         0000000000400394  00000394
       000000000000001c  0000000000000000   A       0     0     1
...
  [11] .init             PROGBITS         0000000000401000  00001000
       0000000000000017  0000000000000000  AX       0     0     4
  [12] .plt              PROGBITS         0000000000401020  00001020
       0000000000000020  0000000000000010  AX       0     0     16
  [13] .text             PROGBITS         0000000000401040  00001040
       0000000000000106  0000000000000000  AX       0     0     16
  [14] .fini             PROGBITS         0000000000401148  00001148
       0000000000000009  0000000000000000  AX       0     0     4
  [15] .rodata           PROGBITS         0000000000402000  00002000
       0000000000000010  0000000000000000   A       0     0     4
  [16] .eh_frame_hdr     PROGBITS         0000000000402010  00002010
       0000000000000024  0000000000000000   A       0     0     4
  [17] .eh_frame         PROGBITS         0000000000402038  00002038
       0000000000000080  0000000000000000   A       0     0     8
...
  [22] .got              PROGBITS         0000000000403fd8  00002fd8
       0000000000000010  0000000000000008  WA       0     0     8
  [23] .got.plt          PROGBITS         0000000000403fe8  00002fe8
       0000000000000020  0000000000000008  WA       0     0     8
  [24] .data             PROGBITS         0000000000404008  00003008
       0000000000000010  0000000000000000  WA       0     0     8
  [26] .comment          PROGBITS         0000000000000000  00003018
       000000000000001f  0000000000000001  MS       0     0     1

For our operating system, we only need the following sections:

.text

This section holds all the compiled code of a program.

.data

This section holds the initialized data of a program. Since the data are initialized with actual values, gcc allocates the section with actual bytes in the executable binary.

.rodata

This section holds read-only data, such as fixed-size strings in a program, e.g. “Hello World”, and others.

.bss

This section, short for Block Started by Symbol, holds the uninitialized data of a program. Unlike other sections, no space is allocated for this section in the image of the executable binary on disk. The section is allocated only when the program is loaded into main memory.

Other sections are mainly needed for dynamic linking, that is code linking at runtime for sharing between many programs. To enable such a feature, an OS as a runtime environment must be present. Since we run our OS on bare metal, we are effectively creating such an environment. For simplicity, we won’t add dynamic linking to our OS.

SYMTAB and DYNSYM

These sections hold a symbol table. A symbol table is an array of entries that describe symbols in a program. A symbol is a name assigned to an entity in a program. The types of these entities are also the types of symbols, and are listed under the Type field below.

Example 5.6. In the sample output, sections 5 and 27 are symbol tables:

  [Nr] Name              Type             Address           Offset
       Size              EntSize          Flags  Link  Info  Align
  [ 5] .dynsym           DYNSYM           00000000004003d0  000003d0
       0000000000000060  0000000000000018   A       6     1     8
...
  [27] .symtab           SYMTAB           0000000000000000  00003038
       0000000000000330  0000000000000018          28    18     8

To show the symbol table:

$ readelf -W -s hello

(-W keeps readelf from abbreviating long symbol names such as __libc_start_main@GLIBC_2.34 to _[...]@GLIBC_2.34.) The output consists of 2 symbol tables, corresponding to the two sections above, .dynsym and .symtab:

Symbol table '.dynsym' contains 4 entries:
   Num:    Value          Size Type    Bind   Vis      Ndx Name
     0: 0000000000000000     0 NOTYPE  LOCAL  DEFAULT  UND
     1: 0000000000000000     0 FUNC    GLOBAL DEFAULT  UND __libc_start_main@GLIBC_2.34 (2)
     2: 0000000000000000     0 FUNC    GLOBAL DEFAULT  UND puts@GLIBC_2.2.5 (3)
     3: 0000000000000000     0 NOTYPE  WEAK   DEFAULT  UND __gmon_start__

Symbol table '.symtab' contains 34 entries:
   Num:    Value          Size Type    Bind   Vis      Ndx Name
......... output omitted .........
    27: 0000000000404020     0 NOTYPE  GLOBAL DEFAULT   25 _end
    28: 0000000000401070     1 FUNC    GLOBAL HIDDEN    13 _dl_relocate_static_pie
    29: 0000000000401040    34 FUNC    GLOBAL DEFAULT   13 _start
    30: 0000000000404018     0 NOTYPE  GLOBAL DEFAULT   25 __bss_start
    31: 0000000000401126    32 FUNC    GLOBAL DEFAULT   13 main
    32: 0000000000404018     0 OBJECT  GLOBAL HIDDEN    24 __TMC_END__
    33: 0000000000401000     0 FUNC    GLOBAL HIDDEN    11 _init
Num

is the index of an entry in a table.

Value

is the virtual memory address where the symbol is located.

Size

is the size of the entity associated with a symbol.

Type

is a symbol type according to this table:

Symbol Types
Type Description
NOTYPE The type of the symbol is not specified.
OBJECT The symbol is associated with a data object. In C, any variable definition is of OBJECT type.
FUNC The symbol is associated with a function or other executable code.
SECTION The symbol is associated with a section, and exists primarily for relocation.
FILE The symbol is the name of a source file associated with an executable binary.
COMMON The symbol labels an uninitialized variable. That is, when a variable in C is defined as a global variable without an initial value, or as an external variable using the extern keyword. In other words, these variables stay in the .bss section.
TLS The symbol is associated with a Thread-Local Storage entity.
Bind

is the scope of a symbol.

LOCAL

are symbols that are only visible in the object files that defined them. In C, the static modifier marks a symbol (e.g. a variable/function) as local to only the file that defines it.

Example 5.7. If we define variables and functions with the static modifier:

hello.c

static int global_static_var = 0;

static void local_func() {
}

int main(int argc, char *argv[])
{
    static int local_static_var = 0;

    return 0;
}

Then we get the static variables listed as local symbols after compiling:

$ gcc $BOOKFLAGS hello.c -o hello
$ readelf -W -s hello

Symbol table '.dynsym' contains 4 entries:
   Num:    Value  Size Type    Bind   Vis      Ndx Name
     0: 00000000     0 NOTYPE  LOCAL  DEFAULT  UND
     1: 00000000     0 FUNC    GLOBAL DEFAULT  UND __libc_start_main@GLIBC_2.34 (2)
     2: 00000000     0 NOTYPE  WEAK   DEFAULT  UND __gmon_start__
     3: 0804a004     4 OBJECT  GLOBAL DEFAULT   14 _IO_stdin_used

Symbol table '.symtab' contains 39 entries:
   Num:    Value  Size Type    Bind   Vis      Ndx Name
     0: 00000000     0 NOTYPE  LOCAL  DEFAULT  UND
......... output omitted .........
    13: 0804c010     4 OBJECT  LOCAL  DEFAULT   24 global_static_var
    14: 08049156     6 FUNC    LOCAL  DEFAULT   12 local_func
    15: 0804c014     4 OBJECT  LOCAL  DEFAULT   24 local_static_var.0
......... output omitted .........

Notice the suffix .0 that gcc appended to local_static_var: a static variable local to a function has no name visible outside that function, so gcc numbers it to keep two functions that each declare a local_static_var apart.

GLOBAL

are symbols that are accessible by other object files when linking together. These symbols are primarily non-static functions and non-static global data. The extern modifier marks a symbol as externally defined elsewhere but accessible in the final executable binary, so an extern variable is also considered GLOBAL.

Example 5.8. Similar to the LOCAL example above, the output lists many GLOBAL symbols such as main:

   Num:    Value  Size Type    Bind   Vis      Ndx Name
......... output omitted .........
    36: 0804915c    10 FUNC    GLOBAL DEFAULT   12 main
......... output omitted .........
WEAK

are symbols whose definitions can be redefined. Normally, a symbol with multiple definitions is reported as an error by a compiler. However, this constraint is lax when a definition is explicitly marked as weak, which means the default implementation can be replaced by a different definition at link time.

Example 5.9. Suppose we have a default implementation of the function add:

hello.c

#include <stdio.h>

__attribute__((weak)) int add(int a, int b) {
    printf("warning: function is not implemented.\n");
    return 0;
}

int main(int argc, char *argv[])
{
    printf("add(1,2) is %d\n", add(1,2));
    return 0;
}

__attribute__((weak)) is a function attribute. A function attribute is extra information for a compiler to handle a function differently from a normal function. In this example, the weak attribute makes the function add a weak function, which means the default implementation can be replaced by a different definition at link time. Function attributes are a feature of a compiler, not standard C.

If we do not supply a different function definition in a different file (must be in a different file, otherwise gcc reports an error), then the default implementation is applied. When the function add is called, it only prints the message "warning: function is not implemented." and returns 0:

$ gcc $BOOKFLAGS hello.c -o hello
$ ./hello
warning: function is not implemented.
add(1,2) is 0

However, if we supply a different definition in another file e.g. math.c:

math.c

int add(int a, int b) {
    return a + b;
}

and compile the two files together:

$ gcc $BOOKFLAGS math.c hello.c -o hello
$ ./hello
add(1,2) is 3

Then, when running hello, no warning message is printed and the correct value is returned.

A weak symbol is a mechanism to provide a default implementation, replaceable when a better implementation is available (e.g. more specialized and optimized) at link-time.

Vis

is the visibility of a symbol. The following values are available:

Symbol Visibility
Value Description
DEFAULT

The visibility is specified by the binding type of a symbol.

  • Global and weak symbols are visible outside of their defining component (executable file or shared object).

  • Local symbols are hidden. See HIDDEN below.

HIDDEN A symbol is hidden when the name is not visible to any other program outside of its running program.
PROTECTED A symbol is protected when it is shared outside of its running program or shared library and cannot be overridden. That is, there can only be one definition for this symbol across running programs that use it. No program can define its own definition of the same symbol.
INTERNAL Visibility is processor-specific and is defined by the processor-specific ABI.
Ndx

is the index of the section that the symbol is in. Aside from fixed index numbers that represent section indexes, the index has these special values:

Symbol Index
Value Description
ABS The index will not be changed by any symbol relocation.
COM The index refers to an unallocated common block.
UND The symbol is undefined in the current object file, which means the symbol depends on the actual definition in another file. Undefined symbols appear when the object file refers to symbols that are available at runtime, from a shared library.

LORESERVE

HIRESERVE

LORESERVE is the lower boundary of the reserved indexes. Its value is 0xff00.

HIRESERVE is the upper boundary of the reserved indexes. Its value is 0xffff.

The operating system reserves exclusive indexes between LORESERVE and HIRESERVE, which do not map to any actual section header.

XINDEX The index is larger than LORESERVE. The actual value will be contained in the section SYMTAB_SHNDX, where each entry is a mapping between a symbol, whose Ndx field is a XINDEX value, and the actual index value.
Others Sometimes, values such as ANSI_COM, LARGE_COM, SCOM, SUND appear. This means that the index is processor-specific.
Name

is the symbol name.

Example 5.10. A C application program always starts from the symbol main. The entry for main in the symbol table in the .symtab section is:

   Num:    Value          Size Type    Bind   Vis      Ndx Name
    31: 0000000000401126    32 FUNC    GLOBAL DEFAULT   13 main

The entry shows that:

  • main is the 31st entry in the table.

  • main starts at address 0x0000000000401126.

  • main consumes 32 bytes.

  • main is a function.

  • main is in global scope.

  • main is visible to other object files that use it.

  • main is inside the 13th section, which is .text. This is logical, since .text holds all program code.

STRTAB

holds a table of null-terminated strings, called a string table. The first and last byte of this section is always a NULL character. A string table section exists because a string can be reused by more than one section to represent symbol and section names, so a program like readelf or objdump can display various objects in a program, e.g. variables, functions, section names, in human-readable text instead of its raw hex address.

Example 5.11. In the sample output, sections 6, 28 and 29 are of STRTAB type:

  [Nr] Name              Type             Address           Offset
       Size              EntSize          Flags  Link  Info  Align
  [ 6] .dynstr           STRTAB           0000000000400430  00000430
       0000000000000048  0000000000000000   A       0     0     1
...
  [28] .strtab           STRTAB           0000000000000000  00003368
       00000000000001a1  0000000000000000           0     0     1
  [29] .shstrtab         STRTAB           0000000000000000  00003509
       0000000000000116  0000000000000000           0     0     1
.dynstr

holds the names of the symbols in .dynsym, the ones resolved at runtime by the dynamic linker.

.shstrtab

holds all the section names.

.strtab

holds the symbols e.g. variable names, function names, struct names, etc., in a C program, but not fixed-size null-terminated C strings; the C strings are kept in the .rodata section.

Example 5.12. Strings in those sections can be inspected with the command:

$ readelf -p 29 hello

The output shows all the section names, with the offset (also the string index) into .shstrtab to the left:

String dump of section '.shstrtab':
  [     1]  .symtab
  [     9]  .strtab
  [    11]  .shstrtab
  [    1b]  .note.gnu.property
  [    2e]  .note.gnu.build-id
  [    41]  .interp
  [    49]  .gnu.hash
  [    53]  .dynsym
  [    5b]  .dynstr
  [    63]  .gnu.version
  [    70]  .gnu.version_r
  [    7f]  .rela.dyn
  [    89]  .rela.plt
  [    93]  .init
  [    99]  .text
  [    9f]  .fini
  [    a5]  .rodata
  [    ad]  .eh_frame_hdr
  [    bb]  .eh_frame
  [    c5]  .note.ABI-tag
  [    d3]  .init_array
  [    df]  .fini_array
  [    eb]  .dynamic
  [    f4]  .got
  [    f9]  .got.plt
  [   102]  .data
  [   108]  .bss
  [   10d]  .comment

The actual implementation of a string table is a contiguous array of null-terminated strings. The index of a string is the position of its first character in the array. For example, in the above string table, .symtab is at index 1 in the array (the NULL character is at index 0). The length of .symtab is 7, plus the NULL character, which takes 8 bytes in total. So, .strtab starts at index 9, .shstrtab at index 0x11, and so on:

String table in memory of .shstrtab. A character in bold is the first character of a string; its column number plus the row offset is the index of that string: .symtab at 0x01, .strtab at 0x09, .shstrtab at 0x11 and .note.gnu.property at 0x1b.
Offset 00 01 02 03 04 05 06 07 08 09 0a 0b 0c 0d 0e 0f
00000000 \0 . s y m t a b \0 . s t r t a b
00000010 \0 . s h s t r t a b \0 . n o t e
… and so on

Similarly, the output of .strtab:

String dump of section '.strtab':
  [     1]  crt1.o
  [     8]  __abi_tag
  [    12]  crtstuff.c
  [    1d]  deregister_tm_clones
  [    32]  __do_global_dtors_aux
  [    48]  completed.0
  [    54]  __do_global_dtors_aux_fini_array_entry
  [    7b]  frame_dummy
  [    87]  __frame_dummy_init_array_entry
  [    a6]  hello.c
  [    ae]  __FRAME_END__
  [    bc]  _DYNAMIC
  [    c5]  __GNU_EH_FRAME_HDR
  [    d8]  _GLOBAL_OFFSET_TABLE_
  [    ee]  __libc_start_main@GLIBC_2.34
  [   10b]  puts@GLIBC_2.2.5
  [   11c]  _edata
  [   123]  _fini
  [   129]  __data_start
  [   136]  __gmon_start__
  [   145]  __dso_handle
  [   152]  _IO_stdin_used
  [   161]  _end
  [   166]  _dl_relocate_static_pie
  [   17e]  __bss_start
  [   18a]  main
  [   18f]  __TMC_END__
  [   19b]  _init

Among the names of the C library’s start-up code, you can recognize the two that come from our source: hello.c and main.

HASH and GNU_HASH

hold a symbol hash table, which supports symbol table access.

DYNAMIC

holds information for dynamic linking.

NOBITS

is similar to PROGBITS but occupies no space.

Example 5.13. The .bss section holds uninitialized data, which means the bytes in the section can have any value. Until an operating system actually loads the section into main memory, there is no need to allocate space for it in the binary image on disk, which reduces the size of the binary file. Here are the details of .bss from the example output:

  [Nr] Name              Type             Address           Offset
       Size              EntSize          Flags  Link  Info  Align
  [25] .bss              NOBITS           0000000000404018  00003018
       0000000000000008  0000000000000000  WA       0     0     1
  [26] .comment          PROGBITS         0000000000000000  00003018
       000000000000001f  0000000000000001  MS       0     0     1

In the above output, the size of the .bss section is 0x8, 8 bytes, while the offsets of both .bss and the section that follows it, .comment, are the same, 0x3018. That is, .bss consumes no byte of the executable binary on disk.

Notice that the .comment section has no starting address. This means that this section is discarded when the executable binary is loaded into memory.

REL

holds relocation entries without explicit addends. This type will be explained in detail in chapter 8, Linking and loading on bare metal.

RELA

holds relocation entries with explicit addends. This type will be explained in detail in chapter 8, Linking and loading on bare metal.

INIT_ARRAY

is an array of function pointers for program initialization. When an application program runs, before getting to main(), initialization code in .init and this section are executed first. The first element in this array is an ignored function pointer.

It might not make sense when we can include initialization code in the main() function. However, for shared object files where there is no main(), this section ensures that the initialization code from an object file executes before any other code to ensure a proper environment for the main code to run properly. It also makes an object file more modular, as the main application code need not be responsible for initializing a proper environment for using a particular object file, but the object file itself. Such a clear division makes code cleaner.

However, we will not use any .init and INIT_ARRAY sections in our operating system, for simplicity, as initializing an environment is part of the operating-system domain.

Example 5.14. To use the INIT_ARRAY, we simply mark a function with the attribute constructor:

hello.c

#include <stdio.h>

__attribute__((constructor)) static void init1(){
    printf("%s\n", __FUNCTION__);
}

__attribute__((constructor)) static void init2(){
    printf("%s\n", __FUNCTION__);
}

int main(int argc, char *argv[])
{
    printf("hello world\n");

    return 0;
}

The program automatically calls the constructors without explicitly invoking them:

$ gcc $BOOKFLAGS hello.c -o hello
$ ./hello
init1
init2
hello world

Example 5.15. Optionally, a constructor can be assigned a priority from 101 onward. The priorities from 0 to 100 are reserved for gcc. If we want init2 to run before init1, we give it a higher priority:

hello.c

#include <stdio.h>

__attribute__((constructor(102))) static void init1(){
    printf("%s\n", __FUNCTION__);
}

__attribute__((constructor(101))) static void init2(){
    printf("%s\n", __FUNCTION__);
}

int main(int argc, char *argv[])
{
    printf("hello world\n");

    return 0;
}

The call order should be exactly as specified:

$ gcc $BOOKFLAGS hello.c -o hello
$ ./hello
init2
init1
hello world

Example 5.16. We can add initialization functions using another method:

hello.c

#include <stdio.h>

void init1() {
    printf("%s\n", __FUNCTION__);
}

void init2() {
    printf("%s\n", __FUNCTION__);
}

/* Without typedef, init is a definition of a function pointer.
   With typedef, init is a declaration of a type.*/
typedef void (*init)();

__attribute__((section(".init_array"))) init init_arr[2] = {init1, init2};

int main(int argc, char *argv[])
{
    printf("hello world!\n");

    return 0;
}

The attribute section("...") puts a variable or a function into a particular section rather than the default (.data or .text). In this example, it is .init_array. The section name is not necessarily the same as a standard section in an ELF file (such as .text or .init_array), but can be anything. Non-standard section names are often used for controlling the final binary layout of a compiled program. We will explore this technique in more detail when learning the GNU ld linker and the linking process. Again, the program automatically calls the constructors without explicitly invoking them:

$ gcc $BOOKFLAGS hello.c -o hello
$ ./hello
init1
init2
hello world!
FINI_ARRAY

is an array of function pointers for program termination, called after exiting main(). If the application terminates abnormally, such as through an abort() call or a crash, the .fini_array is ignored.

Example 5.17. A destructor is automatically called after exiting main(), if one or more are available:

hello.c

#include <stdio.h>

__attribute__((destructor)) static void destructor(){
    printf("%s\n", __FUNCTION__);
}

int main(int argc, char *argv[])
{
    printf("hello world\n");

    return 0;
}
$ gcc $BOOKFLAGS hello.c -o hello
$ ./hello
hello world
destructor
PREINIT_ARRAY

is an array of function pointers that are invoked before all other initialization functions in INIT_ARRAY.

Example 5.18. To use the .preinit_array, the only way to put functions into this section is to use the attribute section():

hello.c

#include <stdio.h>

void preinit1() {
    printf("%s\n", __FUNCTION__);
}

void preinit2() {
    printf("%s\n", __FUNCTION__);
}

void init1() {
    printf("%s\n", __FUNCTION__);
}

void init2() {
    printf("%s\n", __FUNCTION__);
}

typedef void (*preinit)();
typedef void (*init)();

__attribute__((section(".preinit_array"))) preinit preinit_arr[2] = {preinit1, preinit2};
__attribute__((section(".init_array"))) init init_arr[2] = {init1, init2};

int main(int argc, char *argv[])
{
    printf("hello world!\n");

    return 0;
}
$ gcc $BOOKFLAGS hello.c -o hello
$ ./hello
preinit1
preinit2
init1
init2
hello world!
GROUP

defines a section group, which is the same section that appears in different object files but when merged into the final executable binary file, only one copy is kept and the rest in other object files are discarded. This section is only relevant in C++ object files, so we will not examine it further.

SYMTAB_SHNDX

is a section containing extended section indexes, that are associated with a symbol table. This section only appears when the Ndx value of an entry in the symbol table exceeds the LORESERVE value. This section then maps between a symbol and an actual index value of a section header.

Upon understanding section types, we can understand the numbers in the Link and Info fields:

The meanings of Link and Info depend on the section type.
Type Link Info
DYNAMIC Entries in this section use the section index of the dynamic string table. 0

HASH

GNU_HASH

The section index of the symbol table to which the hash table applies. 0

REL

RELA

The section index of the associated symbol table. The section index to which the relocation applies.

SYMTAB

DYNSYM

The section index of the associated string table. One greater than the symbol table index of the last local symbol.
GROUP The section index of the associated symbol table. The symbol index of an entry in the associated symbol table. The name of the specified symbol table entry provides a signature for the section group.
SYMTAB_SHNDX The section header index of the associated symbol table.

Exercise 5.1. Verify that the value of the Link field of a SYMTAB section is the index of a STRTAB section.

Exercise 5.2. Verify that the value of the Info field of a SYMTAB section is the index of the last local symbol + 1. It means, in the symbol table, from the index listed by the Info field onward, no local symbol appears.

Exercise 5.3. Verify that the value of the Link field of a REL section is the index of the SYMTAB section.

Exercise 5.4. Verify that the value of the Info field of a REL section is the index of the section where the relocation is applied. For example, if the section is .rel.text, then the relocated section should be .text.

5.5 Program header table

A program header table is an array of program headers that defines the memory layout of a program at runtime.

A program header is a description of a program segment.

A program segment is a collection of related sections. A segment contains zero or more sections. An operating system, when loading a program, only uses segments, not sections. To see the information of a program header table, we use the -l option with readelf:

$ readelf -l <binary file>

Similar to a section, a program header also has types:

PHDR

specifies the location and size of the program header table itself, both in the file and in the memory image of the program.

INTERP

specifies the location and size of a null-terminated path name to invoke as an interpreter for linking runtime libraries.

LOAD

specifies a loadable segment. That is, this segment is loaded into main memory.

DYNAMIC

specifies dynamic linking information.

NOTE

specifies the location and size of auxiliary information.

TLS

specifies the Thread-Local Storage template, which is formed from the combination of all sections with the flag TLS.

GNU_STACK

indicates whether the program’s stack should be made executable or not. The Linux kernel uses this type.

GNU_RELRO

marks a region that the dynamic linker makes read-only once it has finished relocating the program (relocation read-only). A GNU extension, like the two below.

GNU_EH_FRAME

gives the location of the .eh_frame_hdr section, a lookup table into .eh_frame for stack unwinding.

GNU_PROPERTY

gives the location of the .note.gnu.property section, where gcc records which processor features the code relies on, such as control-flow enforcement.

A segment also has permissions, which are a combination of these 3 values:

Segment Permissions
Permission Description
R Readable
W Writable
E Executable

Example 5.19. The command to get the program header table:

$ readelf -l hello

Output:

Elf file type is EXEC (Executable file)
Entry point 0x401040
There are 14 program headers, starting at offset 64

Program Headers:
  Type           Offset             VirtAddr           PhysAddr
                 FileSiz            MemSiz              Flags  Align
  PHDR           0x0000000000000040 0x0000000000400040 0x0000000000400040
                 0x0000000000000310 0x0000000000000310  R      0x8
  INTERP         0x0000000000000394 0x0000000000400394 0x0000000000400394
                 0x000000000000001c 0x000000000000001c  R      0x1
      [Requesting program interpreter: /lib64/ld-linux-x86-64.so.2]
  LOAD           0x0000000000000000 0x0000000000400000 0x0000000000400000
                 0x00000000000004f8 0x00000000000004f8  R      0x1000
  LOAD           0x0000000000001000 0x0000000000401000 0x0000000000401000
                 0x0000000000000151 0x0000000000000151  R E    0x1000
  LOAD           0x0000000000002000 0x0000000000402000 0x0000000000402000
                 0x00000000000000d8 0x00000000000000d8  R      0x1000
  LOAD           0x0000000000002df8 0x0000000000403df8 0x0000000000403df8
                 0x0000000000000220 0x0000000000000228  RW     0x1000
  DYNAMIC        0x0000000000002e08 0x0000000000403e08 0x0000000000403e08
                 0x00000000000001d0 0x00000000000001d0  RW     0x8
  NOTE           0x0000000000000350 0x0000000000400350 0x0000000000400350
                 0x0000000000000020 0x0000000000000020  R      0x8
  NOTE           0x0000000000000370 0x0000000000400370 0x0000000000400370
                 0x0000000000000024 0x0000000000000024  R      0x4
  NOTE           0x00000000000020b8 0x00000000004020b8 0x00000000004020b8
                 0x0000000000000020 0x0000000000000020  R      0x4
  GNU_PROPERTY   0x0000000000000350 0x0000000000400350 0x0000000000400350
                 0x0000000000000020 0x0000000000000020  R      0x8
  GNU_EH_FRAME   0x0000000000002010 0x0000000000402010 0x0000000000402010
                 0x0000000000000024 0x0000000000000024  R      0x4
  GNU_STACK      0x0000000000000000 0x0000000000000000 0x0000000000000000
                 0x0000000000000000 0x0000000000000000  RW     0x10
  GNU_RELRO      0x0000000000002df8 0x0000000000403df8 0x0000000000403df8
                 0x0000000000000208 0x0000000000000208  R      0x1

 Section to Segment mapping:
  Segment Sections...
   00
   01     .interp
   02     .note.gnu.property .note.gnu.build-id .interp .gnu.hash .dynsym .dynstr .gnu.version .gnu.version_r .rela.dyn .rela.plt
   03     .init .plt .text .fini
   04     .rodata .eh_frame_hdr .eh_frame .note.ABI-tag
   05     .init_array .fini_array .dynamic .got .got.plt .data .bss
   06     .dynamic
   07     .note.gnu.property
   08     .note.gnu.build-id
   09     .note.ABI-tag
   10     .note.gnu.property
   11     .eh_frame_hdr
   12
   13     .init_array .fini_array .dynamic .got

In the sample output, the LOAD segment appears four times:

  LOAD           0x0000000000000000 0x0000000000400000 0x0000000000400000
                 0x00000000000004f8 0x00000000000004f8  R      0x1000
  LOAD           0x0000000000001000 0x0000000000401000 0x0000000000401000
                 0x0000000000000151 0x0000000000000151  R E    0x1000
  LOAD           0x0000000000002000 0x0000000000402000 0x0000000000402000
                 0x00000000000000d8 0x00000000000000d8  R      0x1000
  LOAD           0x0000000000002df8 0x0000000000403df8 0x0000000000403df8
                 0x0000000000000220 0x0000000000000228  RW     0x1000

Why? Notice the permissions:

  • the first LOAD has Read permission only. It holds the ELF header, the program header table and the tables used for dynamic linking: data that the loader reads, but that the program never executes or modifies.

  • the second LOAD has Read and Execute permission. This is a text segment. A text segment contains read-only instructions.

  • the third LOAD has Read permission only again. It holds the read-only data of the program, such as the "Hello World" string in .rodata.

  • the fourth LOAD has Read and Write permission. This is a data segment. It means that this segment can be read and written to, but is not allowed to be used as executable code, for security reasons.

Older linkers produced only two LOAD segments, a Read and Execute one with everything from the ELF header to .rodata, and a Read and Write one with the data. The GNU linker now separates the code from everything else by default (its -z separate-code option), so that no byte that is not an instruction is ever mapped executable. Each segment then starts on its own 4 KB page, hence the Align of 0x1000 and the page-aligned addresses 0x401000 and 0x402000.

Then, the LOAD segments contain the following sections:

   02     .note.gnu.property .note.gnu.build-id .interp .gnu.hash .dynsym .dynstr .gnu.version .gnu.version_r .rela.dyn .rela.plt
   03     .init .plt .text .fini
   04     .rodata .eh_frame_hdr .eh_frame .note.ABI-tag
   05     .init_array .fini_array .dynamic .got .got.plt .data .bss

The first number is the index of a program header in the program header table, and the remaining text is the list of all sections within a segment. Unfortunately, the list of program headers above does not print the indexes, so a user needs to keep track manually of which segment is of which index. The first segment starts at index 0, the second at index 1 and so on. LOAD are the segments at index 2, 3, 4 and 5. As can be seen from the four lists of sections, most sections are loadable and are available at runtime.

5.6 Segments vs sections

As mentioned earlier, an operating system loads program segments, not sections. However, a question arises: Why doesn’t the operating system use sections instead? After all, a section also contains similar information to a program segment, such as the type, the virtual memory address to be loaded, the size, the attributes, the flags and align. As explained before, a segment is the perspective of an operating system, while a section is the perspective of a linker. To understand why, looking into the structure of a segment, we can easily see:

To see the last point more clearly, consider an example of linking two object files. Suppose we have two source files, the hello.c of this chapter:

hello.c

#include <stdio.h>

int main(int argc, char *argv[])
{
    printf("Hello World\n");
    return 0;
}

and:

math.c

int add(int a, int b) {
    return a + b;
}

Now, compile the two source files as object files, 32-bit ones from here on:

$ gcc $BOOKFLAGS -c math.c
$ gcc $BOOKFLAGS -c hello.c

Then, we check the sections of math.o:

$ readelf -S math.o
There are 9 section headers, starting at offset 0xe8:

Section Headers:
  [Nr] Name              Type            Addr     Off    Size   ES Flg Lk Inf Al
  [ 0]                   NULL            00000000 000000 000000 00      0   0  0
  [ 1] .text             PROGBITS        00000000 000034 00000d 00  AX  0   0  1
  [ 2] .data             PROGBITS        00000000 000041 000000 00  WA  0   0  1
  [ 3] .bss              NOBITS          00000000 000041 000000 00  WA  0   0  1
  [ 4] .comment          PROGBITS        00000000 000041 000020 01  MS  0   0  1
  [ 5] .note.GNU-stack   PROGBITS        00000000 000061 000000 00      0   0  1
  [ 6] .symtab           SYMTAB          00000000 000064 000030 10      7   2  4
  [ 7] .strtab           STRTAB          00000000 000094 00000c 00      0   0  1
  [ 8] .shstrtab         STRTAB          00000000 0000a0 000045 00      0   0  1
Key to Flags:
  W (write), A (alloc), X (execute), M (merge), S (strings), I (info),
  L (link order), O (extra OS processing required), G (group), T (TLS),
  C (compressed), x (unknown), o (OS specific), E (exclude),
  D (mbind), p (processor specific)

As shown in the output, the virtual memory addresses of every section are set to 0. At this stage, each object file is simply a block of binary that contains code and data. Its existence is to serve as a material container for the final product, which is the executable binary. As such, the virtual addresses in math.o are all zeroes. (The format of the table is also different from the one we saw earlier: for a 32-bit file, a section fits on one line.)

No segment exists at this stage:

$ readelf -l math.o
There are no program headers in this file.

The same happens to the other object file:

$ readelf -S hello.o
There are 11 section headers, starting at offset 0x158:

Section Headers:
  [Nr] Name              Type            Addr     Off    Size   ES Flg Lk Inf Al
  [ 0]                   NULL            00000000 000000 000000 00      0   0  0
  [ 1] .text             PROGBITS        00000000 000034 00002e 00  AX  0   0  1
  [ 2] .rel.text         REL             00000000 0000f4 000010 08   I  8   1  4
  [ 3] .data             PROGBITS        00000000 000062 000000 00  WA  0   0  1
  [ 4] .bss              NOBITS          00000000 000062 000000 00  WA  0   0  1
  [ 5] .rodata           PROGBITS        00000000 000062 00000c 00   A  0   0  1
  [ 6] .comment          PROGBITS        00000000 00006e 000020 01  MS  0   0  1
  [ 7] .note.GNU-stack   PROGBITS        00000000 00008e 000000 00      0   0  1
  [ 8] .symtab           SYMTAB          00000000 000090 000050 10      9   3  4
  [ 9] .strtab           STRTAB          00000000 0000e0 000013 00      0   0  1
  [10] .shstrtab         STRTAB          00000000 000104 000051 00      0   0  1
Key to Flags:
  W (write), A (alloc), X (execute), M (merge), S (strings), I (info),
  L (link order), O (extra OS processing required), G (group), T (TLS),
  C (compressed), x (unknown), o (OS specific), E (exclude),
  D (mbind), p (processor specific)
$ readelf -l hello.o
There are no program headers in this file.

Only when object files are combined into a final executable binary are sections fully realized:

$ gcc $BOOKFLAGS math.o hello.o -o hello
$ readelf -S hello
There are 29 section headers, starting at offset 0x3570:

Section Headers:
  [Nr] Name              Type            Addr     Off    Size   ES Flg Lk Inf Al
  [ 0]                   NULL            00000000 000000 000000 00      0   0  0
  [ 1] .note.gnu.bu[...] NOTE            080481b4 0001b4 000024 00   A  0   0  4
  [ 2] .interp           PROGBITS        080481d8 0001d8 000013 00   A  0   0  1
  [ 3] .gnu.hash         GNU_HASH        080481ec 0001ec 000020 04   A  4   0  4
  [ 4] .dynsym           DYNSYM          0804820c 00020c 000050 10   A  5   1  4
  [ 5] .dynstr           STRTAB          0804825c 00025c 000055 00   A  0   0  1
  [ 6] .gnu.version      VERSYM          080482b2 0002b2 00000a 02   A  4   0  2
  [ 7] .gnu.version_r    VERNEED         080482bc 0002bc 000030 00   A  5   1  4
  [ 8] .rel.dyn          REL             080482ec 0002ec 000008 08   A  4   0  4
  [ 9] .rel.plt          REL             080482f4 0002f4 000010 08  AI  4  22  4
  [10] .init             PROGBITS        08049000 001000 000020 00  AX  0   0  4
  [11] .plt              PROGBITS        08049020 001020 000030 04  AX  0   0 16
  [12] .text             PROGBITS        08049050 001050 000151 00  AX  0   0 16
  [13] .fini             PROGBITS        080491a4 0011a4 000014 00  AX  0   0  4
  [14] .rodata           PROGBITS        0804a000 002000 000014 00   A  0   0  4
  [15] .eh_frame_hdr     PROGBITS        0804a014 002014 000024 00   A  0   0  4
  [16] .eh_frame         PROGBITS        0804a038 002038 000080 00   A  0   0  4
  [17] .note.ABI-tag     NOTE            0804a0b8 0020b8 000020 00   A  0   0  4
  [18] .init_array       INIT_ARRAY      0804bf00 002f00 000004 04  WA  0   0  4
  [19] .fini_array       FINI_ARRAY      0804bf04 002f04 000004 04  WA  0   0  4
  [20] .dynamic          DYNAMIC         0804bf08 002f08 0000e8 08  WA  5   0  4
  [21] .got              PROGBITS        0804bff0 002ff0 000004 04  WA  0   0  4
  [22] .got.plt          PROGBITS        0804bff4 002ff4 000014 04  WA  0   0  4
  [23] .data             PROGBITS        0804c008 003008 000008 00  WA  0   0  4
  [24] .bss              NOBITS          0804c010 003010 000004 00  WA  0   0  1
  [25] .comment          PROGBITS        00000000 003010 00001f 01  MS  0   0  1
  [26] .symtab           SYMTAB          00000000 003030 000270 10     27  20  4
  [27] .strtab           STRTAB          00000000 0032a0 0001ce 00      0   0  1
  [28] .shstrtab         STRTAB          00000000 00346e 000101 00      0   0  1
Key to Flags:
  W (write), A (alloc), X (execute), M (merge), S (strings), I (info),
  L (link order), O (extra OS processing required), G (group), T (TLS),
  C (compressed), x (unknown), o (OS specific), E (exclude),
  D (mbind), p (processor specific)

Every loadable section is now assigned an address, in the Addr column. The reason each section got its own address is that in reality, gcc does not combine the objects by itself, but invokes the linker ld. The linker ld uses the default script that it can find in the system to build the executable binary. In the default script, the first segment is assigned the starting address 0x8048000, and sections belong to it. Then:

Indeed, the end address of a segment is also the end address of its final section. We can see this by listing all the segments:

$ readelf -l hello

And check, for example, the first LOAD segment, which starts at 0x08048000 and ends at 0x08048000 + 0x304 = 0x08048304:

Elf file type is EXEC (Executable file)
Entry point 0x8049050
There are 12 program headers, starting at offset 52

Program Headers:
  Type           Offset   VirtAddr   PhysAddr   FileSiz MemSiz  Flg Align
  PHDR           0x000034 0x08048034 0x08048034 0x00180 0x00180 R   0x4
  INTERP         0x0001d8 0x080481d8 0x080481d8 0x00013 0x00013 R   0x1
      [Requesting program interpreter: /lib/ld-linux.so.2]
  LOAD           0x000000 0x08048000 0x08048000 0x00304 0x00304 R   0x1000
  LOAD           0x001000 0x08049000 0x08049000 0x001b8 0x001b8 R E 0x1000
  LOAD           0x002000 0x0804a000 0x0804a000 0x000d8 0x000d8 R   0x1000
  LOAD           0x002f00 0x0804bf00 0x0804bf00 0x00110 0x00114 RW  0x1000
  DYNAMIC        0x002f08 0x0804bf08 0x0804bf08 0x000e8 0x000e8 RW  0x4
  NOTE           0x0001b4 0x080481b4 0x080481b4 0x00024 0x00024 R   0x4
  NOTE           0x0020b8 0x0804a0b8 0x0804a0b8 0x00020 0x00020 R   0x4
  GNU_EH_FRAME   0x002014 0x0804a014 0x0804a014 0x00024 0x00024 R   0x4
  GNU_STACK      0x000000 0x00000000 0x00000000 0x00000 0x00000 RW  0x10
  GNU_RELRO      0x002f00 0x0804bf00 0x0804bf00 0x00100 0x00100 R   0x1

 Section to Segment mapping:
  Segment Sections...
   00
   01     .interp
   02     .note.gnu.build-id .interp .gnu.hash .dynsym .dynstr .gnu.version .gnu.version_r .rel.dyn .rel.plt
   03     .init .plt .text .fini
   04     .rodata .eh_frame_hdr .eh_frame .note.ABI-tag
   05     .init_array .fini_array .dynamic .got .got.plt .data .bss
   06     .dynamic
   07     .note.gnu.build-id
   08     .note.ABI-tag
   09     .eh_frame_hdr
   10
   11     .init_array .fini_array .dynamic .got

The last section in the first LOAD segment is .rel.plt. The .rel.plt section starts at 0x080482f4 because the start address is 0x08048000 and the offset into the file is 0x2f4. The end address of .rel.plt should be 0x08048000 + 0x2f4 + 0x10 = 0x08048304, because the section size is 0x10. This is exactly the same as the end address of the first LOAD segment above: 0x08048000 + 0x304 = 0x08048304.

The next segment does not continue where this one ends: the linker starts the code segment on a fresh page, at file offset 0x1000 and address 0x08049000, so that .init, .plt and .text are the only sections mapped executable. This is where the addresses 0x0804xxxx of the objdump listings of chapter 4 come from. In general, the address of a section is the address of its segment plus the distance of the section from the start of the segment in the file; the data segment, for example, starts at address 0x0804bf00 for file offset 0x2f00.

Chapter 8, Linking and loading on bare metal, will explore this whole process in detail.

5.7 Check your understanding

  1. An ELF executable carries both a program header table and a section header table. Why two tables that describe the same bytes? What would still work, and what would stop working, if the section header table of hello were erased?

  2. In the 32-bit hello, .bss has a size of 4 but the section that follows it, .comment, starts at the same file offset. What would change in the file, and what in memory, if the variable that lives in .bss were initialized to 1 instead of 0?

  3. .symtab, .strtab and .shstrtab have the address 0 and no A flag, while .dynsym and .dynstr have an address. What is the difference between the two groups, in terms of who needs them and when?

  4. The Name field of a section header is a 4-byte number, sh_name, not a string. Why is the name not stored in the header itself, and what does readelf do to print .text?

  5. hello.o has a .rel.text section and the linked hello does not, while hello has .rel.dyn and .rel.plt that hello.o lacks. What happened to the first, and why do the other two exist in the executable at all?

  6. The linker produces four LOAD segments with different permissions rather than one segment covering the whole file with all permissions. What does the separation cost, and what kind of bug does it catch at run time that a single segment would hide?

  7. The entry point is _start, not main. What would happen if hello were linked with -e main, so that the operating system jumped to main directly?

  8. Within a LOAD segment, the address of a section is the address of the segment plus the distance of the section from the start of the segment in the file. Why must the layout in memory mirror the layout in the file, and why does each LOAD segment after the first start at a file offset that is a multiple of 0x1000?