Bridging the gap: from object file to process

Executing a raw object file is not the same as running a normal executable. A compiler and linker have already done a lot of work by the time a typical binary lands on disk: they have laid out sections, resolved relocations, and stamped in an ELF header that the kernel knows how to parse. An object file is missing all of that infrastructure. It is a container for code and data that still needs to be wired up before the CPU can touch it.

The first obstacle is the format itself. When you write ld to link an object file into an executable, the linker performs several passes. It concatenates sections with the same flags, assigns virtual addresses, and then fixes up every relocation entry by patching the target locations with the final addresses. None of that has happened in a .o file.

What the loader would normally do

The kernel's ELF loader handles a dynamically linked executable in a well-defined sequence:

  1. Parse the ELF header and program headers to find PT_LOAD segments.
  2. Map those segments into memory at their specified virtual addresses, applying page alignment and permissions.
  3. Zero-fill the .bss region, which has no file content but needs anonymous memory.
  4. For a dynamic binary, locate the interpreter (.interp section), map it, and hand control over to the dynamic linker.
  5. Let the dynamic linker process relocations, resolve symbols against shared libraries, run constructors, and finally jump to the entry point.

An object file has no PT_LOAD segments at all. It only has section headers, which the kernel loader ignores. So even if you mmap the file manually, you get a flat blob of bytes with no assigned addresses and no permission split between code and data. The sections that would normally become separate memory regions are all contiguously packed in the file. You have to impose structure yourself.

Assigning addresses by hand

The strategy for a self-contained object file loader is to decide where each section will live at runtime, then treat the file as a template. You walk the section headers, pick an address for each section that requires allocation (SHF_ALLOC), and build a memory image that mirrors a linked binary. Code sections need read and execute permissions; data sections need read and write. The .bss section is special: it has SHF_ALLOC but sh_size describing memory usage while sh_offset points into a zero-length portion of the file. You allocate and zero that space without copying any file bytes.

Alignment is the detail most easily skipped and most quickly punished. Each section header carries an sh_addralign field. If you map the file at an arbitrary heap address, local symbols inside the file may no longer have the alignment the compiler assumed. For example, an SSE-aligned array expects a 16-byte boundary. The loader must check that the base address for every allocated section is congruent to zero modulo its alignment requirement.

Relocations: the real work

Once sections are in place, every relocation entry in rela or rel sections must be applied. On x86-64, the RELA format uses an explicit addend. Each entry names a symbol and a target location. To process it, you need the symbol's value: the sum of the section it lives in plus its offset within that section, plus the final load bias of the file. The relocation types then decide the arithmetic:

  • R_X86_64_PC32 — Store S + A - P, a 32-bit PC-relative offset.
  • R_X86_64_PLT32 — Similar, used when the compiler emitted a call through the PLT.
  • R_X86_64_64 — Store the full 64-bit absolute value of S + A.

R_X86_64_PLT32 deserves a warning. When code is compiled with -fno-plt or when relocations target shared libraries, this type can still show up in object files. It is not safe to assume every call site has already been converted to a direct branch. The loader must treat a missing PLT relocation as an error unless it has explicitly stubbed the lazy-binding machinery.

Patching and executing

After relocations are applied, the code itself often still contains references to addresses that were never touched by relocation entries. A static binary produced by a normal toolchain would have had the linker resolve these by rewriting the instruction stream. With a hand-rolled loader, you have to decide before execution whether the process is allowed to write code as it runs, which means mprotect must flip the page permissions between relocation and execution. The standard pattern is to map memory with PROT_READ|PROT_WRITE, apply all fixes, then call mprotect to set the final PROT_EXEC on code pages and strip write permission from them.

Finally, you need an entry point. The object file's e_entry is meaningless because there is no ELF header to hold it. The convention is to look for a symbol named _start or main in the symbol table, compute its address by resolving it against the section it belongs to, and set up the register state as the ABI requires. For a simple test, that means clearing %rbp and calling the address as a function pointer.

That entire sequence — mapping sections, applying relocations, fixing permissions, and jumping to a resolved symbol — is enough to run a single-object compilation that does no dynamic linking and no calls into libc. It is a minimal loader, but it proves the point: the difference between an object and a process is not magic, just bookkeeping that someone has to do.

When Functions Call Each Other

In the previous post, we built a small loader capable of importing and executing functions from an object file. The example functions were self-contained, but real code often depends on other code. If we rewrite add10 to call add5, the loader begins to fail — or may even crash. The problem is not in our loader logic, but in how the compiler emits machine code when it cannot know the final layout of all functions.

Disassembling the object file with objdump reveals why. The two instructions that call add5 are encoded as e8 00 00 00 00, a near relative call. The five-byte instruction consists of the 0xe8 opcode and a 32-bit signed displacement that should hold the distance from the next instruction to the target function. The compiler, however, has left that field as zero because both functions have external linkage — the default in C — and it cannot assume where add5 will ultimately reside in memory.

Manually patching a copy of the loaded .text section with the correct displacements confirms this diagnosis. For the first call at offset 0x1f, the next instruction begins at 0x24, and the targeted add5 starts at 0x0; the relative offset is 0x0 - 0x24 = -0x24, expressed as the two’s complement value 0xffffffdc. Patching both call sites yields correct results from add10.

The compiler does emit these values when the function is declared static (internal linkage), proving it can calculate the offset when the layout is known. It also demonstrates that static is not a security boundary: the loader can still call such functions by their symbol table offsets.

Relocations: The Compiler’s Instructions

Reverting add5 to external linkage and inspecting the object file with readelf uncovers a new section, .rela.text. Its entries are the instructions the linker would normally follow to patch the code. Each relocation specifies:

  • Offset — where in .text the patch is needed; these match our manual patch locations.
  • Info — the upper bits index a symbol table entry (in this case, add5); the lower bits give the relocation type.
  • Sym. Value — the symbol offset within its section.
  • Addend — a constant used in the relocation formula.

The relocation type here is R_X86_64_PLT32, which the x86-64 ELF specification defines as producing a 32-bit result using the formula L + A - P, where L is the symbol’s runtime address, A the addend, and P the address of the location being patched. Implementing this in the loader — computing L + A - P against actual runtime addresses — produces the same values as our manual patch and makes add10 execute correctly.

Adding Data Sections

Functions that return constant strings or access global variables introduce two new section types: .rodata for read‑only constants and .data for mutable globals. They also create new relocations in .text, of type R_X86_64_PC32, whose formula S + A - P is functionally the same for our purposes. The symbol table entries here reference section indices rather than named symbols; they point to .data and .rodata respectively.

Mapping each section with a separate mmap call is risky: the kernel may place the mappings far apart in the address space, and a 32-bit displacement can only cover roughly ±2 GiB. The fix is to reserve a single block of memory with one mmap invocation and partition it among the sections, applying the appropriate protection flags per chunk.

Once the loader knows the runtime address for every section, the relocation handler can resolve offsets that point into .data and .rodata just as it resolves those pointing to functions. With this in place, the loader correctly handles code that depends on static data and mutable globals, and the imported functions can manipulate the object’s internal state through the provided accessors.

The next step is importing object code with references to external libraries — a problem the loader will need to solve with dynamic linking.