The object file as a deployable unit
Most compiled languages follow the same pipeline: a compiler turns source into assembly, an assembler turns that into object files, and a linker merges those object files into an executable while resolving external references. But there are cases where skipping the linker is useful — for example, when a library must be built with a different compiler than the rest of the program. Executing an object file directly is also a common technique in malware analysis. Previous posts in this series covered the x86 approach; this one applies the same idea to aarch64, where the overall mechanics are identical but the ELF details differ.
The individual stages are easy to isolate with GCC. Starting from a single main.c file, the compiler produces assembly (main.s), the assembler produces an object file (main.o), and the linker produces the final binary (main). Both main.o and main are ELF files, but only the linked binary contains everything needed to run.
#include <stdio.h>
int main(void)
{
puts("Hello, world!");
return 0;
}
Compiler output — main.s:
$ gcc -S main.c
$ ls
main.c main.s
Assembler output — main.o:
$ gcc -c main.s -o main.o
$ ls
main.c main.o main.s
Linker output — main:
$ gcc main.o -o main
$ ls
main main.c main.o main.s
$ ./main
Hello, world!
All examples assume native aarch64 GCC, or cross-compilation flags for other hosts.
$ file main.o
main.o: ELF 64-bit LSB relocatable, ARM aarch64, version 1 (SYSV), not stripped
$ file main
main: ELF 64-bit LSB pie executable, ARM aarch64, version 1 (SYSV), dynamically
linked, interpreter /lib/ld-linux-aarch64.so.1,
BuildID[sha1]=d3ecd2f8ac3b2dec11ed4cc424f15b3e1f130dd4, for GNU/Linux 3.7.0, not stripped
Anatomy of an ELF object
An ELF file begins with a fixed ELF header that records the target architecture, the entry point, and the locations of the other tables that follow. Two of those tables matter here: the program header and the section header.
The loader uses the program header to map segments — groups of sections with similar permissions and purposes — into memory. In the aarch64 output, the loader is identified as /lib/ld-linux-aarch64.so.1. Segment granularity keeps loading fast, while section granularity keeps permissions precise: code sections are executable but not writable, and data sections are writable but not executable.
$ readelf -h main
ELF Header:
Magic: 7f 45 4c 46 02 01 01 00 00 00 00 00 00 00 00 00
Class: ELF64
Data: 2's complement, little endian
Version: 1 (current)
OS/ABI: UNIX - System V
ABI Version: 0
Type: DYN (Position-Independent Executable file)
Machine: AArch64
Version: 0x1
Entry point address: 0x640
Start of program headers: 64 (bytes into file)
Start of section headers: 68576 (bytes into file)
Flags: 0x0
Size of this header: 64 (bytes)
Size of program headers: 56 (bytes)
Number of program headers: 9
Size of section headers: 64 (bytes)
Number of section headers: 29
Section header string table index: 28
$ readelf -Wl main
Elf file type is DYN (Position-Independent Executable file)
Entry point 0x640
There are 9 program headers, starting at offset 64
Program Headers:
Type Offset VirtAddr PhysAddr FileSiz MemSiz Flg Align
PHDR 0x000040 0x0000000000000040 0x0000000000000040 0x0001f8 0x0001f8 R 0x8
INTERP 0x000238 0x0000000000000238 0x0000000000000238 0x00001b 0x00001b R 0x1
[Requesting program interpreter: /lib/ld-linux-aarch64.so.1]
LOAD 0x000000 0x0000000000000000 0x0000000000000000 0x00088c 0x00088c R E 0x10000
LOAD 0x00fdc8 0x000000000001fdc8 0x000000000001fdc8 0x000270 0x000278 RW 0x10000
DYNAMIC 0x00fdd8 0x000000000001fdd8 0x000000000001fdd8 0x0001e0 0x0001e0 RW 0x8
NOTE 0x000254 0x0000000000000254 0x0000000000000254 0x000044 0x000044 R 0x4
GNU_EH_FRAME 0x0007a0 0x00000000000007a0 0x00000000000007a0 0x00003c 0x00003c R 0x4
GNU_STACK 0x000000 0x0000000000000000 0x0000000000000000 0x000000 0x000000 RW 0x10
GNU_RELRO 0x00fdc8 0x000000000001fdc8 0x000000000001fdc8 0x000238 0x000238 R 0x1
Section to Segment mapping:
Segment Sections...
00
01 .interp
02 .interp .note.gnu.build-id .note.ABI-tag .gnu.hash .dynsym .dynstr .gnu.version .gnu.version_r .rela.dyn .rela.plt .init .plt .text .fini .rodata .eh_frame_hdr .eh_frame
03 .init_array .fini_array .dynamic .got .got.plt .data .bss
04 .dynamic
05 .note.gnu.build-id .note.ABI-tag
06 .eh_frame_hdr
07
08 .init_array .fini_array .dynamic .got
$ readelf -SW main
There are 29 section headers, starting at offset 0x10be0:
Section Headers:
[Nr] Name Type Address Off Size ES Flg Lk Inf Al
[ 0] NULL 0000000000000000 000000 000000 00 0 0 0
[ 1] .interp PROGBITS 0000000000000238 000238 00001b 00 A 0 0 1
[ 2] .note.gnu.build-id NOTE 0000000000000254 000254 000024 00 A 0 0 4
[ 3] .note.ABI-tag NOTE 0000000000000278 000278 000020 00 A 0 0 4
[ 4] .gnu.hash GNU_HASH 0000000000000298 000298 00001c 00 A 5 0 8
[ 5] .dynsym DYNSYM 00000000000002b8 0002b8 0000f0 18 A 6 3 8
[ 6] .dynstr STRTAB 00000000000003a8 0003a8 000092 00 A 0 0 1
[ 7] .gnu.version VERSYM 000000000000043a 00043a 000014 02 A 5 0 2
[ 8] .gnu.version_r VERNEED 0000000000000450 000450 000030 00 A 6 1 8
[ 9] .rela.dyn RELA 0000000000000480 000480 0000c0 18 A 5 0 8
[10] .rela.plt RELA 0000000000000540 000540 000078 18 AI 5 22 8
[11] .init PROGBITS 00000000000005b8 0005b8 000018 00 AX 0 0 4
[12] .plt PROGBITS 00000000000005d0 0005d0 000070 00 AX 0 0 16
[13] .text PROGBITS 0000000000000640 000640 000134 00 AX 0 0 64
[14] .fini PROGBITS 0000000000000774 000774 000014 00 AX 0 0 4
[15] .rodata PROGBITS 0000000000000788 000788 000016 00 A 0 0 8
[16] .eh_frame_hdr PROGBITS 00000000000007a0 0007a0 00003c 00 A 0 0 4
[17] .eh_frame PROGBITS 00000000000007e0 0007e0 0000ac 00 A 0 0 8
[18] .init_array INIT_ARRAY 000000000001fdc8 00fdc8 000008 08 WA 0 0 8
[19] .fini_array FINI_ARRAY 000000000001fdd0 00fdd0 000008 08 WA 0 0 8
[20] .dynamic DYNAMIC 000000000001fdd8 00fdd8 0001e0 10 WA 6 0 8
[21] .got PROGBITS 000000000001ffb8 00ffb8 000030 08 WA 0 0 8
[22] .got.plt PROGBITS 000000000001ffe8 00ffe8 000040 08 WA 0 0 8
[23] .data PROGBITS 0000000000020028 010028 000010 00 WA 0 0 8
[24] .bss NOBITS 0000000000020038 010038 000008 00 WA 0 0 1
[25] .comment PROGBITS 0000000000000000 010038 00001f 01 MS 0 0 1
[26] .symtab SYMTAB 0000000000000000 010058 000858 18 27 66 8
[27] .strtab STRTAB 0000000000000000 0108b0 00022c 00 0 0 1
[28] .shstrtab STRTAB 0000000000000000 010adc 000103 00 0 0 1
Key to Flags:
W (write), A (alloc), X (execute), M (merge), S (strings), I (info),
L (link order), O (extra OS processing required), G (group), T (TLS),
C (compressed), x (unknown), o (OS specific), E (exclude),
D (mbind), p (processor specific)
Finding functions without names
The section header lists sections but, for space efficiency, does not embed full names. Instead, names live in a dedicated .shstrtab string table, where each entry is null-terminated. Other sections reference their names by offset into this table. The ELF header gives a direct pointer to .shstrtab, which resolves the otherwise circular lookup.
The same pattern applies to symbol data: .symtab holds symbol entries and .strtab holds their names. To load add5 and add10 from obj.o, the steps are straightforward:
- Locate the two functions in the
.textsection. - Copy them into executable memory.
- Return their addresses to the calling program.
The code from the x86 post runs on aarch64 without modification, which shows how portable the ELF format and the underlying technique are across architectures.
Branching Out: AArch64 Relocations in Practice
When we modified add10 in Part 2 to call add5 rather than computing everything internally, we introduced the first real-world dependency: relocations. These occur whenever a symbol—be it a function, variable, or constant—lives outside the current compilation unit. The AArch64 version of our loader broke exactly as the x86 one did. The assembly listing reveals why:
$ objdump --disassemble --section=.text obj.o
obj.o: file format elf64-littleaarch64
Disassembly of section .text:
0000000000000000 <add5>:
0: d10043ff sub sp, sp, #0x10
4: b9000fe0 str w0, [sp, #12]
8: b9400fe0 ldr w0, [sp, #12]
c: 11001400 add w0, w0, #0x5
10: 910043ff add sp, sp, #0x10
14: d65f03c0 ret
0000000000000018 <add10>:
18: a9be7bfd stp x29, x30, [sp, #-32]!
1c: 910003fd mov x29, sp
20: b9001fe0 str w0, [sp, #28]
24: b9401fe0 ldr w0, [sp, #28]
28: 94000000 bl 0 <add5>
2c: b9001fe0 str w0, [sp, #28]
30: b9401fe0 ldr w0, [sp, #28]
34: 94000000 bl 0 <add5>
38: a8c27bfd ldp x29, x30, [sp], #32
3c: d65f03c0 ret
Notice that every hex value in the second column is uniform in length—a stark contrast to the variable-length x86 instructions from Part 2. AArch64 instructions are fixed at 32 bits. Because immediate values often don't fit in that space alongside opcode and register fields, complex operations may need multiple instructions. Our focus is on the bl (branch with link) instructions at rows 28 and 34. These jumps save the return address in the link register (lr) before branching, enabling the callee to recover the caller's address. While the opcode takes up bits [31:26], bl needs no register operands, freeing all remaining 26 bits for the immediate branch offset. That offset is a signed value relative to the current pc, giving a range of roughly ±128 MB after accounting for the required 4-byte instruction alignment.
Our bl instructions currently contain all zeros in their immediate fields—effectively jumping nowhere. To call add5 we need offsets of -0x28 and -0x34 from the pc. Dividing by 4 yields -0xA and -0xD, which in two's complement fill the 26-bit immediate field. Folding those back into the opcode gives the final instructions 0x97FFFFF6 and 0x97FFFFF3. A key detail: these offsets are computed from the address of the bl instruction itself, not the next instruction as in x86.
Hardcoding those values gets us a working binary:
...
static void parse_obj(void)
{
...
/* copy the contents of `.text` section from the ELF file */
memcpy(text_runtime_base, obj.base + text_hdr->sh_offset, text_hdr->sh_size);
*((uint32_t *)(text_runtime_base + 0x28)) = 0x97FFFFF6;
*((uint32_t *)(text_runtime_base + 0x34)) = 0x97FFFFF3;
/* make the `.text` copy readonly and executable */
if (mprotect(text_runtime_base, page_align(text_hdr->sh_size), PROT_READ | PROT_EXEC)) {
...
$ gcc -o loader loader.c
$ ./loader
Executing add5...
add5(42) = 47
Executing add10...
add10(42) = 52
That confirms the math, but it's not how a real linker operates. Linkers consult relocation types and their associated formulas. Our relocation is R_AARCH64_CALL26. Its formula—where S is the symbol address, A the addend, and P the place being relocated—guides the computation:
|
ELF64 Code |
Name |
Operation |
|
283 |
R_<CLS>_CALL26 |
S + A - P |
Implementing that in our loader is straightforward:
/* Replace `#define R_X86_64_PLT32 4` with our Type */
#define R_AARCH64_CALL26 283
...
static void do_text_relocations(void)
{
...
uint32_t val;
switch (type)
{
case R_AARCH64_CALL26:
/* The mask separates opcode (6 bits) and the immediate value */
uint32_t mask_bl = (0xffffffff << 26);
/* S+A-P, divided by 4 */
val = (symbol_address + relocations[i].r_addend - patch_offset) >> 2;
/* Concatenate opcode and value to get final instruction */
*((uint32_t *)patch_offset) &= mask_bl;
val &= ~mask_bl;
*((uint32_t *)patch_offset) |= val;
break;
}
...
}
$ gcc -o loader loader.c
$ ./loader
Calculated relocation: 0x97fffff6
Calculated relocation: 0x97fffff3
Executing add5...
add5(42) = 47
Executing add10...
add10(42) = 52
Handling Data: Addressing the Page Problem
With function calls working, the next step is introducing global variables and constants. This surfaces two additional relocation types unused by x86 code on AArch64: R_AARCH64_ADR_PREL_PG_HI21 and R_AARCH64_ADD_ABS_LO12_NC. Their combined effect is visible in the assembly:
$ objdump --disassemble --section=.text obj.o
obj.o: file format elf64-littleaarch64
Disassembly of section .text:
0000000000000000 <get_hello>:
0: 90000000 adrp x0, 0 <get_hello>
4: 91000000 add x0, x0, #0x0
8: d65f03c0 ret
000000000000000c <get_var>:
c: 90000000 adrp x0, 0 <get_hello>
10: 91000000 add x0, x0, #0x0
14: b9400000 ldr w0, [x0]
18: d65f03c0 ret
000000000000001c <set_var>:
1c: d10043ff sub sp, sp, #0x10
20: b9000fe0 str w0, [sp, #12]
24: 90000000 adrp x0, 0 <get_hello>
28: 91000000 add x0, x0, #0x0
2c: b9400fe1 ldr w1, [sp, #12]
30: b9000001 str w1, [x0]
34: d503201f nop
38: 910043ff add sp, sp, #0x10
3c: d65f03c0 ret
The pattern is clear: every adrp instruction is paired with a subsequent add. The reason is the 12-bit immediate field of the add instruction, which alone can only address a tiny offset from a base register. The adrp instruction provides that base by computing a PC-relative address to the start of the current 4KB page, using a 21-bit immediate shifted left by 12 bits to cover up to ±1 GB range. The relationship is the key:
Page(expr), as used in the relocation formulas, is simply the expression's address masked at 4KB boundaries:(expr & ~0xFFF).R_AARCH64_ADR_PREL_PG_HI21computes the page difference between the symbolSand the relocation siteP, encoding that result into theadrpinstruction.R_AARCH64_ADD_ABS_LO12_NCsupplies the low 12 bits of the symbol's address—the page offset that the shift masked out—for the followingaddinstruction to incorporate.
|
ELF64 Code |
Name |
Operation |
|
275 |
R_<CLS>_ ADR_PREL_PG_HI21 |
Page(S+A) - Page(P) |
|
277 |
R_<CLS>_ ADD_ABS_LO12_NC |
S + A |
Only when both relocations are resolved does a full symbol address become available. Encoding the adrp immediate is itself non-trivial: the 21 bits are split, with the low 2 bits occupying positions [30:29] and the high 19 bits in positions [23:5]. Applying the formulas and instruction layouts to update our loader yields:
#define R_AARCH64_CALL26 283
#define R_AARCH64_ADD_ABS_LO12_NC 277
#define R_AARCH64_ADR_PREL_PG_HI21 275
...
{
case R_AARCH64_CALL26:
/* The mask separates opcode (6 bits) and the immediate value */
uint32_t mask_bl = (0xffffffff << 26);
/* S+A-P, divided by 4 */
val = (symbol_address + relocations[i].r_addend - patch_offset) >> 2;
/* Concatenate opcode and value to get final instruction */
*((uint32_t *)patch_offset) &= mask_bl;
val &= ~mask_bl;
*((uint32_t *)patch_offset) |= val;
break;
case R_AARCH64_ADD_ABS_LO12_NC:
/* The mask of `add` instruction to separate
* opcode, registers and calculated value
*/
uint32_t mask_add = 0b11111111110000000000001111111111;
/* S + A */
uint32_t val = *(symbol_address + relocations[i].r_addend);
val &= ~mask_add;
*((uint32_t *)patch_offset) &= mask_add;
/* Final instruction */
*((uint32_t *)patch_offset) |= val;
case R_AARCH64_ADR_PREL_PG_HI21:
/* Page(S+A)-Page(P), Page(expr) is defined as (expr & ~0xFFF) */
val = (((uint64_t)(symbol_address + relocations[i].r_addend)) & ~0xFFF) - (((uint64_t)patch_offset) & ~0xFFF);
/* Shift right the calculated value by 12 bits.
* During decoding it will be shifted left as described above,
* so we do the opposite.
*/
val >>= 12;
/* Separate the lower and upper bits to place them in different positions */
uint32_t immlo = (val & (0xf >> 2)) << 29 ;
uint32_t immhi = (val & ((0xffffff >> 13) << 2)) << 22;
*((uint32_t *)patch_offset) |= immlo;
*((uint32_t *)patch_offset) |= immhi;
break;
}
$ gcc -o loader loader.c
$ ./loader
Executing add5...
add5(42) = 47
Executing add10...
add10(42) = 52
Executing get_hello...
get_hello() = Hello, world!
Executing get_var...
get_var() = 5
Executing set_var(42)...
Executing get_var again...
get_var() = 42
The final, complete loader is available in the accompanying repository.
From Relative Branches to Register Jumps
The PLT/GOT mechanism solves a practical problem: system libraries should be shareable across processes, and we only want to resolve the symbols we actually use. Instead of resolving every external function up front, the dynamic loader creates small stubs — puts@plt, for example — that lazily look up the real address on first call and cache the result in the GOT.
Our Part 3 implementation on x86 simplified this by replacing the PLT with a jump table and the GOT with embedded assembly instructions. Each stub loads the resolved address and jumps to it. For AArch64, the same idea applies, but the mechanics differ because we increasingly deal with 64-bit addresses and a different instruction set. The biggest change: instead of relative branches within a small address range, we need to load a full 64-bit address into a register and use a register-indirect branch (br or blr).
Reading the Native Resolution Pattern
The compiler’s own output for an unresolved call shows how the platform expects this to work:
$ objdump --disassemble --section=.text loader
...
1d2c: 97fffb45 bl a40 <puts@plt>
1d30: f94017e0 ldr x0, [sp, #40]
1d34: d63f0000 blr x0
...
The sequence is: bl jumps to the puts@plt stub, ldr loads a value from the stack into x0 (each function keeps local state in its own stack frame), and blr branches to the address now held in that register. AArch64 register naming reflects operand width: x0–x30 hold 64-bit values, while w0–w30 use only the lower 32 bits and zero the upper half.
We only need a plain br in our stubs: the eventual return goes back to say_hello in obj.c, so there is no need to save a link register. A minimal C function shows how the compiler materializes an arbitrary address:
#include <stdint.h>
void say_hello(void)
{
uint64_t reg = 0x555555550c14;
}
$ gcc -c hello.c
$ objdump --disassemble --section=.text hello.o
hello.o: file format elf64-littleaarch64
Disassembly of section .text:
0000000000000000 <say_hello>:
0: d10043ff sub sp, sp, #0x10
4: d2818280 mov x0, #0xc14 // #3092
8: f2aaaaa0 movk x0, #0x5555, lsl #16
c: f2caaaa0 movk x0, #0x5555, lsl #32
10: f90007e0 str x0, [sp, #8]
14: d503201f nop
18: 910043ff add sp, sp, #0x10
1c: d65f03c0 ret
The hex value 0x555555550c14 is the address returned by lookup_ext_function — printed here as an example, though any valid 48-bit address would do. The compiler splits this value across three instructions: one mov and two movk, each carrying 16 bits of immediate data with a left shift (lsl).
Register choice matters. x0 is reserved by the AArch64 procedure call standard for passing function arguments, and x0–x7 are caller-saved. Our stub should use x9 instead so it doesn’t clobber arguments or interfere with the caller’s expectations.
Building the AArch64 Jump Table
The jump table entry expands from one instruction to four:
...
struct ext_jump {
uint32_t instr[4];
};
...
Three instructions load the 64-bit address (16 bits at a time via mov/movk), and the fourth, br, jumps to that register. No stack frame is needed since we’re not preserving anything — just loading an address and branching. But we cannot write assembly mnemonics directly into the table; we need raw machine code. A short assembly file gives us the hex encodings:
.global _start
_start: mov x9, #0xc14
movk x9, #0x5555, lsl #16
movk x9, #0x5555, lsl #32
br x9
$ as -o hw.o hw.s
$ objdump --disassemble --section=.text hw.o
hw.o: file format elf64-littleaarch64
Disassembly of section .text:
0000000000000000 <_start>:
0: d2818289 mov x9, #0xc14 // #3092
4: f2aaaaa9 movk x9, #0x5555, lsl #16
8: f2caaaa9 movk x9, #0x5555, lsl #32
c: d61f0120 br x9
Patching for ASLR
A static address won’t work under ASLR. The address returned by lookup_ext_function changes on every run, so we have to patch the mov and movk instructions with the actual value at load time. The address is split into three 16-bit chunks and written into the corresponding immediate fields, just as the compiler does when it generates code:
if (symbols[symbol_idx].st_shndx == SHN_UNDEF) {
static int curr_jmp_idx = 0;
uint64_t addr = lookup_ext_function(strtab + symbols[symbol_idx].st_name);
uint32_t mov = 0b11010010100000000000000000001001 | ((addr << 48) >> 43);
uint32_t movk1 = 0b11110010101000000000000000001001 | (((addr >> 16) << 48) >> 43);
uint32_t movk2 = 0b11110010110000000000000000001001 | (((addr >> 32) << 48) >> 43);
jumptable[curr_jmp_idx].instr[0] = mov; // mov x9, #0x0c14
jumptable[curr_jmp_idx].instr[1] = movk1; // movk x9, #0x5555, lsl #16
jumptable[curr_jmp_idx].instr[2] = movk2; // movk x9, #0x5555, lsl #32
jumptable[curr_jmp_idx].instr[3] = 0xd61f0120; // br x9
symbol_address = (uint8_t *)(&jumptable[curr_jmp_idx].instr[0]);
curr_jmp_idx++;
} else {
symbol_address = section_runtime_base(§ions[symbols[symbol_idx].st_shndx]) + symbols[symbol_idx].st_value;
}
uint32_t val;
switch (type)
{
case R_AARCH64_CALL26:
/* The mask separates opcode (6 bits) and the immediate value */
uint32_t mask_bl = (0xffffffff << 26);
/* S+A-P, divided by 4 */
val = (symbol_address + relocations[i].r_addend - patch_offset) >> 2;
/* Concatenate opcode and value to get final instruction */
*((uint32_t *)patch_offset) &= mask_bl;
val &= ~mask_bl;
*((uint32_t *)patch_offset) |= val;
break;
...
In the code above, &jumptable[curr_jmp_idx].instr[0] becomes the target for the bl instruction, since the relocation type remains R_AARCH64_CALL26. Execution flows through the entire stub, ending with blr jumping to the resolved wrapper:
$ gcc -o loader loader.c
$ ./loader
Executing add5...
add5(42) = 47
Executing add10...
add10(42) = 52
Executing get_hello...
get_hello() = Hello, world!
Executing get_var...
get_var() = 5
Executing set_var(42)...
Executing get_var again...
get_var() = 42
Executing say_hello...
my_puts executed
Hello, world!
Cross-Architecture Takeaways
This series walked through the full path from an object file on disk to running code: linking, relocations, dependency resolution, and the low-level mechanics of jumping between modules. The AArch64 port highlights how much the implementation depends on the instruction set — relative branches within range for x86 become register loads plus indirect branches on 64-bit ARM, and every immediate must be patched to survive ASLR. We also saw how stub code can intercept a function call and redirect it to a custom wrapper, which is a useful hooking technique in its own right.
None of this should be treated as production code. The examples deliberately omit bounds and integrity checks to stay short. Loading and executing arbitrary external input is inherently dangerous — the security warning from the first post in this series still applies. The code is here for learning, and nothing more.



