Many function-call explanations still use a 32-bit x86 diagram: arguments are pushed right to left, and ebp+8 holds the first argument. That model does not describe a typical Linux x86-64 program. Under the SysV AMD64 ABI, the first six integer or pointer arguments normally travel in registers, and rbp does not have to be a frame pointer.

This article covers the Linux user-space SysV AMD64 ABI only. It does not mix in 32-bit cdecl, the Windows x64 ABI, or the Linux system-call ABI. The goal is to prove four things in a debugger: where arguments enter a function, what call saves, how GDB reconstructs a backtrace, and why optimization makes source variables or entire frames disappear.

1. Build a reproducible program

Save this as stack_demo.c:

#include <stdio.h>

__attribute__((noinline))
long add(long left, long right)
{
    long result = left + right;
    return result;
}

__attribute__((noinline))
long calculate(long input)
{
    long doubled = input * 2;
    return add(doubled, 7);
}

int main(void)
{
    volatile long input = 5;
    printf("%ld\n", calculate(input));
    return 0;
}

Build a version that is easy to inspect:

gcc -std=c11 -Wall -Wextra -g3 -O0 \
  -fno-omit-frame-pointer -fno-pie -no-pie \
  stack_demo.c -o stack_demo

./stack_demo
# 17

-O0 reduces optimizer rewrites, -fno-omit-frame-pointer encourages a visible frame chain, and -fno-pie -no-pie makes code addresses easier to read during the exercise. These are not production hardening flags. In particular, do not disable PIE in a release build merely to simplify debugging.

2. How SysV AMD64 passes arguments

The common locations for integer and pointer arguments are:

Value Location
Argument 1 rdi
Argument 2 rsi
Argument 3 rdx
Argument 4 rcx
Argument 5 r8
Argument 6 r9
Argument 7 and later Stack argument area
Integer return value rax

For add(doubled, 7), doubled enters rdi and 7 enters rsi. The compiler may spill either value into the current stack frame for debugging or register allocation, but that is a code-generation choice. It does not mean the ABI requires every argument to be pushed first.

General-purpose registers also have preservation rules:

  • Caller-saved: rax, rcx, rdx, rsi, rdi, and r8-r11. A caller that needs one of these values after a call must preserve it.
  • Callee-saved: rbx, rbp, and r12-r15. A callee that changes one of these registers must restore it before returning.

Stack alignment is part of the ABI too. Immediately before call, rsp is normally aligned to 16 bytes. After call pushes an eight-byte return address, the callee enters with (rsp + 8) % 16 == 0. A common push rbp prologue restores 16-byte alignment before the function makes another call.

3. What call, ret, and a frame actually do

When calculate calls add, two things happen:

  1. call add pushes the address of the next instruction onto the stack.
  2. The processor transfers control to add.

When add finishes, ret pops that address into the instruction pointer. call and ret do not save local variables or preserve every register. The compiler emits those operations according to the ABI and the needs of the current function.

An unoptimized function that keeps rbp often starts with:

push rbp
mov  rbp, rsp
sub  rsp, 0x10

The resulting frame can be simplified as:

Higher addresses
+----------------------+
| Caller's frame       |
+----------------------+
| Return address       |  [rbp + 8]
+----------------------+
| Caller's saved rbp   |  [rbp]
+----------------------+
| Locals and spills    |  [rbp - offset]
+----------------------+  <- rsp
Lower addresses

That diagram describes this build, not a permanent memory layout. Optimized code may omit rbp, allocate no stack space, or use the ABI’s 128-byte red zone. Treating one frame diagram as universal is a common source of bad debugging conclusions.

4. Walk through the disassembly

Inspect the two functions separately:

objdump -d -M intel --disassemble=calculate stack_demo
objdump -d -M intel --disassemble=add stack_demo

Exact instructions vary with the GCC release, distribution hardening defaults, and target options. A typical calculate body may look similar to this:

calculate:
    push   rbp
    mov    rbp, rsp
    sub    rsp, 0x10
    mov    QWORD PTR [rbp-0x8], rdi
    mov    rax, QWORD PTR [rbp-0x8]
    add    rax, rax
    mov    QWORD PTR [rbp-0x10], rax
    mov    rax, QWORD PTR [rbp-0x10]
    mov    esi, 0x7
    mov    rdi, rax
    call   add
    leave
    ret

Follow the data flow instead of memorizing offsets:

  • input is already in rdi when calculate starts.
  • The result of input * 2 eventually moves into rdi as the first add argument.
  • 7 moves into rsi as the second argument.
  • call add writes the return address at the top of the stack.
  • add returns through rax, first to calculate, then to main.

An endbr64 instruction, stack-canary checks, different local offsets, or pop rbp instead of leave do not indicate a broken program. Record the complete compiler command before interpreting the output.

5. Prove the registers and return address in GDB

Start GDB and break on the first instruction of both functions:

gdb -q ./stack_demo
set disassembly-flavor intel
set pagination off
break *calculate
break *add
run

At the first stop, on entry to calculate:

info registers rip rsp rbp rdi
x/i $rip
x/gx $rsp

You should observe that:

  • rdi contains 5, the first argument to calculate(input).
  • $rsp points to the address used to return from calculate to main.
  • The prologue has not run yet, so rbp cannot be assumed to describe a calculate frame.

Use display/i $pc and stepi until execution passes mov rbp, rsp, then inspect the frame:

info registers rsp rbp
x/2gx $rbp
info frame

In this build, [$rbp] is the saved old rbp, and [$rbp+8] is the return address. Continue to add:

continue
info registers rdi rsi rsp
x/gx $rsp
bt
finish
p/d $rax

At add, expect rdi=10 and rsi=7. After finish, rax should contain 17. Absolute addresses may change between runs; the useful evidence is the register values and their relationships, not a fixed hexadecimal address copied from a screenshot.

6. A backtrace is not just an rbp linked list

With frame pointers enabled, saved rbp values often form an easy-to-follow chain. Modern GDB can also use DWARF debug data, call-frame information (CFI), and generated machine code to recover caller state. Both of these claims are therefore wrong:

  • “No rbp means no backtrace.”
  • “A successful bt proves the physical frame chain is intact.”

Useful commands include:

bt
bt full
frame 1
info args
info locals
thread apply all bt full

A backtrace can still be incomplete when:

  • Stack memory or a return address is corrupted.
  • The binary was stripped and matching separate debug symbols are unavailable.
  • The production executable or shared libraries do not match the debugger’s files.
  • JIT code, handwritten assembly, or a third-party library lacks correct unwind information.
  • Optimization introduces inlining, tail calls, or <optimized out> values.

For a core dump, verify the executable, shared libraries, build IDs, and symbol packages before explaining a missing frame. A mismatched binary can still produce plausible-looking function names and line numbers.

7. What optimization changes

Build an optimized copy:

gcc -std=c11 -Wall -Wextra -g3 -O2 \
  stack_demo.c -o stack_demo_O2

objdump -d -M intel --disassemble=calculate stack_demo_O2

Possible differences include:

  • A local value remains only in a register and has no fixed stack slot.
  • calculate tail-calls add, so its frame is not retained.
  • The compiler omits rbp and relies on rsp plus unwind information.
  • Several source expressions collapse into the same instructions.
  • GDB reports parameters or locals as <optimized out>.

-fno-omit-frame-pointer can improve some profiling and incident-debugging workflows, but GCC explicitly notes that it does not guarantee a frame pointer in every function on every target. Decide whether to enable it in production using the profiler, CPU cost, and unwind support you actually operate; it is not a universal backtrace repair switch.

8. Resolve addresses in PIE executables

Check the ELF type first:

readelf -hW ./stack_demo | grep 'Type:'
readelf -hW ./stack_demo_O2 | grep 'Type:'

The teaching build uses -no-pie and normally reports EXEC. Distribution-default builds often report DYN, meaning PIE. For a live process or core dump, let GDB use the loaded-module information whenever possible:

info files
info sharedlibrary
info proc mappings
info symbol $pc
list *$pc

If a log contains only a runtime address, first identify the owning module, then subtract that module’s load base to obtain a module-relative address:

addr2line -e ./your-app -f -C -i 0x<module-relative-address>

An address inside a shared library must be resolved against the matching .so and debug symbols, not against the main executable. If the wrong library may have loaded, use LD_DEBUG=libs on a trusted program to inspect loader decisions. Do not execute an unknown binary merely to investigate its symbols.

9. Common wrong assumptions

Assumption Better check
All x86-64 arguments are on the stack Check ABI registers first; a stack copy may only be a spill.
rbp is always a frame pointer Inspect disassembly and unwind data; optimized code may omit it.
bt only walks an rbp chain GDB also uses DWARF, CFI, and instruction analysis.
main is the process’s first code _start and the C runtime initialize the process before calling main.
The same C source always produces the same assembly Compiler, flags, target CPU, and hardening defaults all matter.
A function name proves symbols match Verify build IDs and exact executable and library revisions.

10. Stack layout and security boundaries

The return address commonly sits above local stack data, so an out-of-bounds write may corrupt control data. That does not make the layout dependable: optimization, alignment, stack protectors, ASLR, PIE, and control-flow protection can all change the result.

Make memory bugs fail early during testing:

gcc -std=c11 -Wall -Wextra -g3 -O1 \
  -fsanitize=address,undefined -fno-omit-frame-pointer \
  stack_demo.c -o stack_demo_asan

Apply the toolchain and distribution hardening baseline to release builds, for example:

gcc -O2 -g -fstack-protector-strong -D_FORTIFY_SOURCE=2 \
  -fPIE -pie -Wl,-z,relro,-z,now \
  stack_demo.c -o stack_demo_hardened

These options provide defense in depth. They do not replace bounds checks, length validation, or memory-safe design. Sanitizers belong primarily in test and pre-production builds rather than a default production binary.

11. A practical debugging order

  1. Record the architecture, ABI, compiler release, and complete build flags.
  2. Preserve the exact unstripped executable, shared libraries, build IDs, and debug symbols.
  3. Read the generated instructions with objdump; do not guess argument locations from C source.
  4. At function entry, inspect argument registers, rsp, and the return address in GDB.
  5. Compare -O0 with the production optimization level to identify inlining, tail calls, and frame-pointer changes.
  6. Convert PIE or shared-library addresses to module-relative addresses before using addr2line.
  7. When unwinding looks wrong, rule out build mismatches and stack corruption before blaming GDB.

The reliable model is the intersection of the ABI, generated instructions, and runtime state. A single memorized frame diagram stops being useful as soon as an optimized x86-64 binary differs from it.

References