Linux Memory Allocation: malloc, brk, mmap, and RSS
Use a reproducible program, /proc, and strace to inspect how glibc malloc uses brk, mmap, and madvise, and why RSS may stay high after free.
malloc() is not a system call. It is a user-space allocator interface. glibc first searches chunks it already manages through tcache, bins, and arenas. Only when that space is insufficient may it request more virtual address space through brk or mmap.
At least four different quantities appear in discussions about “memory usage”:
- Requested bytes: the size passed to
malloc. - Virtual address space: ranges represented by
VmSize, VMAs, and page tables. - Resident memory: pages currently backed in RAM, usually observed through
VmRSSorsmaps. - Allocator-retained memory: chunks already passed to
freebut still held by glibc for reuse.
Mixing those values produces claims such as “free does nothing,” “anything above 128KB always uses mmap,” or “a large allocation requires physically contiguous RAM.” None is a reliable general rule.
1. The real path from malloc to a physical page
A common path looks like this:
malloc(size)
|
+-- reusable chunk exists in glibc tcache / bins / arena
| `-- return it without a system call
|
`-- existing space is insufficient
+-- extend an arena with brk or mmap
`-- possibly mmap a large request directly
|
`-- kernel creates or extends a virtual mapping
|
`-- first write to each page triggers a fault
`-- kernel allocates and maps a physical page
There is no one-to-one relationship between malloc and a system call. One brk or mmap operation can supply many later allocations, while a free call may remain entirely in user space.
brk and mmap primarily change virtual address space. Anonymous pages are commonly backed on first access. A contiguous user-space virtual range does not require contiguous physical pages.
2. Observe allocation, page touching, and release
The following program accepts a byte count, calls malloc, writes one byte per page, and pauses twice. The first pause lets you inspect the live allocation. The second follows free and malloc_trim.
Save it as malloc_probe.c:
#define _GNU_SOURCE
#include <errno.h>
#include <malloc.h>
#include <stdint.h>
#include <stdio.h>
#include <stdlib.h>
#include <unistd.h>
static size_t parse_size(const char *text)
{
char *end = NULL;
errno = 0;
unsigned long long value = strtoull(text, &end, 10);
if (errno != 0 || end == text || *end != '\0' ||
value == 0 || value > SIZE_MAX) {
fprintf(stderr, "invalid byte count: %s\n", text);
exit(EXIT_FAILURE);
}
return (size_t)value;
}
int main(int argc, char **argv)
{
if (argc != 2) {
fprintf(stderr, "usage: %s BYTES\n", argv[0]);
return EXIT_FAILURE;
}
const size_t bytes = parse_size(argv[1]);
const long page_size_value = sysconf(_SC_PAGESIZE);
if (page_size_value <= 0) {
perror("sysconf");
return EXIT_FAILURE;
}
const size_t page_size = (size_t)page_size_value;
unsigned char *buffer = malloc(bytes);
if (buffer == NULL) {
perror("malloc");
return EXIT_FAILURE;
}
for (size_t offset = 0; offset < bytes; ) {
buffer[offset] = 0x5a;
if (bytes - offset <= page_size) {
break;
}
offset += page_size;
}
buffer[bytes - 1] = 0x5a;
printf("pid=%ld ptr=%p bytes=%zu\n",
(long)getpid(), (void *)buffer, bytes);
puts("inspect allocation, then press Enter");
fflush(stdout);
getchar();
free(buffer);
const int trimmed = malloc_trim(0);
printf("free complete; malloc_trim=%d\n", trimmed);
puts("inspect again, then press Enter");
fflush(stdout);
getchar();
return EXIT_SUCCESS;
}
Compile it and trace memory-related system calls:
gcc -std=c11 -O2 -g -Wall -Wextra malloc_probe.c -o malloc_probe
strace -f -e trace=brk,mmap,munmap,madvise \
./malloc_probe $((64 * 1024 * 1024))
At the first pause, open another terminal and use the printed PID:
PID=<pid-from-program>
grep -E '\[heap\]|rw-p' /proc/$PID/maps
grep -E 'Vm(Size|RSS|Data)' /proc/$PID/status
cat /proc/$PID/smaps_rollup
Press Enter so the program calls free and malloc_trim(0), then inspect the same files again. Observe the result instead of assuming it:
- The pointer may fall inside
[heap]or a separate anonymous mapping. - RSS normally grows after the program writes to every page.
freemay causemunmapormadvise, or no relevant system call at all.- A
malloc_trimreturn value of1means some memory was released; it does not promise that RSS returns to its original value.
Run both a smaller and a larger request:
./malloc_probe $((64 * 1024))
./malloc_probe $((64 * 1024 * 1024))
Do not infer a permanent threshold from those two runs. glibc can change its policy using allocation history, environment variables, tunables, architecture, and release-specific behavior.
3. The boundary around brk and sbrk
The program break marks the end of the traditional process heap. Linux’s brk system call attempts to move that boundary. sbrk is a historical libc interface that applies a relative increment; it is not a separate Linux system call.
The main arena can extend the program break when conditions allow. Growth occurs at the top, so a freed chunk in the middle cannot be returned by lowering the break alone. The allocator usually keeps it in a bin for reuse or applies another page-reclamation mechanism when whole pages become available.
Do not move brk manually inside a normal program that still uses malloc, printf, threading libraries, or other libc components. The program break is process-wide state. Bypassing the allocator corrupts glibc’s view of heap boundaries and chunks. Use strace, /proc/<pid>/maps, and allocator statistics for observation instead of implementing half of an allocator inside the process.
4. mmap creates independent virtual mappings
Anonymous mmap can create virtual ranges separate from the traditional [heap]. glibc commonly uses it in two situations:
- A large allocation receives a direct mapping.
- An additional arena receives heap space to reduce multithreaded contention.
A direct mapping can usually be removed as a unit with munmap, so virtual size and RSS are more likely to fall on release. The tradeoff includes more VMAs, page-table work, TLB effects, and system calls. An allocator balances reuse against prompt return to the kernel.
An mmap range can be virtually contiguous while its physical pages remain scattered. A user process requesting 64MB of anonymous memory does not require the buddy allocator to find one contiguous 64MB physical block. That constraint belongs to some kernel allocation APIs, not ordinary user-space malloc.
5. Modern glibc is more than brk plus a fixed threshold
The rule “below 128KB uses brk; above 128KB uses mmap” is not accurate enough for current glibc. The decision also depends on:
- Whether the current thread’s tcache already contains a suitable chunk.
- Availability in the arena’s fastbins, small bins, large bins, unsorted bin, and top chunk.
- Which arena serves the thread and whether that arena must grow.
- Tunables such as
mmap_threshold,trim_threshold,top_pad, andarena_max. - glibc release, architecture, and prior allocation and release behavior.
glibc documentation still lists 128KiB as the initial default mmap_threshold, but dynamic adjustment is enabled by default. Setting glibc.malloc.mmap_threshold or the corresponding mallopt parameter also changes that behavior. These predictions are therefore unreliable:
- “129KB must produce one
mmapcall.” - “64KB always comes from
[heap].” - “The same binary must take the same path on another host.”
Use GLIBC_TUNABLES for controlled experiments, such as limiting the arena count:
GLIBC_TUNABLES=glibc.malloc.arena_max=2 \
./malloc_probe $((64 * 1024 * 1024))
This is not a universal production setting. Fewer arenas can reduce virtual memory and fragmentation, but can also increase lock contention under concurrency. Establish a baseline with real workload, RSS, latency, and allocation profiles before changing it.
6. Why RSS may stay high after free
free(pointer) directly returns a chunk to the allocator. What happens next depends on where the chunk came from and what surrounds it:
- Tcache or a bin receives the chunk: no system call; RSS normally stays unchanged.
- A directly mapped large allocation is released: glibc commonly calls
munmap, making a drop more likely. - A large free range reaches the top of an arena: the allocator may lower the program break.
- Whole pages inside an arena become free: the allocator may use
madvisewhile retaining the virtual range. - Live chunks separate free ranges: fragmentation prevents returning one continuous region.
malloc_trim(0) is a best-effort reclamation request. Neither its return value nor RSS alone proves or disproves a leak. A useful investigation order is:
- Check whether object counts or heap profiles continue to grow.
- Compare anonymous private pages in
smaps_rollup, not only total RSS. - Account for multiple arenas, thread stacks, JIT memory, file mappings, and shared memory.
- Test whether later allocations reuse the retained memory.
High RSS after free may be allocator retention, or it may be a real leak. The two cases require different evidence; one top snapshot cannot distinguish them.
7. Page faults, overcommit, and container OOM
A non-null result from malloc does not mean physical RAM has already been reserved for every byte. A common anonymous-memory sequence is:
- The allocator obtains virtual addresses.
- The process writes to a page for the first time.
- The CPU raises a page fault.
- The kernel allocates, zeroes, and maps a physical page.
As a result, a process that reserves but does not touch memory can have a large VmSize and a small RSS. Do not test the shared zero page by reading uninitialized malloc memory; that read has no dependable meaning in C. Write each page as the example does, or experiment directly with anonymous mmap.
Linux overcommit behavior is controlled by settings including vm.overcommit_memory:
- Mode
0: heuristic overcommit. - Mode
1: always overcommit. - Mode
2: enforce a commit limit.
Even when the host allows overcommit, a container may hit its cgroup limit first. In a cgroup v2 environment, inspect at least:
cat /sys/fs/cgroup/memory.current
cat /sys/fs/cgroup/memory.max
cat /sys/fs/cgroup/memory.events
Increasing oom or oom_kill counters in memory.events identifies cgroup memory pressure. Free memory on the host does not rule out a container OOM.
8. A sufficient observation toolkit
Do not rely on one metric. Each group of tools answers a different question.
System calls: did the allocator grow or release mappings?
strace -f -e trace=brk,mmap,munmap,madvise ./malloc_probe 67108864
This shows address-space changes, but not reuse inside tcache or bins.
Address space: which mapping owns the pointer?
cat /proc/$PID/maps
pmap -x $PID
[heap] is the traditional program-break heap. An unlabeled anonymous mapping may belong to an additional arena, a direct mmap allocation, a thread stack, or another runtime component. One maps line is not enough to identify its source.
Resident pages: what is actually counted in RSS?
cat /proc/$PID/smaps_rollup
grep -E 'Vm(Size|RSS|Data|Swap)' /proc/$PID/status
Inspect Rss, Pss, Private_Dirty, Anonymous, and Swap. Shared libraries, file-backed cache, and anonymous heap pages have different ownership and reclamation behavior.
Page faults: what did first touch cost?
perf stat -e page-faults,minor-faults,major-faults \
./malloc_probe 67108864
First writes to anonymous memory normally produce minor faults. Major faults commonly involve file I/O or swap. Interpret the counters together with the mapping type.
9. Do not confuse malloc with kernel allocators
A user process cannot call kmalloc, slab, or vmalloc directly. Those are kernel-internal facilities:
kmallocserves kernel objects with particular physical-contiguity requirements.- slab/SLUB caches repeatedly allocated kernel object types.
vmalloccreates contiguous kernel virtual addresses backed by noncontiguous physical pages.- The page allocator supplies physical pages, and page tables map them into user virtual space.
User-space malloc manages virtual addresses through brk and mmap; the kernel then resolves pages on fault. Seeing malloc(1GB) in application code does not imply that the kernel is requesting one contiguous 1GB physical block.
Kernel structure layouts change between releases. For user-space RSS incidents, stable evidence comes from system calls, /proc, cgroup files, and allocator documentation—not a copied struct mm_struct definition from an old kernel tree.
10. Five common questions
Q1: RSS remains high after free. Is that a leak?
Not necessarily. Verify whether live objects or heap profiles keep growing, then separate tcache/bin retention, arena fragmentation, thread stacks, and other anonymous mappings.
Q2: How can I tell whether an allocation used brk or mmap?
Use strace for system calls and /proc/<pid>/maps for the pointer’s mapping. If the allocator reused an existing chunk, the allocation may show neither call.
Q3: Does anything above 128KB always use mmap?
No. 128KiB is glibc’s documented initial default threshold, not a fixed boundary. Dynamic adjustment, tunables, arena state, and glibc release all affect the choice.
Q4: Why did malloc succeed, but the process died while writing memory?
Overcommit may allow address-space reservation before first touch creates real memory pressure. A container may also hit memory.max and trigger a cgroup OOM.
Q5: Why is [heap] absent or barely growing in maps?
The process may rely mostly on direct mmap, additional arenas, another allocator, or a program break that has not grown. [heap] is not a summary of all dynamic process memory.
11. Investigation order
- Decide whether the signal is requested bytes, virtual size, RSS, PSS, or cgroup usage.
- Use
straceto detectbrk,mmap,munmap, andmadviseactivity. - Use
mapsto locate the address andsmaps_rollupto break down resident pages. - Inside a container, inspect
memory.current,memory.max, andmemory.eventstogether. - Prove a leak with a heap profiler or live-object counts; do not equate retained RSS with a leak.
- Change glibc arena, threshold, or trim tunables only after a baseline identifies them as the bottleneck.
- Retest with pinned glibc, kernel, and container-runtime versions; allocation policy is not constant across releases.
The useful model is not a memorized threshold. It is one evidence chain connecting allocator state, virtual mappings, page faults, and cgroup limits.