Memory 01 - The virtual memory of a user process in Linux
Linux Process Memory Learning · Next: Diagnosing process memory use in Linux
The Slovak original of this document: Memory 01 - Virtuálna pamäť užívateľského procesu v OS Linux (slovensky).
1 Intro
1.1 Physical memory
Physical memory is a hardware storage device. More detail -> https://en.wikipedia.org/wiki/Computer_memory
1.2 Virtual memory
Virtual memory provides software-controlled management of memory addresses, letting every program/process have its own unique view of the computer's memory.
On a multi-tasking operating system every process runs in its own memory sandbox, called the virtual address space. On a 32-bit platform that is a 4 GB block of addresses.
One way an OS manages memory is paging, which lets a program's physical address space be discontiguous. Linux uses paging, as do other operating systems. Under paging, the virtual address space is mapped onto physical memory through page tables the kernel maintains.
1.3 Paging
Under paging, both physical and logical memory are divided into blocks of the same size: for physical memory we speak of frames, for logical memory of pages. Since the blocks are of constant size, it is enough to number the pages and it is not necessary to record whole page addresses. The block diagram below shows physical addresses being mapped onto virtual ones by paging.
+----------+
0| |
+----------+ +---+ +----------+
| page 0 | 0| 1 | 1| page 0 |
+----------+ +---+ +----------+
| page 1 | 1| 5 | 2| page 2 |
+----------+ +---+ +----------+
| page 2 | 2| 2 | 3| |
+----------+ +---+ +----------+
| page 3 | 3| 8 | 4| |
+----------+ +---+ +----------+
| ... | 4| | 5| page 1 |
+----------+ +---+ +----------+
| page N | 5| | 6| |
+----------+ +---+ +----------+
7| |
+----------+
8| page 3 |
+----------+
logical memory page table physical memory1.4 The linear address space
From user space the address space is a flat linear one, but the kernel sees it differently: it is split in two, the user address space and the kernel's own. Where the split falls is set by "PAGE_OFFSET", which on x86 puts it at address "0xc0000000". So 1 GiB is always mapped by the kernel and the remaining 3 GiB are available to user processes.
1.5 Address space management in Linux
The address space a Linux process can use is managed by the data structure "mm_struct". Every address space is made of some number of memory regions, which never overlap. A memory region may be the process's heap, where "malloc()" allocates; a file mapped into memory, such as a shared library; or anonymous memory allocated with "mmap()".
1.6 The memory regions of a Linux process
A process rarely uses the whole of its address space; usually only parts of the memory regions are in use. Each region is represented by the structure "vm\_area\_struct" and, as said above, regions never overlap. The list of every memory region of a Linux process can be read through the PROC interface, in "/proc/[pid]/maps", where [pid] is the process identifier.
1.7 The system calls that act on a Linux process's memory regions
- fork() - Creates a new process with a new address space. Every page is marked COW (copy on write) and shared between the two processes until a page fault makes private copies of them.
- clone() - Creates a new process, letting it share part of its parent's context. This is how threads are implemented on Linux.
- mmap() - Creates a new memory region in the process's address space.
- mremap() - Remaps a memory region, or changes its size.
- munmap() - Removes part or all of a memory region.
- shmat() - Attaches a shared memory segment to the process's address space.
- shmdt() - Detaches a shared memory segment from the process's address space.
- execve() - Loads a new executable, overwriting the current address space.
- exit() - Tears down the address space and every memory region of the process.
2 The virtual memory of a Linux program or process
!!! Warning !!! Some of what follows applies to kernel 3.10. On other kernel versions it may differ.
2.1 A block diagram of a Linux process's memory
The block diagram below shows how the memory of a Linux process written in C is laid out.
Higher memory addresses
0xffffffff -------> =============================
| Kernel / System | User programs may neither read nor write these addresses
| | An attempt to read or write ends in a segmentation fault
| | (Segmentation Fault).
0xc0000000 -------> =============================
|###########################|
|###########################| Empty memory space
|###########################| Random stack offset
|###########################|
start_stack -----------> ============================= . . . . . . . . . . . . . . . . . . . . ==========================
direction | | Stack | | STACK segment |
of growth | | | | |
| | env | | |
| | argv | | |
| | argc | | |
| ----------------------------- | |
| | automatic variables of | | |
| | the function "main()" | | |
| ----------------------------- | |
| | automatic variables of | | |
v | the function "func()" | | |
stack_pointer -----------> ============================= . . . . . . . . . . . . . . . . . . . . ==========================
(points at the top of the stack)|###########################|
|###########################| Empty memory space
|###########################| Available for stack growth
|###########################|
mmap_base -----------> ============================= . . . . . . . . . . . . . . . . . . . . ==========================
direction | | malloc.o (lib*.so) | Mapped files (library functions when | Memory mapping segment |
of growth | | | linked dynamically), or | |
v | printf.o (lib*.so) | anonymous mappings | |
------------ ============================= . . . . . . . . . . . . . . . . . . . . ==========================
|###########################|
|###########################| Empty memory space
|###########################| Available for heap growth
program break |###########################|
brk -----------> ============================= . . . . . . . . . . . . . . . . . . . . ==========================
^ | Heap | | HEAP segment |
| | | | |
direction | | malloc() | | |
of growth | | calloc() | | |
| new | | |
start_brk -----------> ============================= . . . . . . . . . . . . . . . . . . . . ==========================
|###########################|
|###########################| Empty memory space
|###########################| Random brk offset
|###########################|
-----------> ============================= . . . . . . . . . . . . . . . . . . . . ==========================
end_bss | Global variables | Uninitialised data | BSS segment |
| | (initialised to zero) | (.bss) |
| char *s; | | |
start_bss -----------> |---------------------------| . . . . . . . . . . . . . . . . . . . . ==========================
end_data | int n = 10; | Initialised data | DATA segment |
| | (variables the programmer initialised) | (.data) |
| char *s = "String"; | | |
start_data -----------> ============================= . . . . . . . . . . . . . . . . . . . . ==========================
end_code | malloc.o (lib*.a) | Library functions | TEXT / CODE (ELF) |
| | when linked | segment |
| printf.o (lib*.a) | statically | |
============================= . . . . . . . . . . . . . . . . . . . . | |
| program.o | | Compiled code |
| | | (program.out) |
----------------------------- | |
| main.o | | |
| | | Process binary image |
| func() | <--- Return address | (/bin/program) |
============================= | |
| | | |
| crt0.o (start-up routine) | | |
| | | |
start_code ------------ ============================= ==========================
Lower memory addresses2.2 The memory descriptor of a Linux process
The Linux kernel represents a process's address space with a data structure called the memory descriptor. It holds everything to do with the process's address space. On Linux (kernel 3.10) it is the structure "mm\_struct", defined in the header "<linux/mm\_types.h>" (in older versions, in <linux/sched.h>). There is exactly one "mm\_struct" per process, shared by the process's threads.
struct mm_struct {
struct vm_area_struct *mmap; /* Pointer to the head of the list of memory area objects (Virtual Memory Areas). */
/* Put another way, it is the list of memory areas (Virtual Memory Areas). */
struct rb_root mm_rb; /* Pointer to the root of the red-black tree of memory area objects. */
struct vm_area_struct *mmap_cache; /* Pointer to the most recently used (referenced) memory area. */
unsigned long mmap_base; /* Base address of the memory mapping (mmap) areas. */
unsigned long task_size; /* Total size of the process address space. */
unsigned long cached_hole_size; /* If non-zero, it holds the size of the largest free memory hole that lies */
/* below the address "free_area_cache". */
unsigned long free_area_cache; /* First address pointing at a free hole of size "cached_hole_size" or larger. */
/* The address from which the kernel starts looking for a free range of linear addresses in the */
/* process address space. */
unsigned long highest_vm_end; /* Highest end/last address of a virtual memory area. */
pgd_t *pgd; /* Pointer to the so-called Page Global Directory. Every process has this pointer set */
/* to its own PGD, which in fact is a physical page frame. */
atomic_t mm_users; /* Secondary usage counter. Number of processes sharing the "mm_struct" data structure. */
atomic_t mm_count; /* Primary usage counter. */
int map_count; /* Number of memory areas. */
spinlock_t page_table_lock; /* Page tables lock. Protects the page tables and some counters. */
struct rw_semaphore mmap_sem; /* Semaphore of the memory areas (read/write semaphore). */
struct list_head mmlist; /* List of all "mm_struct" structures. */
unsigned long hiwater_rss; /* High-watermark of the RSS (Resident Set Size) figure. */
unsigned long hiwater_vm; /* High-watermark of the total size of the process address space (number of */
/* pages of the process). */
unsigned long total_vm; /* Total size of the process address space (number of pages of the process). */
unsigned long locked_vm; /* Number of locked pages, which cannot be moved to SWAP space. */
/* These are the pages that have the "PG_mlocked" flag set. */
unsigned long pinned_vm; /* Number of pages of the process address space that are permanently pinned in memory. */
/* TODO -> find out what exactly is meant by this pinning. */
unsigned long shared_vm; /* Number of shared pages (files). */
unsigned long exec_vm; /* Number of pages of the process address space with the permissions VM_EXEC & ~VM_WRITE set */
unsigned long stack_vm; /* Number of pages of the process address space that belong to the stack. VM_GROWSUP/DOWN */
unsigned long def_flags; /* Default access permissions (flags) of the memory areas. */
unsigned long nr_ptes; /* Page table pages. */
unsigned long start_code; /* Start address of the code segment. */
unsigned long end_code; /* End address of the code segment. */
unsigned long start_data; /* Start address of the data segment. */
unsigned long end_data; /* End address of the data segment. */
unsigned long start_brk; /* Start address of the heap segment. */
unsigned long brk; /* End address of the heap segment. */
unsigned long start_stack; /* Start address of the stack segment. */
unsigned long arg_start; /* Start address of the program arguments. */
unsigned long arg_end; /* End address of the program arguments. */
unsigned long env_start; /* Start address of the environment variables. */
unsigned long env_end; /* End address of the environment variables. */
/* Special counters, which in some configurations are protected by "page_table_lock" and in other configurations by the fact that the */
/* operations are atomic. */
struct mm_rss_stat rss_stat;
struct linux_binfmt *binfmt;
cpumask_var_t cpu_vm_mask_var; /* Bit mask for the lazy TLB (Translation Lookaside Buffer) switch. */
/* The TLB (Translation Lookaside Buffer) is a cache on the processor (in the MMU unit) that is */
/* used to reduce the time needed to access a user memory location. */
/* In other words, the TLB stores recent translations of virtual addresses to physical addresses. */
mm_context_t context; /* Architecture-specific data/context. */
...
};The "mm\_struct" structure has more members than these, but for memory work this definition is enough. The whole of it for kernel 3.10 is at -> http://lxr.linux.no/linux+v3.10/include/linux/mm_types.h#L325
3 Memory information about a Linux program or process
3.1 Memory statistics for a Linux process
The basic memory figures for a program or process, measured in pages, are read like this.
# cat /proc/[pid]/statm -------------------------------------------------------------------------------- size resident shared text lib data dt -------------------------------------------------------------------------------- 40938 1257 894 261 0 408 0
- size/VmSize - Total program size in memory
- resident/VmRSS - Resident set size
- shared/RssFile+RssShmem - Resident shared pages
- text/code - Text (Code)
- lib - Library (unused since 2.6; always 0)
- data - Data + Stack
- dt - Dirty pages (unused since 2.6; always 0)
3.2 The memory regions of a Linux process, described
The virtual memory regions of a program, process or thread are listed like this.
# cat /proc/[pid]/maps -------------------------------------------------------------------------------- address perm offset dev inode path -------------------------------------------------------------------------------- 00400000-00505000 r-xp 00000000 fd:00 50818536 /usr/bin/mc 00705000-0070a000 r--p 00105000 fd:00 50818536 /usr/bin/mc 0070a000-0070f000 rw-p 0010a000 fd:00 50818536 /usr/bin/mc 0070f000-00747000 rw-p 00000000 00:00 0 01e5f000-01f10000 rw-p 00000000 00:00 0 [heap]
- address - The first and last address of the memory region in the address space of the program or process.
- perm - How the pages of the region may be reached:
r = read [reading allowed] w = write [writing allowed] x = execute [execution allowed] s = shared [shared virtual memory] p = private (copy on write) [private virtual memory]
Note: the permissions can be changed with the "mprotect()" system call.
- offset - If the memory region is mapped from a file with "mmap()", this is the offset in that file where the mapping starts. For a region not mapped from a file the offset is zero.
- device - If the memory region is mapped from a file, this is the major and minor number of the device the file is on.
- inode - If the memory region is mapped from a file, this is the file's inode number.
- path - If the memory region is mapped from a file, this is the path to it. It can be empty for an anonymous mapping.
There are special memory regions, such as: [stack] = the stack of the program/process [heap] = the heap of the program/process [vdso] = virtual dynamic shared object
4 Reading the memory regions of a Linux process
Linux exposes a process's virtual memory through the "/proc" pseudo filesystem, in the pseudo file "/proc/[pid]/mem". That file shows the contents of the process's memory regions exactly as they are mapped in the process itself: the byte at offset X in "/proc/[pid]/mem" is the byte at address X in the process. If the address is unmapped in the process, reading at that offset fails with an input/output error ("EIO", "Input/Output Error"). Since nothing is usually mapped on a process's first page, reading the first page fails that way. You can test this by trying to read a process's memory with "cat /proc/[pid]/mem" as "root".
Not every memory region can be read. A process's memory regions are recorded in the "/proc" filesystem, in the pseudo text file "/proc/[pid]/maps". Basically it is the memory map of the process. To read another process's memory you first read the description of its regions from "/proc/[pid]/maps", which also says how each region may be reached (read, write and so on). Knowing which regions allow access, you can read and copy them out of the pseudo file "/proc/[pid]/mem" with "read()" and "mmap()", but reading the memory is not that straightforward and certain conditions have to be met.
For one process (the reader) to read the memory regions of another (the target), certain conditions have to be met:
- A process that wants to read another process's memory regions from "/proc/[pid]/mem" has to attach to that process with "ptrace()" and the flag "PTRACE\_ATTACH". When it has finished reading it should detach again with "ptrace()" and "PTRACE\_DETACH". This is how debuggers attach to a process. A reader running as "root" does not have to attach, but the target process still has to be stopped.
- The target process, whose memory is being read, has to be stopped. A running process allocates memory as it goes and so changes its own memory regions, which means the reader may try to read a region (or pages) that has already been unmapped, and get "EIO", "Input/Output Error" back, see the beginning of this chapter above.
"ptrace()" with "PTRACE\_ATTACH" stops the target process by sending it a STOP signal. Signal delivery is asynchronous, so the reader has to wait a moment for the target to change state; it does so by calling "waitpid()", which suspends the reader until the process named by the PID changes state.
A piece of C that attaches to a target process by PID and reads one of its memory regions:
/* Keep the path to the PROC pseudo file in "mem_file_name". */ /* E.g.: mem_file_name = "/proc/15482/mem". */ sprintf(mem_file_name, "/proc/%d/mem", pid); /* Open the PROC pseudo file for reading (O_RDONLY). */ mem_fd = open(mem_file_name, O_RDONLY); /* Attach to the target process with the "ptrace()" system call */ /* with the flag "PTRACE_ATTACH" set. */ ptrace(PTRACE_ATTACH, pid, NULL, NULL); /* Since "ptrace()" is asynchronous, the reader has to wait a moment. */ waitpid(pid, NULL, 0); /* In the open PROC pseudo file, set the offset to the region */ /* we want to read. */ lseek(mem_fd, offset, SEEK_SET); /* Read a page of memory from the open PROC pseudo file (mem_fd) */ read(mem_fd, buf, sysconf(_SC_PAGE_SIZE)); /* Detach from the target process with the "ptrace()" system call */ /* with the flag "PTRACE_DETACH" set. */ ptrace(PTRACE_DETACH, pid, NULL, NULL); /* Close the PROC pseudo file. */ close(mem_fd);
5 Tools for reading, dumping and searching the memory of a process - Linux
5.1 Volatility Framework
The Volatility Framework is open source and written in Python. Releases are available in zip and tar archives, Python module installers, and standalone executables.
- [http://www.volatilityfoundation.org](http://www.volatilityfoundation.org "test")
- https://github.com/volatilityfoundation/volatility/wiki
- https://www.aldeid.com/wiki/Volatility
- https://www.aldeid.com/wiki/Volatility/Retrieve-hostname
- https://www.aldeid.com/wiki/Volatility/Retrieve-password
5.2 memdump
A utility to dump memory of unixy processes
5.3 The Coroner's Toolkit (TCT) - memdump (Wietse Venema)
5.4 memgrep (Matt Miller a.k.a skape)
5.5 Memfetch (Michal Zalewski a.k.a lcamtuf)
Memfetch, a simple utility to take non-destructive snapshots of process address space.
5.6 fmem - kernel driver, that creates /dev/fmem device (Ivo Kollar a.k.a niekt0)
/dev/fmem behave in same way that /dev/mem (direct access to physical memory), but does not have limits that /dev/mem have. It is possible to dump whole physical memory through /dev/fmem. (Alternative to /dev/crash)
5.7 Foriana - FOrensic Ram Image ANAlyzer (Ivo Kollar a.k.a niekt0)
Version 1.0 can list processes and modules from memory dump of i386/x86_64/arm linux/bsd kernels, and provide option for reading linear memory from dumps. Theory is described in my master thesis (english).
- http://hysteria.cz/niekt0/
- http://hysteria.cz/niekt0/foriana/foriana_current.tgz
- http://hysteria.cz/niekt0/foriana/doc/foriana.pdf
5.8 LiME ~ Linux Memory Extractor
A Loadable Kernel Module (LKM) which allows for volatile memory acquisition from Linux and Linux-based devices, such as Android. This makes LiME unique as it is the first tool that allows for full memory captures on Android devices. It also minimizes its interaction between user and kernel space processes during acquisition, which allows it to produce memory captures that are more forensically sound than those of other tools designed for Linux memory acquisition.
5.9 BASH memory dumper
A small script I wrote for the BASH shell.
6 Tools for reading, dumping and searching the memory of a process - Windows
6.1 memgrep - Memory grep
A memory searching utility across multiple processes on Windows platform.
- Opens each process
- Works out the valid memory pages
- Search for ascii and unicode incarnation of the string
- https://github.com/nccgroup/memgrep
- http://blog.wirhabenstil.de/2016/01/19/searching-greping-inside-live-process-memory/
6.2 FireEye - Memoryze
7 memdump - installing the tool and examples of its use [5.2]
7.1 Downloading and unpacking memdump
# cd /install # wget https://github.com/bitw1ze/memdump/archive/master.zip # mv ./master.zip ./memdump.zip # unzip ./memdump.zip
7.2 Compiling memdump
# cd /install/memdump-master # gcc main.c memdump.c -o memdump
7.3 Running memdump's help
# cd /install/memdump-master
# ./memdump -h
----------------------------------------------------------------------------------------------------------------
Usage: ./memdump <segment(s)> [opts] -p <pid>
Options:
-A dump all segments
-D dump data segments
-S dump the stack
-H dump the heap
-d [dir] save dumps to custom directory [dir]
-p [pid] pid of the process to dump
-v verbose
-h this menu7.4 Dumping selected memory regions of a process
-S - dump the stack segment -H - dump the heap segment -p - dump the process whose PID=102347 -d - write the output into the directory mc.dump ---------------------------------------------------------------------------------------------------------------- # cd /install/memdump-master # mkdir dumps # ./memdump -d ./dumps/mc.dump -S -H -p 102347
The command wrote the process's memory regions to these files:
# ls -lh /install/memdump-master/dumps/mc.dump ---------------------------------------------------------------------------------------------------------------- -rw-r--r--. 1 root root 528K Dec 2 11:23 0000000002715000-0000000002799000.dump -rw-r--r--. 1 root root 132K Dec 2 11:23 00007ffd2f019000-00007ffd2f03a000.dump -rw-r--r--. 1 root root 11K Dec 2 11:23 maps
7.5 Identifying the files of the stack and heap regions - using the "maps" file
# cd /install/memdump-master/dumps/mc.dump # cat ./maps | grep 'heap\|stack' ---------------------------------------------------------------------------------------------------------------- 0000000002715000-0000000002799000 rw-p 0000000000000000 00:00 0 [heap] 00007ffd2f019000-00007ffd2f03a000 rw-p 0000000000000000 00:00 0 [stack]
From that output it is clear that:
- the stack region of the "mc" process was written to the file -> 00007ffd2f019000-00007ffd2f03a000.dump
- the heap segment of the "mc" process was written to the file -> 0000000002715000-0000000002799000.dump
7.6 Searching the contents of the memory regions
[1] Write every occurrence of a string found in the heap into the file "mc-heap-strings" [2] Write every occurrence of a string found in the stack into the file "mc-stack-strings" ---------------------------------------------------------------------------------------------------------------- # cd /install/memdump-master/dumps/mc.dump [1]# strings 0000000002715000-0000000002799000.dump > ./mc-heap-strings [2]# strings 00007ffd2f019000-00007ffd2f03a000.dump > ./mc-stack-strings
7.7 Dumping every memory region of a process
-A - dump every memory region -p - dump the process whose PID=102347 -d - write the output into the directory "mc.dump.all" ---------------------------------------------------------------------------------------------------------------- ./memdump -d ./dumps/mc.dump.all -A -p 102347
The files "memdump" wrote every memory region of the process into are:
# ls -lh /install/memdump-master/dumps/mc.dump.all ---------------------------------------------------------------------------------------------------------------- -rw-r--r--. 1 root root 1.1M Dec 2 11:58 0000000000400000-0000000000505000.dump -rw-r--r--. 1 root root 20K Dec 2 11:58 0000000000705000-000000000070a000.dump -rw-r--r--. 1 root root 20K Dec 2 11:58 000000000070a000-000000000070f000.dump -rw-r--r--. 1 root root 224K Dec 2 11:58 000000000070f000-0000000000747000.dump -rw-r--r--. 1 root root 528K Dec 2 11:58 0000000002715000-0000000002799000.dump -rw-r--r--. 1 root root 24K Dec 2 11:58 00007f4677010000-00007f4677016000.dump -rw-r--r--. 1 root root 102M Dec 2 11:58 00007f4677016000-00007f467d53f000.dump -rw-r--r--. 1 root root 8.0K Dec 2 11:58 00007f467d9c5000-00007f467d9c7000.dump -rw-r--r--. 1 root root 8.0K Dec 2 11:58 00007f467dbdf000-00007f467dbe1000.dump -rw-r--r--. 1 root root 4.0K Dec 2 11:58 00007f467e225000-00007f467e226000.dump -rw-r--r--. 1 root root 16K Dec 2 11:58 00007f467ef5b000-00007f467ef5f000.dump -rw-r--r--. 1 root root 20K Dec 2 11:58 00007f467fa8f000-00007f467fa94000.dump -rw-r--r--. 1 root root 16K Dec 2 11:58 00007f467fcac000-00007f467fcb0000.dump -rw-r--r--. 1 root root 4.0K Dec 2 11:58 00007f467ffe6000-00007f467ffe7000.dump -rw-r--r--. 1 root root 400K Dec 2 11:58 00007f4680930000-00007f4680994000.dump -rw-r--r--. 1 root root 44K Dec 2 11:58 00007f4680b9e000-00007f4680ba9000.dump -rw-r--r--. 1 root root 4.0K Dec 2 11:58 00007f4680baa000-00007f4680bab000.dump -rw-r--r--. 1 root root 28K Dec 2 11:58 00007f4680bab000-00007f4680bb2000.dump -rw-r--r--. 1 root root 4.0K Dec 2 11:58 00007f4680bb2000-00007f4680bb3000.dump -rw-r--r--. 1 root root 4.0K Dec 2 11:58 00007f4680bb5000-00007f4680bb6000.dump -rw-r--r--. 1 root root 132K Dec 2 11:58 00007ffd2f019000-00007ffd2f03a000.dump -rw-r--r--. 1 root root 8.0K Dec 2 11:58 00007ffd2f03f000-00007ffd2f041000.dump -rw-r--r--. 1 root root 4.0K Dec 2 11:58 ffffffffff600000-ffffffffff601000.dump -rw-r--r--. 1 root root 11K Dec 2 11:58 maps
8 memdump (Wietse Venema) - installing the tool and examples of its use [5.3]
This software is by the well-known Dutch programmer and physicist Wietse Venema, author of the even better-known Postfix mail server, of TCP wrapper, and of the digital-forensics toolkit TCT (The Coroner's Toolkit). memdump comes from TCT.
8.1 Downloading and unpacking memdump (Wietse Venema)
I do not know why the archive has a ".gz" extension when it is not a gzipped TAR archive. It is a plain POSIX TAR archive, so "-x" on its own is enough.
# cd /install # wget http://www.porcupine.org/forensics/memdump-1.01.tar.gz # tar -xvf ./memdump-1.01.tar.gz
8.2 Compiling memdump
This software was written for Linux kernel "2" or "2.4", and the system I tested on is RHEL 7.3 running kernel 3.10, so we will not waste time adapting it for newer versions of Linux, because adequate replacements exist. The advantage of this software is that it can save memory regions over the network to a remote server, which is very useful in digital forensics (we do not overwrite fragments of the evidence).
9 memgrep - installing the tool and examples of its use [5.4]
Warning: this software was written when the 64-bit platform was not yet widespread, so I recommend using it only on a 32-bit one.
9.1 Downloading and unpacking memgrep
# cd /install # wget http://www.hick.org/code/skape/memgrep/memgrep-0.8.0.tar.gz # tar -xzvf memgrep-0.8.0.tar.gz # cd memgrep-0.8.0
9.2 Getting memgrep to compile
"memgrep" has a few ailments that we have to solve before the compilation itself. They come from the fact that this software is a bit older and was written for older versions of the Linux kernel. During the life of the Linux kernel changes were made in the memory subsystem and in other parts of its code, so without editing the code of memgrep itself this software cannot be compiled.
9.2.1 Failure to include the header "memgrep.h"
In the source file "memgrep.c" the include of the header "memgrep.h" fails. There are two ways to fix it:
- A> by editing the source file "memgrep.c", see below
# vi /install/memgrep-0.8.0/src/memgrep.c ---------------------------------------------------------------------------------------------------------------- Wrong: #include "memgrep.h" Correct: #include "../include/memgrep.h"
- B> by putting the file "include/memgrep.h" into the "src/" directory, see below
# cp /install/memgrep-0.8.0/include/memgrep.h /install/memgrep-0.8.0/src/
9.2.2 Failure to include the header "<linux/user.h>"
The source file "memgrep.c" tries to include "<linux/user.h>", where the structure "user\_regs\_struct" used to be defined. It now lives in a different header, "<sys/user.h>". A small edit fixes the mismatch.
# vi /install/memgrep-0.8.0/src/memgrep.c ---------------------------------------------------------------------------------------------------------------- Wrong: #include <linux/user.h> Correct: #include <sys/user.h>
9.2.3 Failure to include the header "<sys/ptrace.h>"
The source file "memgrep.c" uses the flags of the "ptrace()" system call (PTRACE\_ATTACH, PTRACE\_DETACH, PTRACE\_GETREGS and so on), but the Linux part of the code never includes the right header, "<sys/ptrace.h>". This header is included only in the FreeBSD part of the code. A small edit fixes it: after the include of "<sys/user.h>", add one for "<sys/ptrace.h>".
# vi /install/memgrep-0.8.0/src/memgrep.c ---------------------------------------------------------------------------------------------------------------- Original: #include <sys/user.h> ... Modified: #include <sys/user.h> #include <sys/ptrace.h>
9.2.4 Error in the explicit declaration of the function "ptrace()"
The source file "memgrep.c" declares "ptrace()" explicitly, and that declaration conflicts with the definition in "<sys/ptrace.h>". Comment the explicit declaration out.
# vi /install/memgrep-0.8.0/src/memgrep.c ---------------------------------------------------------------------------------------------------------------- Original: extern long int ptrace (unsigned long int cmd, unsigned long int pid, void *param, unsigned long int data); Modified: /* LH OFF - remove the explicit declaration of "ptrace()" extern long int ptrace (unsigned long int cmd, unsigned long int pid, void *param, unsigned long int data); */
9.2.5 The structure "user\_regs\_struct" has no member "esp" on a 64-bit platform
The source file "memgrep.c" uses the structure "user\_regs\_struct", which on 64-bit platforms has no "esp" member. We solve this by compiling "memgrep" for the 32-bit platform. We tell the compiler (gcc) that we want the 32-bit version (the -m32 switch), see the compilation in [9.3]. To build a 32-bit binary on a 64-bit system we will need the extra packages of 32-bit libraries, so install them.
# yum install glibc-devel.i686 # yum install libgcc-4.8.5-11.el7.i686
9.3 Compiling memgrep
Once every problem described in [9.2.1] through [9.2.5] is solved, we can finally get to compiling.
# cd /install/memgrep-0.8.0/src # gcc -m32 -Wall -O3 memgrep.c -o memgrep
9.4 Running memgrep's help
# cd /install/memgrep-0.8.0/src/
# ./memgrep -h
----------------------------------------------------------------------------------------------------------------
memgrep -- Run-time/core-time memory searching, dumping and modifying utility.
Usage: ./memgrep [-p pid] [-o core] [-T] [-d] [-r] [-s] [-e] [-a addr1,addr2,bss,addr3] [-l length]
[-f fmt,search data] [-t fmt,replace data] [-b pad] [-m minimum size]
[-F fmt] [-L] [-v] [-h]
-p [pid] The process id to operate on.
-o [core] The core file to operate on.
-T Build a referential tree for the given address(es).
-d Dump memory from the specified address(es) for the given length (-l).
-r Replace memory at the specified address(es). If -s is also specified.
only memory that matches the search criteria will be replaced.
-s Search memory at the specified address(es).
-e Enumerate the heap.
-a [addr] The address(es) to operate on seperated by commas. Addresses can be
in the following format:
0x821c4ac
821c4ac
Also, the following keywords can be used:
bss -> Uses the VMA associated with the .bss section (uninit global vars, heap data).
rodata -> Uses the VMA associated with the .rodata section (read-only data, ie, static text).
data -> Uses the VMA associated with the .data section (data, ie, global variables).
text -> Uses the VMA associated with the .text section (text, ie, executable code).
stack -> Dynamically determines the current stack pointer.
all -> Uses bss, stack, rodata, data, text. This is the only keyword that can be used
when operating on core files.
-l [len] The length to use when searching or dumping. A length of 0 means search
till end-of-memory.
-f [data] This specifies the search criteria. Multiple formats are accepted for ease
of use. Below are accepted formats and their examples:
s -> String format (Ex: 's,Testing')
x -> Hex format (Ex: 'x,00414100AB')
i -> Integer format (Ex: 'i,4724')
-t [data] This specifies the replace data. The same formats used with the -f parameter
are valid for the -t parameter.
-m [minsz] The minimum size of a heap allocation for use when enumerating.
-b [pad] Number of bytes of padding to use around dump addresses (default is 0).
-F [fmt] The format to use when dumping memory, can be one of the following:
hexint -> Four byte hexi-decimal integers.
hexshort -> Two byte hexi-decimal shorts.
hexbyte -> One byte hexi-decimal characters.
decint -> Four byte decimal integers.
decshort -> Two byte decimal shorts.
decbyte -> One byte decimal characters.
printable -> Printable characters.
-L List memory segments of a process or core file.
-v Version information.
-h Help.
Example search (search for 'Jane' in .bss):
./memgrep -p 1335 -s -a bss -f s,Jane
Example replace (replace memory at 0x8423143 and 0x8443147 with 0x00ff0041):
./memgrep -p 1335 -r -a 0x8423143,0x8443147 -t x,00ff0041
Example search/replace (Replace 'Test' with 'Rest' in .bss and .rodata):
./memgrep -p 1335 -s -r -a bss,rodata -f s,Test -t s,Rest
Example dump (Dump memory starting at 0x8422113 for 16 bytes):
./memgrep -p 1335 -d -a 0x8422113 -l 16TODO
TODO
99 Links
99.1 Talks and slides on the Linux memory system
- [1] 2016 - Alan Ott - Virtual Memory and Linux (****)
- [2] 2015 - Memory Subsystem and Data Types in the Linux Kernel - Praktikum Kernel Programming (****)
- [3] Free Electrons - Complete training materials (*****) - A lot of interesting information about low level things, kernel development and so on !!!
99.2 The Linux memory system in detail
These works are older (for kernels 2.4, 2.6), but quite a lot of the areas/terms still apply. I wish the fourth edition of the book Linux Device Drivers was already out.
- [1] 2007 - Ulrich Drepper - What Every Programmer Should Know About Memory (*****)
- [2] 2004 - Understanding the Linux Virtual Memory Manager (*****)
- [3] 2005 - Linux Device Drivers, Third Edition - Memory Mapping and DMA
99.3 Further documents on the Linux memory system
An excellent document on the changes to the memory management subsystem in Linux, covering the kernel from 2.6.32 to 4.0-rc4.
- [2] 2016 - An adaptive approach for Linux memory analysis based on kernel code reconstruction
- [3] 2015 - Deep dive into linux memory management
- [4] COMPILER, ASSEMBLER, LINKER AND LOADER: A BRIEF STORY (****)
- [5] 2005 - What is linux-gate.so.1?
99.4 Tutorial - Intersec Techtalk (****)
- [1] 2013 - Intersec Techtalk - Memory – Part 1: Memory Types
- [2] 2013 - Intersec Techtalk - Memory – Part 2: Understanding Process memory
- [3] 2013 - Intersec Techtalk - Memory – Part 3: Managing memory
- [4] 2014 - Intersec Techtalk - More about locality
- [5] 2013 - Intersec Techtalk - Memory – Part 4: Intersec’s custom allocators
- [6] 2013 - Intersec Techtalk - Memory – Part 5: Debugging Tools
- [7] 2014 - Intersec Techtalk - Memory – Part 6: Optimizing the FIFO and Stack allocators
99.5 Tutorial - Gustavo Duarte - Software Illustrated (****)
- [1] 2009 - Getting Physical With Memory
- [2] 2009 - Anatomy of a Program in Memory
- [3] 2009 - How the Kernel Manages Your Memory
- [4] 2009 - Page Cache, the Affair Between Memory and Files
- [5] 2010 - Journey to the Stack, Part I
- [6] 2010 - Epilogues, Canaries, and Buffer Overflows
99.6 Older material on the Linux memory system
- [1] 2003 - KernelAnalysis-HOWTO - 7. Linux Memory Management
- [2] 2003 - The Linux Kernel - 9. Memory
- [3] 1999 - The Linux Kernel - Chapter 3 Memory Management
Current practice (checked 2026-10)
mm_structfields: the listing in 2.2 is the kernel 3.10 layout. The VMA list headmmapand the red-black treemm_rbwere replaced in Linux 6.1 by one maple tree (struct maple_tree mm_mt), andhighest_vm_endwent with them.mmap_cachewas removed in 3.15,free_area_cacheandcached_hole_sizein 3.11,shared_vmin 4.5 (there isdata_vmnow),nr_ptesbecamepgtables_bytesin 4.15,mmap_semwas renamedmmap_lockin 5.8, andrss_statis an array of per-CPU counters. Read the structure from the source of the kernel you run.- Address space layout: the 4 GB space with the 3 GiB / 1 GiB split at
0xc0000000describes 32-bit x86. On x86-64 a process has about 128 TB of user address space with 4-level page tables and about 64 PB with 5-level page tables, and the stack, mmap and heap bases are randomised (ASLR). - Reading
/proc/[pid]/mem: chapter 4 says the reader has to attach withPTRACE_ATTACHand the target has to be stopped. Current kernels require neither: proc_pid_mem(5) says access is governed by a ptrace access mode check (PTRACE_MODE_ATTACH_FSCREDS), so a process that would be allowed to attach can open and read the file. Stopping the target is still sensible if you want a consistent snapshot.process_vm_readv()(since Linux 3.2) reads another process's memory under a similar ptrace access check without going through/procat all. - Yama
ptrace_scope: the article does not mention it, but it decides whether the above works for a non-root user. Withkernel.yama.ptrace_scope0 any process of the same UID can attach, with 1 only an ancestor of the target can, with 2 only a process withCAP_SYS_PTRACE, and 3 disables attaching until reboot. Check the value before concluding that a dump tool is broken. /dev/memand fmem: on x86 kernels built withCONFIG_STRICT_DEVMEMRAM cannot be read through/dev/mem, and kernel lockdown (Linux 5.4 and later), where it is enabled, blocks/dev/mem,/dev/kmemand/dev/kcoreand only loads signed modules. That rules out an out-of-tree module such as fmem on a locked-down system.- Dumping one process: instead of compiling the small dump tools from 5.2 to 5.5 as root, use
gcorefrom the gdb package, which writes a core file of a running process that gdb and other tools can read. - Forensic tools: the Volatility 2 repository linked in 5.1 was archived on 16 May 2025 and needs Python 2; its successor is Volatility 3 (Python 3, a rewrite), which for Linux images needs a symbol table generated for the exact kernel. LiME has moved to
github.com/jtsylve/LiME(the old URL redirects). AVML from Microsoft acquires memory from userland without a kernel module, using/dev/crash,/proc/kcoreor/dev/mem. - memgrep and TCT memdump: the article already notes that these do not build cleanly on 64-bit systems. Nothing has changed there; treat them as history.
$ cat /proc/sys/kernel/yama/ptrace_scope $ gcore -o /tmp/mc 102347
Sources: