systemd Journal Forensics: Files, Seqnum Gaps, .journal~
systemd journal forensics without journalctl: the file format, trusted fields, compression, sequence-number gaps, dirty .journal~ files and time fields.
TL;DR. A journal file is a binary, append-only database with the signature LPKSHHRH. Every entry carries a sequence number that journald increments across all files of a machine, so a missing range of numbers is visible even after the entries are gone. Read the files themselves, not only a journalctl export: headers, sequence numbers, .journal~ files and user journals are where deletion shows.
Where the files are
| Path | Content |
|---|---|
/var/log/journal/<machine-id>/system.journal | Active system journal (persistent storage) |
/var/log/journal/<machine-id>/system@<id>-<seqnum>-<time>.journal | Archived (rotated) system journals |
/var/log/journal/<machine-id>/user-<UID>.journal | Per-user journal for regular UIDs (default SplitMode=uid) |
*.journal~ | Files renamed after an unclean journald stop or detected corruption |
/run/log/journal/<machine-id>/ | Volatile journal, lost at reboot |
<machine-id> matches /etc/machine-id. Whether anything is written to /var depends on Storage= in journald.conf: read the configuration on the image rather than assuming.
Why the journal is worth the trouble
The systemd journal stores each entry as FIELD=value pairs. The fields that start with an underscore are added by journald from the kernel's view of the sender, so the sending process cannot forge them:
_PID,_UID,_GID,_COMM,_EXE,_CMDLINE: who really sent the message._SYSTEMD_UNIT,_SYSTEMD_USER_UNIT: which service or user service it came from._BOOT_ID: which boot; reboots become explicit boundaries._AUDIT_SESSION,_AUDIT_LOGINUID: the audit session and login UID, which link the entry to auditdsesand auid._TRANSPORT:syslog,journal,stdout,kernel,audit,driver.
MESSAGE, SYSLOG_IDENTIFIER and PRIORITY are supplied by the client. Anyone can write sshd: Accepted password ... with logger; the underscore fields will show it came from logger under that user's UID.
The file format in brief
- Header. Signature
LPKSHHRH, compatible and incompatible feature flags,state(offline, online, archived),machine_id,seqnum_id,head_entry_seqnumandtail_entry_seqnum, first and last realtime timestamps, tail boot ID, and object counts. Current systemd writes a 272-byte header; readers must use the header size stored in the file, since older files have shorter headers. - Objects. After the header come typed objects:
DATA(one field=value payload, deduplicated by hash),FIELD,ENTRY(the list of DATA objects making up one entry, with seqnum, realtime, monotonic and boot ID),ENTRY_ARRAY, hash tables and, with Forward Secure Sealing,TAG. - Compression. Large
DATApayloads may be compressed with XZ, LZ4 or ZSTD, flagged per object. The LZ4 variant stores the uncompressed size as a 64-bit little-endian integer before the LZ4 block. - Compact mode. Since systemd 252, files may use the compact layout, with 32-bit offsets instead of 64-bit ones. Older journalctl versions refuse such files, which is one reason to use a recent reader.
A parser that walks the objects directly can read entries that journalctl would skip, for example from a file whose header was not updated after a crash.
Sequence numbers: the built-in tamper evidence
Each entry has a seqnum. journald increments one counter per machine (identified by seqnum_id) and uses it across all files it writes, system and user alike. Sort every entry from every file by seqnum within one seqnum_id, and the numbers should be continuous. A journal seqnum gap means entries that once existed are not in the files you have.
Benign causes come first:
- User journals not collected. Entries written to
user-1001.journalconsume numbers too. If onlysystem.journalwas collected, every user entry appears as a gap. - Vacuumed archives.
journalctl --vacuum-*and size-based rotation delete old archived files; the gap then sits at the beginning of the retained range. - Rotation boundaries between an archive you have and one you do not.
The interesting case is a gap inside the incident window, in a set where all user journals are present, or a gap that lines up with a user journal deleted by hand. In the tool's sample, the attacker removes user-1001.journal: the system journal keeps its numbers, but the entries that went to the user file are missing, and the user-manager lines they contained survive only in /var/log/syslog (rsyslog received them too).
Dirty and online files
- A
.journal~file was renamed because journald did not close it cleanly or found it corrupt. It often contains the last minutes before a crash or a hard power-off. Parse it; the tail may be partially written. - A header
stateof online in a collected file means journald had it open. That is normal forsystem.journalcopied from a live host, but the last entries may be incomplete. - A file whose header says it contains more entries than can be read, or whose tail is damaged, is worth noting in the report, not dropping.
Time fields
| Field | Meaning |
|---|---|
__REALTIME_TIMESTAMP | When journald received the entry, microseconds since the epoch (UTC) |
_SOURCE_REALTIME_TIMESTAMP | Earliest trusted time of the message at its source, when it differs from reception; journalctl displays this one when present |
__MONOTONIC_TIMESTAMP | Microseconds since boot, meaningful with _BOOT_ID |
Realtime values follow the system clock. A clock set backwards shows up as realtime going down while seqnum goes up; within one boot, the monotonic clock gives the true order.
Reading exports
If raw files are not available, journalctl -o export is lossless per entry (binary values are length-prefixed) and -o json is convenient. Both lose the file headers and are limited to what journalctl could read. See the collection guide.
journalctl -D /mnt/evidence/var/log/journal --header # file ranges, state, seqnum ids
journalctl -D /mnt/evidence/var/log/journal --list-boots --utc
journalctl -D /mnt/evidence/var/log/journal -o export > journal.export
In the browser
Linux Log Parser parses regular and compact journal files, XZ/LZ4/ZSTD-compressed fields, .journal~ files and exports, reports sequence-number gaps (and warns when no user journal was loaded), flags dirty and online files, and marks auth.log lines that the journal also holds. Auth lines present in the journal but missing from auth.log are reported as a separate finding, covered in the tampering guide.
The systemd journal cheat sheet on linuxforensics.app lists the retention settings and journalctl commands in detail.