Skip to content

systemd Journal Forensics: Files, Seqnum Gaps, .journal~

systemd journal forensics without journalctl: the file format, trusted fields, compression, sequence-number gaps, dirty .journal~ files and time fields.

Published on 5 min read

TL;DR. A journal file is a binary, append-only database with the signature LPKSHHRH. Every entry carries a sequence number that journald increments across all files of a machine, so a missing range of numbers is visible even after the entries are gone. Read the files themselves, not only a journalctl export: headers, sequence numbers, .journal~ files and user journals are where deletion shows.

Where the files are

PathContent
/var/log/journal/<machine-id>/system.journalActive system journal (persistent storage)
/var/log/journal/<machine-id>/system@<id>-<seqnum>-<time>.journalArchived (rotated) system journals
/var/log/journal/<machine-id>/user-<UID>.journalPer-user journal for regular UIDs (default SplitMode=uid)
*.journal~Files renamed after an unclean journald stop or detected corruption
/run/log/journal/<machine-id>/Volatile journal, lost at reboot

<machine-id> matches /etc/machine-id. Whether anything is written to /var depends on Storage= in journald.conf: read the configuration on the image rather than assuming.

Why the journal is worth the trouble

The systemd journal stores each entry as FIELD=value pairs. The fields that start with an underscore are added by journald from the kernel's view of the sender, so the sending process cannot forge them:

  • _PID, _UID, _GID, _COMM, _EXE, _CMDLINE: who really sent the message.
  • _SYSTEMD_UNIT, _SYSTEMD_USER_UNIT: which service or user service it came from.
  • _BOOT_ID: which boot; reboots become explicit boundaries.
  • _AUDIT_SESSION, _AUDIT_LOGINUID: the audit session and login UID, which link the entry to auditd ses and auid.
  • _TRANSPORT: syslog, journal, stdout, kernel, audit, driver.

MESSAGE, SYSLOG_IDENTIFIER and PRIORITY are supplied by the client. Anyone can write sshd: Accepted password ... with logger; the underscore fields will show it came from logger under that user's UID.

The file format in brief

  • Header. Signature LPKSHHRH, compatible and incompatible feature flags, state (offline, online, archived), machine_id, seqnum_id, head_entry_seqnum and tail_entry_seqnum, first and last realtime timestamps, tail boot ID, and object counts. Current systemd writes a 272-byte header; readers must use the header size stored in the file, since older files have shorter headers.
  • Objects. After the header come typed objects: DATA (one field=value payload, deduplicated by hash), FIELD, ENTRY (the list of DATA objects making up one entry, with seqnum, realtime, monotonic and boot ID), ENTRY_ARRAY, hash tables and, with Forward Secure Sealing, TAG.
  • Compression. Large DATA payloads may be compressed with XZ, LZ4 or ZSTD, flagged per object. The LZ4 variant stores the uncompressed size as a 64-bit little-endian integer before the LZ4 block.
  • Compact mode. Since systemd 252, files may use the compact layout, with 32-bit offsets instead of 64-bit ones. Older journalctl versions refuse such files, which is one reason to use a recent reader.

A parser that walks the objects directly can read entries that journalctl would skip, for example from a file whose header was not updated after a crash.

Sequence numbers: the built-in tamper evidence

Each entry has a seqnum. journald increments one counter per machine (identified by seqnum_id) and uses it across all files it writes, system and user alike. Sort every entry from every file by seqnum within one seqnum_id, and the numbers should be continuous. A journal seqnum gap means entries that once existed are not in the files you have.

Benign causes come first:

  • User journals not collected. Entries written to user-1001.journal consume numbers too. If only system.journal was collected, every user entry appears as a gap.
  • Vacuumed archives. journalctl --vacuum-* and size-based rotation delete old archived files; the gap then sits at the beginning of the retained range.
  • Rotation boundaries between an archive you have and one you do not.

The interesting case is a gap inside the incident window, in a set where all user journals are present, or a gap that lines up with a user journal deleted by hand. In the tool's sample, the attacker removes user-1001.journal: the system journal keeps its numbers, but the entries that went to the user file are missing, and the user-manager lines they contained survive only in /var/log/syslog (rsyslog received them too).

Dirty and online files

  • A .journal~ file was renamed because journald did not close it cleanly or found it corrupt. It often contains the last minutes before a crash or a hard power-off. Parse it; the tail may be partially written.
  • A header state of online in a collected file means journald had it open. That is normal for system.journal copied from a live host, but the last entries may be incomplete.
  • A file whose header says it contains more entries than can be read, or whose tail is damaged, is worth noting in the report, not dropping.

Time fields

FieldMeaning
__REALTIME_TIMESTAMPWhen journald received the entry, microseconds since the epoch (UTC)
_SOURCE_REALTIME_TIMESTAMPEarliest trusted time of the message at its source, when it differs from reception; journalctl displays this one when present
__MONOTONIC_TIMESTAMPMicroseconds since boot, meaningful with _BOOT_ID

Realtime values follow the system clock. A clock set backwards shows up as realtime going down while seqnum goes up; within one boot, the monotonic clock gives the true order.

Reading exports

If raw files are not available, journalctl -o export is lossless per entry (binary values are length-prefixed) and -o json is convenient. Both lose the file headers and are limited to what journalctl could read. See the collection guide.

journalctl -D /mnt/evidence/var/log/journal --header       # file ranges, state, seqnum ids
journalctl -D /mnt/evidence/var/log/journal --list-boots --utc
journalctl -D /mnt/evidence/var/log/journal -o export > journal.export

In the browser

Linux Log Parser parses regular and compact journal files, XZ/LZ4/ZSTD-compressed fields, .journal~ files and exports, reports sequence-number gaps (and warns when no user journal was loaded), flags dirty and online files, and marks auth.log lines that the journal also holds. Auth lines present in the journal but missing from auth.log are reported as a separate finding, covered in the tampering guide.

The systemd journal cheat sheet on linuxforensics.app lists the retention settings and journalctl commands in detail.

Related articles

How to prove Linux logs were edited or deleted: journal seqnum gaps, lines missing from auth.log, blanked wtmp records, stopped daemons and clearing commands.
What to collect for a Linux log investigation and how: one tar command on a live host, journalctl export, ausearch, UAC, Velociraptor, or a mounted disk image.
A worked Linux log investigation on a synthetic jump host: SSH password guessing, a login from a new IP, sudo -i, cron and systemd persistence, log tampering.