Skip to content

Collect Linux Logs for Forensics: tar, UAC, Velociraptor

What to collect for a Linux log investigation and how: one tar command on a live host, journalctl export, ausearch, UAC, Velociraptor, or a mounted disk image.

Published on 5 min read

TL;DR. As root, archive /var/log (text logs, journal, audit, wtmp/btmp/lastlog), the volatile journal in /run/log/journal, and four context files: /etc/localtime, /etc/timezone, /etc/passwd and /etc/hostname. Keep every rotation and every .journal~ file, keep the folder layout, hash the archive. One tar command does it; UAC and Velociraptor do it at scale.

The file list

PathContentWhy
/var/log/auth.log*, /var/log/syslog*Debian, Ubuntu text logsSSH, sudo, PAM, cron, services
/var/log/secure*, /var/log/messages*RHEL, Rocky, Alma, Fedora text logsSame, with dated rotations (secure-20260913)
/var/log/journal/<machine-id>/*.journal, *.journal~Persistent systemd journalTrusted fields, microseconds, sequence numbers
/run/log/journal/<machine-id>/*.journalVolatile journalLost at reboot; the only journal on some hosts
/var/log/audit/audit.log*auditdLogins, sessions, executed programs
/var/log/wtmp*, /var/log/btmp*, /var/log/lastlogLogin recordsSessions, failed logins, last login per UID
/run/utmpCurrent sessionsLive host only
/var/lib/wtmpdb/wtmp.db, /var/log/wtmp.dbwtmpdbopenSUSE, Debian 13
/var/lib/lastlog/lastlog2.dblastlog2util-linux 2.40+
/etc/localtime, /etc/timezoneTime zoneFor syslog lines without an offset
/etc/passwd, /etc/hostnameUID to name, host nameAttribution

The time zone files matter more than they look. A traditional RFC 3164 line such as Sep 14 10:02:11 has no year and no zone: without /etc/localtime every conversion to UTC is a guess.

Live host: one tar command

As root on the host:

sudo tar -C / -czf /tmp/linux-logs.tar.gz --ignore-failed-read --sparse \
  var/log run/log/journal etc/localtime etc/timezone etc/passwd etc/hostname

Then copy the archive to your workstation:

scp user@host:/tmp/linux-logs.tar.gz .

Notes on the flags:

  • --ignore-failed-read skips paths that do not exist on this distribution (no /etc/timezone on RHEL, no /run/log/journal on hosts with persistent storage) with a warning instead of failing.
  • --sparse matters for lastlog: it is a sparse file indexed by UID and can look many gigabytes large when a high UID has logged in. --sparse stores only the used blocks.
  • Root is required: auth.log, btmp, audit.log and the journal are not readable by ordinary users.

Writing the archive to /tmp changes the target's disk. If that is unacceptable, write to mounted external media or stream the archive over SSH instead.

If you can only run journalctl

journalctl's export format is lossless and keeps every field:

sudo journalctl -o export > journal.export
sudo journalctl -D /mnt/evidence/var/log/journal -o export > journal.export   # from a mounted image

-o json works too. Raw journal files are still better: the export is what journalctl could read, while the files keep sequence numbers, file headers and damaged tails, which is where deleted entries show up. Always pass -D when reading evidence: without it journalctl reads the analysis host's own journal.

All audit logs in one file

sudo ausearch --input-logs --raw > audit.log

--input-logs reads the log location from auditd.conf, including rotated audit.log.N files, and --raw keeps the original records, including the ENRICHED fields after the 0x1D separator. On a live host also save the loaded rules: auditctl -l and auditctl -s. Loaded rules can differ from /etc/audit/rules.d/.

UAC

UAC (Unix-like Artifacts Collector) is a shell script that runs on almost any Unix. With only the log artifacts:

sudo ./uac -a files/logs/var_log.yaml,files/logs/run_log.yaml,files/system/etc.yaml,files/system/utmp.yaml /tmp

-a takes a comma-separated list of artifact files. The output is uac-<host>-linux-<date>.tar.gz, which Linux Log Parser opens directly. Or use the standard triage profile, which includes the same logs and much more:

sudo ./uac -p ir_triage /tmp

UAC ships the profiles ir_triage, full, offline and offline_ir_triage. The offline ones are meant for a mounted image rather than a running system.

Velociraptor

Collect the raw files with Generic.Collectors.File, rooted at /, from the GUI, a hunt, an offline collector or the command line:

velociraptor artifacts collect Generic.Collectors.File \
  --args Root=/ --args collectionSpec=$'Glob\nvar/log/**\nrun/log/journal/**\netc/localtime\netc/timezone\netc/passwd\netc/hostname' \
  --output linux-logs.zip

Drop the resulting ZIP on the tool as it is. Velociraptor also has parsing artifacts such as Linux.Forensics.Journal, Linux.Sys.LastUserLogin, Linux.Syslog.SSHLogin and Linux.Ssh.AuthorizedKeys. They are useful for hunting across a fleet, but they return rows, not files, so they cannot be re-parsed or checked for sequence gaps later. Collect the files as well.

Dead box or disk image

Mount the root file system read-only (ideally without replaying the ext4 journal), then archive the same paths:

sudo mount -o ro /dev/sdX1 /mnt/evidence        # or an image mounted read-only
tar -C /mnt/evidence -czf ~/linux-logs.tar.gz --ignore-failed-read --sparse \
  var/log etc/localtime etc/timezone etc/passwd etc/hostname

The volatile journal in /run/log/journal and /run/utmp only exist on a running system. If the host was powered off and journald was volatile, the journal is gone; memory acquisition is the only other route.

Gotchas

  • Take every rotation. auth.log.1, auth.log.2.gz, secure-20260913, wtmp.1, audit.log.4: the older history is where the first access usually is.
  • Take the .journal~ files. They are journals renamed after an unclean stop or detected corruption, and often cover the incident.
  • Collect user journals. journald numbers entries across all files of a machine, so user-<UID>.journal files that are left behind look like deleted entries.
  • Preserve modification times. The year of traditional syslog lines is inferred from them. tar keeps them; a copy through some file-sharing tools does not.
  • Note the time. logrotate may run while you collect. Hash the archive and record when it was made.
  • Check the architecture. wtmp records are 384 bytes on x86_64 but 400 bytes on aarch64; note which one the host was.

Next step

Drop the archive, the UAC tarball, the Velociraptor ZIP or the folder on Linux Log Parser. The pillar guide explains what each source contributes, and the walkthrough shows a full analysis. The journal cheat sheet and auth.log cheat sheet list per-distribution paths in more detail.

Related articles

How to prove Linux logs were edited or deleted: journal seqnum gaps, lines missing from auth.log, blanked wtmp records, stopped daemons and clearing commands.
How the four Linux log families fit together in an investigation, what each one proves, where they disagree, and how to merge them into one timeline.
systemd journal forensics without journalctl: the file format, trusted fields, compression, sequence-number gaps, dirty .journal~ files and time fields.