File Integrity Monitor
I built a dependency-free integrity application with two deliberate operating modes: Integrity Desk for visual local review and a command-line interface for automation. Both paths use the same SHA-256 baseline engine and produce deterministic evidence for added, modified, deleted, moved, and unreadable files.
Why I built it
A file-integrity tool should answer more than whether two hashes differ. An operator needs to know what baseline was trusted, what was added or removed, which metadata changed, whether the scan completed, and whether the evidence itself was protected from casual alteration.
This project keeps the engine small enough to inspect, uses only the Python standard library, and exposes the same behavior through a CLI and Integrity Desk, a local browser interface for people who do not want to manage every scan from a terminal.
Baseline lifecycle
The first scan creates a deterministic inventory of the selected tree. Each record binds a normalized relative path to file type, size, timestamps where appropriate, and SHA-256 content evidence. Later scans compare current observations with the chosen baseline and classify additions, removals, content changes, and metadata changes.
Baseline creation is intentionally separate from accepting a new trusted state. If every change automatically becomes the new baseline, an unauthorized modification can erase its own evidence during the next run.
Interface and reports
Integrity Desk lets the user select a directory, create or load a baseline, run a comparison, review grouped changes, and export a report. The UI calls the same core functions as the CLI, which reduces the chance that automation and the visible result disagree.
Reports are deterministic so equivalent input produces comparable output. Clear statuses distinguish a completed clean scan from an error, skipped path, permission problem, or interrupted run.
Failure cases
The scanner handles unreadable files, disappearing paths, symbolic-link decisions, permission errors, malformed baselines, large trees, and files that change during observation. A result is not labelled clean when important paths could not be examined.
The project also documents an essential limitation: if an attacker can modify both monitored files and the baseline or reporting environment, hashing alone cannot establish truth. Baselines need separate access control, backups, or signed external storage depending on the threat model.
Verification and tradeoffs
In a controlled 500-file fixture set, the monitor detected 45 of 45 expected changes: 20 modifications, 10 deletions, 10 additions, and 5 moves, with zero scan errors.
The project has nine passing automated tests covering deterministic scanning, change classification, same-size content tampering, rename inference, saved dashboard evidence, required-path validation, reports, error states, and the shared interface/core behavior.
Using only the standard library keeps installation and review simple, but it means the project does not depend on a database, background task queue, kernel event feed, or real-time filesystem watcher. It favours reproducible point-in-time evidence over continuous monitoring.
Lessons and next work
The main lesson is that integrity is a lifecycle, not a hash function. Trust in the initial state, protection of the baseline, completeness of observation, review of changes, and controlled acceptance of a new state all matter.
Possible next steps include signed baseline manifests, scheduled scans, stronger platform-specific metadata, protected remote evidence, performance profiles for larger trees, and documented recovery actions for each class of finding.