Wildlife Monitor

Status: working  ·  hardware-software  ·  View on GitHub →

#wildlife #ai #computer-vision #nas #raspberry-pi

What It Does

Automated wildlife detection pipeline for security camera and camera trap footage. Videos are pulled from a network-attached storage (NAS), run through AI detection (MegaDetector) and species identification (SpeciesNet), and browsed through a local web dashboard. Footage with detections gets archived in an organised structure; blank footage is purged on a configurable schedule.

Supports dual-lens cameras (fixed wide + telephoto) with synchronised playback and automatic lens pairing.

Hardware

Component Notes
Linux machine (Ubuntu 20.04+) Runs the detection pipeline and dashboard
Network-attached storage (NAS) Accessible via NFS or SMB — stores raw and archived footage
Security cameras / camera traps Any that write video to the NAS
GPU (optional) NVIDIA CUDA 11/12 — ~10x faster than CPU

How It Works

  1. nas_sync.sh pulls recent video from the NAS to local staging
  2. MegaDetector V6 scans each frame for animals, people, and vehicles
  3. SpeciesNet identifies species (2,000+ species, 65M training images) with state/province-level geo-filtering to cut impossible IDs
  4. Each animal crop is scored for image quality (sharpness, brightness, contrast)
  5. Kept footage is archived back to the NAS under camera/year/month/day/
  6. Blank footage is purged on a configurable retention schedule
  7. Web dashboard (FastAPI) lets you browse species, gallery crops, videos, and activity trends

Runs automatically at 6 AM daily via systemd. Dashboard starts on boot.

Installation

See GitHub repo for full instructions.

Quick start:

chmod +x setup.sh nas_connect.sh nas_sync.sh
./setup.sh
./nas_connect.sh   # interactive NAS setup wizard
./nas_sync.sh --then-process --country US --admin1-region US-UT

Build Log


2026-08-20 — Phase 13: a preview feature nobody could reach, and a text harness that already knew

Two small, unrelated fixes to the raw-cleanup subsystem this phase, both closing gaps left over from Phase 6. The interesting part wasn’t either fix — it was what got found while making them.

The first bug had a shape worth naming on its own: a preview feature that technically existed but was functionally unreachable. nas_sync.sh --dry-run looked, from the flag name, like exactly what an operator would want before trusting an automated deletion sweep — a look-before-you-leap check. But all of the raw-cleanup preview logic was nested inside the --then-process branch, so --dry-run used alone just synced videos and exited, having previewed nothing. The only way to actually see what a cleanup sweep would touch was to hand-set an undocumented environment variable, WM_RAW_CLEANUP_DRY_RUN, inside a full --then-process invocation — something no operator would discover without reading the source. A safety feature that only its author knows how to reach isn’t really a safety feature; it’s a checkbox that happens to compile. The fix pulls the preview into a standalone branch reachable from --dry-run alone, with no --then-process and no environment variable required. That’s also a deliberate behavior change, not a strict superset of the old behavior: --dry-run by itself used to sync-then-exit, and now it previews-then-exits instead. Preserving the old sync-then-exit behavior alongside the new preview was considered and rejected — a preview-only invocation gains nothing from copying gigabytes of video to local staging first, and keeping that dead weight around would have just meant two ways to do a no-op.

The more interesting find, though, was archaeological. Planning for this phase assumed nas_sync.sh had no automated coverage — it’s a bash script, and this project’s test story has always been Python-only stdlib harnesses. That assumption was wrong. scripts/verify_raw_cleanup.py has been parsing nas_sync.sh as text since Phase 6, encoding structural invariants across roughly twenty cases — not by running the script, but by asserting things about its literal source: that certain sentinel comments occur exactly once, that a code block’s extent runs from one divider string to another, that an exit guard is keyed on finding a particular variable name. Every one of those techniques has the same failure mode once you add a second, near-identical block of code to the file: silent mistargeting, not a loud error. A first-occurrence scan for a marker comment doesn’t know which occurrence it’s supposed to mean; it just returns the first one, confidently, whether or not that’s the block you actually inserted. Inserting a second preview-flavored copy of the raw-cleanup verification logic without accounting for this would have meant existing test cases quietly started asserting things about the wrong block — passing while checking nothing meaningful. The fix was to give the new branch its own distinct sentinel comments, so every text-scan in the harness has an unambiguous target, and to add one more case on top: a drift guard (PRV8) that diffs the duplicated verification functions against the canonical copy and fails if they ever diverge, since the two copies exist for placement reasons, not because they’re supposed to evolve independently. The general lesson: text-assertion harnesses are cheap, durable, and genuinely useful, but they fail by going quiet, not by getting loud — worth remembering the next time a “just add a similar block” change looks free.

One smaller trap, set -euo pipefail flavored: NAS_ARCHIVE_ROOT and NAS_BLANK_ROOT are only assigned inside the existing --then-process block, further down the script than where the new standalone preview branch needed to sit. The obvious insertion point — right after the mount check, before the --then-process block runs — would have hit an unbound-variable abort the moment the preview tried to read either path. The new branch derives both locally instead of depending on assignments from a block it may never reach.

The second fix, CLEANUP-05, was smaller in every sense: the Settings page’s raw-retention warning previously only checked raw retention against blank-video retention, and said nothing if raw retention outlived kept-video retention instead — a real gap, since a kept archive copy purged before its raw source was ever cross-verified would strand data with no warning at all. The design call worth recording is that this shipped as a second, independent warning stacked alongside the original one, rather than folding both conditions into a single merged message. Each message names the specific archive type actually at risk, and — more importantly — the original blank-retention warning survives byte-for-byte unchanged, which the harness pins directly as a no-regression guarantee rather than trusting a code review to notice a wording drift.

Deployed and verified live on production hardware against the real NAS and a five-figure raw_recordings directory: the preview correctly reported real skip reasons (no_raw, ok) with nothing copied and nothing deleted — the only file-count movement during the test window came from live camera motion events landing mid-check, confirmed by timestamp, not from the preview doing anything. Both retention warnings behaved exactly as designed across all five tested configurations, stacked correctly, and never blocked a save.

2026-08-16 — Phase 12: two species-correction write paths, one that never left the room

Closed out the last of the three open decision items from the v1.3 milestone’s “observability, UX, and monitoring” bucket, and the one that turned out to be a bigger job than the roadmap line implied.

The starting question was small: a gallery tile for a human-corrected species still shows the AI’s original confidence percentage, which is misleading once a human has overridden the label. Fixing the badge meant finding every place a corrected species gets displayed — and that’s where the actual gap surfaced. This codebase has had two independent species-correction write paths since Phase 3: the Gallery popover’s correct_species(), which writes species.user_common_name directly onto the detection row, and the video player’s per-crop editor, save_video_correction(), which writes a separate video_corrections row keyed by (video_id, original_label). Only one of them was ever readable outside the view that wrote it. The Gallery popover’s corrections propagate everywhere, because the shared DISPLAY_COMMON SQL expression reads user_common_name directly. The video player’s corrections were visible only inside that one video’s own detail page — forever, by construction, since nothing else queried video_corrections at all. Nobody noticed for months, because the two correction UIs look and behave identically from the operator’s chair. It took a Phase 10 UAT session, re-verifying an unrelated fix on a video that happened to carry an old video-player correction, for the mismatch to become visible.

The fix followed a pattern this project had already used once: Phase 10’s NOT_EFFECTIVELY_UNKNOWN made the unknown-species suppression filter correction-aware with a correlated subquery instead of a per-caller Python overlay. Display correctness got the same treatment here — a SQL expression that composes into every existing interpolation site — rather than bolting apply_corrections_to_species()-style post-processing onto six separate reader functions individually.

There was a boundary drawn on purpose, though, and it’s worth explaining why. Making the displayed name correction-aware is safe. Making the grouping correction-aware is not — not without a decision this project hasn’t made yet. If a species drilldown groups on the raw AI label but selects a corrected display name within that group, SQLite hands back an arbitrary member row’s value whenever a group is only partially corrected, which means a species tile could render a different name on every page load. That’s a worse failure than the stale-but-stable behavior it would replace. Fixing it for real means changing what a species drilldown, a filter dropdown, and a chart series are keyed on — not just what they display — and that’s a decision that needed to go to the operator rather than get smuggled into a display-layer bug fix.

The badge itself came down to a small but deliberate ordering choice: a stale AI confidence score attached to a label a human has since overridden is worse than no score at all, so the corrected state replaces the percentage rather than annotating it, and the corrected-check runs before the null-confidence guard — so a corrected crop that was never scored by SpeciesNet in the first place still shows the “corrected” signal instead of falling through to a blank badge.

The third item on this phase’s list, NOTIFY-01, closed out as an honest dead end rather than a fix. Five different failure-injection strategies were tried against a copy of the pipeline to simulate a corrupted-video partial-run failure, and OpenCV/ffmpeg absorbed every one of them as a benign “0 frames extracted” rather than the hard error the alert path is built to catch. The alert code stays in production, live and armed — it’s just permanently unverifiable by any means available, which is now written down as an accepted limitation instead of an item that looks perpetually half-done.

Production deployment surfaced nothing new: full harness suite green on the real runtime, timing deltas all within a few percent of baseline, and the operator’s browser walkthrough confirmed the badge, the propagation fix, and the one v1.2-era view that’s supposed to stay unchanged (the video detail page itself) all behaved as designed. The operator also flagged the grouping gap as worth a follow-up — and, digging into why two correction mechanisms exist in the first place, asked whether they should really be two mechanisms at all. Both questions are now filed as their own todo rather than re-discovered the next time a corrected species shows up somewhere it shouldn’t.

2026-08-16 — FIX-02: 14,354 stale file-path references repaired in production

Closed out a data-quality gap that had been sitting quietly in the database since before this project started tracking its own history: about 22% of crops.crop_path and videos.thumbnail_path references — 14,354 rows out of roughly 65,000 — pointed at a /home/nash/... directory prefix that hadn’t existed on disk in years. The box’s home directory was renamed to /home/twostar/ at some point during this project’s early life, and whatever did that rename never touched the database rows that stored full paths under the old name. Nobody had noticed, because nothing had gone looking. It surfaced by accident, mid-rehearsal, during Phase 9’s dedup backfill audit — a completely unrelated operation that happened to read every crop and thumbnail path on its way through. Confirmed unrelated to and unaffected by that backfill before either was allowed near production.

The actual fix is a one-line idea: replace a stale string prefix with the current one, on two columns, no row deletion, no schema change. That idea got four plans and its own phase anyway, because this project’s standing rule is that every production write earns a harness first, full stop — dry-run-first, snapshot-before-write, an operator go/no-go checkpoint, and post-write reconciliation against a baseline captured before the transaction opened. The cost of that discipline on a change this small was low: one script, a 37-case fixture harness, a rehearsal against a live copy of the data. The cost of skipping it once, on the day the “obvious” one-liner turns out not to be so obvious, is unbounded. That asymmetry is the whole argument.

And it did turn out not to be quite so obvious. The first-instinct implementation — a SQL UPDATE ... WHERE crop_path LIKE '/home/nash/%' doing a REPLACE on the matched rows — is subtly wrong. The LIKE clause controls which rows get touched, but SQL’s substring replace doesn’t know or care where in the string the substitution lands. A path containing /home/nash/ a second time later in its own value — plausible if a filename or a nested directory ever echoed the old username — would get corrupted in the tail, not just repaired at the head. The fix that avoids the failure mode entirely: compute the new value in Python as a leading-prefix slice, so only the first, known-position occurrence is ever touched, structurally. The fixture harness carries a dedicated case asserting exactly this — a path with the stale string appearing twice, confirmed rewritten only at the front.

The roadmap’s Success Criterion 3 — “a sample of previously-broken image URLs resolves correctly in the dashboard after migration” — assumed the stale paths meant broken images in the browser. Reading web_app.py‘s serving code before writing anything showed that assumption was shaky: the dashboard’s media routes reduce every stored path to its basename before serving the file, so the directory prefix never actually reached them. The pre-migration baseline measurement confirmed it empirically — all five sampled dashboard URLs (3 crops, 2 thumbnails) already returned HTTP 200 before the migration ran. The dashboard was never broken by this bug. The real damage was elsewhere: the full-path consumers, wildlife_processor.py and backfill_dedup_videos.py, which do use the complete stored path and would have failed to find any of these 14,354 files on disk. Reporting the criterion as “passed, but for a boring reason — no regression, not a visible repair” felt more honest than quietly checking the box.

The production write itself, once rehearsed, was exactly as uneventful as the plan wanted it to be: 14,354 rows rewritten (4,274 crops, 10,080 thumbnails) in a single invocation, no service stopped — this is a two-column string UPDATE, not a DELETE-based consolidation, so there was never a window where a row briefly didn’t exist. Post-write, every figure reconciled exactly against the pre-write baseline: row counts unchanged (17,298 crops, 48,464 videos), both non-path-column fingerprints byte-identical before and after, zero foreign-key violations, and all 14,354 rewritten paths individually confirmed to resolve against the live filesystem — not sampled, checked. A snapshot backup was taken before the transaction opened regardless, cheap insurance on a 94MB database, matching this project’s practice for every prior production write.

2026-08-10 — Historical dedup backfill: 19,289 duplicate video rows consolidated in production

Closed out the last and highest-risk phase of the v1.2 milestone: a one-time data-surgery script to clean up ~19,291 duplicate videos rows left behind by an old archive-collision bug (fixed at the source back in Phase 5, but the historical mess was explicitly deferred until now). Deliberately sequenced last — irreversible production writes only happen after everything else has stabilized.

The plan itself was built in layers on purpose: a read-only audit against real production data first (before any delete logic existed), then a tracer proving one full consolidation path end-to-end on a fixture, then widening that tracer to handle every real shape the audit found — operator corrections, empty winners, a hazard where crops silently migrate to the wrong sibling via INSERT OR REPLACE, dual-lens pairings. Two hazards surfaced just from reading the existing code before any of this ran: nearly every duplicate group (98.7%, it turned out) shares one identical thumbnail file with no uniquifier, so a naive delete would blank the surviving row’s gallery tile; and crops could end up orphaned on the wrong sibling entirely. Both got dedicated reference-guarded handling instead of being discovered the hard way in production.

The fixture harness — 53 cases across 10 suites, all synthetic — passed clean the whole way through. It proved the logic was right. It said nothing about whether the logic was fast enough, and that gap showed up hard during the full-scale rehearsal against a real production copy: a first dry-run attempt was killed after 33+ minutes of runaway CPU, root-caused to an N+1 query pattern invisible at fixture scale (six rows) but catastrophic at real scale (19,291 groups, ~64,000 rows). Fixed with a bulk-prefetch cache. Then the same shape of bug reappeared in the actual write path — an unindexed DELETE ... WHERE detection_id IN (subquery) — fixed by deleting on primary key instead. And then a third layer: even with the query-level fixes, SQLite’s own foreign-key enforcement was still doing full-table scans on every delete, because three FK columns (two on crops/species, one a self-referencing paired_video_id on videos) had never been indexed. That one required stepping outside the phase’s own files into the shared database.py schema — flagged for explicit sign-off rather than treated as a same-file auto-fix, given the blast radius. Three schema-level fixes later, the full 19,291-group pass went from a projected ~2.3 hours down to 42 seconds.

The rehearsal earned its keep: none of those three bugs were reachable from a fixture database small enough to write by hand. Real data volume is a different animal, and “the tests pass” turned out to be a necessary but insufficient claim for anything touching the live database.

The actual production run was almost anticlimactic by comparison — 41 seconds, services stopped for a few minutes total, every post-run check clean (zero FK violations, zero broken pairings, deltas reconciling exactly). 19,289 of 19,291 groups consolidated; the remaining 2 were left alone on purpose, reported by name rather than force-resolved, because both plausible automatic answers would have either destroyed a video’s only crops or duplicated its detections. Quarantine directory ended up empty — turns out almost nothing was actually safe to delete once the shared-thumbnail guard was accounted for, which says something about how deep that particular bug’s fingerprints ran through the data.

2026-05-07 — Project initialized and milestone planned

Started using a proper planning workflow for the first time on this project, which has been running and evolving organically for a while now. The codebase reached a point where it works well enough to trust the output but has enough rough edges that fixing one thing breaks another — the classic “it works on my machine, in this exact order” problem.

The planning pass surfaced six silent failure modes that have been lurking in the code. The most insidious: a NameError: log in database.py that crashes any startup against a database that needs a filepath migration. It never showed up during normal operation — only when something changed in the environment. Similar story with SETTINGS_FILE using a global path instead of honoring --data-dir, meaning settings saved in the dashboard silently went to the wrong place when the service was running with a custom data directory.

The dual-lens sync situation turned out to be three separate bugs stacked on each other: a SQL LIKE wildcard matching _ as any character (so camera_01 could pair with camera_X01), a JavaScript isSyncing guard that resets synchronously and immediately lets the next timeupdate event fire (feedback loop), and no database migration to repair the pairs that got wrongly matched by the SQL bug. All three need to be fixed in the right order.

Planning broke into four phases: fix the six blocking bugs first (Phase 1), then run monitoring + email alerts (Phase 2), scheduling UI + gallery UX improvements (Phase 3), and the dual-lens overhaul (Phase 4). Phases 2–4 can run in parallel once Phase 1 clears. The constraint on dependencies is worth noting: the dual-lens database fix must land before the JavaScript player rewrite, and Phase 1 must land before anything else, because until those six bugs are fixed the codebase isn’t safe to extend.

Key architectural constraint I locked in early: no new pip dependencies. The project already has a heavy ML stack (MegaDetector + SpeciesNet), and adding smtplib-style packages that are already in stdlib, or <dialog> elements that are now baseline browser-native, keeps the installation story clean. Everything new goes into Python stdlib or browser APIs.

Total scope: 26 requirements across the four phases.

Scroll to Top