No description
  • C++ 89.6%
  • Python 3.4%
  • Makefile 2.9%
  • Shell 1.9%
  • CMake 1.6%
  • Other 0.6%
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
2026-06-09 03:39:28 -07:00
scripts Initial commit 2026-06-09 03:39:28 -07:00
src Initial commit 2026-06-09 03:39:28 -07:00
.gitignore Initial commit 2026-06-09 03:39:28 -07:00
Makefile Initial commit 2026-06-09 03:39:28 -07:00
readaccel.c Initial commit 2026-06-09 03:39:28 -07:00
README.md Initial commit 2026-06-09 03:39:28 -07:00
run.sh Initial commit 2026-06-09 03:39:28 -07:00
STATE.md Initial commit 2026-06-09 03:39:28 -07:00
zones.conf.example Initial commit 2026-06-09 03:39:28 -07:00

kinectsentry

A room security / activity monitor built on the Xbox 360 Kinect (v1). It captures RGB + depth continuously, learns the static room, detects and tracks the people in it, builds a 3D scan of each person, logs what they do, and serves a live web monitor.

Built from handoff.md; reuses the Kinect capture and web-streaming code from the sibling kinectosc project.

Status

Implemented — all eight handoff stages:

  1. Capture + web — Kinect RGB+depth, web monitor at :8080.
  2. Foreground segmentation — farthest-surface background model; the foreground mask is viewable at /mask.
  3. People — depth-blob clustering with optional YOLO confirmation; stable per-person IDs; boxed and labelled in /live.
  4. Events + log — person_entered / person_left / person_moved / dwell to events.jsonl and kinectsentry.log.
  5. Reconstruction — optional Open3D GPU SLAM backend for the handheld room scan: RGB-D odometry tracks the camera, a CUDA TSDF voxel grid accumulates the room. Off by default; see Stage 5 / Open3D below.
  6. Person 3D model — each tracked person's foreground points are registered into a skeleton-based body-local frame (rigid torso frame) and accumulated into a colour voxel grid; snapshotted to subject_<id>.ply when they leave.
  7. 3D /room page — the learned room, moving people, and the per-person subject scans, all in one interactive WebGL view.
  8. Polish — user-defined room zones (zone_entered/zone_exited), --armed motion events, size-based log rotation.

RoomModel is the stable §4 interface; it has two backends — the fixed-camera background model (default) and the Open3D SLAM backend (stage 5, opt-in). Both expose the static room as a depth+colour image registered to the live view, so foreground segmentation is identical for either.

Note: the subject scan (§6) and the body frame depend on the YOLO skeleton, so it is active only when --model is supplied.

Build

sudo apt install build-essential libfreenect-dev libjpeg-dev curl \
                 cmake nvidia-cuda-toolkit ffmpeg        # Debian
make                                                     # builds everything

A single make builds the whole project — it fetches ONNX Runtime, clones and CUDA-builds Open3D, exports the YOLO11-pose model, and builds the kinectsentry binary. Every step is cached, so re-running make is cheap. The first build is long (the Open3D CUDA compile takes tens of minutes).

The Open3D and model steps are non-fatal: if the CUDA toolkit or python is missing, make still produces a working core binary (fixed-camera, no YOLO).

Narrower targets:

Command Builds
make everything — binary + Open3D SLAM + YOLO model
make core just the kinectsentry binary
make slam just the Open3D GPU backend (long; needs the CUDA toolkit)
make model just the YOLO11-pose model (needs python3)

The Open3D SLAM backend (handoff §5) switches on automatically once make slam has built it — no flag to pass. It targets the RTX 4070 (sm_89); with it, RoomModel uses Open3D RGB-D odometry + a CUDA TSDF voxel grid, so the camera can be carried around to scan the room. The detector is yolo11x-pose at 640px — the largest YOLO11 pose model.

Run

./run.sh                                       # depth-blob detection only
./run.sh --model models/yolo11x-pose.onnx      # + YOLO + 3D subject scan
./run.sh --model models/yolo11x-pose.onnx --zones zones.conf --armed
./run.sh --port 9000

run.sh just launches the binary, auto-selecting an exported YOLO model; it does not touch CUDA libraries, so the system loader path must already provide CUDA 12.x + cuDNN 9 for the GPU provider (ONNX Runtime falls back to CPU otherwise). Flags: --port, --model, --threads, --zones FILE, --armed (see --help).

Then open http://<laptop>:8080/:

Page Shows
/live RGB with people boxed (cyan = YOLO-confirmed, yellow = blob only) and ID-labelled
/depth depth map
/mask foreground mask
/room interactive 3D point cloud — the learned room with moving people in red, plus each person's accumulated 3D subject scan in a row off to one side
/log the event log, newest first

Leave the room empty for the first few seconds so the static background learns cleanly before anyone enters.

Zones

Copy zones.conf.example to zones.conf, edit the axis-aligned boxes to your room (metres, sensor space), and pass --zones zones.conf. A person's centroid crossing a box edge emits zone_entered / zone_exited.

Logs and artifacts

events.jsonl (machine-readable) and kinectsentry.log (text) are append-only and flushed on every event; each rotates to .1 past 16 MB. Event JSON includes a free-form detail field — YOLO confidence on person_entered, delta/total/speed on person_moved, duration+path on person_left, etc.

Recordings

Whenever a YOLO-confirmed person is in view, kinectsentry records H.264 video of the RGB stream to recordings/<isoTimestamp>.mp4 (via ffmpeg), and keeps rolling for a few seconds after they leave. recording_started and recording_stopped events log the filename in events.jsonl. Pass --no-record to disable. The recorder is a no-op (with a stderr warning) if ffmpeg is not on PATH.

A finished subject scan is written to subject_<id>.ply when that person leaves.

Hardware

Kinect v1 needs its 12 V power adapter and a udev rule for non-root USB access — see PI_SETUP.md / 51-kinect.rules in the kinectosc repo.

Layout

src/
  capture/   kinect_camera.*   Kinect RGB+depth      (from kinectosc)
  room/      room_model.*      §4 RoomModel interface + backend factory
             fixed_room_model.*  fixed-camera background model (default)
             slam_room_model.*   Open3D GPU SLAM backend (stage 5, USE_OPEN3D)
  detect/    foreground.*      motion segmentation   (§5)
             people.*          blob clustering + tracking + zones (§6/§8)
             pose_model.*      multi-person YOLO detector (from kinectosc)
             subject_scan.*    per-person 3D body scan (§6)
  log/       event_log.*       events.jsonl + text log, rotation (§7)
  web/       web_server.*      HTTP/stream server    (from kinectosc)
             monitor.*         the monitor pages     (§8)
  sentry.*   the capture -> process loop
  main.cpp