Reference8 min read

Design invariants

The boundaries that stop one phase leaking into another, and why capture happens before rendering.

This page is the one to read if you are about to change something, or if you want to know whether a behaviour is deliberate. Everything here is a boundary that was chosen rather than discovered.

Capture into a raw frame, and render late

Capture into a raw frame, and render late.

There is no PNG, JPEG, base64, JSON, or filesystem access anywhere in the capture path. The backend produces a Frame — pixels and source geometry — and exits the picture.

  1. capture backendReads pixels from X11 into a Frame. Does nothing else.
  2. compareTwo Frames become changed-pixel counts and a bounding box. No encoder.
  3. observeCaptures repeatedly until the screen changes and settles. Encodes once, at the end.
  4. sessionKeeps a bounded history of Frames. Later requests read from memory.
  5. presentationChooses the size and format of Frames that already exist.
  6. encode + outputThe only place bytes become an image and are written.

Three invariants fall out of this, and they are worth stating separately because each one is load-bearing:

  • The capture backend produces only a raw Frame. It never encodes.
  • Comparison consumes only raw frames. It never sees a format, and therefore cannot depend on one.
  • The temporal state machine consumes only raw frames. It has no encoder in scope, so encoding cannot leak into a sampling window.

Rendering late also means one capture serves many views

Because the raw frame survives, several renderings can come from one capture. Nine views over one frame hold one frame’s pixels — SessionFrame::frame is an Arc<Frame>, and the presentation layer never constructs a Frame. So cropping allocates no new frame ID, and a multi-view response costs one frame’s memory rather than nine.

Why the service polls for new connections instead of blocking

A blocking accept cannot be interrupted into noticing a shutdown flag: the standard library retries it across EINTR, so a signal handler that only sets a flag leaves the loop parked and the process appears hung when asked to stop.

The obvious workaround — a non-blocking listener retried in a loop with a short sleep — is worse than the problem, and measurably so. Every incoming connection then waits up to the whole sleep interval to be accepted, which turned a trivially cheap request into a multi-millisecond one and made the persistent path slower than starting a fresh process. That is the exact opposite of the point.

poll blocks until a connection arrives or the timeout elapses, so a connection is accepted immediately while shutdown is still noticed within one interval. The wakeup cost falls on an idle service, where it does not matter, instead of on every request.

Why each connection gets its own thread

This is correctness, not throughput. A single-threaded loop makes one long observation block every other session, because the next connection’s request cannot even be read until the observation finishes. That is a global lock in effect, and it would make the documented session_busy refusal unreachable — a second observation would be silently queued behind the first instead of being told to wait.

Concurrency is safe because the locking is already per-session. Two connections touching different sessions never contend, and two touching the same session serialize on that session’s own mutex. No lock is taken across sessions.

Why comparison is a single pass with no mask

A temporal observation loop may evaluate a comparison dozens of times per second, so compare_frames computes its counts and bounding box in one traversal and allocates nothing. It does not build a per-pixel change mask, and it does not stop early when the area threshold is satisfied — because the counts and the bounding box must be exact regardless of the verdict.

Why observation has a Clock trait

Temporal logic is a state machine over time, and a state machine tested against the real clock can only be tested slowly and unreliably. The two-method Clock trait has a SystemClock for the CLI and a ManualClock whose time advances only when a test says so.

Every state-machine scenario — thirty polls, a timeout, a stability window that must not complete early — runs in microseconds with exact assertions rather than approximated ones. Only two real-time tests exist, with generous margins, so the system clock path is exercised at least once.

Why one engine, not three loops

wait-change, wait-stable, and observe share capture scheduling, timeout handling, target-consistency checking, and timing accumulation. They differ in one thing: which frames each comparison pairs up.

That difference is stated explicitly at each call site rather than hidden behind a shared abstraction, because getting it backwards produces a plausible but wrong answer — and a wrong answer is worse than a failure.

Why the JPEG encoder is written by hand

To keep the build dependency-free on any machine: no libjpeg, no cc, no system image libraries. It is a baseline 4:4:4 encoder with the standard Annex K Huffman tables, which suits screenshots full of small text.

Never guess on the caller's behalf

The second rule, applied wherever a silent choice was available. The transform is always present, even as the identity. Errors carry classes rather than prose. Regions are refused rather than clipped. Applied settings are reported separately from requested ones. A timeout is a distinct outcome rather than exit 1.

Each of these costs a little verbosity and buys the property that a caller can reason about the tool without experimenting on it.

the shape of the guarantee
every response says   where the pixels came from
every view says       how to map its pixels back
every failure says    what class of problem it is
every adjustment says what changed, from what, to what

Next: the numbers behind all of this — the performance model.