Performance model
Where the time goes during a capture, and which stages actually dominate.
Measure, do not assume
--time reports each stage separately, which is the whole point — the answer to “why is this slow” is a different lever depending on which number is large.
eensh capture --display :1 --time screenshot.pngHow much time each stage takes
Two stages dominate, and which one dominates depends entirely on what you asked for.
| Stage | Scales with | Reduced by |
|---|---|---|
| capture | Pixels transferred from the X server — roughly four bytes per pixel at depth 24 | --region only |
| encode | Pixels encoded, and the effort level or quality | --region, --width, --compression, --format |
| resize | Pixels resized | Nothing — it is on the way to reducing encode |
| base64 | Bytes encoded | Everything above, but it is rarely significant |
The four options that affect encode cost, ranked
| Lever | Reduces capture | Reduces encode | Notes |
|---|---|---|---|
| --region | yes | yes | The strongest lever for both, and the only one that helps capture. |
| --format jpeg | no | yes | Also shrinks the payload several-fold, which matters when the result is going into a model’s context. |
| --width | no | yes | Linear in pixels: halving the width quarters the encode cost. |
| --compression fast | no | yes | A modest multiple, and the right default for most uses. |
Why PNG is slow and JPEG is not
PNG encoding has two stages, and the first is the one people forget:
RGB8 pixels ──▶ filter selection ──▶ deflate ──▶ bytes
(per scanline, (the part everyone
over full width) thinks is all of it)Filter selection evaluates candidate filters per scanline across the entire row width, so it is on the order of several times the raw pixel traffic before compression even starts. JPEG has no equivalent search — just DCT, quantisation, and Huffman — which is the structural reason the two encoders behave so differently on the same image.
Why capture time barely changes when you shrink the region
On a given server, a small region and a full-screen capture cost surprisingly similar amounts, because a capture opens an X11 connection and the handshake dominates. On local X.Org or Xvfb that handshake is sub-millisecond; some Xvfb builds (snap-packaged ones, notably) spend tens of milliseconds there, which affects every X client equally.
That is what the persistent session exists to amortise — and it is why the saving is honestly described as modest for one-shot captures. The handshake is a small part of a round trip that also includes process startup, and the CLI is one process per invocation.
A real timing breakdown, from a four-frame observation
An observation of four captures with PNG presentation held fixed:
Capture plus sleep account for the window to within microseconds, and the encodes are reported outside it entirely. That gap is also whyage_us is not near zero — it is the encoding time that elapsed after the last sample was taken.
Next: what is deliberately not here — scope.