Getting started6 min read

Coordinates and the transform

How the transform maps image pixels back to screen pixels, worked through in both directions.

This is the part that trips people up, and it is worth five minutes because it is the difference between a tool you trust and one you re-derive the arithmetic for every single time.

There are two coordinate spaces, and a capture almost always involves both:

  • Source space — pixels on the screen. This is where windows live, where your region was specified, and where a click has to land.
  • Image space — pixels in the image you received. This is the only space the bytes know about.

They are the same only when you captured without resizing and without a region offset. In every other case they differ, and the transform is what reconciles them.

Source space and image space, with the transform between themregion at (640, 200)(0, 0)source space — the screen×2.0(0, 0)image space — what you get backoffset_x = 640, scale_x = 2.0
Image pixel (0, 0) is source pixel (640, 200). That offset is the offset_ fields; the size ratio is the scale_ fields.

The rule

source from image
source_x = offset_x + image_x * scale_x
source_y = offset_y + image_y * scale_y
offset_x / offset_y

Where image pixel (0, 0) sits in source space. Non-zero whenever you captured a region or a window away from the origin.

scale_x / scale_y

How many source pixels one image pixel covers. 2.0 means the image is half the source width; 0.5 means it was scaled up.

origin is always top-left: x increases to the right and y increases downwards. It is present so that a reader never has to guess the convention.

Worked example, forwards

A 640×480 region at (640, 200) on the screen, resized to 320×240:

transform
{
  "origin": "top-left",
  "offset_x": 640, "offset_y": 200,
  "scale_x": 2.0, "scale_y": 2.0
}

Image pixel (100, 50) is therefore source pixel (840, 300):

text
source_x = 640 + 100 * 2.0 = 840
source_y = 200 +  50 * 2.0 = 300

Read it as: move 100 image pixels right, which is 200 screen pixels, and you have moved from 640 to 840.

Worked example, backwards

The other direction is the one a click needs — you know where you want to click on the screen and need to locate it in the image you already have:

image from source
image_x = (source_x - offset_x) / scale_x
image_y = (source_y - offset_y) / scale_y

Source (840, 300) with the same transform:

text
image_x = (840 - 640) / 2.0 = 100
image_y = (300 - 200) / 2.0 =  50

Try it

scale

1.33

offset_x

640

source_x

773

source_x = 640 + 100 * 1.33 = 773

Why the source rectangle is reported, not the local one

A crop is a rectangle within a frame, and it would be natural to report it that way — as an offset from the frame's own origin. eensh reports it in source coordinates instead, and the reason is that a caller should never have to hold two frames in mind to interpret one response.

So a region at screen (100, 200) reports source_rect of (100, 200), even though there is obviously a frame-local (0, 0) somewhere. The response describes the screen, because that is what the caller is reasoning about.

The three ways to get this wrong

  1. Using the image width as the source width. Correct right up until you resize, which is exactly when you need it.
  2. Applying the transform in the wrong direction. The transform maps image → source. Converting a click needs the inverse, which is the section above.
  3. Assuming the offset is zero. It is zero for a full-screen capture and non-zero for everything else — including a window that starts away from the screen origin.

All three produce a wrong coordinate that is plausible rather than obviously broken, which is why the tool states the transform instead of letting you infer it.

Next: what happens when the capture does not succeed at all — errors that guide you.