How Transduce Resizes (and Upscales) Your Footage¶
Overview¶
Every clip that lands in Transduce showed up a different size. A generator handed you 1024×1024. Your delivery spec wants 3840×2160. Somewhere in between, two very different questions have to get answered: what shape does the frame become, and where do the new pixels actually come from? Get the first one wrong and your subject's head is cropped off. Get the second one wrong and a clean 2× blowup looks like it was decoded through a screen door.
This page covers both — the deterministic geometry (crop vs. letterbox, which filter, why two separate code paths exist for preview vs. export) and the part with actual machine learning in it (the SPAN upscale cascade). Same spirit as the colour science page: real numbers, no hand-waving, but hopefully a little more fun to read.
Two decisions, not one¶
Transduce treats "resize" as two independent questions, and it's worth keeping them separate in your head because they're solved by completely different code:
- Shape — how does a source frame that isn't your target's aspect
ratio get reconciled with a frame that is? This is
aspect_handling. - Pixels — once the scale factor is known, how do the actual sample
values get computed at the new resolution? This is
resize_filter(deterministic) orupscale_mode(deterministic or AI, your call).
Nothing about picking Fill & Crop vs. Fit changes which resampling kernel runs, and nothing about Fast vs. Enhanced changes whether you get black bars or a crop. They're orthogonal knobs.
Deciding the shape¶
Transduce supports exactly two aspect-handling modes, and deliberately no third:
- Fill & Crop (default) — scale the source until the short side matches the target, then crop the excess off the long side. Nothing letterboxes; the frame is filled edge to edge. This is the "Instagram crop" behavior — you lose some picture at the edges, but you never see black bars.
- Fit — the mirror image: scale until the long side matches the target, then pad the short side with black. Nothing gets cropped; the full source frame is always visible, letterboxed or pillarboxed as needed.
Both modes always preserve the source aspect ratio — there is no "stretch to fill" option, on purpose. Squashing or stretching footage is treated as a deliberate creative distortion, not a resize operation, so if you actually want that look, it lives in the Effects module system instead of hiding inside the resize logic.
When you're in Fill & Crop, the crop window doesn't have to sit dead
centre — nudging it (arrow keys in the Viewer) shifts offset_x/
offset_y in scaled pixel space, clamped so you can never nudge the
crop window past the edge of the scaled source. The clamp is a genuinely
small piece of math but worth seeing once:
scale = max(target_w / src_w, target_h / src_h)
scaled_w = round(src_w * scale)
scaled_h = round(src_h * scale)
max_x = max(0, (scaled_w - target_w) // 2)
max_y = max(0, (scaled_h - target_h) // 2)
That max() of the two scale ratios is the whole trick behind Fill &
Crop — Fit uses the same formula with min() instead. One line decides
whether you crop or letterbox.
The Fast path — deterministic resampling¶
"Fast" is the default upscale mode and the only mode used for downscales, previews, and anyone without a GPU to spare. There are five named filters in the UI — Lanczos, Mitchell, Bicubic, Gaussian, Linear — and, honestly, worth being upfront about: they don't all map to five distinct algorithms.
Transduce actually runs two different resampling backends depending on what kind of frame is moving through it:
| Backend | Data | Used by |
|---|---|---|
Pillow (PIL.Image.resize) |
uint8 RGB |
Viewer preview — fast, interactive, no HDR precision to protect |
scipy.ndimage.zoom |
float32 ACEScg |
Export pipeline — never clamps, preserves small negative ACEScg values |
And here's the part that's actually interesting: on the export
path, Lanczos, Mitchell, and Bicubic all resolve to the exact same
operation — a cubic spline, order=3 — because scipy.ndimage doesn't
have a native Lanczos or Mitchell kernel to call. On the preview
path, Mitchell and Bicubic collapse to the same thing for the same
reason (Pillow has no native Mitchell filter either), but Lanczos really
is its own distinct algorithm there. Practically: the three "sharp"
filters will look identical to each other in the exported frame today,
but Lanczos in the live Viewer is doing something the other two aren't.
If you're chasing a specific look, Gaussian and Linear are the ones that
reliably behave differently from the rest.
When downscaling with the Gaussian filter, a real anti-alias pass runs first — a small Gaussian blur sized to the scale ratio, so high-frequency detail doesn't fold back into moiré when the resolution drops:
sigma = max(0.0, 0.5 / scale - 0.5)
The smaller the target relative to the source, the more aggressive the pre-blur. This blur only exists on the float32 export path — the fast uint8 preview skips it, so a heavily-downscaled Gaussian export can look very slightly softer than what you saw in the Viewer a moment earlier. That's a known, accepted gap, not a bug: the Viewer optimizes for interactivity, the export pipeline optimizes for correctness.
If scipy genuinely isn't installed, there's a pure-NumPy bilinear
fallback that still preserves float32 precision (including negative
ACEScg values) — lower quality than the cubic path, but it will never
silently corrupt HDR data or crash the export.
The Enhanced path — an actual neural network¶
Flip the toolbar's Resize mode from Fast to Enhanced and, only when you're upscaling, Transduce routes the frame through SPAN — a compact super-resolution network — instead of a math kernel. Pure downscales get zero benefit from a super-resolution model (there's nothing to hallucinate detail into), so Enhanced mode silently uses the Fast path for those; it only actually engages for ratio ≥ 1.
The cascade¶
SPAN is a fixed 2× model — it doesn't know how to do 1.4× or 3.7×, only double. So getting from your source to an arbitrary target size is a short pipeline, not a single inference call:
ratio < 1 → pure downscale, no AI pass at all
1 ≤ ratio < 2 → one 2× SPAN pass, then a mathematical trim to target
2 ≤ ratio ≤ 5 → two 2× SPAN passes (→ 4× total), then trim to target
ratio > 5 → clamp to two passes and warn (4× is the ceiling)
The Quality setting controls the second half of that: Best runs the full cascade above; Medium always caps at a single SPAN pass and lets a plain bilinear step cover whatever scale remains, trading a little quality for meaningfully less GPU time. Either way, the final step back down to your exact target resolution is always the same mathematical resize used by the Fast path — SPAN only ever produces clean multiples of 2×, so something has to trim the excess off to hit 1920×1080 on the nose.
Borrowing a colorist's monitor for a second¶
Here's the part that actually matters if you care about the pipeline: SPAN was trained on ordinary display-referred sRGB images, not scene-linear ACEScg. Feeding it linear light data directly would produce garbage — the model has never seen a value like that. So every SPAN pass does a full round trip: ACEScg → sRGB linear → sRGB gamma-encoded → inference → sRGB linear → back to ACEScg (the exact matrices are the same ACEScg⇄Rec.709 pair documented on the colour science page). Wide-gamut values outside the sRGB ceiling get clamped just for the inference step — for AI-generated footage this is rarely a real loss, since it's uncommon for that footage to carry meaningful energy outside [0,1] to begin with — and everything downstream of the upscale still only ever sees linear ACEScg. The model briefly borrows a colorist's Rec.709 monitor, does its job, and hands the frame back exactly the way it found it.
Why your GPU doesn't melt on a 4K frame¶
A single SPAN inference call on a full 4K frame would either blow your VRAM budget or just fail outright, so frames are tiled: cut into 512×512 chunks with a 32-pixel overlap, each tile inferred separately, then stitched back together. The seam between tiles is hidden with a cosine-shaped blend ramp — weight fades smoothly to zero across the overlap region on each side, tiles get summed with those weights, and the result is normalized by the total weight per pixel. It's the same math as feathered alpha compositing, just applied to hide a grid instead of a matte edge. 512px was chosen specifically to be safe on 8GB of VRAM regardless of scale factor — a deliberately conservative number so Enhanced mode doesn't become a crash-roulette on modest hardware.
One more small, very deliberate engineering choice: every SPANUpscaler
instance in the app shares a single ONNX Runtime session behind the
scenes, rather than each render worker loading its own. That's not a
memory optimization — it's a stability one. Two concurrent CoreML
sessions on the same model, on macOS, will crash the process with no
Python traceback to explain why. One shared, thread-safe session sidesteps
that failure mode entirely.
If the model isn't installed¶
The SPAN weights are a separate, opt-in download
(scripts/export_span_onnx.py --auto-download), not bundled with the
app. If you toggle Enhanced mode without the model present, Transduce
doesn't error and doesn't block your export — it logs a warning and
quietly falls back to the Fast path for that render. Enhanced mode is
additive, never a hard requirement.
Fast vs. Enhanced — the honest tradeoff¶
Fast is deterministic, instant, needs no GPU, and is always available — it's also what every downscale uses regardless of which mode you've selected. Enhanced only ever helps on upscales, costs real GPU time (tiled inference isn't free), and depends on a model file being present on disk. Neither is "better" in the abstract — Fast is correct and predictable for a plate that's already close to your delivery size; Enhanced earns its cost specifically when you're pushing a generator's native resolution meaningfully upward and want more than a mathematical guess at the missing detail.
Where the target size comes from¶
You rarely type raw pixel dimensions by hand. The toolbar's Resize bar and Project Settings' Resize page both edit the same underlying fields — width, height, pixel aspect ratio, aspect handling, upscale mode and quality, resize filter — and either can drive a set of named presets. Built-in presets cover the obvious targets:
| Preset | Resolution |
|---|---|
| 720p HD | 1280 × 720 |
| 1080p HD | 1920 × 1080 |
| 2K DCI | 2048 × 1080 |
| 4K UHD | 3840 × 2160 |
| 4K DCI | 4096 × 2160 |
| Vertical HD | 1080 × 1920 |
| Vertical 4K | 2160 × 3840 |
Beyond the built-ins, any combination of fields can be saved as a named, user-defined preset — including ones derived automatically from a Delivery Profile — so a recurring delivery spec is a dropdown selection, not a set of fields you re-enter every project. Pixel aspect ratio has its own small set of options (Square 1.0, Anamorphic 1.333, Scope 2.0) for the rarer case where your delivery format itself expects non-square pixels.
Preview matches delivery¶
The Viewer keeps entirely separate caches for Fast and Enhanced preview
frames, specifically so switching modes to compare them doesn't thrash a
single cache or force a full re-render just to look back and forth. That
same upscale_mode value is what the export pipeline reads when it
actually renders your final frames — so what you're comparing in the
Viewer is the same decision the exporter makes, not a preview
approximation of it.