Engines
One platform, three engines. They answer the same question — which of these files are redundant, and which only look redundant — for three kinds of input.
| Engine | Input | Matched by | Status |
|---|---|---|---|
| Health | DICOM, PNG, JPEG, HEIC, TIFF, or a .zip | Perceptual signature, DICOM-aware | Available |
| Curate | PNG, JPEG, WebP, TIFF, HEIC, or a .zip | Perceptual signature | Available |
| Drive | .las / .laz point clouds | Rigid geometric registration | Built, not launched |
Health and Curate run the same engine. That is not a simplification for the docs — a CT slice and a product photograph are matched the same way, and the difference between the two products is the vocabulary, the reporting, and the DICOM-specific handling described below. Drive is the only one that is different code.
How matching works
Every image is reduced to a compact perceptual signature derived from its visual structure rather than from its bytes or its metadata. Two images are near-duplicates when their signatures differ by less than the configured threshold.
What a match survives
The signature is computed from normalised structure, so all of the following still match the original:
- Re-encoding, and any change in compression quality.
- Resizing, rescaling and format conversion.
- Renaming, and stripped, rewritten or absent metadata.
- For DICOM: a different display window applied to the same pixel data.
This is what makes it metadata-independent. A checksum sees four different files here; the engine sees one image four times.
What it will not match
The signature answers "is this the same image", not "is this the same subject". It will not match two photographs of one object taken from different angles, and it is deliberately not asked to — that is the job of the separate embedding recall layer, which returns candidates for review rather than duplicates.
How sure you can be
Matching is a threshold, not a verdict, and the threshold is yours to set. Every report names the level it ran at and lists the evidence for each cluster, so a result can be audited rather than trusted. Voxelion never deletes anything: it returns a keep/drop manifest and the decision stays with you.
Throughput
| Signature | Per comparison | Per core |
|---|---|---|
| Standard (2D images) | ~0.6 µs | ~1.6M comparisons/sec |
| Volumetric (DICOM scans) | ~2.8 µs | ~0.36M comparisons/sec |
Measured on a single core, and stated because a number is worth more than an adjective. Comparison is not the bottleneck in practice — decoding is, and large corpora are indexed rather than compared pairwise, so a lookup touches a small candidate set instead of the whole library. Plan capacity from the per-GB rate on the pricing page, which is what you are actually billed on.
Volumetric mode (DICOM)
A CT/MR scan is an ordered stack of slices, not a pile of images. Adjacent slices in one scan are genuinely near-identical — measured on real CT, neighbouring slices match even at the strictest setting — so slice-level matching reports a clean scan as ~99% self-duplicate. When a batch contains slices of a TOMOGRAPHIC series, Voxelion automatically switches to volumetric mode.
A batch of DICOM slices carrying a SeriesInstanceUID is a scan,
not a pile of images. Treated as loose images, a CT series reports as almost entirely
self-duplicate — neighbouring slices genuinely do look alike — which is both useless and
alarming. Volumetric mode is detected automatically rather than opted into, precisely because
the people most exposed to that failure are the ones who would not know to ask for it.
In volumetric mode the signature widens substantially. The standard width is measurably too narrow at clinical scale — distinct patients collide — and at the wider setting the same populations separate cleanly. The threshold you choose is rescaled to match, so a level means the same thing in both modes.
Modality matters. Read from DICOM tag (0008,0060) — never guessed, never asked for. Only CT/MR/PT/NM form spatial stacks and get grouped. Ultrasound, x-ray (CR/DX), mammography and angiography carry a SeriesInstanceUID too, but their series are sets of INDEPENDENT captures: two near-identical ultrasound frames are a genuine duplicate, so grouping them would hide it. Those fall back to slice-level dedup. Per-modality rules are configurable (GET/PUT /api/settings -> config.modalities).
Windowing. CT is always rendered at a fixed 0/2000 HU window; the file's own WindowCenter/WindowWidth is ignored. Hounsfield units are absolute, so a CT's window is a display preference rather than part of the image — hashing through it made 70% of same-slice re-renders (measured over five real series) fall outside the match cutoff, with a soft-tissue re-window landing further from its own source than a different patient's anatomy. Those were silent misses. Only CT: MR/US/CR values are not calibrated to an absolute scale. Note this rescues DICOM-to-DICOM only — a scan already exported as a windowed 8-bit image has lost the out-of-range values for good.
The unit of removal becomes the scan. A cleaned export drops every slice of a redundant volume and keeps the most complete copy — never N−1 slices of a series, which would leave a corrupted study behind.
Choosing an engine
Each account has a default engine, set at sign-up and changeable from Profile. API calls may
override it per request with an engine field. The engine is recorded on the report,
so a run remains attributable long after the upload is deleted.
The engine also selects the rate — see Pricing.
Reference generated from live API (https://api.voxelion.ai) on 2026-08-30. The endpoint list, error codes and limits on this page are produced from the API's own route table — if an endpoint is not listed here, it is not enabled on production.