# Full-scroll surface predictions: m7, Socratic, HercUNet

Yupeng Tang (Georgia Tech), for Paul Henderson. Run on the Georgia Tech PACE cluster.

Three surface models over two scrolls (test-time augmentation on for m7 and Socratic, off for
HercUNet, see its section), saved as probability volumes so you can pick your own threshold, plus
binarised masks at three thresholds.

## Volumes

| Scroll | Volume | Level-0 shape | Voxel |
|---|---|---|---|
| PHerc1447 | `20250521151220-8.640um-1.2m-116keV-masked.zarr` | 24297 x 8343 x 8343 | 8.640 um |
| PHerc0814 | `20250804134230-9.362um-1.2m-113keV-masked.zarr` | 19390 x 5784 x 5784 | 9.362 um |

Both were streamed from `s3://vesuvius-challenge-open-data`. PHerc0814 is the 9.362 um scan, not
the 2.399 um or 1.129 um ones.

## What is here

For each model and scroll:

- `<model>_<scroll>/` is the deliverable: foreground probability as `uint8`, 0 to 255.
- `<model>_<scroll>_th0.2.zarr`, `_th0.3.zarr`, `_th0.5.zarr` are binarised masks, 0 or 255.

All stores are OME-NGFF 0.4, Zarr v2, `128^3` chunks, blosc/zstd level 3, in the scroll's own
coordinates, with a 6-level max-pooled pyramid (max, not mean, so a thin sheet survives
downsampling). They are sparse: only chunks holding predictions exist, so a store spans the whole
scroll while holding only the predicted region, exactly like the published M7 surface stores.

To threshold the probability volume yourself, the cut in `uint8` is `ceil(t * 255)`:
0.2 is 51, 0.3 is 77, 0.5 is 128.

## m7 (nnU-Net, `scrollprize/surface_m7_nnunet`, fold 0)

Run through villa's own `vesuvius.models.run.inference`, so the recipe is the library's:
`192^3` patches at overlap 0.5, 8-way mirror TTA, Gaussian blending with sigma = patch / 8, then
softmax over the blended logits.

**Normalisation, which you asked me to check:** the log line is
`Using model-declared normalization 'ct' instead of CLI/default 'instance_zscore'`, and the applied
values are the ones in the model's `plans.json` (clip 0 to 212, mean 87.5442, std 47.7438). So the
current villa code takes the plan's CT normalisation over the CLI default. On a checkout from before
villa PR #1386 this would have been instance z-score, which is the failure you warned about.

Because per-patch logits for a whole scroll are about 20 TB, I blend in memory and keep only the
blended probability. To show that this changes nothing, I ran villa's own `predict` then
`blend_logits` then `finalize_outputs` on a 384^3 region and compared voxel by voxel: maximum
difference 1 of 255, 0.26 percent of voxels differ at all (all by exactly 1, which is float
rounding), and Dice at thresholds 0.2, 0.3 and 0.5 is at least 0.9999.

| Scroll | Blocks | Patches | GPU time |
|---|---|---|---|
| PHerc1447 | 446 | 698,298 | 61.8 h |
| PHerc0814 | 228 | 375,228 | 37.2 h |

Complete: every block has an output and no block came back empty.

## Socratic (`Hob3rMallow/socratic_method`, c3-f0-200k)

Run with its own shipped `socratic-predict`, so every pinned parameter is the release's: 8-way
mirror TTA, BF16, `128^3` cubes with a 32-voxel halo (`192^3` context), operating threshold 0.30.
The checkpoint's SHA-256 verified as `54db9a59cb602ff08e3dda0acc1195ad980ab05fb660332d32fb436a0d49a175`.
Its per-cube float16 probabilities are what I converted to `uint8`, so `th0.3` here is the shipped
operating point (77 in `uint8`).

| Scroll | Cubes predicted |
|---|---|
| PHerc1447 | 256,616 |
| PHerc0814 | 135,072 |

**One gap to know about.** `socratic-predict` predicts a cube only when all 26 neighbour cubes are
present, so one 128-voxel cube layer on each of the six faces is never predicted: 246 blocks on
PHerc1447 and 8 on PHerc0814 came back with no cubes for that reason. It matters at the z ends of
PHerc1447, where the outer 256 voxels are 28 percent material measured on pyramid level 5; the y and
x faces are under 0.21 percent, and PHerc0814 is 0.26 percent overall. m7 and HercUNet have no such
restriction, so when comparing the three, either exclude that shell or treat Socratic as having no
prediction there.

**Training provenance, since it affects how you read the numbers.** PHerc1447 was deliberately held
out of F0's training corpus. PHerc0814 was one of its two validation scrolls, so it is not unseen
data for this model.

## HercUNet (`jimmylomro/hercUNet`, `jimmylomro/hercunet-v0`)

4 Jacobi passes at overlap 0.25, which is the published recipe, but **without TTA**, unlike m7 and
Socratic here. I measured what it would cost on our hardware before deciding: a plain pass is about
32 GPU-hours on PHerc1447 and 14 on PHerc0814 using 4 H100s, and TTA multiplies the pass it is
applied to by about 6 (measured: 8.4 windows/s plain against 1.4 with `--tta all`). Running
`--tta all --tta-passes last` on both scrolls would therefore have cost roughly 300 GPU-hours more
than the plain run, against about 76 for the plain one.

I chose to deliver the complete three-model comparison rather than spend that, since the gain from
TTA here is unmeasured and HercUNet is in any case not directly comparable to the other two. Adding
it later is cheap and does not need a rerun of the earlier passes: the last pass alone on PHerc0814
is about 70 GPU-hours more, on PHerc1447 about 190. Say the word and I will run whichever you want.

**HercUNet predicts the medial crest of a sheet rather than the face m7 and Socratic label**, so a
plain Dice against face labels understates it. Its author makes the same point.

One practical note in case you hit it: `--scroll PHerc0814` resolves through its data layer to the
2.399 um scan (74568 x 23891 x 23891), not the 9.362 um one. I passed the volume explicitly with
`--local-vol` so both scrolls use the scans listed above.

## Reproducing

The pipeline scripts are small and I can share them as a repo if useful:

- `m7_blocks.py`: block-wise m7 with villa's `Inferer` and an in-memory Gaussian blend.
- `soc_blocks.py`: `socratic-predict` per 512^3 region, cubes copied into one store.
- `herc_full.sbatch` / `herc_multi.sbatch`: whole-scroll HercUNet, with the explicit volume.
- `derive_masks.py`: thresholds and the max-pooled pyramid.

Environment: torch 2.14 with CUDA 13.2 for m7 and HercUNet, torch 2.11 with CUDA 12.8 for Socratic
(its pinned requirement), zarr v2 output throughout.

Happy to rerun anything with different settings, or to add other thresholds; the probability volumes
make that cheap.
