feat(seg): the fast segmentation engine, non-rigid drift, and a test suite that can fail - #122
Open
CSSFrancis wants to merge 53 commits into
Open
feat(seg): the fast segmentation engine, non-rigid drift, and a test suite that can fail#122CSSFrancis wants to merge 53 commits into
CSSFrancis wants to merge 53 commits into
Conversation
…surement Steps A1 and B1/B5/B6 of DRIFT_AND_PARTICLES_PLAN.md. spyde/drift/ — rigid translation by FFT phase correlation with a running Fourier average reference and Guizar-Sicairos matrix-multiply DFT upsampling. Streams one frame at a time; a solve returns an (N, 2) DriftModel, never an aligned copy of the movie. The running reference is accumulated in FOURIER space via a phase ramp, which is exact for sub-pixel shifts and avoids the resample blur that would accumulate over thousands of frames. spyde/particles/ — the classical engine (ParticleSpy's segptcls vocabulary) plus split_instances(), the instance-split shared by all three planned engines, and calibrated regionprops measurement. spyde/signals/particles.py — SpyDEParticles, ragged per-frame CSR storage mirroring SpyDEDiffractionVectors. Outlines are quantised int16 contours (~120 MB at 1.5M particles, vs 770 MB for bbox bitmaps); full-frame label images are never stored. Three things measurement corrected: - _upsampled_dft ignored its upsample argument, so the refinement ran at 1/upsample of the intended resolution. It still found a peak, so every recovered shift merely quantised to 1/8 px. - The accuracy gate was vacuous: every truth shift was a multiple of 1/upsample and so exactly representable. Off-grid truth gives 0.065 px at u=8, and would have caught the bug above. - torch-CPU is 7.7x faster than numpy for per-frame FFT (139 vs 18 frames/s on 512^2) because np.fft is single-threaded. Backend order is now cuda > mps > torch-cpu > numpy; numpy stays as the parity reference. ParticleSpy's watershed_size filters markers by AREA, which erases a 3x3 particle's one-pixel local-maximum marker — the sensitivity failure plan Section 0.9 exists to prevent. Replaced by min_separation + marker_smooth. 103 tests; numbers in benchmarks.md.
…gs it found Step 0 of DRIFT_AND_PARTICLES_PLAN.md: the fixture every later step is graded against. `spyde.data.synthetic.particle_movie` gives a 24-frame in-situ movie with nine particles on a drifting speckled support film, and stamps its whole motion model as ground truth: per-frame drift, radii, and the nucleation (8), dissolution (16) and merge (14) frames. `particle_truth_at` evaluates that model, so no consumer re-derives it and no test can pass by repeating the generator's own mistake. Reachable in the app as `load_test_data_particles`, lazy at one frame per chunk like a real .mrc. It earned its place immediately by finding two defects in the drift solver: 1. A FULL HANN WINDOW DESTROYS THE REGISTRATION. With apodize=1.0 the solve returned a 25 px error on a 6 px drift -- worse than not correcting. The spurious peak at (-19, 19) scores 0.121 against the true peak's 0.088, because a full-frame window reweights the two frames differently once the drift is large. skimage's phase_cross_correlation returns the SAME wrong answer on the same windowed input, so this is a property of full-frame windowing, not of either implementation. The window is now a Tukey taper (alpha 0.25, edge only): 0.124 px. 2. THE RUNNING REFERENCE WAS NOT ROBUST TO A BAD FRAME, despite the docstring saying so. A frame of pure noise has a broadband spectrum, so after phase normalisation it contributed as much to the accumulated reference as a good frame and dragged the next two registrations 3.9 px off. The per-frame peak sharpness was already computed and unused; it now gates entry to the reference. Both constants are measured, not chosen: worst natural frame sits at 0.388 of the running median, a noise frame at 0.007, so the threshold is 0.25. A 3-sample warm-up was tried and let the bad frame through on a short stack; windowing the median made no difference at N=3, 5 or unbounded. Also floored the phase-normalisation divisor at 100*eps instead of adding 1e-12, matching skimage -- correct in principle, though it was not the bug. Every claim above is pinned by a test, including two that assert the failure mode still exists when the fix is disabled, so they cannot quietly go vacuous. 189 tests.
…rticles `scripts/verify_drift_particles.py` runs the python suites, the frontend typecheck, the Playwright specs and the benchmarks from one place, because the feature spans two languages, three test tiers and a separate repo. Every stage is independent and the exit code is the worst result, so one broken tier still reports the state of the others -- it is a status board, not a fail-fast gate. Suites and specs that a plan step has not reached yet report SKIP rather than failing, so the same command is useful from the first step to the last. `scripts/bench_drift_particles.py` prints the numbers every recorded decision rests on, in a form that pastes into benchmarks.md. Deliberately not a pytest: these are machine-dependent, and a benchmark that fails CI because a runner was busy teaches nothing. Tests assert correctness; this reports cost. Recorded in benchmarks.md, including one number that points at future work: measure_frame costs 3.7x segment_frame, which is the wrong way round. The segmenter is vectorised over pixels while measurement still loops over particles to crop each bbox and trace each contour. Harmless at fixture scale (52 s for 3000 small frames) but it is the first place to look if the real 2048-4096 px target misses the plan's "minutes" budget.
…t stream Wave C1/C2. One `linear_sum_assignment` per frame pair over a distance cost gated by `max_dist`, with `memory=k` gap closing; the assignment's leftovers carry the events. Distances are in the particles' CALIBRATED units because the measured centroids are, and `sample_frame_positions` is the single seam where DriftModel's pixels cross over. The gate uses a rectangular matrix with a derived sentinel for infeasible pairs rather than the padded square Jaqaman formulation -- same optimum, a quarter the matrix at 500 particles/frame. The property-similarity penalty deliberately cannot move the gate; it only reorders pairs that are already admissible. Verified independently of the agent that wrote it: no link exceeds `max_dist` at any `max_dist` from 1 to 40 px, and the identity `Dcount = births + splits - deaths - merges` holds on every frame of the fixture. Trajectories exact (mover is ONE track across all 24 frames, 0.175 px max error); one birth at frame 8; one death at frame 16; drift correction takes the static anchors from a 9.3 px lab excursion down to 0.3 px. THREE CORRECTIONS TO THE PLAN, all found by building this: 1. C1 said the linker runs on "raw minus tree.drift". That is `to_lab_frame` -- the INVERSE -- and doubles the drift instead of removing it. Exactly the sign trap drift/model.py's docstring warns about, written into the plan anyway. It is `to_sample_frame`, which ADDS the shifts. 2. C2 claimed the unmatched rows and columns ARE the event stream. Birth and death, yes. Merge and split, no: a one-to-one assignment cannot express two-to-one, so they need an explicit post-pass with its own radius and its own failure modes. Now documented, along with those failure modes. 3. "Recovers the merge exactly" was not achievable and is now reworded. The merge FRAME is a segmentation property: geometric contact is frame 14, but the segmenter resolves one region at 18 with watershed on and at 12 with it off -- the truth is bracketed on one boolean. The gate now asserts what belongs to the linker, and a separate test asserts the bracketing so the offset is provably segmentation, not a linker bug. 102 tests.
A dead track kept painting its head dot for as long as its trajectory intersected the trailing window. The dot means 'the particle is HERE NOW', so on a dead track it is a lie -- it reads as a real particle the segmenter has stopped filling. Same for a track inside its memory gap. Recorded with the fix. The count lane was drawn as a straight interpolation between frames, which puts the visual transition half a frame early: nucleation at frame 8 looked like 7. Integer lanes are step plots; continuous ones stay lines. Neither is visible in a passing test -- both came from looking at the pixels, which is what CLAUDE.md's verification standard is about.
Wave B2/B3. `features.py` builds a 36-channel stack (gaussian, DoG, rank median/min/max, Sobel, Hessian eigenvalues, Laplacian, optional membrane projections) batched in torch; `scribble.py` trains a small MLP head on pixels the user painted and returns a per-frame foreground probability that `split_instances` turns into instances. torch is imported lazily, and every call site takes `accelerator_lock` -- test_device_lock.py gains a class asserting that, since it is the shared-lock registry. Measured: train + apply on a 96x112 frame is 0.492 s, inside the ~1 s interaction budget, and the fit is FIXED cost (1.5 ms/step at any thread count from 1 to 24 -- pure dispatch, not arithmetic), so it will not degrade with frame size. IoU 0.941 against the sklearn RandomForest reference on identical labels and identical channels. Faint probes come back at 0.983 and 0.9997. Sharing intermediates across the feature family is worth 156 -> 54 ms; the rank family alone 205 -> 33 ms by reusing one unfold, and max_pool2d(stride=1) measured 4x SLOWER than unfold. THREE FINDINGS THAT CHANGE THE PLAN: 1. Section 0.9 needs at least one FAINT scribble. Trained on bright particles only, the head finds at most 1 of 2 faint probes and the forest finds 0 of 2 -- exactly 0.0, which is structural rather than noise, because a tree cannot predict outside the leaves it was shown. One seven-pixel dab fixes it. So the caret's per-class pixel counts are load-bearing, not decoration: under-training a class is THE failure mode, and no amount of sensitivity slider substitutes for a missing example. Recorded as its own test class rather than worked around. 2. "Fine scales are mandatory" was right for the wrong reason. A coarse (4, 8) stack still DETECTS both faint probes; what it loses is their SIZE -- radius error 13% -> 26% overall, and -12% -> -44% on the r=3 probe. The floor stays <=1 px for measurement fidelity, not detection. 3. Importing spyde.particles costs ~6 s, and torch is not why. measure.py reads the column schema from spyde.signals.particles, and importing anything under spyde.signals executes its __init__, which pulls hyperspy (5.5 s of the 6.3). Free in the app -- the backend loads hyperspy at startup anyway -- but it means a script that only wants segmentation pays for a signal framework it never touches, against the "constructible standalone" contract. Left alone deliberately: the fix is in a file other waves share. Also corrected an over-confident docstring: the bright-only faint-probe counts depend on where the background scribbles land. An independent probe with different dabs got 0 of 2, not 1 of 2 -- same conclusion, different numbers. The assertion was already the robust `< 2`; the prose now says so. 150 + 16 tests.
…ze=0 is a footgun Adding one faint scribble to bright-only labels takes detection from 7/9 to 8/9 true particles -- and costs 25 spurious instances, because teaching the classifier faint contrast necessarily teaches it to fire on film speckle of similar contrast (foreground pixels 977 -> 1334). min_size=10 removes 24 of those 25. So the classifier is not what buys specificity; the instance-split's size filter is. Three consequences for the caret: a user who zeroes min_size to catch the small ones gets the opposite of what they want; the live preview must report the count AFTER the size filter or the number moves invisibly; and sensitivity and min_size are coupled, so they belong adjacent rather than in separate tabs. Found by looking at the rendered probability map. The tests were green and asserted only on the probes' own probabilities, which is exactly the blind spot CLAUDE.md's 'look at the pixels' rule exists for.
Plan step 5, the architectural half. Segmentation spawns a NEW SignalTree rather
than decorating the movie it came from -- a segmentation is a derived dataset,
like a strain map, not a property of its source.
ParticleTree
root lazy LABEL MOVIE, one frame per chunk, painted from contours
.particles the SpyDEParticles store
.source_node the signal it was computed from
.nav_map particle frame -> source navigation index
.nav_traces count(t) / mean size(t) / event lanes
Three things this buys that an attribute would not. It answers Wave D by
construction: particles found on a 4D-STEM virtual image record which node they
came from and which parent nav positions each covers, so "the mean diffraction
pattern for this particle" is a slice rather than a guess -- `particle_nav_positions`
is that seam, and it distinguishes the movie case (a particle's pixels are SIGNAL
coordinates, so the only nav index is the frame) from the virtual-image case
(they ARE nav positions). The label movie is a real dataset, so scrubbing,
saving, the report builder and the movie editor need no special case. And
re-segmenting produces a sibling to compare against instead of destroying the
previous result.
The label movie stays lazy at one frame per chunk, and a test asserts that
rendering one frame does not compute the stack -- the same patch.object guard
find_vectors uses. A materialised label movie is 64 MB PER FRAME at 4096".
Wave 0 alongside it:
* `requires_particles` in BOTH toolbar filter paths. A test greps the module
to enforce that, because a gate added to only one path renders a button that
never dispatches -- actions/README.md section 6 records that bug already
happening once with requires_vectors.
* `lifecycle.wait_for_particles` + `seg_batch_running`. Deliberately no
`strict` switch: wait_for_vectors needs one because vectors attach to the
tree the user clicked, so an any-tree fallback could re-dispatch forever into
a tree-specific gate. Particles live on their OWN tree, so there is no
ambiguity to resolve and therefore no knob. Timeout is 600 s not 300 --
segmenting thousands of frames is the stated target and a legitimate
eight-minute run must not be abandoned at five.
* `particles` registered as a hyperspy signal type (ParticleMap /
LazyParticleMap), so downstream actions gate by type instead of hunting up
the parent chain -- the same job `insitu` does for Play / Fast-Forward.
22 tests.
The one genuinely new renderer primitive this feature needs. There was no table
and no sorting anywhere in the app, and no virtualisation library -- react-rnd is
a declared dependency that two files explicitly reject in favour of a hand-rolled
Pointer-Capture gesture, so this follows that grain and adds no dependency.
DataTable is data-agnostic: the backend supplies the column set, so showing
tracks instead of particles is a column change rather than a new panel. Sortable
headers, single/multi selection, swatch cells, inline units, tabular-nums for
numerics. Virtualisation is a scrollTop slice plus two spacer divs, with
scrollTop QUANTISED to the row grid so a wheel gesture re-renders about once per
row instead of once per event. Verified at 5000 rows: ~30 DOM rows, and scrolling
to row 2000 swaps the window without growing the DOM.
Two traps the codebase had already paid for and documented:
* `sendAction` is recreated on every provider render, so listing it in a
dependency array re-runs the effect on every state update -- and if that
effect requests data whose reply is state, it loops. Routed through a ref
(ConsoleBar.tsx:226 records this as the "flashing preview" bug).
* A Dropdown's menu is absolutely positioned, so one rendered inside the
table's `overflow: auto` body would be clipped. Selects stay in the header;
a comment marks the exact spot where a future one must not go.
The dock is capped at 50% of the window: the whole bottom stack is
`flexShrink: 0`, so LogPanel + this + ConsoleBar + StatusBar could otherwise
squeeze MDIArea to nothing. Screenshot 07 confirms the MDI keeps real height at
full extension.
Visibility had to go through SpyDEContext rather than App state, because there
was no View menu at all and MenuBar reads only the context -- so this adds View,
following the existing dialog open/close triple, plus a StatusBar toggle beside
Log.
typecheck clean, build clean, 12/12 in tests/data_table.spec.ts, and 23/23 across
the shell-adjacent specs (app_log, examples_menu, update_gpu_dialogs, ui_fixes).
Screenshots in electron/data_table_shots were looked at, not just captured.
Known gap until the backend lands: `particles_query` has no handler yet, so
opening the dock logs one Unknown-action warning.
…osed Both wizards per actions/README.md: preview on the CURRENT frame only, progressive + cancellable runs, latest-wins generation guards, schemas declared in both the controller and registry._WIZARD_SCHEMAS. seg_: open/close/set_method/tune/paint/train/run/commit. The caret gets seg_state (classes with per-class labelled-pixel counts -- load-bearing, since under-training a class is THE failure mode), seg_preview (count AFTER the size filter, plus the effective min_size so the caret never shows a number different from what ran) and seg_trained. drift_: open/close/set_method/tune/run/commit, with the before/after check window registered as a bare-figure controller. drift_tune re-solves the FIRST PAIR only -- 2 FFTs -- so tuning stays interactive on a long movie. TWO FIXES TO MY OWN EARLIER CODE, both found by building against it: 1. `solve_translation` could not stream its shift trace, so the drift caret could show a progress bar but not a curve. `progress` carries only a count and the shift array is solver-local until the return; chunking the solve would change the ANSWER, because the running Fourier reference accumulates across the whole stack. Added an `on_shift(i, dy, dx, sharpness)` callback -- the only way to stream the trace without altering the result. 2. `open_particle_tree` published `tree.particles` at construction, which contradicts the plan and unlocks `requires_particles` against an empty store -- the user could click Track on zero particles. Now takes `attach=False`. The placeholder is still handed in and mutated IN PLACE at finalize, because the lazy label movie closes over that exact object; swapping in a fresh store would leave the already-open window rendering zeros forever. That subtlety is now in the docstring rather than in a caller's workaround. Three plan gaps recorded rather than papered over: `rigid+affine` has no engine (spyde/drift/ has only translation.py, so A4's affine search does not exist) and stays a stub that keeps DriftModel.kind honest; a ONE-frame particle tree is not constructible (a size-1 hyperspy nav axis leaves MultiplotManager without a selector_type), so the single-image shape routes through commit_result_tree instead; and the eraser rasterises through a scratch LabelStore so brush and eraser geometry cannot drift apart. 159 tests across the wizard, drift, tree, schema and double-fire suites. Per CLAUDE.md this is built + headless-tested and NOT yet verified in the app -- there is no caret to drive until the renderer side lands.
…exposed SegmentWizard (330 px, 2-column), DriftWizard, and a floating ClassStrip, registered in FloatingToolbar. 7 message types typed in protocol.ts and added to the SpyDEContext re-broadcast list -- none were there. The caret follows the layout the plan fixed, and the two decisions that came out of MEASUREMENT are visible in it: sensitivity and min_size sit adjacent because they are coupled (one faint scribble buys +1 true particle and 25 spurious ones; min_size removes 24), and typing min_size=0 shows the effective 10 with the reason inline, so the caret can never display a number different from what ran. Per-class labelled-pixel counts are rendered prominently and a class with zero pixels is dimmed and flagged -- under-training a class is THE failure mode and the counts are how a user notices. TWO BACKEND BUGS, both found by driving the real UI and invisible to any headless test: 1. `seg_run` never emitted a terminal progress, so `state.loading.busy` never cleared: the StatusBar spinner spun forever and -- because it prefers loading.text while busy -- "Segmenting (33%)" permanently masked the "Found N particles" line emitted right after it. 2. The label-movie window stayed BLACK after finalize. Clearing the stale cached dask array makes the next read correct but nothing triggers a read, so the result looked empty until the user happened to scrub. Now repaints the displayed frame. Also wired drift_run to the on_shift callback added earlier, so the trace streams while solving instead of appearing all at once at the end. Batched 16 frames per message: one per frame would flood the PLOTAPP line protocol at thousands of frames for a curve the eye cannot follow that finely. One renderer bug the screenshots caught that headless could not: mixing a `border` shorthand with a `borderColor` longhand across React inline style objects left a deactivated class row with a stale white border, so two rows read as selected. Fixed, with a computed-style assertion in the spec. typecheck clean, 7/7 across segment_wizard.spec.ts + drift_wizard.spec.ts on real Dask and real data, 143 python tests. Screenshots in electron/segment_wizard_shots and drift_wizard_shots were looked at.
`solve_translation(..., roi=(y0, x0, h, w))` measures the shift on a sub-region while the returned shifts still apply to the whole frame -- a translation is a translation regardless of the window you measured it in. This is not a speed switch, it is often the more CORRECT answer. Whole-frame correlation averages over everything that moved, so on an in-situ movie where the sample is genuinely evolving -- particles growing, drifting, appearing -- the sample's own motion contaminates the estimate of the stage's. Restricting to a static feature-rich landmark measures the stage and nothing else. A test builds exactly that adversarial case: a bright square tracking the other way outside the ROI drags the whole-frame solve while the ROI solve correctly reports ~zero, and it asserts BOTH halves so the fixture cannot quietly stop demonstrating the point. The ROI is fixed in frame coordinates, so the landmark drifts within it; that is fine while the drift is small against the box, and it is why the box wants to be comfortably larger than the excursion. Documented, and the forthcoming caret preview exists so a user judges it by eye instead of guessing. Out-of-bounds and too-small ROIs RAISE rather than clamp. A silently shrunk box would correlate on a region the user never dragged, and the resulting drift curve would be wrong in a way nothing on screen could explain. 49 tests.
Added after reviewing the first carets: 'way too complicated. Too many options.
Information overload.' A fair verdict on a Segment caret carrying ~15 visible
controls, and a drift from section 0.9's own instruction -- 'expose one
sensitivity control, not independent knobs' -- that happened one
reasonable-looking addition at a time.
The rule now applies to every action in this feature: the default face carries
the TASK, not the algorithm; everything else sits behind a collapsed Advanced,
including parameters that matter but that nobody should normally touch; a warning
belongs beside the control it is about rather than on the front; and buttons are
named for the job ('Find in all frames', not 'Run All'). Nothing is deleted --
the Python API and the provenance keep it all.
Drift is the same rule applied harder, and gets a better shape than it had: its
parameters have one right answer we already know, so the caret becomes a button,
a progress bar and 2-3 toggles -- one being 'use ROI for alignment'. The dx/dy
curve stops being caret furniture and becomes its own plot window filled as the
solve runs. And discovery precedes commitment: a draggable ROI with a live
drift-corrected sum over ~20 frames, so a user SEES whether alignment works on a
subset before paying for the whole movie.
Plan B9 + C2/C3. Track-coloured 25% fills with 1 px outlines on the label movie, labels on SELECTION only, trails as a fading line plus a head dot, click / track / rubber-band selection, and delete / merge / split recorded on the tree so a re-run cannot silently discard a correction. The two decisions that came out of looking at renders are implemented and pinned. A dead track draws NO head dot -- the dot means "the particle is here now", so on a dead track it reads as a real particle the segmenter stopped filling. And the count lane is a genuine steps-post staircase built into the DATA, because anyplotlib's plot has no drawstyle: verified that a change at frame 2 draws its transition AT frame 2, not at 1. FOUR BUGS FOUND BY DRIVING THE APP, none visible headless: 1. A `parameters:` block on a toolbar TOGGLE makes the renderer show a form with its own Run button -- so the overlay never drew and clicking only opened a panel. The knobs belong in the caret schema, not on the toggle. 2. FloatingToolbar sends `set_action_active(false)` but never a second `toolbar_action` for an action it believes is live, so a self-toggling function can be switched ON and never OFF. Returning a handle with `active_children` + `close()` hands teardown to `_track_action_artifacts`, which is the framework's own answer. 3. anyplotlib's Axes has no set_title/set_ylim (the existing `_stack_navigators` silently loses its titles the same way), so the event lane autoscaled to an invisible baseline and drew nothing. 4. The selection readout was drawn ACROSS the particle it described. Stacked-navigator machinery is reused wholesale except the figure builder, which draws every lane as a plain line -- count must be a staircase and events are markers, not a trace. Known limits, all documented rather than worked around: the lanes button cannot FORCE the stack on screen (WindowContent only stacks when its own chip selection has >=2 entries and no backend message can set that), so it registers the lanes as named navigators and the user shift-clicks; `tree.particle_edits` is recorded but nothing consumes it yet -- `pending_edits(tree)` is the seam for seg_run; and the birth/death badge on the frame (plan C2's third surface) is carried but undrawn. 88 tests. Full migrated suite 2961 passed.
…e movie
BaseSignalTree.close() tears down by iterating hard-coded attribute NAME LISTS,
so anything nobody remembers to add is silently exempt -- no error, no warning,
just state that outlives its tree. Three omissions from this feature:
* `_seg_wizard` / `_drift_wizard` were absent from the wizard list, so closing
a tree left a live wizard controller holding its windows and generation state.
* `particles`, `_seg_pending_particles`, `particle_events`, `particle_edits`,
`nav_traces`, `drift` and `nav_map` were absent from the results list.
* Worst: `source_node` and `source_tree`. A particle tree back-references the
movie it was segmented FROM, so closing the particle tree kept the source's
lazy multi-GB array reachable -- and closing the source movie freed nothing.
test_particle_lifecycle.py closes things and asserts on what is left, including a
weakref test proving the source signal really is collectable afterwards, plus its
non-vacuity partner showing the source is still pinned WITHOUT close -- otherwise
that test would only be demonstrating that a local went out of scope.
Because the real failure mode is an attribute nobody thought to LIST, one class
reads close()'s source and fails on a missing name. A purely behavioural test
only catches the attributes you remembered to set, which is exactly the blind
spot that caused the bug.
Also: the verification runner was never actually in the repo. .gitignore carries
a `verify_*.py` rule for throwaway ad-hoc scripts and it silently swallowed
scripts/verify_drift_particles.py -- so an earlier commit claimed to add it while
only the benchmark landed. Renamed to scripts/check_drift_particles.py rather
than punching a hole in a shared ignore rule.
22 lifecycle tests.
Per plan section 0.9a, after "way too complicated. Too many options. Information
overload." The Classical default face goes from 17 interactive elements to 7 --
and from 13 actual knobs to THREE: the sensitivity slider, "Find in all frames",
and a collapsed Advanced. The spec pins the 7 and asserts 17 named testids are
absent from the DOM, so it cannot quietly drift back.
Everything moved, nothing deleted: min_size (with its floor warning now directly
under the field it explains rather than shouting from the front), max_size,
split/separation/smoothing, threshold, pre-blur, rolling ball, local window,
dark-particles, store-outlines, link-tracks, the histogram, and Commit Frame.
Grouped under quiet section labels. params() is byte-identical and no Python
changed.
Two further removals, made deliberately rather than demoted:
* `+ add class` is GONE. It was permanently disabled with "not wired yet" on
it -- pure noise on a face this redesign had just emptied, advertising
something that cannot happen. It comes back WITH a seg_add_class verb.
* The floating brush strip is now scoped to the Scribble tab. It floats OVER
the image, so on Classical -- where there is nothing to paint -- it was
chrome covering the data for no reason. The spec asserts it is absent on
Classical, appears on Scribble, and goes away again on the way back.
A real bug fixed on the way (predates this change, but the redesign makes the
disclosure the main interaction so it would have bitten immediately): the
caret's placement layout-effect only ran when FloatingToolbar itself re-rendered,
but the caret's height changes from the WIZARD's own state, and a child's
setState does not re-render its parent. So placement went stale and the caret
jumped on some later unrelated render -- landing between a mousedown and a
mouseup, which makes the browser emit no click at all. The symptom was "every
other click on the caret does nothing". Fixed with a ResizeObserver on the caret
box.
typecheck + build clean; 5/5 segment_wizard.spec.ts; fit_wizard, drift_wizard and
play_no_caret_loops all still pass.
… toggles The drift redesign, per plan section 0.9a and the shape asked for in review: "the dx and dy should be a new plot filled in as it is computed... a discovery feature... a draggable ROI, a drift corrected sum of said ROI... then mostly 2-3 togglable options, one of them use ROI for alignment." DISCOVERY BEFORE COMMITMENT. A draggable rectangle on the movie plus a live drift-corrected sum of just that box over ~20 frames, with a gain number (gradient energy of the aligned sum over the raw sum, on the SAME pixels) so "is this landmark any good?" has an answer that is not purely visual. A good box sums sharp, a bad one blurs. The preview frames are sampled EVENLY over the whole movie rather than the first 20 in a row: the question is whether the landmark survives the full excursion, and 20 consecutive frames of a long movie drift by almost nothing, so a contiguous window would answer "looks fine" for every box including the useless ones. THE CARET IS TWO TOGGLES AND A BUTTON. Use ROI for alignment, ignore bad frames, Correct Drift. Reference mode, sub-pixel factor, max shift, interpolation order and the model tabs are all still there, behind a collapsed Advanced -- including the affine/non-rigid stubs, which were three visible tabs for two choices that do not work. THE dy/dx CURVE IS ITS OWN WINDOW, opened when the solve starts and filled from the solver's on_shift callback, batched by frame count OR a 0.15 s interval so a slow solve still animates. Geometry is image pixels end to end -- anyplotlib widgets report x/y/w/h in image pixels and solve_translation's roi is in pixels, so they meet with no conversion. Nothing materialises the movie: the solve streams, the check sums stream over a bounded subset, and the preview keeps only the crop under a byte cap. Verified against the fixture's ground truth: the recovered curve has dy rising monotonically to 6 px and dx swinging negative then back -- the fixture's exact shape -- and the status line reports max shift 6.00 px against a fixture built with drift_amplitude=6.0. 73 python tests, 4/4 drift_wizard.spec.ts, typecheck and build clean. Written by an agent that hit its session limit mid-verification; the code was complete and the constants it was mid-sentence about (_TRACE_MAX_INTERVAL, the time import) were already in place. I ran the verification it did not reach.
Reported as "hanging or really slow with a 4k x 4k movie". It was not hanging.
Measured, one frame, classical preview path (segment + measure):
256^2 0.27 s
1024^2 0.57 s
2048^2 1.69 s
4096^2 8.36 s <-- 7.6 s of it is segment_frame alone
The caret re-previews on every sensitivity nudge, so a real 4k in-situ movie
meant 8 seconds of work per keystroke. Plan section 0.8 said "preview the CURRENT
frame only" and never bounded the frame's SIZE.
Above a megapixel the preview now runs on a centred crop at FULL resolution:
408 ms on a 4096^2 frame, a 20x speedup, and frames already under the budget
take the untouched fast path (box=None, zero added cost).
CROP, NOT DOWNSAMPLE, and that is the whole point. Downsampling would be cheaper
again but it makes the preview a DIFFERENT computation from the run it exists to
predict -- section 0.9 records that the fine feature scales are what find small
faint particles, so a preview that silently detects a different population is
worse than a slow one. A full-resolution crop runs the identical algorithm on
identical pixels.
The caret says so rather than quietly lying: the count reads "N particles in this
region" instead of "on this frame", with the window size beneath it and the full
reasoning in the tooltip. "12 particles on this frame" would understate a 4096^2
frame by fifteen sixteenths. The overlay is placed back at its offset so it draws
over the region it describes instead of in the corner.
Also adds `tutorial_particles` to the user-facing Tutorial Data menu: the same
24 x 96x112 synthetic movie the tests use, with its ground truth stamped in
metadata. Asked for directly -- a dataset where the whole loop is interactive
while the large-data path is still being worked on -- and its docstring carries
the numbers above so the size choice is not mysterious.
91 python tests, 5/5 segment_wizard.spec.ts, typecheck and build clean.
Reported as "I still can't scribble... pretty sure you wired that wrong". It was wrong, in two independent ways, and neither was visible to any test: 1. NOTHING EVER CREATED A BRUSH WIDGET. `add_brush_widget` was called nowhere in the codebase, so Shift+drag had nothing to hit. The widget existed in anyplotlib and no code asked for one. 2. THE CARET LISTENED ON A PATH THE BRUSH CANNOT REACH. It waited for a renderer-side `spyde:figure_event` carrying a points array, but an anyplotlib stroke travels to PYTHON (event_json -> Figure._dispatch_event -> Widget._update_from_js -> plot.callbacks.fire). The renderer never sees it, so even with a brush present no stroke could have arrived. The existing tests passed throughout because they posted a synthetic seg_paint payload -- which exercises the rasteriser and proves nothing about whether a real stroke can get there. The new tests drive the WIDGET: 0 -> 51 labelled pixels from a 4-point stroke. The brush is now created, owned and read in Python, armed only on the Scribble tab (it floats over the image, so on Classical it would be a paint cursor over data with nothing to paint into) and detached on leaving it or closing the caret. Only unconsumed strokes are applied -- the widget accumulates for its whole life, so replaying the list every event would re-paint everything and grow the class counts quadratically. seg_paint survives as the programmatic/test door, and both paths now share ONE rasteriser (`_paint_stroke`) so brush and eraser cannot disagree about which pixels a path covers. INSTRUCTIONS, which were also missing and were asked for: arming now emits "Shift+drag on the image to paint labels · plain drag still pans · pick the class and brush size on the strip beside the plot". A brush that arms silently is undiscoverable -- no cursor change, no affordance -- so the honest outcome was a user dragging at a picture that never responds. And when the brush is genuinely unavailable it SAYS SO with the version numbers, instead of leaving the image unresponsive: the widget needs anyplotlib >= 0.5.0 (PR CSSFrancis/anyplotlib#47), and SpyDE's floor is still 0.4.2. Also puts the small dataset where it was asked for: Examples -> Dummy Data -> "Particles & Drift (small)". 47 wizard tests, typecheck and build clean.
Two user-visible bugs, one root cause: "I can only scribble one colour" and
"delete doesn't work".
The ClassStrip's onSelect/onEraser set REACT state and nothing sent it to Python.
Meanwhile the backend read `active_class` and `erase` out of `wiz.params` --
parameters that were never declared in DEFAULTS and never set by any action. So
`params.get("active_class", 0)` was permanently 0 and `params.get("erase", False)`
permanently False. Every stroke painted class 0; the eraser was a no-op.
Fixed at all three layers, because any one alone leaves it broken:
* DEFAULTS declares `active_class` and `erase`, so _coerce types them and they
are real parameters rather than silent fallbacks.
* The caret's params() sends them, and the strip's handlers call tune()
instead of only setting local state (brush SIZE already did this, which is
why size worked and the other two did not).
* seg_tune pushes them onto the LIVE WIDGET via _sync_brush. This is the part
that actually matters: the widget tags each stroke with its own class at
paint time in JS, so a class change that stops at wiz.params paints the
previous colour forever, and `erase` is a widget MODE the handler cannot
apply after the fact because the stroke arrives already tagged.
Verified end to end rather than by inspection: classes go
{0:27} -> {0:27,1:27} -> {0:27,1:27,2:27} across two switches, and erasing over
the class-0 path takes it 27 -> 0 while classes 1 and 2 are untouched.
The new tests assert on the WIDGET's attributes and on what landed in the store,
never on the params dict -- params was already "correct" the whole time the paint
was wrong, so a params-level test would have passed throughout.
51 wizard tests.
…to 2.8s Reported as "the classical segmentation is just too slow for a 4k x 4k image". Profiled rather than guessed, and the answer was not where intuition points: `distance_transform_edt` is 61% of a 4096^2 frame (3.93 s of 6.40 s) while watershed itself is only 9%. Optimising watershed would have been wasted work. The EDT exists to seed markers and to give watershed an elevation, and neither needs full resolution. Above ~2 MP both now run on a decimated grid with the elevation bilinearly upsampled and rescaled back into pixel units. DETECTION IS UNTOUCHED, and the distinction is the whole justification: the threshold still runs at full resolution, so WHICH bodies are found is unchanged and section 0.9's faint-particle sensitivity is unaffected. Only the cut BETWEEN two touching bodies moves, by about `factor` px. `min_separation` is divided by the factor so a 3 px separation does not silently become 6 and merge two markers. Measured, and the accuracy claim is checked rather than asserted: 4096^2 touching 7.09 s -> 2.82 s (2.5x) 162 = 162 particles, area 4762 = 4762 4096^2 isolated 8.14 s -> 2.80 s (2.9x) 81 = 81 particles 1024^2 0.34 s -> 0.30 s unchanged (below the threshold) Identical counts and identical median areas at 0.0% difference on both touching and isolated fields. Also worth knowing and now recorded in benchmarks.md: turning "Split touching" OFF is a further 1.9x (2.80 -> 1.45 s) and is simply correct when particles are isolated, since watershed then has nothing to do and the EDT is skipped entirely. Decimating past 2 buys almost nothing once the EDT is no longer dominant. Next levers, unmeasured, recorded with reasoning: tile+thread (scipy's EDT and watershed are single-threaded and release the GIL -- the region_sum.py precedent got 6.6x from that shape), GPU EDT, and skipping the split for components whose area already matches a single-body prior.
"Still can't swap between labels for the painting" -- while "delete works". That pair is the diagnosis: both flow through the SAME handler on the SAME stroke, but `erase` was read from `wiz.params` and the class from the widget's `stroke_classes`. The one with a Python authority worked; the one with a JS authority did not. `Figure._push_widget` sends a targeted update that never writes `panel_<id>_json`, so a Python-side `brush.class_id = 1` does not reliably reach the widget before the next stroke -- anyplotlib's own brush test has to press Shift BEFORE pushing the class to dodge exactly this. So `_on_stroke` now takes the class from `wiz.params["active_class"]`, like erase already did. Nothing is lost: a stroke cannot change class midway, so the active class when it completes IS its class. The widget push stays, but only so the stroke DRAWS in the right colour while painting -- cosmetic, no longer load-bearing. MY EARLIER TEST WAS THE PROBLEM. It set `brush.class_id` and read it back, both on the Python side, so it passed while the app was broken -- it never modelled the JS widget being out of sync, which is the entire failure. The replacement forces the widget's class STALE and wrong and asserts the stroke still lands where the caret asked. That change also broke `test_a_second_stroke_only_paints_its_own_points`, which had encoded the old wrong authority by handing the widget different stroke_classes. Updated to switch class through seg_tune, the way the strip does. 52 wizard tests.
Snapshot of the uncommitted particles/drift work so the rebase onto main has a safe point to fall back to. Contains the vectorised measure_frame (props/hull/ contours/intensity), the batch fan-out, the CNN scribble prototype and its benchmark, plus the segmentation overlay and wizard changes. Four files that had been duplicated here were dropped first -- profile_backend, benchmark_nav_fill_dispatch, test_chunk_dispatch_guard and test_live_fill_poller are byte-identical to the copies that merged to main via #95, so they come from there now.
…olver (A2-A5) Rigid translation removes what moves the whole frame. What is left has two physical causes, both real, so the model is SELECTABLE rather than assumed: * Scan distortion -- a scanned frame is not acquired instantaneously, so each ROW is displaced by a different amount. The distortion is a function of the SLOW scan coordinate, which is why the scan-knot model gives one displacement per row (Bezier-smoothed) and not a general 2-D field. A free per-row fit would have one parameter per row and would happily absorb sample motion and noise into the "scan" term, which is the failure this parameterisation exists to avoid. * Sample deformation -- the specimen bends and parts of the field move independently. No function of the scan coordinate can express that, hence the dense control-point field with bending-energy regularisation. They share the warp, the solver and the regularisation; only the parameter->displacement map differs, which is what makes building both affordable rather than twice the work. GATHER, not the KDE scatter the plan sketched. quantem scatters because it is BUILDING a reconstruction from many scans, where several source pixels land on one output pixel and must accumulate. Here the job is the inverse -- one frame resampled onto the reference grid -- which is exactly grid_sample: differentiable, fused, no normalisation pass, and no holes where nothing landed. The scatter formulation is still right for multi-scan reconstruction; it is not needed to correct a movie. Regularisation penalises CURVATURE, not magnitude: a uniform or linearly-varying displacement is what real drift looks like and must not be taxed, while a field that folds between neighbouring control points is not a physical deformation. The temporal term is what lets a noisy frame borrow support from its neighbours instead of acquiring its own wild field. Windows CUDA-autograd mitigations are both applied and both load-bearing: a warm-up backward on the calling thread before the loop, and multithreaded autograd disabled around it. Yields happen INSIDE the step loop (every ~12), not per stage, and always sync the device before releasing it -- handing off with kernels in flight is the MPS race the device lock exists to prevent. Edge policy is the locked one: out-of-bounds samples are NaN, never zero. Zero-filling is what nucleates a spurious edge "particle" downstream. Verified against the plan's acceptance criterion -- recovers a synthetic KNOWN warp, both parameterisations. Ground truth is built with the same resampler the solver uses, so the test measures the solver and not a difference between two interpolators, and the residual is scored as a RATIO of the distorted-vs- reference residual so no unjustifiable absolute tolerance is baked in. The recovery test was checked to FAIL on a solver that does not fit: steps= 1 ratio 0.791 -> assertion FAILS steps=220 ratio 0.118 -> assertion PASSES (~88% of the warp removed) Not yet wired to an action or caret; solve_nonrigid is importable and tested. 13 tests, and the whole drift suite (135) still passes.
…x 300
`nonrigid` was a stub in _UNAVAILABLE that silently reverted to rigid. It is
now the real solve: Advanced gains "Non-rigid field" (scan_knot | dense) and
"Non-rigid steps", and drift_run fits the field AFTER the rigid pass -- on top
of it, not instead of it, because rigid has already removed everything that
moves the whole frame, so what is left to fit is scan distortion or sample
deformation rather than a mixture of those and the stage.
The fit CANNOT see the movie at full size: 300 x 4096^2 float32 is 20.1 GB.
_decimated_stack reads frames one at a time and strides each immediately, so
the stack is never resident. Strided rather than area-averaged on purpose --
the fit needs crisp gradients to correlate and a box mean blurs exactly those
(same reasoning as the navigator's base-frame subsample). The fitted parameters
are resolution-independent, which is what makes fit-small/apply-large correct
rather than a shortcut.
A non-rigid failure keeps the RIGID model and says so. torch missing, CUDA OOM,
a device that will not take the graph -- none of those should throw away a good
rigid answer the user already waited for.
Measured, 300 frames, CUDA, 120 steps (benchmarks.md):
FIT, whole movie at once scan-knot dense(6x6)
128^2 (32x decimation) 2.73 s 2.29 s
256^2 (16x) 2.41 s 7.71 s
512^2 (8x) 9.08 s 28.39 s
APPLY, 4096^2, PER FRAME 385 ms 432 ms
-> over 300 frames 115 s 130 s
The apply is the expensive half and it is the one paid per frame -- 385 ms is
~23x the 60 fps budget, so a non-rigid corrected movie cannot be scrubbed the
way a rigid one can (rigid is an np.roll for integer shifts and preserves dtype;
non-rigid resamples every pixel). Fine for export at ~2 min/300 frames; for
interactive display the corrected node needs the tiered async read or a
decimated view while scrubbing. The FIT is not worth optimising: even the
slowest one measured is a quarter of the apply cost over the same movie.
test_drift_wizard's stub test asserted non-rigid reverts to rigid, which was
correct for a stub and is now wrong. Replaced with the real contract: the
selection STICKS (a caret that quietly reverted would put the wrong `kind` in
provenance), both parameterisations are selectable, and an unknown one falls
back rather than raising. 137 drift tests pass.
The benchmark also had to flush stdout before `os._exit` -- `_exit` skips stdio
flushing, so with output redirected every print was discarded and the first run
looked like it produced nothing.
…4096^2
`apply_nonrigid` had no `device` argument and always ran on the host, on a
machine whose GPU the fit was already using. That was the whole of the 385 ms
reported in the previous commit.
Profiling the stages first, rather than optimising the obvious-looking one:
CPU build field 68 ms | warp 262 ms | total 392 ms
CUDA warp with the frame already resident 7.9 ms
+ both host<->device copies 41 ms
The warp is 7.9 ms and the other ~33 ms is PCIe, so this is TRANSFER-bound and
micro-optimising the resample would have bought nothing. Two consequences drove
the fix:
* Build the field ON the device from the fitted parameters (a few hundred bytes)
instead of building it host-side and shipping it -- that would have added
134 MB to the cost that already dominates.
* Dropped a trailing `.copy()`: `.to("cpu")` already returns a tensor owning its
memory and `.numpy()` keeps that storage alive, so the copy was 67 MB of pure
waste -- a fifth of the GPU path's total.
CPU 392 -> 278 ms/frame (83.5 s per 300 frames)
CUDA 47.8 ms/frame (14.4 s per 300) 5.8x, and the default
The number worth remembering for later: a batch path that keeps frames RESIDENT
pays only the 7.9 ms, i.e. ~2.4 s for a 300-frame movie. Per-frame calls from
host memory can never beat the copy, so that is the shape a whole-movie GPU
correction should take.
CPU/CUDA agree to 4.2e-04 with identical NaN masks -- float32 grid_sample
kernels differ between backends, so this is close but NOT bit-identical, unlike
the region integrator's exact contract. Do not build an equality test on it.
MPS is not selected by device=None: the win is one fused kernel, and an
unsolicited MPS submission contends for the shared device lock the neural and
scribble paths take. Opt in with device="mps".
…ture unreachable The solve landed and every headless test passed, but a user could not select it. `drift_action._UNAVAILABLE` is DUPLICATED as `UNAVAILABLE` in DriftWizard.tsx, so removing non-rigid from the backend list left the tab `disabled` in the UI. Nothing headless can see a disabled tab: tsc was clean, 137 drift tests were green, and the finished feature was unreachable. Found by running the app, which is the only thing that could have found it. Renderer now: * drops non-rigid from its UNAVAILABLE copy (with a note naming the trap, since the duplication itself is the hazard and is staying), * renders "Field" (Scan distortion | Sample deformation) and "Fit steps" ONLY when non-rigid is selected -- they mean nothing under a rigid solve, and §0.9a is that the caret shows the task, * threads nonrigid_model / nonrigid_steps through vals.current and params(), without which the controls would have moved nothing. TabRow marked the active tab by inline STYLE alone. Added role="tablist"/"tab" and aria-selected: the state was invisible to a screen reader, and a test would have had to assert on a style object, which is a check that passes for the wrong reason. The backend now logs the fit and its RESULTING KIND. drift_run deliberately falls back to rigid when the fit cannot run (no torch, OOM, a device that will not take the graph) -- so without this there is no way, from a log or a test or the UI, to tell a real non-rigid solve from a quiet fallback, and "the caret said non-rigid" is exactly the claim that must not be trusted. e2e: the old test asserted the tab STAYS locked, which was right for a stub and is now wrong. Replaced with the real contract, plus a test that drives an actual solve and asserts the backend reports kind=dense. 6 tests pass, and the screenshots show it: Non-rigid active, "Fitting the non-rigid field (dense)…" in the status bar, the dy/dx window filling, GPU at 94%, no duplicate windows. Also renumbered the new screenshots to 10-12; they were written to 06/07, which the trace test already owns -- two tests sharing a path means the later one silently destroys the earlier one's evidence.
…nverting Found reviewing the branch. `classical._prepare` fills NaN with the finite MINIMUM before filtering -- correct, and the docstring called it "the one value guaranteed not to threshold as a particle". That guarantee is false whenever `invert` is set. `invert` (dark particles on a bright background) maps x -> -x AFTER the fill, so the minimum becomes the MAXIMUM: the drift-padded border ends up the brightest region in the image that is actually thresholded. Measured on a 96x96 frame with a 12 px NaN border: the border reads as the frame maximum and `segment_frame` puts a 240 px instance on it. This is precisely the failure `spyde.drift.warp` names as the most likely integration bug in the feature -- "a threshold applied to NaN ... invents a large 'particle' along the edge that then nucleates a spurious track" -- so a drift-corrected in-situ movie of dark particles would have grown a spurious edge track through the linker, in every frame, silently. The fill cannot simply move after the invert: it has to happen BEFORE filtering because skimage filters propagate NaN outward and would erase a band of real data. So the fill takes the polarity of the thresholded image -- finite MAXIMUM when inverting, finite minimum otherwise -- which leaves the padding at the minimum after inversion, exactly what the rest of the function already assumes. Note the asymmetry that hid this: the scribble and CNN engines both force invalid pixels to zero probability in every class (`proba[:, ~prepared.valid] = 0`), so neither could label the padding. Only the classical path relied on the fill VALUE alone, and only in one polarity. Regression test covers both polarities: the padding must not be the brightest region, and no instance may be found on it.
…und trips -> 2
Reported as "why is the computing navigator using a single worker": the
dashboard showed the distributed backend live and the cluster idle, tasks
arriving one by one, and it got worse with dataset size (an 800 GB 4D-STEM scan
and a long movie both).
Not placement. `dispatch_chunks` tops up on EVERY completion, so on the UNPINNED
lane `lane_cap - outstanding` is 1 in steady state and
n = min(submit_batch, len(pending), lane_cap - outstanding[lane])
collapses to n = 1. `submit_batch=8` only ever applied to the first fill. Every
chunk after that was its own blocking scheduler round trip with the GIL held in
the client process -- precisely the cost #95 existed to remove, reintroduced
through the back door by me, and scaling with chunk count. Hence "only big
datasets".
The window was never justified on this lane:
* it did not measure as backpressure -- benchmarks.md has bounded at 46-50 s
against unbounded at 50.2 s on the same 977-chunk movie, within noise;
* distributed >= 2022.3 queues root tasks at the scheduler, which is this
window's stated job done in the right process without our GIL;
* there is no placement decision to make when nothing is pinned, so there is
nothing for a window to balance.
Prime with one small batch, then send the rest in a single submit. 977 chunks:
one-at-a-time (the bug) 970 submits first 656 ms 12.40 s
all-at-once 1 submit first 1292 ms 5.03 s
prime + bulk (shipped) 2 submits first 45 ms 5.28 s
485x fewer round trips, 14.6x faster to first paint, 2.3x faster overall. The
middle row is why this is two submits and not one: all-at-once is fastest in
total but DOUBLES time-to-first-chunk, because the client serialises the whole
graph before anything returns -- and the progressive fill exists so the
navigator starts filling immediately. Measuring only wall-clock would have
shipped the wrong variant.
THE DUAL-LANE PATH IS UNCHANGED and keeps its window: there a completion
genuinely pulls the next chunk so a ~30x-faster GPU lane and the CPU lane drain
one pool and finish together. That is real work stealing the scheduler cannot
do, and it is why this module exists.
`batch_unpinned=False` restores the old behaviour per call, which is what the
benchmark's control arm uses.
test_chunk_dispatch_guard asserted `max_in_flight <= 4` as backpressure -- a
fence I put there myself. Retired with the argument against it written out in
the test, and replaced by an assertion on the submit COUNT; the scheduler owns
memory backpressure now. 377 dispatch/nav/progressive tests pass.
`_start_progressive_nav_compute` picks its path from `self.client` ONCE. The cluster takes ~10 s, so a file opened right after launch finds None and takes the threaded branch -- one background thread, one chunk at a time. That choice was then PERMANENT: however many workers registered a second later, the entire fill ran single-threaded. On a 977-frame movie that is the difference between seconds and minutes, and on the dashboard it looks like an idle cluster. Why it presented as a movie-only bug: a 4D-STEM scan goes through the nav-shape prompt, which is a human round trip, so its cluster is always up by the time the tree is built. A movie opens straight through and loses the race. Same code, different timing. The loop now re-checks each chunk and re-enters the dispatcher path once a client exists. Re-entering rather than switching in place is deliberate: the distributed branch owns cancellation, the sidecar save and the final repaint, and duplicating any of that in the threaded loop is how two paths drift apart. Handing over recomputes the chunks already painted, hence _NAV_HANDOVER_MIN_CHUNKS = 8 -- with only a handful left the recompute costs more than it saves. The check runs per chunk, so in practice it fires within the first few and almost nothing is redone. Tests cover the decision (hands over on the first chunk after the cluster registers; does NOT near the end; never without a cluster) plus a wiring guard, because the decision logic is only meaningful if the real loop still makes it.
Reported: "the 1k x 1k outline only shows in the classical; if you move to a different mode it continues to show, but toggle out/back in and it won't show up again." Three symptoms, one cause. The box was drawn only as a SIDE EFFECT of a successful segmentation, and `_preview` returns early when there is no engine -- an untrained Scribble, or Prompt before anything is prompted -- painting nothing and clearing nothing: * Classical is the one engine that always has a solver, so only it drew the box. * Switching to an untrained engine left the OLD box on screen, because the early return cleared nothing. It looked correct and was stale. * Toggling the caret ran _drop_overlay, and reopening in an untrained engine early-returned again, so it never came back. The box documents WHERE the 1-megapixel preview budget looks. That is true whenever the caret is open and has nothing to do with whether an engine has been trained -- arguably it matters MOST before training, when the user is deciding where to scribble. `show_preview_window()` now draws it (and clears any previous engine's outlines, so they cannot linger looking like the new engine's answer) on the no-engine path. It pushes only when the box or the cleared-state CHANGES. `seg_tune` fires on every slider tick and lands here whenever there is no engine, and `_push_groups` falls back to `MarkerGroup.set`, which re-serialises the whole panel -- so an unconditional push would put a full serialisation on every tick of a drag to redraw a rectangle that had not moved. That is what `test_an_unrelated_tune_does_NOT_force_a_panel_push` caught when I first wrote this without the guard, which is exactly what that test is for. 664 particle/seg/scribble tests pass.
…y engine Reported with a screenshot: 547 instances in the preview window where ~30 are real, and no obvious way to cut them down. The Advanced block has six knobs -- min size, max size, split touching, min separation, marker smoothing, drop edge -- and not one of them says "fewer particles". Worse, the size/shape ones CANNOT fix this failure: over-split support-film texture is often small AND round, which is exactly what a size or circularity filter keeps. Adds a per-instance confidence score and a single slider that filters on it. WHAT THE SCORE IS. Contrast-to-noise against the instance's own dilated background ring, |intensity_mean - background| / spread, squashed to [0, 1] by cnr/(cnr+1) so the control is a plain 0-100% with no dataset-dependent range to explain. That statistic is the one that separates these two populations: a real particle sits well away from its immediate surroundings, while a fragment of textured film is BY CONSTRUCTION the same brightness as the texture around it. Measured on a fixture built to reproduce the failure (30 real particles, 300 labelled fragments of a noisy film): real median 0.947 (p10 0.945) against texture median 0.090 (p90 0.221) -- a clean gap, and a 0.94 cut keeps 30/30 real with zero texture. That fixture is synthetic and built to have the property, so it validates the MECHANISM, not anyone's data; the test asserts a separating threshold EXISTS rather than any particular number. WHY IT IS FAST AND UNIFORM. The score is derived from intensity columns already measured, so scoring costs no extra pass, and `filter_by_score` is a numpy mask over a few hundred rows. Dragging re-filters an existing result and never re-segments. It also means the same thing on Classical, Scribble and Prompt because it acts on the measured OUTPUT rather than on any one method's parameters -- which is why the slider is on the default face for all three. min_score=0 is a no-op returning the inputs unchanged, so the default behaves exactly as before. An UNMEASURABLE instance (no background ring) scores 1.0, not 0.0. Absent evidence is not evidence of a bad particle, and a "hide the marginal ones" control must not silently delete things it knows nothing about. FORMAT_VERSION 1 -> 2, with MIGRATION rather than rejection. `score` is appended at the end, which is the documented way to extend this layout, so an older file now loads with the column padded (1.0, per above). Refusing it would make every previously-saved particle result unopenable in order to add one derived number -- and it is derived, so it never needed to have been stored. The layout guard still rejects a genuinely reordered/renamed layout; its test now checks that case instead of the version, which is what it was always for. 686 particle/seg/scribble/track tests pass.
…gons "A lot of particles overlaid starts to make everything slow." Every contour is a path the renderer re-transforms on each pan/zoom frame, so several hundred of them cost real interactivity. Above _RASTER_ABOVE (100) the overlay switches to a single mask; below it the vector outlines stay, because they are crisp at any zoom and each is an object the UI can hover. Only the DRAWING changes -- the particle count and every measurement are untouched. Two things had to be got right, and I had the first one wrong going in. 1. WHICH PRIMITIVE. `Plot2D` has no `add_raster` at all (only `Plot1D` does), so the marker-layer raster I had planned does not exist on the plot type the signal uses. The right primitive is `Plot2D.set_overlay_mask`, which composites client-side onto the transparent 2-D canvas that sits ABOVE the WebGPU canvas -- so it works on a GPU-rendered base. My note that "raster masks are invisible under GPU tile mode" was too broad: the mask layer is fine, and the real constraint is (2). 2. THE TILING RESOLUTION. In tile mode the renderer checks `bytes.length === iw * ih` with `iw = base_width || image_width` -- the OVERVIEW size, not the native frame. A native-resolution mask fails that check and is dropped SILENTLY: no error, no overlay, nothing in the log to say why. On a 4096² movie (always tiled) the feature would simply have drawn nothing, and looked like it was not wired up. The mask is therefore built at frame size and reduced to `base_width x base_height` whenever it is set. The reduction is a block ANY, not a subsample: particles are often a few pixels across, and striding a 4096² mask down to 1024² drops three quarters of them at random. Mask colour is the same green as the outlines so crossing the threshold does not read as a mode change, and the raster is cleared when the count drops back below it and on caret teardown, so the two can never double up. Tests cover both resolutions (native when untiled, overview when tiled, with the failing case spelled out), that small particles survive the reduction, and that a 3-particle frame is NOT rastered. 690 tests pass.
… pausing workers
Reported from a real batch run: dask workers pausing at 80% of a 9.24 GiB limit,
restarting at 95%, "Unmanaged memory: 6.47 GiB".
The dask graph was never the problem -- segment_movie is a plain map_blocks over
time chunks and nothing computes the movie, so it is genuinely embarrassingly
parallel. ONE unit of that work was simply enormous. Measured at 4096²,
classical:
input frame 64 MB
_prepare 24 MB 0.04 s
threshold 7 MB 0.02 s
split_instances (watershed) 546 MB 4.15 s <-- 64% of the peak
whole frame 852 MB 4.4 s
With threads_per_worker=4 that is ~3.4 GB of concurrent peak per worker before
frames in flight. "Unmanaged" is exactly right: the rasters are ours, inside the
task, so dask can neither account for them nor spill them.
Fix: watershed each connected component inside its own bbox. EXACT, not an
approximation -- a component is surrounded by background by definition, so a
1 px pad holds every pixel the distance transform, the markers and the watershed
can depend on, and no watershed flows between components that do not touch. Same
shape measure.py already uses for contours.
4096², 400 overlapping discs: whole-frame 538 MB / 1.59 s
per-component 272 MB / 1.09 s
2x less memory AND 1.5x faster -- the distance transform now runs over the
particles' bboxes instead of 16.7 M pixels of mostly background.
ONE BEHAVIOUR CHANGE, and it is an improvement. Counts differ on large frames
(395 vs 388) because _split_factor decimates the whole-frame split geometry 4x
at 4096² (the earlier 7.1 -> 2.8 s optimisation) while a per-component crop is a
few hundred px and gets factor 1, i.e. full-resolution markers. At 1024², where
neither route decimates, the two agree PIXEL FOR PIXEL -- that is what the
parity test asserts, and it is why the difference is decimation rather than a
cropping error.
Gated at _COMPONENT_ROUTE_PX (4 MP); below it the bookkeeping costs more than it
saves. Tests: exact parity where neither decimates, a long diagonal particle is
not broken at a crop edge, and labels stay globally unique across components (an
off-by-one in the offset would silently merge two particles). 693 tests pass.
Reported on a real in-situ movie: "14028 particles in this region", the preview window a flat sheet of green, the renderer hung, and the Scribble tab unusable because the frame you have to paint on was under that sheet. Four defects, one screenshot, and every test green through all of them. THE RASTER OVERLAY WAS UNREACHABLE ABOVE 1024 px -- i.e. on exactly the frames it was added for. The renderer sizes the mask against `base_width || image_width` (the tile OVERVIEW grid) while tile mode sets `image_width` to the FULL native frame. `_set_raster_overlay` reduced to the overview grid, which is what the renderer wants and what `set_overlay_mask` then rejected against the image shape; the ValueError landed in a bare `except` and was logged at DEBUG, so every large-frame preview fell back to one filled polygon per instance. At 14028 instances that is the hang. Verified both directions on a 4096^2 tiled plot: the overview-sized mask raised, and a full-resolution one encoded 22.4 MB that the renderer silently discards. anyplotlib now owns the reduction and accepts either shape (CSSFrancis/anyplotlib fix/overlay-mask-tile-mode), so this hands over the native mask and the arithmetic has ONE owner that cannot disagree with the shape check beside it. The failure log is now WARNING: the fallback from there is worse than the failure. NOTHING CAPPED THAT FALLBACK. `_MAX_OUTLINE_POLYS` is the seatbelt, independent of `_RASTER_ABOVE` -- reaching it means the raster was unavailable, and the honest answer is then "too many to draw" rather than thousands of paths the renderer re-transforms every pan frame. The count still reports every instance. `show_preview_window` CLEARED THE OUTLINES BUT NOT THE RASTER, so above 100 instances the previous engine's result survived a switch to an untrained Scribble -- covering the image you have to paint on. Its own docstring says it prevents exactly that; it only covered one of the two drawing routes. A FAILED THRESHOLD IS NOT A RESULT. Otsu on a low-contrast frame has no bimodal histogram to find, lands inside the noise, and the split shatters the support film. Measured on a stand-in (noisy film, 8 faint particles, 1024^2): defaults give 4873 instances at 39% coverage; min_size=2000 gives 17 at 7.2%; and gaussian=2 + min_size=200 + no watershed gives 8 -- which looks like the 8 real particles until you see it covers 52% of the frame, because those 8 bodies ARE the film. So `_threshold_failed` tests count AND coverage, neither being diagnostic alone, and the caret names the failure and points at Scribble (plan 0.9) rather than reporting a number. It does not silently re-tune. Also: contours are ~half a preview's cost at high instance counts and are discarded above the draw cap anyway, so `measure_frame(want_contours=False)` skips them there -- 1513 -> 969 ms at n=4873 with the rows BIT-IDENTICAL, so count, histogram, median and the confidence filter are untouched. That breaks the one-contour-per-row correspondence `SpyDEParticles.from_frames` needs, so `commit()` refuses such a preview and says why instead of building a store whose outlines do not match its rows. `set_overlay` now takes the instance count explicitly, since `len(contours)` is no longer it. The test for the tiling trap FAKED `set_overlay_mask` and asserted the mask SpyDE built, so it never enforced the contract that failed and was green throughout. It uses a real Plot2D now and asserts the bytes that SHIP. 692 particle/seg tests pass; 3415 in the full suite.
`ee74ad6` swapped the Confidence slider for two nanometre sliders and left four things broken at once. None was visible to pytest or tsc, and `segment_wizard.spec.ts` -- which asserts the caret is visible on its first line -- would have caught the worst of them on its own. It had not been run. `Field` WAS USED BUT NEVER IMPORTED, so the caret threw on mount and the window came up blank. One line; it was already exported from WizardShell. CONFIDENCE WENT NOWHERE. The commit's comment says it moved to Advanced. It did not: `min_score` kept its state, its payload field and a working backend filter with no control anywhere able to move it off 0. Restored under Advanced, which is where that comment always said it was -- and it matters, because it is the only control that cuts over-split film texture, which is small AND round and so survives every size and shape filter in there. THE nm SLIDERS ARE ONLY nm ON A LENGTH AXIS. They divide by the signal's scale, and a reciprocal-space signal reports a perfectly healthy positive scale in nm^-1 -- so the conversion ran and produced a merge radius wrong by the camera length while the caret still read "nm". The label now follows `face_units`: nm where the axis is a real-space length (via `_NM_PER`, so um and A work too), px where it is not. ADVANCED DID NOT FIT. Stacked in one column it measured 907 px in an 805 px MDI area, and the caret has no scroller of its own (the Threshold menu is absolutely positioned, so an overflow:auto ancestor clips it) -- so the size histogram and Commit Frame were unreachable, not merely cramped. Two columns is plan B7's own answer and the only one that neither deletes a control nor needs a scroller; the caret widens only while Advanced is open, so the collapsed face stays calm. ...which then exposed a placement bug worth its own paragraph: a side-placed caret anchors its RIGHT edge to the window's left, so at 520 px it walked off the left edge of the app entirely and its controls became unclickable. FloatingToolbar clamps side placements into the MDI area now -- overlapping the owning window is recoverable, being off-screen is not -- and the caret's measured width is mirrored into state so the clamp re-runs when a disclosure changes it. When there IS room the arithmetic is identical to before. `expectCaretFits` checks BOTH AXES now. Vertically-only is how the off-screen placement got through: the two-column caret fit vertically while half of it sat outside the viewport. The control-count guard goes 7 -> 9 for the two nm sliders. It stays exact on purpose -- it is what stops the face refilling one reasonable-looking addition at a time -- and the three filters that had no coverage now have some, each polling `data-seq` to prove the control reaches the backend rather than merely existing, which is exactly what a dead slider passes.
The bundled particle fixture is small, clean and high-contrast, so a global threshold works on it and every existing spec stayed green while a real low-contrast in-situ frame produced 14028 instances, a hung renderer and an unusable Scribble tab. A fixture that cannot reproduce the failure cannot guard the fix. `load_test_data_particles` takes `noise` now. Around 0.35 there is no bimodal histogram left for otsu to find, it lands inside the noise, and the split shatters the film exactly as reported; at the default 0.015 nothing changes. `seg_oversegment.spec.ts` pairs that with `size: [1200, 1200]` -- above anyplotlib's 1024 tile threshold -- and that pairing is the whole point, since either alone reproduces nothing. It asserts the fixture really does over-segment first (or the spec proves nothing), then that the caret NAMES the failure rather than counting it, that the raster-failure WARNING and the outline draw cap both stay silent, that the overlay is drawn but does not blanket the window, and that switching to Scribble leaves the frame clean.
benchmarks.md gets the over-segmentation table (otsu on a low-contrast frame is
39-53% coverage at every setting, and the one that yields 8 instances covers
52% -- those 8 bodies are the film) and the preview cost breakdown: the
watershed's 639 ms is inherent to a mask covering 39% of the frame and is NOT
the part to optimise, while the contours are the removable half.
The plan gets both traps as 0/0b/0c, because each cost real time to find:
* ASSERT ON THE BYTES THAT SHIP, never on what the caller built. The test for
the tiling trap faked `set_overlay_mask`, so it checked the mask SpyDE
constructed and never the contract that actually failed.
* A SILENT `except ... log.debug` AROUND A RENDER CALL HIDES A HANG. The
fallback was worse than the failure it was hiding.
* THE CLEAN FIXTURE CANNOT REPRODUCE A THRESHOLD FAILURE, which is why the
suite was green while the app was not.
Screenshots under docs/pr/seg-overlay-seam/ are the before/after pairs, each
captured by reverting exactly one fix and re-running, so they isolate what they
claim to: no overlay at all -> the mask draws; the dead classical result buried
the Scribble frame -> the frame is clean.
Tried on a real 132.8 GB in-situ movie, the classical engine did not work, and
the decision to remove rather than retune is measured:
* NO PARAMETER SETTING RECOVERS THE PARTICLES. A global threshold needs a
bimodal histogram; a low-contrast frame has none, so otsu lands inside the
noise and the split shatters the support film. On a 1024^2 stand-in:
defaults 4873 instances at 39% coverage, min_size=2000 gets it to 17 at
7.2%, and the settings that yield a plausible EIGHT instances cover 52% of
the frame -- those eight bodies ARE the film.
* THE PREVIEW COULD NOT PREDICT THE RUN. The caret previews a centred ~1 MP
crop and the batch segments the whole frame, and otsu is computed from
whatever array it is handed: 120.0 on the crop, 146.0 on the full frame of
the same image. So tuning the preview did not control the run, which is the
"I had it well fit and then it blew up" report.
* Plan 0.9 already made the learned classifier the primary path.
A trained scribble head has neither problem -- it is a fixed function, so a
crop, a tile and the whole frame get the same answer. This removes a class of
bug rather than fixing one, and it removes the per-tile threshold-consistency
problem that progressive preview tiling would otherwise have to solve.
`classical.py` -> `instances.py`, holding only the shared split every engine
funnels into. Six `SegmentParams` fields go with the engine; a `.spyde` saved
with them still loads, because `_coerce` and `_segment_kwargs` name known keys
rather than passing a dict through. `resolve_engine` RAISES for classical
instead of substituting something that behaves differently.
FOUR THINGS THE DELETION EXPOSED, all of which shipped silently:
1. The brush was armed only by `seg_set_method("scribble")` -- the verb the
caret sent when you clicked the Scribble tab. No tab row, nothing sends it,
and the caret opened telling you to paint with nothing to paint with.
`seg_open` arms it now, which is the right trigger: there is one engine and
painting is the only way to teach it.
2. `_emit` addressed messages with a window id CAPTURED at construction, itself
copied from `Plot.window_id`, itself copied from `plot_window` when the plot
was built. A plot that acquired its window later carried None through both
and every message went to nobody -- the renderer filters `seg_state` by
window id, so the caret showed an empty class list and no brush strip. It
resolves from the live plot at send time now, and warns rather than dropping
silently.
3. Every preview became a torch submission. That was harmless while classical
(pure scipy) was the default; afterwards one preview per tune and per
navigator move went at the device with nothing bounding them. The preview
now takes the same device semaphore as the batch -- and deliberately NOT
`_cap_torch_threads`, which is documented as a no-op off a worker because a
single latency-sensitive call should keep the whole machine.
4. `show_preview_window` cleared the vector outlines but not the RASTER, so a
dead result stayed on the image.
GPU FEEDERS ARE NOW CAPPED, which is the reported freeze. `band_budget_bytes`
claims 25% of FREE VRAM per submission, read at call time, so concurrent
feeders each claim their share of a figure that ignores the others. The count
landed on four by construction -- `SPYDE_FV_GPU_CONC`=2 per process, a "one"
lane, and an app process that is NOT in the lane at all -- i.e. 100% of free
VRAM. Observed at feeders=4: six busy workers and 11.5/12.0 GB, where the card
stops responding.
* `PARTICLE_DEVICE_CONC` = 1, not find-vectors' 2. Free: batch.py's own lane
measurement is FLAT (1/2/4 feeders = 0.270/0.252/0.260 frames/s), so the
concurrency that bought no throughput was buying a VRAM multiplier.
* `PARTICLE_GPU_LANE_MAX` = 2 caps the number of PROCESSES even when the
shared setting asks for more -- the one place this path does not obey it.
At "all" the split returns NO lane, `_dispatch` hands every worker chunks,
and the per-worker gate says yes to each; the clamped mode now goes to both
so they cannot disagree. Segmentation needs this where find-vectors does
not: its device cost is a torch head plus a VRAM-proportional band, not a
fixed buffer.
* The invariant, asserted in a test that reads the real constants:
`(lane + 1) * _GPU_FREE_FRACTION <= 1`, the +1 being the app process.
Also: a coverage verdict so a failed segmentation is NAMED rather than reported
as a count (it reads the measured output, so an under-trained head trips it as
readily as a bad threshold did); a hard cap on outline polygons; the preview
skips contours it cannot draw (1513 -> 969 ms at 4873 instances, rows
bit-identical); and `_report_stage_costs` drains the per-worker stage log at
the end of a run, which every worker already recorded and only the benchmark
ever read -- so "it's slow" now arrives with a per-stage number attached.
`seg_autolabel` is a TEST DOOR: it paints from a synthetic fixture's stamped
ground truth through the real rasteriser and refuses any signal without it, so
a spec whose subject is the overlay or the batch can get a trained head without
hand-placing strokes.
…annels
Profiled on a 4096^2 frame, CUDA (TITAN X Pascal, 373 GB/s measured), 36
channels, 875 ms wall with the GPU 92% BUSY -- so this is real work, not launch
overhead and not something fusing would fix:
aten::median 226 ms x 32 7046 us each
aten::cudnn_convolution 228 ms x 400 570 us each
aten::copy_ 191 ms x 472 406 us each
aten::amin + aten::amax 15 ms x 32 ~470 us each
The last line is the argument. `minimum` and `maximum` are the SAME unfold over
the SAME windows for twice as many channels and cost 15 ms against the median's
226 -- because a median is a SORT of every window (9 and 25 taps per pixel over
16.7 M pixels) while a min/max is one pass. The transposed-reduction trick
already took it 52 -> 23 ms per band; it is still 15x its own family.
MEASURED COST TO ACCURACY: none.
with median (36 ch) bright 7/7 faint 2/2 fit 1780 ms
without (34 ch) bright 7/7 faint 2/2 fit 466 ms
Both faint probes still found, which is what plan 0.9 makes the priority, and
the fit is 3.8x faster. At 4096^2: featurise+head 1191 -> 763 ms (-36%), whole
frame 3225 -> 2814 ms, and a 7914-frame movie 7.1 -> 6.2 h.
The implementation is UNCHANGED and still parity-tested against scipy -- only
the default moved. A head saved with it keeps working: a saved model carries
its own FeatureSpec and is never re-specced from this default.
DROPPING MORE WAS TRIED AND REVERTED, which the docstring records so nobody
repeats it. Scored by the faint-probe margin over six seeds, dropping sobel and
min/max as well looked STRICTLY better -- 25 channels, 32% faster, and the
highest minimum margin of anything tried (0.990 vs 0.952), nothing lost in 12
runs. It is not free. That metric only asks whether a particle is FOUND, and
those channels also feed the boundary class, i.e. instance SPLITTING. Counting
instances from both split routes over eight seeds on the nine-particle fixture:
25 ch (-sobel -minmax) boundary WRONG 2/8 watershed ok
30 ch (-minmax) boundary ok watershed WRONG 1/8
20 ch (-sobel) boundary WRONG 1/8 watershed ok
34 ch (median off only) ok ok
Sobel is an edge detector and a seam between two touching particles IS an edge,
so a head asked to predict boundaries without it over-splits. Only the median
drop is clean on both routes. Anything further here needs BOTH gates, and a
sigma cut needs a third: the fixture's particles are 3-9 px and the largest
sigma is exactly what a big particle needs, so sigmas=(1,2,4) (22 ch, margin
0.880) is untested for real data and must not be taken on this evidence.
Batching the convolutions was also measured and rejected: padding every kernel
to the largest radius is SLOWER (0.92x, since sigma 0.5 then does 65 taps
instead of 5), and the cascade formulation is 1.31x but only ~5% of the frame
and it perturbs the numerics through five accumulated stages.
The tab row is gone with the classical engine. One shipping engine and one unimplemented one (plan B4's prompt head) is not a choice, and a tab row over a single tab is furniture -- it comes back WITH that head, not before. The consequence is deliberate and is the headline change: the caret now opens having found NOTHING, with the instruction as its primary content, because an untrained classifier has no opinion. The old face showed a result the instant it opened and on real low-contrast data that result was the support film. "Nothing yet, here is what to do" beats a confident wrong answer. So the empty state IS the design: the paint instruction, the class list, Train, and a disabled "Find in all frames" until a head exists. The brush strip is up from the start rather than gated behind a tab, because painting is now the only way to teach the only engine -- hiding it hid the one thing the user must do first. Everything engine-independent stays on the face: the two nanometre filters, the count, the run button. The Advanced disclosure keeps its two columns and loses only the `detection` block, whose six knobs all belonged to the deleted engine. What counts as foreground is the classifier's business now, and a classifier is taught by painting, not by a slider. The coverage verdict survives the deletion and is reworded rather than removed: it reads the measured OUTPUT, so an under-trained head that learnt "film is particle" trips it exactly as a bad threshold did -- and that is a common mistake, one careless dab on the background. `SubWindow` now exposes `data-window-id`. A staged action is routed by it (`window.electron.action(a, p, windowId)`), so a spec that needs to SEND one rather than click the control that sends it had no way to name the target window. It is also the only stable handle across a close -- a `.nth(i)` locator silently re-targets, which cost a debugging run.
Every defect this branch fixed shipped past a green suite, and in each case for
a reason in the TEST, not a gap in coverage. The fixes are as much here as in
the code.
THE HELPERS WERE SUPPLYING THE MISSING CALL. Both copies of `_scribble_wizard`
called `seg_set_method` themselves -- the exact verb whose absence was the
missing-brush bug -- so they armed a widget the caret never asked for. Removed
from both, and `test_a_brush_is_attached_by_seg_open_ALONE` now calls the one
verb the caret sends on mount. Verified non-vacuous: reverting the fix fails it.
THE E2E PAINTED WITHOUT THE WIDGET. It dispatched a synthetic `figure_event`
straight at `seg_paint`, which never touches the brush. There is now a REAL
Shift+drag over the figure canvas -- actual DOM events into the iframe, the only
path a user has -- and it too was checked against the reverted fix.
THE FAKE ENFORCED NOTHING. `TestRasterOverlayAboveThreshold` stubbed
`set_overlay_mask` and asserted the mask SpyDE BUILT, so it never enforced the
contract that failed. It uses a real Plot2D and asserts the bytes that SHIP.
THE FIXTURE COULD NOT FAIL. `load_test_data_particles` is clean and
high-contrast, so a threshold works on it and no spec could reproduce an
over-segmentation. `seg_oversegment.spec.ts` pairs `noise: 0.35` with a tiled
frame; `seg_autolabel {mislabel: true}` swaps the classes to produce the
over-segmenting head that a CORRECTLY trained one does not (measured -- which is
itself an argument for the deletion).
Also new: `seg_second_window.spec.ts` for closing a window and segmenting
another (three variants, all passing -- so it is regression cover, not a
reproduction of the report); the feeder-cap invariant read off the real
constants; and `_labels.py`, which replaces the deleted `segment_frame` for the
~15 tests that only ever needed A label image. It reproduces the old call
exactly (NaN fill polarity, blur, otsu) so their asserted counts do not move,
and it REJECTS the dead parameter names with a message saying what to use --
`gaussian=` is the natural thing to write when porting, and one slipped through.
FOUR TESTS LOST THEIR PREMISE AND WERE REPOINTED, NOT DELETED:
* "leaving scribble detaches the brush" -> CLOSING the caret does; there is
no other engine to leave to.
* "untrained scribble keeps the last preview" -> `prompt` is the unrunnable
engine now.
* The classical baseline -> "no global threshold finds the faint probes",
which is the standing justification for the deletion.
* RandomForest parity -> split. The head finds a held-out faint probe the
forest does not, which is the head being BETTER (a tree cannot predict
outside its leaves; the MLP extrapolates). Asserting agreement where one is
expected to win makes the better model fail, so parity is asserted on
BRIGHT particles and a new test fails only if the forest beats the head.
One false positive fixed while here: `test_an_unrelated_tune_does_NOT_force_a_
panel_push` counts every push to the signal plot, and that plot has an async
fill of its own -- under full-suite load its last push landed inside the window
and the test blamed a slider drag. Confirmed zero pushes in isolation with a
stack-trace spy, then added a settle. A failure that only appears in a full run
AND accuses the wrong subsystem is the worst kind.
`_opened` now autolabels and trains, since "open it and a preview appears" is
no longer reachable; `_opened_untrained` serves the tests that are genuinely
about the empty state. That also took the file from 33 min to 51 s -- the old
time was 35 tests each burning a 60 s timeout waiting for a preview that could
not exist.
Both were borrowing: Drift used `center_zero_beam`, and Segment Particles and Particle Overlay both used `peak_finding`. `peak_finding` says "find the POSITIONS of points" -- crosshairs on dots -- and segmentation does not find points. It finds BODIES and traces their outlines, which is what everything downstream measures. So: irregular masses with traced red edges, and the touching pair actually touches, cut by a red seam. Splitting touching bodies is the hard part of the instance step and the only feature of that action worth spending an icon on. Drift reads as "the same feature recorded twice, pulled back into register": the drifted frame offset from the reference, the same particle in each, and one bold correction arrow between them. House convention kept -- 512^2, white is what you have, red is what the action found. BOTH WERE REVISED ON WHAT THEY LOOK LIKE AT 26 px IN THE REAL TOOLBAR, not at 512 in an editor. Segment's first draft had near-circular blobs that were indistinguishable from peak_finding's dots at that size, which is the exact thing a new icon was for; Drift's had three stacked frames and a thin arrow that merged into mush. Detail that only survives at 512 is detail spent on nothing.
Plan 0.2 now carries the decision and its measurements, so the engine cannot be reintroduced on the reasoning that put it there: no parameter setting recovers the particles on low-contrast data (the settings giving a plausible 8 instances cover 52% of the frame -- those bodies are the film), and the preview's crop threshold (otsu 120.0) is not the run's whole-frame threshold (146.0), so tuning one never controlled the other. Also recorded: the caret losing its tab row and opening untrained is the deliberate consequence, not an oversight; and the coverage verdict survives the deletion because it reads the measured OUTPUT -- it needed a SECOND branch to do so, because the two engines fail with different SHAPES. A threshold shatters the film into thousands of fragments; a bad classifier merges it into a few hundred large blobs. Measured in the app, a mis-trained head returned 228 instances covering ~100% of the frame -- a solid sheet, plainly wrong, and comfortably under a 500-instance bar tuned on the shatter.
Adds the untracked remainder of the fast-engine work: `fast_engine.py` itself, the V2 plan it implements, and the two Playwright specs that drive it (`seg_fast_engine`, `seg_4k_fullframe`). Committed as-is so the branch is self-consistent rather than depending on files that only exist in one working tree. NOTE: the suite is NOT green here -- `test_particles_*` and `test_seg_autolabel` fail with `KeyError: 'contours'` out of `particles_action.py`, i.e. the preview payload this engine returns does not carry the key the wizard reads. That is a real gap in this branch, not a test problem.
… phantom batch
The live preview is mask-only by design, but commit() still read
prev["contours"] -- a KeyError on the ordinary tune-then-Commit path.
Split on commit instead, on a worker, per plan section 3(iii): re-run the
engine on the previewed window (the stored mask has lost its boundary
channel, so splitting it would silently take the watershed route and
disagree with the batch), then measure and commit through the same
one-contour-per-row guard both commit paths now share.
close() never cleared _seg_batch_running despite lifecycle.py's docstring
saying it does, so a closed tree kept reporting a running batch. The test
that should have caught it removed the tree from session.signal_trees
itself and asserted on the emptied session; it now asserts the flag on
the tree after close(), and the close-list guards match quoted tokens
instead of substrings ("particles" was satisfied by _seg_pending_particles).
The wizard tests catch up with the contract they predate: count == -1
means not-counted, the box group ships empty on the mask path, the Commit
tests drive the async split, and the two chip tests that could never fail
(untrained seg_run bails before the chip is raised; del msgs[:] erased
the list they then re-read) now run a real trained batch filtered by the
result window's id. Every seg_train fixture pins device="cpu" --
auto-selection was torch-CUDA inside pytest, the documented Windows trap.
The review's verdicts, applied: - test_particles_batch: force_cpu=True at every in-process resolve_engine call -- EngineSpec carries no device, so loading the head re-selected CUDA under pytest; all 8 reds traced here - test_particles_gpu: prime cuBLASLt before the conv feature stack. On Pascal + cu124 + Windows the FIRST cublasLt initialisation fails with CUBLAS_STATUS_NOT_INITIALIZED when it happens late, established by minimal pair on an idle card -- not contention, not VRAM. Areas now compare within 5 px on matched pairs, per the module's own IoU contract - test_nav_cluster_handover: drive the real _start_progressive_nav_compute with a client that flips mid-fill, instead of a hand-copied replica of the handover decision; handed_over_with is finally asserted, and blocks computed after handover are measured, not inferred - test_chunk_dispatch_guard: a pinned dual-lane test pins the per-lane in-flight window and the per-completion pull -- nothing covered the pinned lane after the unpinned window's retirement - test_particle_overlay: five tests spy the real _push_markers seam; the builder-level asserts stayed green with the push dropped entirely - test_device_lock: FastScribbleClassifier fit/predict join the lock roster -- the shipped engine had no lock coverage, the superseded one did - drift: the superseded-preview test gates a real tune on an Event (red without the generation bump), four sleeps became completion signals, and the expensive nonrigid fits are class-scoped fixtures - across the board: flat sleeps -> real signals, 60-180 s wait ceilings -> 30 s, autouse Session fixtures off the pure-function tests, and the scribble rescue flag accumulates instead of being overwritten (the finding it exposed is documented on its xfail)
… deleted benchmark_particles_batch imported segment_frame and defaulted to a classical EngineSpec -- an ImportError before the first timing, since this branch removed both. Scribble-only now, with training labels from _labels.labels_from, os._exit(0) after the flush like the other three, and --purge-cache for the page-cache trap (an explicit warm-cache caveat when omitted). The other three get the review's methodology nits: a throwaway warm solve before the nonrigid FIT table so cold CUDA init stops landing on the first cell, submit-counting that sees non-list submits plus a guard on the no-chunk case, cuda synchronize around the scribble fit/predict timings, and the same cuBLASLt priming the GPU test needed.
Both sides landed the nav-dispatch wave, so the dispatch/handover test files conflicted as add/add against #115's copies -- resolved to this branch's post-review versions (real-loop handover drive, pinned-lane window test, armed-predicate waits), which supersede them. close()'s clear-list is the union of main's _ipf_window/_ipf_pick_fn and this branch's _seg_batch_running; test_device_lock keeps both new lock classes (particles here, fitting/EBSD from #114); benchmarks.md keeps both appended sections. pyproject drops the committed editable ../anyplotlib override -- the brush widget it existed for landed in 0.7.1, and uv.lock is main's. The one test that needs the still-unreleased overlay-mask-tile fix (test_tiled_mask_SHIPS_at_the_OVERVIEW_size) now gates on the fix's presence (_reduce_mask_any) rather than a version: it runs against the dev checkout and skips on released 0.7.1. That skip is a release gate -- ship anyplotlib fix/overlay-mask-tile-mode and bump the pin to lift it before this feature is called complete.
`_opened` fires the discovery preview, which streams `drift_trace` batches with the same window id and the same leading frame indices as the run that follows. Under CI timing they interleave, and the raw point count read 10 for 8 frames (windows-py3.12, first matrix run). Snapshot the message list before drift_run and assert the run streamed every frame INDEX -- the contract is "every solved frame went out", not "nobody else spoke". No-drop coverage stays with the filled == N_FRAMES wait above it.
0.7.2 ships fix/overlay-mask-tile-mode (set_overlay_mask usable in tile mode, both directions), which was the last unreleased anyplotlib change this branch leaned on. The capability gate on test_tiled_mask_SHIPS_at_the_OVERVIEW_size finds _reduce_mask_any in the published wheel (verified against PyPI, not the dev checkout), so the test now RUNS on CI instead of skipping -- the green-with-skip release gate is closed. Lock updated for anyplotlib only.
numba/llvmlite (0.65.1 and 0.66.0 alike) miscompiles the particle kernels on the cp310/cp311 wheels on every OS, and on macOS on every python: the uninstrumented builds corrupt the heap and the matrix dies with 0xc0000374 / SIGABRT wherever the corruption happens to be DETECTED (hull, intensity, contours, instances). The minimal pair, on a Windows py3.11 venv against this repo: a cold compile crashes in label_contours; the same cold compile under NUMBA_BOUNDSCHECK=1 passes 13/13 with zero bounds violations. In-bounds indexing plus instrumentation-dependent behaviour is a miscompile, not a kernel bug, so the remedy is the instrumented codegen on exactly the affected legs. The shipped app pins Python 3.12 -- a clean leg -- and is unaffected. Remove when a fixed numba lands.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this is
The segmentation/particles wave: the fast scribble engine replaces the classical threshold engine (removed, with the numbers recorded in-repo), with a mask-only live preview that splits instances on demand and on commit rather than per tune. Around it:
requires_particlesgating, live overlay with editing, frame-to-frame linker (birth/death/merge/split), batch segmentation with dual-lane (GPU/CPU) dispatchThe verification wave (last three commits)
A targeted review of the branch's test suite (~11k new test lines) found 10 failures / 12 errors and a set of tests that could not fail. Everything found is fixed:
SegmentWizard.commit()raisedKeyError: 'contours'on the ordinary tune-then-Commit path with the fast engine — commit now re-runs the engine on the previewed window and splits on a worker (plan §3(iii)), sharing the one-contour-per-row guard with the full-preview pathclose()never cleared_seg_batch_running, so a closed tree kept reporting a running batch; the test masking it asserted on a self-emptied sessionresolve_engine/seg_trainpins CPU (auto-selection was torch-CUDA under pytest, the documented Windows trap);FastScribbleClassifierfit/predict joined the device-lock roster (the shipped engine had no lock coverage)_start_progressive_nav_computeinstead of a hand-copied replica; a pinned dual-lane test covers the per-lane in-flight window (nothing did after the unpinned window's retirement); overlay tests spy the real_push_markersseam; the two vacuous chip tests run a real trained batchCUBLAS_STATUS_NOT_INITIALIZEDwhen it happens after cuDNN conv work (minimal pair on an idle card — not contention). The GPU parity driver and both MLP benchmarks prime it earlyTest status
spyde/tests/migrated/suite: 3443 passed, 3 skipped, 1 xfailed, 1 xpassed, 0 failed (15:43 local)test_particles_gpu.pyon real CUDA (subprocess pattern): 7 passed — green for the first time--helpclean;benchmark_particles_batchported off the deleted classical engine (it was an ImportError before the first timing)For review
test_particles_scribble.py: after fixing an accumulation bug,labels_from(sensitivity=0.9)can rescue both faint probes in full-suite order (fg 14.1%, under the flood line) — order-dependent within a process, so whether this weakens the classical-removal rationale is a deliberate call, not a test fix