Marine Debris Detector

U-Net · ResNet34 · MARIDA

Overlays
Opacity
Legend
Patch filter
Patches
Model
U-Net · ResNet34 encoder
11-band Sentinel-2 · params
Test F1 · op-thr

Every overlay, uncertainty map and metric on this page is the real trained model's output, precomputed offline on the MARIDA held-out patches and bundled — no runtime model download.

Sentinel-2 · 10 m/px · 1.0× · drag a box to zoom · drag to pan · scroll to zoom · ⤢ reset
Loading MARIDA patches…
Review Validation Label QA Uncertainty Active Learning Curation Annotators Ops Architecture

Human review

Keyboard-fast triage of model detections vs MARIDA ground truth

Precision
Recall
F1
IoU
0
true pos
0
false pos
0
false neg
0
true neg

Region-level, scored against MARIDA ground truth across the detections you've reviewed on 0 patches. Updates live.

Analyst agreement
Agreement
0
accepted
0
rejected
0
relabelled
0
needs expert
Review queue

Corrections accumulate into a retraining set. Low-confidence detections (< op threshold) route here — never auto-accepted.

Threshold & PR tradeoff

Drag the cutoff — P/R/F1/IoU recompute live on bundled GT + probs

Trained model · held-out test split
Precision
Recall
F1 / IoU

Threshold 0.50
Precision
Recall
F1
IoU

Computed over labeled pixels across the demo patches. Moving the slider also re-thresholds the map overlay in the center. The dot marks the current cutoff on the real PR curve.

Label QA

Cleanlab confident-learning · likely-wrong labels, worst first

Uncertainty

MC-Dropout epistemic uncertainty · 20 stochastic passes

Dropout is kept active at inference and the model is run 20 times per patch. Where those passes disagree, the per-pixel standard deviation is high — the model is unsure there. High-uncertainty regions are exactly what the active-learning queue prioritises.

This patch ⌀ unc
Dataset ⌀ unc
Most uncertain regions here

Confidence calibration: routing on MC-Dropout std is more trustworthy than raw softmax, which is typically over-confident on rare classes like debris.

Active learning

Next-to-label queue ranked by MC-Dropout uncertainty + entropy

Retrain loop

The ranking above is a real uncertainty computation. The retrain step needs a live GPU backend, so its projected F1 lift is a scripted illustration — no training happens in your browser.

Data curation

Encoder-feature embedding · click a point to load that patch

debris present hard negative ring = high uncertainty

Deepest ResNet34 encoder features per patch, reduced to 2D (). Nearby points look alike to the model; outliers & high-uncertainty (amber-ringed) points are the most valuable to label or double-check.

Multi-annotator & agreement

Consensus labelling · inter-annotator agreement (Cohen's κ)

Fleiss' κ (3 raters)

Three simulated annotators (a marine-ecology expert, a trained analyst, and a crowd worker) label this patch's detections. Agreement is computed live; disagreements route to adjudication.

Per-region annotator calls
consensus
disputed
expert-only
3
raters

Operations

Throughput, turnaround & auto-calibrated threshold

0
reviewed (session)
in queue
regions / min
est. to clear
Auto-calibrated confidence threshold → target 95% precision
Suggested cutoff

est. label cost saved
auto-accepted
routed to human
patches in scope

Active-learning routing means only low-confidence detections need a human — the rest auto-accept. Cost figures assume a $0.12/region annotation rate.

Production architecture

Where this graduates from a browser demo to an operational pipeline

Every box is a named open-source or managed component. The browser demo you're using exercises the Model, Uncertainty, Label QA and Curation stages on real bundled data; the ingest, orchestration and live-retrain lanes are what a full deployment adds.

Data: MARIDA (Kikaki et al. 2022, PLoS ONE) · Sentinel-2 L2A · CC BY 4.0. Model, uncertainty, label-QA & embeddings computed by GIA.

← All Labs