Lookout log · OpenCV AI Competition 2026

The camera wall finally points at the smoke.

It is 04:50, and one dispatcher faces a wall of camera tiles — ALERTCalifornia alone runs more than 1,200. PlumeRank scores every camera against its own last 15 minutes and ranks the wall under a hard budget of 2 alerts per camera-window. On 54 held-out fires it caught 27; counting every miss as never, the median time to detection is 38.0 minutes. The same neural network run on every frame caught 2.

git clone https://github.com/edycutjong/plumerank && cd plumerank && make verify
≤3 min · no keys · no account · recomputes real metrics on committed pixels
Real HPWREN frames · Border58 Fire, Otay Mountain West, 22 Jun 2024 (development set).
PlumeRank detected it at +8 min. Neither baseline detected it within the window. results/arm_dev_2026-10-04.json
Status · 10 Oct 2026Held-out result committed: 27 of 54 fires detected, median 38.0 min, read once after the constants were frozen (results/bench_2026-10-10.json). The development numbers appear only where labelled.
Ledger · what the files say

Minutes to a human look, at a fixed budget.

Published smoke detectors report per-frame F1, or time-to-detection with no cap on false alarms. Neither tells a dispatcher which tile to open, or tells an agency whether the system survives a real shift before someone mutes it. PlumeRank reports one thing: minutes after ignition until a human is pointed at the right camera, with every system held to the same alert budget.

27/54
held-out fires caught
results/bench_2026-10-10.json · headline.n_detected
38.0 min
median time to detection · misses count as never
same file · headline.p50_minutes
1.4 vs 3,645
CNN calls per camera-window, PlumeRank vs B1
same file · model_calls_per_camera_night
735
tests · 97% coverage
README.md · Testing · 1,500 property-based cases

Held-out = 54 FIgLib fires no constant was fit on, read once after the constants were frozen. 27 of 54 is exactly half, so the median is fragile: it reads 32–38 min with the ignition clock shifted −5…0 min, and does not exist at +1…+5 (results/sensitivity_2026-10-10.csv). The development set (29 fires, in-sample) gave 17/29 and 32.0 min.

Scene 01 · the wall, shown twice

62 cameras. One list. Nothing hidden.

On the left, the live wall on the AWS box: 29 development fires and 33 fire-free confusion windows, replayed one minute of footage per second, re-ranked on every tick. On the right, the judge command run on a clean checkout. It recomputes the micro metrics from committed pixels, then prints the published held-out headline verbatim, labelled as not recomputed.

44.242.91.201 / the wallopen live ↗
The live PlumeRank wall: 62 HPWREN camera tiles with a ranked watch list on the right; rank 1 and rank 2 carry alert tags and per-statistic scores

Captured from the live box on 10 Oct 2026. The watch list keeps all 62 cameras and only reorders them.

plumerank — zshreplay of a real run
$ .venv/bin/python scripts/bench.py --arm micro
windows unified to 31 frames
CPUCreditBalance before : n/a
CPUCreditBalance after  : n/a

========================================================================
REPRODUCED HERE
========================================================================
  what           : the micro corpus, recomputed live by the real OpenCV 5 pipeline
  corpus         : data/micro/ (10 windows x 31 frames)
  PlumeRank      : p50    n/a min   p95    n/a min   (2/5 detected)
  B1 CNN grid    : p50    n/a min   p95    n/a min   (1/5 detected)      (the SAME onnx, every frame)
  B2 frame-diff  : p50    n/a min   p95    n/a min   (0/5 detected)
  realised rate  : 1.57 alerts/camera-night on the 5 non-fire micro windows (target 2.00, uncapped)
  escalation     : 0.65% of frames touched by ONNX (top-2 of a 10-window roster; the published figure in block 2 is the same rule over a larger one)
  checked against: results/verify_micro.json -- MATCH (byte-for-byte)

========================================================================
PUBLISHED, NOT RECOMPUTED HERE
========================================================================
  source         : results/bench_2026-10-10.json, printed verbatim -- not recomputed here

  PlumeRank -- median time-to-detection @ 2.00 alerts/camera-night
    arm            : hold (54 sequences, 6 excluded from the 60 in manifest_hold.json, see data/excluded.json)
    PlumeRank      : p50   38.0 min   p95    n/a min
    B1 CNN grid    : p50    n/a min   p95    n/a min      (the SAME onnx, every frame)
    B2 frame-diff  : p50    n/a min   p95    n/a min
    realised rate  : 1.99 alerts/camera-night on the confusion set (target 2.00, uncapped)
    escalation     : 1.76% of frames touched by ONNX
    model calls    : PlumeRank 1 / camera-night   |   B1 3645 / camera-night
    t0 sensitivity : advantage not finite at all 11 offsets in [-5,+5]
    background     : mog2
    commit 5df73321059f2797c32d302676b813c81b87fd0c  instance local (not EC2)  wall clock 12m29s  2026-10-10
  -> results/bench_2026-10-10.json

  reproduce with: make fetch && make bench  (about 95 min)

$ curl -s http://44.242.91.201/build-info | python3 -c "import json,sys;print(json.load(sys.stdin)['opencv']['version'])"
5.0.0
The log · every held-out fire, caught or missed

Fifty-four fires. One line each.

A fire lookout writes one line per event: time, bearing, what changed. Here is the held-out set written the same way: fires no constant was fit on, read once. Each row is a real FIgLib ignition, listed with the minute after ignition at which each system first detected it: the camera in its top four, with an alert. All three systems are calibrated to the same budget on the same fire-free windows. Misses are listed too.

T+minfire · cameraPlumeRankB1 CNNB2 diff
+0Junction Fire · 2026 · bm-e-mobo-c+0——
+0Rainbow Fire · 2026 · rm-e-mobo-c+0——
+1Springs Fire · 2025 · lp-n-mobo-c+1+24—
+1Tusil Fire · 2026 · mlo-s-mobo-c+1——
+2Lodge Fire · 2025 · sm-n-mobo-c+2—+28
+3Academy Fire · 2025 · wc-n-mobo-c+3——
+5Tenaja Fire · 2024 · buff-n-mobo-c+5——
+5Bernardo Fire · 2026 · wc-w-mobo-c+5——
+7Beaver Fire · 2026 · lp-w-mobo-c+7——
+8Henshaw Fire · 2025 · mg-n-mobo-c+8——
+8Scissors Fire · 2025 · mp-n-mobo-m+8—+9
+8Creelman Fire · 2026 · cp-w-mobo-c+8——
+9Creelman Fire · 2026 · mg-s-mobo-c+9——
+10Border 65 Fire · 2024 · lp-s-mobo-c+10——
+11Church 2 Fire · 2026 · mlo-s-mobo-c+11——
+12Monte Fire · 2025 · cp-w-mobo-c+12——
+13Posta 3 Fire · 2024 · pi-e-mobo-c+13——
+16Warners Fire · 2024 · sojr-s-mobo-c+16——
+16Bridge Fire · 2024 · starr-n-mobo-c+16——
+20Steele Fire · 2025 · om-n-mobo-c+20——
+24Mission Fire · 2026 · rm-e-mobo-c+24——
+25Junction Fire · 2026 · mg-e-mobo-c+25——
+26Junction Fire · 2026 · vo-w-mobo-c+26——
+28Border 65 Fire · 2024 · om-e-mobo-c+28——
+29Ted Williams Fire · 2025 · wc-w-mobo-m+29——
+35Henderson Fire · 2025 · bh-w-mobo-c+35——
+38Crosley Fire · 2026 · hp-w-mobo-c+38——
missProctor Fire · 2024 · sm-s-mobo-c—+18—
missBorder 65 Fire · 2024 · pi-s-mobo-c———
missPapa Fire · 2024 · rm-w-mobo-c———
missNavajo Fire · 2024 · cp-n-mobo-c———
missBridge Fire · 2024 · marconi-n-mobo-c———
missBridge Fire · 2024 · stgo-e-mobo-c——+37
missSaddleback Fire · 2024 · stgo-w-mobo-c———
missQuarry Fire · 2024 · sm-w-mobo-c———
missPasqual 4 Fire · 2024 · wc-n-mobo-c———
missResort Fire · 2024 · bm-n-mobo-c———
missResort Fire · 2024 · mg-n-mobo-c———
missUnnamed Fire · 2024 · dwpgm-w-mobo-c———
missBernardo Fire · 2025 · ch-s-mobo-c———
missPenny Fire · 2025 · ml-w-mobo-c———
missTrail Fire · 2025 · rm-w-mobo-c———
missFelipe Fire · 2025 · ml-n-mobo-c———
missSteele Fire · 2025 · lp-w-mobo-c———
missBernardo Fire · 2025 · bl-n-mobo-c———
missVail Fire · 2025 · hp-w-mobo-c———
missCoches Fire · 2025 · sm-n-mobo-c———
missRanch Fire · 2025 · wc-w-mobo-m———
missScissors Fire · 2025 · vo-e-mobo-m———
missUnnamed Fire · 2026 · wilson-s-mobo-c———
missCrosley Fire · 2026 · mpo-w-mobo-c———
missJunction Fire · 2026 · vo-n-mobo-c———
missRainbow Fire · 2026 · bh-w-mobo-c———
missSorrento Fire · 2026 · sdsc-e-mobo-c——+21
Same model, both sides

The win is not a bigger network. It is the same one, asked less often.

B1 is PlumeRank's own ONNX tile classifier, run on every frame instead of the ~2% the classical stage escalates (1.76% on the held-out run). Because the model is identical on both sides, the gap comes from the pipeline, not from model size. Every system is calibrated to 2.0 alerts per camera-window on the confusion set; all of them land at 1.99.

PlumeRank · per-camera baseline · median 38.0 min27 / 54
B2 through the same baseline · no median20 / 54
B1 through the same baseline · no median18 / 54
As deployed, no per-camera baseline
B2 · frame differencing · no median4 / 54
B1 · the same CNN on every frame · no median2 / 54

Held-out set, 54 fires, read once. "No median" means the system detected fewer than half the fires; that is the result itself, not missing data. Source: results/bench_2026-10-10.json. On the development set PlumeRank also led, with 17 of 29; there B1 beat B2 (14 vs 10 through the same baseline, 6 vs 3 as deployed), the reverse of the held-out order.

Mechanism · OpenCV 5 does the per-frame work

A camera that runs hot all night is not a fire.

23 OpenCV 5 surfaces across 6 modules, counted from the source by scripts/check_claims.py on every test run. Four classical statistics fold into one hazard per camera. Each camera is then judged only against itself.

05 · the part that carries the ranking

Each camera against its own last 15 minutes

The fused hazard is z-scored against the camera's own trailing 15 minutes, then a solver picks the one threshold that spends exactly the budget on fire-free windows. Without the baseline, the same pipeline catches fewer fires; with it, both baselines improve too, which is why the comparison above runs them through it.

z(cam, t) = (h(t) − mean15 min) / sd15 min
θ = solve_threshold(confusion set, budget = 2.0 / camera-window) → 0.849
rank.Baselinesolve_thresholdCloudWatch drift ±15%
01 · stabilise

Hold the horizon still

Mast sway would read as motion. Sub-pixel phase correlation registers each frame to the last; when the sky is too featureless to correlate, a MAGSAC++ affine fit takes over.

phaseCorrelatecreateHanningWindowestimateAffine2D · USAC_MAGSAC
02 · segment

Find what changed

A Gaussian-mixture background model per camera; connected components become the boxes the statistics read.

createBackgroundSubtractorMOG2connectedComponentsWithStats
03 · four statistics

Measured, then fused

S1base anchoring0.528
S2vertical growth0.50
S3diffuse boundary0.44
S4DIS flow drift0.538

AUC alone, pre- vs post-ignition, on frames with a component. None is a detector by itself, and this page says so.

SobelLaplacianDISOpticalFlow_create
04 · escalate

A CNN for ~2% of frames

75 of 4,266 held-out frames (1.76%) go to a small ONNX tile classifier through cv2.dnn, trained only on Pyro-SDIS (French fire services), never on a Californian pixel. On the development ablation grid, turning it off left the count unchanged in five of six pairs and caught one more fire in the sixth.

dnn.readNetFromONNXENGINE_NEWblobFromImage
Cloud · one box, on purpose

One Graviton instance. Every number committed.

PlumeRank on AWSAn S3 bucket feeds one t4g.small Graviton instance running OpenCV 5, FastAPI and Caddy; CloudWatch watches the alert rate and the heartbeat; judges reach the instance over HTTP. EC2 · t4g.small · arm64 OpenCV 5.0.0 · NEON FastAPI · 4 GET routes Caddy · static wall S3corpus + bundle CloudWatchdrift ±15% · heartbeat Judgebrowser · curl HPWREN FIgLib replay → stabilise → segment → S1–S4 → escalate ~2% → fuse → baseline → solve_threshold → ranked watch list
  • Compute
    One t4g.small (Graviton, arm64) runs OpenCV 5.0.0 built for NEON, the FastAPI service and the static site behind Caddy. /build-info prints the build string.
  • Storage
    One S3 bucket: the DEV + CONFUSE corpus (5,082 frames) and the code bundle the box boots from.
  • Watch
    CloudWatch tracks the realised alert rate on the fire-free cameras, with drift alarms at ±15% of 2.0. A heartbeat-silence alarm reboots the instance if the service stops beating.
  • Not here
    No Lambda, no managed vision API, no second instance, no upload route, no dispatch route. Four GET routes, all tested.
Honest by design · lifted from the README

What it cannot do yet, in its own words.

The median is on a knife edge, and that is part of the result.

27 of 54 is exactly half. Shift the ignition clock a minute later and the median disappears; the detection counts are the steadier comparison.

— README · What it actually measures
Cloud is the adversary — not the fog this project was designed against.

Measured before the per-camera baseline shipped (and not re-measured since): low cloud cost 4.25 and cloud shadow 3.86 alerts per camera-window, against fog's 1.15. A blind census then found the fog class picks marine layer at chance.

— README · Honest limitations, 2
The novelty is a metric reframe, not a new sensor.

The OpenCV operations are well known. What is new is the attention-budget objective they are pointed at. The system is also daylight only, because FIgLib is.

— README · Honest limitations, 1 and 4
Questions a judge would ask

Asked, and answered.

Is the wall actually live, or a recording?

Live. The AWS box replays real HPWREN FIgLib footage through the real pipeline, one minute of footage per second, and re-ranks every camera on every tick over server-sent events. Nothing is pre-rendered. /health shows the tick counter and the measured tick time.

What is mocked?

Nothing on the demo path. There are no offline, mock or dry-run branches in the pipeline, API or scripts. make verify recomputes real metrics on committed pixels and checks them byte-for-byte against results/verify_micro.json. The replay is the one thing that is not live camera input: the footage is historical, because the ground truth is the ignition time recorded in FIgLib.

Why not report F1 like every other smoke detector?

Because a dispatcher does not act on frames. They act on a short list, and they mute any system that interrupts too often. PlumeRank holds every system to the same alert budget and reports minutes after ignition until a human is pointed at the right camera. FIgLib's per-sequence ignition clock means this needs no hand labelling.

What exactly is a "camera-night"?

The code's name for one evaluation window: 81 minutes of footage on one camera, the length of a FIgLib sequence. So the budget is 2 alerts per camera per 81-minute window, about 18 per camera over a 12-hour shift if run continuously. This page calls it a camera-window to avoid the confusion; the windows are daytime.

How was the held-out number protected?

The held-out set is 60 FIgLib fires in data/manifest_hold.json, 6 excluded by a published rule (data/excluded.json), so 54 were measured, and no constant was fit on any of them. The constants were frozen and committed first (results/constants.sha256); the benchmark refuses to read a held-out frame unless that hash matches, checks every frame against the fetch receipt, and logged its single read with the commit (results/hold_access.log). The result, whatever it said, became the headline: 27 of 54, median 38.0 minutes.

Could it dispatch an engine, or hide a camera?

No. It is a ranking aid for a human. There is no dispatch route and no upload route, and tests assert that exactly four documented routes exist. The full roster always stays on screen, every ranked entry carries its evidence crop, and the alert budget is a published, logged parameter with a drift alarm.

04:50 · next sixty seconds

Open the right tile first.