All Case Studies
TestR&DSensor Fusion

Inside the Bench-Rig Validation Program

By Engineering — R&D · July 6, 2026 · 7 min read

An 18-month bench rig, ~3,200 catalogued events, and staged thermal-vs-gas fusion trials — how detection is validated before it ever ships.

Before any vehicle-deck deployment, RoRoSafe's detection platform is validated on a bench rig that has staged roughly 3,200 catalogued thermal events over 18 months. That reproducible event catalogue — not the most recent voyage — is what every algorithm change is measured against, and it is where the staged thermal-versus-gas trials quantified exactly what multi-modal detection buys, and where it does not.

Why a reproducible event catalogue matters

The bench rig is undramatic — a steel frame, a movable thermal source, and twelve sensor cells in fixed positions — but its value is reproducibility. Over 18 months it produced about 3,200 catalogued events spanning early thermal anomalies, frank fires, and fourteen distinct nuisance signatures. Every algorithm change is regressed against that same catalogue, so false-positive rate can be quoted against a fixed dataset rather than against whichever voyage happened last — which is how detection products usually drift into being unmeasurable. The current alarm threshold sits at a 6 °C delta from each vehicle's rolling baseline.

What the bench rig got wrong

The rig also taught its own limits. Three things it missed before the first sea trial: solar gain (there was no solar source on the bench, so the first sailing's sun-warmed deck was a surprise), vehicle motion (engine bays cool unevenly as the ship rolls), and cargo density (close-packed vehicles share thermal mass in ways the bench did not model). All three are now represented in the rig in some form. None of them was obvious beforehand — which is the argument for staged testing that deliberately hunts for what the model leaves out, rather than testing to confirm what it already assumes.

~3,200
staged events catalogued over 18 months on the bench rig
6 °C
current alarm threshold (delta from rolling baseline)
4 of 5
abuse profiles where thermal+H₂ fusion beat thermal alone

What the fusion trials measured

The staged trials put multi-modal detection to the test rather than assuming it. A six-week thermal-plus-hydrogen trial across five cell-abuse profiles found fused detection beat thermal-only on four of five, adding a median of about three minutes of lead — but the slow-thermal-injection case, with no electrochemical fault and therefore no H₂ signature, was carried by thermal alone, showing exactly where the gas layer is and is not load-bearing. A separate four-week thermal-versus-off-gas run put numbers on the Stage-2/Stage-3 split: on venting events, off-gas led visible smoke by roughly nine minutes and thermal by about six, with fusion reaching around twelve — while on non-vent events off-gas never tripped and thermal carried the detection on its own.

Sensor fusion is a hypothesis to test, not a feature to add. The gain has to be measured against the cases that matter — including the ones where one sensor is blind.

Why you benchmark on the hard cases

The consistent lesson across both trials is that a fusion architecture must be benchmarked on the event mix a real fleet sees, not the cases that flatter one sensor. Off-gas detection looks unbeatable if you only test venting events; add the non-vent class and it misses entirely, while thermal holds. Fusion earns its added complexity precisely because it keeps the strength of each layer without losing the cases where one is blind. A detection number quoted only on vent events — or only against the latest voyage — is not a measurement; it is a selection, and it is why like-for-like, catalogue-based benchmarking is the point of the rig.

What it means for owners and class

For shipowners and class surveyors, the value of a bench-rig program is that detection claims arrive with a reproducible provenance: a fixed catalogue, staged abuse profiles, and fusion gains measured against the hard cases rather than the convenient ones. That is the same evidence discipline a class society's type-approval package demands and the same reproducibility an underwriter's evaluation tests for. The bench rig is where a detection figure stops being a demo and becomes a number someone else can independently check.

Sources

  • RoRoSafe bench-rig program records (internal R&D, Chennai): ~3,200 catalogued events over 18 months; 14 distinct nuisance signatures isolated; 6 °C delta-from-baseline alarm threshold; solar-gain, vehicle-motion and cargo-density gaps identified at first sea trial.
  • RoRoSafe staged H₂-fusion trial (six weeks, five cell-abuse profiles): thermal+H₂ fusion beat thermal-only on 4 of 5 profiles, ~3 min median additional lead; slow-thermal-injection case carried by thermal alone.
  • RoRoSafe staged off-gas-vs-thermal trial (four weeks): median lead over visible smoke — off-gas ~9 min, thermal ~6 min, fusion ~12 min on vent events; off-gas 0 min on non-vent events.
  • [VERIFY: all bench figures are internal R&D results and have not been independently published; the staged profiles, thresholds and lead times illustrate the program's methodology and should be treated as internal to RoRoSafe.]
Frequently asked

Questions, answered

What is the RoRoSafe bench rig?+

A fixed test rig — a steel frame, a movable thermal source, and twelve sensor cells in set positions — used to stage and catalogue thermal events for validating detection. Over 18 months it produced about 3,200 catalogued events (early anomalies, fires and nuisances), which serve as the reproducible dataset every algorithm change is regressed against.

Why does a reproducible event catalogue matter for detection?+

Because it lets false-positive and detection performance be quoted against a fixed dataset rather than the most recent voyage. Without a fixed catalogue, a detection product drifts into being unmeasurable — every claim is against different data. The bench catalogue is what makes a RoRoSafe detection figure something a class surveyor or underwriter can independently check.

Does sensor fusion actually improve detection?+

In the staged trials, on most events. Thermal-plus-H₂ fusion beat thermal-only on four of five abuse profiles (~3 min median extra lead), and thermal-plus-off-gas fusion reached about 12 minutes of lead over visible smoke on venting events versus ~9 for off-gas and ~6 for thermal. But on non-vent events the gas layer never trips, so fusion's value is keeping thermal's coverage while adding gas's early lead where it exists.

Why benchmark detection on non-venting events?+

Because they expose where a single sensor is blind. Off-gas detection looks unbeatable if you only test venting events, but on a non-vent fault it misses entirely while thermal still catches it. A real fleet sees a mix, so benchmarking only on vent events flatters the gas layer and overstates the system. The fusion architecture is validated on the full mix for that reason.

Related reading

Continue the thread