Journal
ResearchJuly 2026

Why an ICU Arrhythmia Alarm Is Not Yet a Clinical Event

An ICU alarm is a detector output, not physiological ground truth. Reliable ventricular-tachycardia verification must separate electrical rhythm, mechanical consequence and sensor failure without suppressing true events.

An intensive-care monitor sounds a ventricular-tachycardia alarm.

That alarm matters immediately. It may indicate a lethal rhythm requiring urgent intervention. It may also have been triggered by motion, a loose electrode, pacing, electrical interference or a waveform that the monitor interpreted incorrectly.

The alarm itself does not decide between those possibilities. It reports that a detection rule was satisfied by the measurements available to the device.

This creates three separate questions:

  1. Did the ventricular electrical rhythm actually occur?
  2. If it occurred, what mechanical effect did it have on pressure and blood flow?
  3. What clinical response is warranted in this patient and context?

Those questions are related, but they are not interchangeable. A technically correct rhythm alarm is not automatically haemodynamically catastrophic. A clinically important rhythm can coexist with a corrupted pulse signal. A pattern that resembles ventricular tachycardia in one ECG channel can be an electrode artefact while the patient's actual rhythm remains visible elsewhere.

Reliable alarm verification therefore cannot be reduced to making the ECG classifier more confident. It requires an evidence process that distinguishes the physiological event from the measurements, distinguishes each sensor from its failure modes and remains conservative when the available evidence cannot resolve the case.

An alarm is a claim about measurements

A bedside alarm is produced by a chain of operations. Electrodes or other sensors acquire physical signals. Hardware and software condition those signals. A detection process identifies a pattern or threshold crossing. Alarm logic then determines whether to issue an audible or visual notification.

An error can enter at any point in that chain.

For ECG monitoring, movement can change the electrode-skin interface and create rapid deflections. A partially detached electrode can introduce intermittent noise. Low-amplitude QRS complexes may be missed in one lead. Bundle-branch block, ventricular pacing and other wide-complex rhythms can resemble patterns that a detector associates with ventricular tachycardia. Electrical interference may affect one or several channels.

None of these mechanisms implies that ECG monitoring is unreliable in general. They explain why an alarm should be understood as a test result rather than as direct access to the underlying cardiac state.

The language used to describe alarm performance must also remain precise:

  • A false rhythm alarm means that the rhythm specified by the alarm was not present under the declared adjudication definition.
  • A true but non-actionable alarm means that the rhythm was present, but its duration, physiological consequence or clinical context did not meet the relevant threshold for intervention.
  • A technical alarm identifies a problem with a sensor, connection, signal or monitoring system rather than a patient rhythm.
  • An unresolvable event is one for which the available evidence is insufficient to support a reliable true-or-false determination.

Collapsing these categories into “nuisance alarms” conceals materially different problems. A false alarm calls for better detection or sensor-failure handling. A true but non-actionable alarm may call for different prioritisation or notification policy. A technical alarm may require staff to restore monitoring. An unresolvable event should not be silently converted into a confident suppression decision.

The measured burden is substantial, but it is not one universal number

One of the most detailed observational studies of ICU alarm burden was conducted across five adult intensive-care units at the University of California, San Francisco. The study followed 461 consecutive patients over 31 days in 2013 and recorded 2,558,760 unique audible and inaudible alarms. Of these, 381,560 were audible, corresponding to an average of 187 audible alarms per bed per day.

Clinical experts manually adjudicated 12,671 arrhythmia alarms; 88.8% were classified as false positives. The study also reported that 93% of its 168 true ventricular-tachycardia alarms were not sustained long enough to warrant treatment under the study's clinical criteria. Drew et al., PLOS ONE (2014)

These figures establish the scale and layered nature of the problem in that study. They should not be quoted as a universal false-alarm rate for every contemporary ICU. The data came from one health system, a particular monitor environment and a defined month more than a decade ago. Alarm configurations, software and clinical practice vary between institutions and have changed over time.

The newer VTaC benchmark provides a more diverse view specifically for ventricular-tachycardia alarms. Its current release contains 5,037 adjudicated alarm events from commercial bedside monitors at three major US hospitals and from three monitor manufacturers. Every event includes at least two ECG leads and one or more pulsatile waveforms: photoplethysmography, arterial blood pressure or both. Of the 5,037 events, 1,441—28.61%—were labelled true. Lehman et al., VTaC, PhysioNet (2024)

VTaC also exposes the difficulty of establishing the reference label. At least two experts independently reviewed 5,742 candidate events. Their initial decisions conflicted on 1,208 events, or 21.04%. Some disagreements were adjudicated; events that remained unadjudicated, uncertain or rejected were excluded before the final 5,037-event release.

That disagreement is not evidence that expert review lacks value. It shows that real waveforms can be genuinely ambiguous and that a professional benchmark must make uncertainty visible. A dataset containing only the cases on which a final label was reached cannot, by itself, establish how a deployed system should handle the cases that experts found unreadable or unresolved.

Electrical rhythm and mechanical consequence are different states

ECG, arterial pressure and photoplethysmography do not provide three interchangeable measurements of the same quantity.

  • ECG measures differences in electrical potential at the body surface. It provides evidence about cardiac electrical activation and rhythm.
  • An arterial-pressure waveform is produced by an invasive arterial catheter and pressure-transducer system. It provides beat-to-beat evidence about arterial pressure at the measurement site.
  • Photoplethysmography, including the pulsatile component used by pulse-oximetry sensors, optically measures changes associated with peripheral blood volume at the sensor site.

A real ventricular electrical event should be interpreted together with its expected mechanical consequences, but the relationship is not a simple equality. Mechanical output depends on the rhythm as well as preload, afterload, contractility, vascular state and the location and quality of the mechanical sensor.

This distinction enables several clinically different observations:

  • An ECG channel may display rapid artefact while a clean ECG lead and valid pulsatile channels continue to show the patient's underlying rhythm.
  • Genuine ventricular tachycardia may be accompanied by a changed but persistent arterial pulse.
  • Genuine ventricular tachycardia may produce severely impaired or absent effective circulation.
  • A PPG waveform may disappear because of peripheral vasoconstriction, motion or loss of sensor contact rather than because cardiac mechanical output has ceased.
  • An arterial-pressure waveform may be unavailable or technically distorted even when the cardiac event is real.

The UCSF alarm study published concrete examples of this reasoning. In one false VT alarm, six ECG leads contained motion artefact while another lead continued to show sinus rhythm; arterial-pressure and oxygen-saturation pulse waveforms also continued at the underlying sinus rate. In a true VT example, the arterial-pressure waveform fell to near zero during the arrhythmia. These were illustrative cases, not universal decision rules. Drew et al., PLOS ONE (2014)

The defensible engineering principle is therefore conditional physiological corroboration:

Use signals produced by different physical processes to test whether the proposed event is jointly compatible with the available evidence, but only after establishing that each contributing signal is valid enough for that test.

This is not majority voting. Two pulsatile channels are not two independent witnesses when both depend on the same failing circulation. Nor can a normal-looking pulse channel automatically overrule a well-supported electrical rhythm. The evidential value of each channel depends on what it measures, how it can fail and whether it is usable at that moment.

More sensors do not automatically create more certainty

Multimodal monitoring is valuable because different sensors expose different parts of the physiological process. It also introduces additional failure modes.

An ECG lead can detach. A PPG sensor can move or lose peripheral signal. An arterial line can be damped, flushed, disconnected or otherwise distorted. Channels can have different sample rates, delays and data gaps. Patient movement can contaminate more than one sensor. A downstream communication or integration fault can make an otherwise valid waveform unavailable to the alarm-verification layer.

Simply concatenating every available waveform does not resolve these problems. A model may learn to rely on the channel that is easiest in a development dataset, even if that channel is frequently missing or differently configured at another hospital. It may interpret absence of a signal as physiological absence when the actual cause is sensor failure. It may also inherit manufacturer-specific processing characteristics that look predictive in one data source but do not transfer.

A deployable verification system therefore needs separate evidence about:

  • whether each expected channel is present;
  • whether its signal quality supports the proposed use;
  • whether its timing is adequate for cross-channel comparison;
  • whether a contradiction is physiological or technical;
  • whether the remaining valid channels make the event identifiable; and
  • whether the system must retain the alarm because uncertainty is unresolved.

These requirements should be tested explicitly. “Multimodal” is a description of the inputs, not evidence of robustness.

Alarm correctness and clinical actionability must not be conflated

The VTaC reference task asks whether an alarmed VT rhythm was true or false under a waveform definition. It follows the PhysioNet Challenge convention of five or more consecutive ventricular beats at a rate above 100 beats per minute.

That is a legitimate and reproducible benchmark target. It is not a complete clinical decision.

Clinical significance may depend on duration, perfusion, symptoms, underlying disease, treatment status and other information absent from an alarm waveform window. A short self-terminating run and sustained haemodynamically unstable VT can both satisfy a rhythm definition while requiring very different responses. Conversely, an unreadable mechanical channel does not make a genuine electrical event clinically unimportant.

This leads to a necessary separation of outputs:

  1. Rhythm evidence: what electrical event is supported?
  2. Perfusion evidence: what mechanical circulation is supported during the event?
  3. Sensor evidence: which channels are valid, failed or contradictory?
  4. Decision sufficiency: is there enough evidence to support retention, prioritisation or possible suppression under a declared policy?

A technical system can assist these questions without pretending to make the entire clinical decision. Treatment remains a clinical responsibility governed by the patient's condition and the applicable care protocol.

Headline accuracy is not the operating requirement

Alarm verification is a high-consequence, asymmetric decision problem. Incorrectly retaining a false alarm adds burden. Incorrectly suppressing a true life-threatening alarm can cause far greater harm.

An aggregate accuracy or area under the receiver-operating-characteristic curve averages across decisions and thresholds. A deployed system operates at a specific threshold, with a particular case mix, alarm prevalence and cost of error.

At minimum, evaluation should report:

  • True-alarm sensitivity: the proportion of adjudicated true alarms that the system retains.
  • False-alarm suppression rate: the proportion of adjudicated false alarms that it suppresses.
  • Positive predictive value after verification: the proportion of retained alarms that are true in the evaluated setting.
  • Coverage: the proportion of alarms for which the system issues a determinate verification result rather than abstaining.
  • False alarms per patient-hour or bed-day: the operational burden remaining after verification.
  • Decision latency: the time required to reach the result relative to the clinical response window.
  • Uncertainty intervals: the statistical precision of every headline operating metric.

Positive predictive value depends on prevalence. A system tested on a balanced research set can have the same sensitivity and specificity yet produce a different proportion of correct retained alarms in an ICU where true VT is much rarer. Performance must therefore be reported on a population and sampling process that match the intended use or be recalculated under clearly stated assumptions.

The evaluation unit also matters. VTaC consists of alarms already triggered by bedside monitors. Success on that dataset supports a claim about verification of triggered alarms. It does not establish sensitivity for finding all VT episodes in continuous ICU monitoring, because unalarmed episodes are outside the benchmark.

Similarly, a random alarm-level split can leak patient, hospital or device characteristics across development and test data. When multiple alarms come from the same patient record, those events are not independent examples of deployment to a new patient. Patient-disjoint evaluation is the minimum credible boundary. Hospital-, manufacturer- and time-disjoint tests answer additional questions that a random split cannot.

A defensible validation programme

No single retrospective score is sufficient for a monitor-level safety claim. Validation should advance through stages with different purposes.

1. Reproducible retrospective benchmarking

The first stage should use a fixed, documented public split where available and prevent patient overlap between development and evaluation. Results should include the selected operating threshold, confidence intervals, coverage and full error counts—not only AUROC.

VTaC is particularly useful because it provides multiple hospitals, manufacturers, ECG leads and pulsatile channels. Its limitations must remain attached to every claim: the released dataset is centred on triggered VT alarms, excludes unresolved events from the final labelled set and does not contain the detailed clinical information needed to establish actionability or patient outcome.

2. Deliberate failure-condition testing

The system should be tested with conditions that matter operationally:

  • missing PPG;
  • missing arterial pressure;
  • both mechanical channels absent;
  • paced and wide-complex rhythms;
  • low-perfusion states;
  • motion and electrode artefact;
  • unreadable or conflicting channels;
  • differences in monitor manufacturer and signal configuration; and
  • repeated alarms from the same patient.

These are not optional subgroup analyses after the headline result. They determine whether the system fails safely.

3. External and temporal evaluation

A hospital held out from all model development tests institutional transfer. A monitor manufacturer held out from development tests dependence on device-specific signal processing. A later time period tests whether changes in practice, hardware or population degrade performance.

Results should be published at the same prespecified operating threshold where technically possible. Retuning the threshold separately on every test site changes the claim from transfer to local recalibration.

4. Prospective silent-mode validation

Before influencing alarm delivery, the system should run prospectively without changing clinical care. Its outputs can then be compared with blinded expert adjudication, actual channel availability, alarm burden and workflow.

Silent-mode validation tests real-time data acquisition, timing, missingness and implementation behaviour that retrospective waveform files cannot reproduce. It should also measure how often the system returns “insufficient evidence,” because abstention observed only after deployment is not a validated safety behaviour.

5. Controlled clinical and human-factors evaluation

Only a subsequent study can determine whether changing alarm presentation reduces workload without delaying or preventing response to genuine events. Relevant outcomes include retained true alarms, false alarms delivered to staff, response times, missed or delayed events, clinician overrides and unintended changes in alarm-management behaviour.

The hardware and software context matters here. Alarm systems in medical electrical equipment are covered by the FDA-recognised IEC 60601-1-8 consensus standard, which addresses the safety and essential performance of alarm systems and signals. Meeting an algorithmic benchmark does not replace device-level risk management, verification, validation or human-factors work. US FDA, Recognized Consensus Standard IEC 60601-1-8

Safe uncertainty is a required output

A forced binary answer makes every alarm look resolvable. Real monitoring data is not.

The system may lack a usable pulsatile channel. ECG leads may disagree while several are contaminated. Timing may be unreliable. The waveform may fall outside the states represented in validation. The available measurements may establish a rapid electrical rhythm without resolving whether effective circulation persists.

In such cases, “cannot determine” is not a model failure to be hidden from the performance table. It is an operational outcome whose frequency, causes and consequences must be measured.

A professional output should make clear:

  • which alarm hypothesis was evaluated;
  • which channels supported it;
  • which channels contradicted it;
  • which channels were unavailable or invalid;
  • whether the result concerns electrical rhythm, perfusion or both;
  • whether the evidence was sufficient for the configured action; and
  • why the original alarm was retained when uncertainty remained.

This level of explanation does not require exposing proprietary implementation. It requires the system to communicate the evidential status of its decision in terms that can be audited.

What the public evidence proves—and what it does not

The public literature establishes several defensible conclusions.

False and non-actionable arrhythmia alarms can create substantial ICU burden. VT alarms are difficult because the system must preserve very high sensitivity while improving positive predictive value. Multiple ECG leads and pulsatile waveforms can provide evidence that is not present in one ECG channel. Cross-hospital and cross-manufacturer data is available for retrospective development and comparison.

The literature does not establish that every false alarm can be removed safely. It does not show that a particular retrospective model will transfer to a new monitor fleet or patient population. It does not turn PPG or arterial pressure into infallible reference signals. It does not prove that fewer alarms improve patient outcomes. It also does not establish the performance of any Mondren system.

The 2015 PhysioNet/Computing in Cardiology Challenge was explicit about the correct objective: reduce false arrhythmia alarms with minimal or no effect on true alarms. Its dataset and subsequent algorithms advanced the field, but the challenge data was smaller and less diverse than VTaC. The existence of a stronger benchmark improves the feasibility of evaluation; it does not remove the need for prospective proof. PhysioNet/Computing in Cardiology Challenge (2015)

The Mondren perspective

At Mondren, cardiac monitoring is a frontier research area. Our interest is not in producing another alarm classifier whose main result is a higher average score.

The public engineering position is more specific:

A monitor alarm should be evaluated as a physical claim assembled from partial, fallible measurements. Electrical rhythm, mechanical circulation and sensor integrity should remain distinguishable, and the system should retain the alarm when those states cannot be resolved safely.

This article reports no Mondren benchmark, clinical validation, regulatory status or deployment-readiness result. Every numerical result cited here comes from external public research.

Any future performance claim from us should state the dataset, hospitals, monitor manufacturers, patient and event counts, split boundaries, alarm definition, operating threshold, sensitivity, false-alarm suppression, abstention rate and relevant failure-condition results. A prospective claim should additionally state whether the system operated silently or affected care, and which clinical or workflow outcomes were measured.

The problem, evidence requirements and validation principles can be public. Our model construction, internal state representation, signal-quality calculations, cross-channel logic, timing rules, parameter values and implementation remain proprietary.

The commercial implication

The commercial value of alarm verification is not an abstract improvement in classification. It is a measurable reduction in avoidable alarm burden without a clinically unacceptable loss or delay of genuine alarms.

For a hospital, the evidence package should address alarms delivered per bed-day, staff response, override behaviour, unresolved cases and performance across the institution's actual devices and units. For a monitor manufacturer, it should address integration, real-time latency, missing channels, device-specific processing, fail-safe behaviour and the regulatory evidence required for the intended use.

The safest product boundary may initially be a verification and prioritisation layer rather than autonomous suppression. That boundary is not a rhetorical precaution; it determines the hazards, validation endpoints, user interface and regulatory claims that must be addressed.

A useful system would give clinicians fewer unsupported alarms and more information about the events that remain. A trustworthy system would also state when it lacks enough evidence to do that.

The distinction is fundamental: a detector can recognise a pattern in a waveform. A clinical monitoring system must establish what that pattern means across electrical activity, mechanical consequence, sensor integrity and uncertainty.

This article discusses clinical-monitoring research and engineering evaluation. It is not medical advice and does not make a claim of clinical safety, efficacy, regulatory clearance or patient benefit.