Killing the Alarms

Moving from reactive noise to autonomous action.

Your best engineers are working as network janitors. They watch dashboards, correlate alarms by hand, and clean up configuration errors a machine should never have allowed. None of that is the work you hired them for, and none of it gets easier as the network grows.

Why doesn’t hiring solve this?

Because the economics underneath have stopped working. Bandwidth demand rises on an exponential curve. Revenue per bit does not. Every operator therefore faces the same question: how do you cover a materially larger network without a proportionally larger operations team?

The conventional answers have run out. You can buy another monitoring tool, but tool count is already part of the problem. You can hire, except the workforce that understands both legacy transport and modern fibre plant is retiring faster than it is being replaced, and a new hire spends most of a year getting useful. You can accept longer restoration times, and that shows up in churn.

The remaining answer is to change what a human is for.

What is actually wrong with forty-nine alarms?

Nothing, individually. Each one is accurate. An OLT line card fails, every ONT behind it loses GPON signal, and each reports independently. The network is telling the truth forty-nine times.

The problem is that answering the only question that matters — how many customers are down and which ones — requires a human to read forty-nine rows, recognise they describe one event, find the parent cause, and work out who sits behind it. Under a manual model that is thirty to sixty minutes before anyone knows the customer impact. Not because the engineer is slow, but because the work is serialised through a person who has to reconstruct topology from memory.

Every layer of the operation repeats the pattern. Alarm volume grows with the network. Support tickets grow with the subscriber base. Integration work grows with every vendor added. Headcount is the only lever most operators have against any of it.

Why can’t your data lake do this?

Most operators are drowning in telemetry and starving for action. The usual architecture routes everything into a single store — which is, in practice, where telemetry goes to die. Query latency measured in seconds to minutes is fine for a quarterly capacity review and useless for an event in progress.

The two workloads have incompatible requirements. Correlating an event as it unfolds needs sub-second reads against current state. Capacity modelling needs deep, cheap scans across months. Force one store to serve both and the urgent job is the one that loses.

Separating them is the foundation everything else rests on: a hot path holding current state where correlation runs while the event is still happening, and a cold path holding history for trending, capacity work, and audit — depth that never competes with real-time processing for the same resources. The platform page covers how that is built.

What makes correlation actually work?

Topology, not timing windows and not alarm-text pattern matching. Forty-nine alarms collapse into one service event because the system knows those ONTs sit behind that card — and knows which subscribers, on which service tiers, with which SLA commitments, sit behind them.

That has an honest consequence, and it is the first thing a serious evaluator should test: correlation quality is bounded by topology quality. If your inventory is wrong, root-cause identification will be wrong in exactly the same places. Operators with poor topology data should expect a reconciliation phase and should budget for it. Any vendor telling you otherwise is selling the demo rather than the deployment.

What does an agent do that automation doesn’t?

Automation executes a rule somebody wrote in advance. It handles the case it was written for and nothing else, and it needs maintaining every time the network changes.

An agent owns a job. It has a defined mandate, acts inside that mandate without being asked, and hands off with context attached when the work needs a human. Support tickets get answered from documented procedure rather than queued. Customers get notified from the correlated event rather than after someone remembers. The supervisor picture stays current instead of being reconstructed at shift change. The six agents are here.

The distinction that matters commercially is grounding. Every agent decision anchors to retrieved documentation with a confidence score attached, and you set the thresholds that map confidence to action. If the answer is not in your documentation, no agent invents one — it escalates. The failure mode of AI in operations is confident, wrong, and fast; addressing it structurally rather than with a better prompt is the difference between a system you can deploy and a demo.

What happens when the platform itself is under load?

An assurance platform that degrades during a mass-outage event is worse than no platform, because by then the operations team has reorganised around it.

Two mechanisms answer that. Worker capacity scales against real-time demand rather than sitting provisioned for the average, so a storm gets processing capacity rather than a queue. And the architecture runs Active-Active across datacentres, so losing one does not create a surveillance gap. It is designed to a 99.999% availability target — stated as a target, because availability achieved in any given deployment depends on the environment it lands in.

Where does this argument break down?

In four places, and the paper says so at length rather than hoping you do not notice.

Autonomy has a blast radius — an agent that can execute a change can execute a bad one, which is a legitimate reason to phase adoption slowly. The labour argument gets oversold; reassigning operations staff is a multi-year organisational change, and any model showing headcount savings arriving on the same schedule as the software is a model to distrust, including one presented by a vendor. Topology quality bounds everything, as above. And Rapax is a young platform from a small company — the counterweight is a team that built and sold assurance software before, but that is track record rather than scale.

Get the white paper

The full paper adds the architecture in detail — the hot and cold path design, deployment and on-premise options, the resilience model — along with the complete counterpoint section and three things worth doing regardless of which vendor you choose.

ContactUs KillingAlarmsWP

About Rapax

Rapax is an AI-native network service assurance and automation platform for telecom operators, bringing fault, performance, topology, and service management into one system worked by six AI agents. It deploys on Kubernetes in cloud or on-premise environments, including local model inference. Rapax is a business unit of Citus Technologies, LLC.

Shawn Ennis is the Founder & CEO of Citus Technologies and the founder of Rapax. He spent 25 years in telecom operations, holds 12 patents in network management and service assurance, and previously founded Assure1 — acquired by Oracle in 2021.

Not ready to download? Book fifteen minutes — no prep, no deck. cal.com/shawn-ennis · sales@rapax.app · More white papers