13 min read

Telecom Circuit Uptime: What SLA Metrics to Verify on Every WAN Link

What the availability number on a telecom SLA really commits to: uptime, MTTR, latency, jitter, packet loss, measurement points, and the credit mechanics to verify on every WAN link.

ByAndré Ribeiro· Founder, Obelinf
Telecom Circuit Uptime: What SLA Metrics to Verify on Every WAN Link
Telecom Circuit Uptime: What SLA Metrics to Verify on Every WAN Link · August 20, 2026
On this page

Every business circuit is sold with the same reassuring number. The quote says 99.99 percent availability, procurement files it away, and nobody reads the contract again until the night the link drops. When that happens, the discovery begins: the availability percentage is not a promise about your experience, it is a carefully defined claim about a specific measurement, taken at a specific point, over a specific period, excluding a specific list of events. The value of an SLA is decided almost entirely in those definitions, and most teams never look at them until they are already arguing with the carrier who holds the outage records.

This guide covers the SLA metrics to verify on every WAN link, whether the service is MPLS, dedicated internet access, metro Ethernet, or dark fiber: what an uptime figure actually commits to, how the repair clock runs, the latency, jitter, and packet loss thresholds that degrade applications long before a circuit counts as down, the bandwidth guarantee that congestion can silently erode, and the measurement boundaries that decide who wins a dispute. By the end you will know which numbers to check with a calculator and which clauses to check with a lawyer.

What an Availability Percentage Commits To

Your circuit quote leads with a percentage, usually 99.9 or 99.99, and that number is where most SLA conversations begin and end. Availability is best understood as a downtime budget disguised as a compliment. On a 30 day month, 99.9 percent leaves about 43 minutes of allowed unavailability and 99.99 percent leaves just over 4 minutes; spread across a year the same two figures allow 8.8 hours and 52.6 minutes respectively. That is a wide enough spread that the first question to ask any carrier is which tier you are actually buying, because the sales deck and the contract do not always agree.

Yearly downtime budget by availability tier, from 99 percent to 99.999 percent Five nines Four nines Three nines Two nines 99.999% = 5.3 min a year 99.99% = 52.6 min a year 99.9% = 8.8 hours a year 99% = 87.7 hours a year Each extra nine cuts the yearly downtime budget by roughly ten times. Typical contracts commit to 99.9 or 99.99 percent, not five nines.

The budget only means something if you agree on what consumes it. Most contracts define unavailability as a complete loss of service at the demarcation point, which sounds precise and is actually narrow. A link that is up but degraded, congested, or dropping a meaningful percentage of packets does not breach the availability commitment in most contracts, because availability in this sense is a connectivity measurement, not a performance measurement. Latency, jitter, and packet loss live in separate clauses with their own thresholds, which is why a circuit can post a perfect 99.99 percent month while the applications running over it feel noticeably worse.

The exclusions list quietly shrinks the number further. Scheduled maintenance windows with proper notice are usually removed from the calculation entirely, as are force majeure events, power failures at your site, and faults in customer owned equipment. That is reasonable in isolation, but the exclusions accumulate: a month with no outages at all still cannot reach four nines if a two hour maintenance window was subtracted from the denominator. Ask the carrier to show you how the accounting works before you treat the headline number as a floor.

MTTR and MTBF: The Repair Clock and the Relapse Rate

Availability tells you how much downtime is allowed; MTTR, mean time to restore, tells you how long a single incident can run. The clause is usually written as a maximum, four hours for a metro fiber circuit being typical and eight hours for less critical links, with response commitments layered on top, such as acknowledging a ticket within 30 minutes and dispatching a technician by a set deadline. The fine print that matters more than the hours is when the clock starts and whether it runs continuously.

Carriers frame that clock differently. The version most favorable to them starts at dispatch, when a technician accepts the case, which ignores everything between your call and that moment. The version most favorable to you starts at detection, in the carrier’s network management system or at your own ticket open time, whichever is earlier, and it is a clause worth negotiating on any circuit that carries revenue. Equally important is the distinction between business hours and 24/7 restoration. A four hour MTTR with the clock running nine to five can mean a circuit stays down overnight and through the next morning before the carrier’s obligation resets, which is why the contract must state the schedule explicitly.

MTBF, mean time between failures, is the quieter of the two metrics and often the more telling. It describes how often outages happen rather than how long they last, and together the two numbers describe the shape of your risk. A circuit with one 43 minute outage has the same monthly availability as a circuit with ten four minute outages, but completely different operational meaning for a site running time sensitive applications. Ask the carrier for historical outage frequency on the specific path, not just the product brochure, because MTBF is where the real world shows up: the difference between a rural last mile that fails quarterly and a well maintained fiber path that has not blinked in years.

Latency, Jitter, and Packet Loss: The Degradation Trio

When a link stops working you notice immediately; when it degrades you notice slowly, through support tickets about slow pages, poor call quality, and unexplained application timeouts. The three numbers that govern that degradation are latency, jitter, and packet loss, and each has a specific contract question attached. For latency, ask how it is measured and between which points. A commitment of 15 milliseconds to the carrier’s nearest point of presence is not the same as a commitment of 15 milliseconds end to end between your sites, and the headline number on the datasheet is almost always the shorter of the two.

Jitter is the variation in latency between consecutive packets, and it is the metric voice and video feel first. A few milliseconds of variation is normal on a healthy carrier network, commitments of 5 milliseconds or less are common on premium products, and once the variation climbs past roughly 10 to 15 milliseconds, voice quality visibly suffers even when average latency looks fine. Packet loss is the third leg, typically committed at 0.1 percent or better on business circuits, and its real world impact depends on how the metric is sampled more than on the value itself, because a monthly average smooths away the short bursts that actually break sessions.

Five short outages inside one month that the monthly availability average hides Carrier monthly report 99.99% available Your link logs 5 separate drops 12 s 18 s 14 s 9 s 25 s Five outages totaling 78 seconds fit inside a 99.99 percent monthly budget, yet each one kills an active session. The monthly average stays green while the applications see five disconnects. Verify per interval, not per month.

The same logic applies to every averaging scheme in the contract. A committed loss of 0.1 percent can be met with a flawless month punctuated by a lost burst every time a session needs to cross the network, and a latency commitment framed as a monthly mean can be met while individual minutes are far outside budget. Ask whether each metric is evaluated as a monthly average or in shorter intervals, five minute buckets are the practical standard, and whether percentile reporting such as p95 or p99 is available. An SLA you can only read as a monthly average is an SLA you cannot verify weekly, and verifiability is the entire point.

The Bandwidth Promise: CIR and the Silent Brownout

On Ethernet and carrier Ethernet services, the port speed and the committed information rate are different numbers, and the gap between them is where a carrier can meet every availability target while still shortchanging you. A circuit with a gigabit port and a 100 megabit CIR is fully available every day of the month while delivering nothing above the committed rate, and congestion on the carrier’s network silently turns that gap into reality. If the contract contains no throughput commitment at all, a permanently congested handoff can satisfy every other clause while your users feel the squeeze.

Verify the bandwidth commitment at installation and again at every renewal. Service activation tests based on ITU Y.1564 exercise the circuit at its full committed rate and measure throughput, frame loss, and delay in both directions, which is the same methodology the carrier’s own installers use, so asking for the results is a fair request rather than an adversarial one. A circuit that cannot pass its own turn up test at the committed rate is a circuit you should document now, while the goodwill around the sale still exists, because the same clause that let it underperform at install will cover it for the life of the contract.

Ask explicitly whether the SLA governs throughput at all. Many product SLAs cover connectivity only, with no commitment about how much traffic the link can carry, which makes your own utilization records the only real evidence you can lean on from the first billing month. Capture them consistently, because at renewal time a circuit running against its CIR supports a request for a higher committed rate or a discounted price, while a claim made without utilization data reads as anecdote no matter how correct it is.

Where the SLA Is Measured Decides Who Wins

Circuit path from customer edge to core backbone, with your measurement point and the carrier's test point marked your probe sees the whole path carrier tests stop at the POP Customer edge Demarc / NID Local loop Carrier POP Core backbone where you measure the ownership line fails most often where carrier tests shared by many Your scope ends at the demarc Carrier scope: everything behind the handoff A carrier that measures at its own edge never counts local loop outages, the segment that fails most often, against your availability.

The single most consequential location in an SLA is not a percentage, it is a coordinate. Everything the carrier commits is measured at a specific point on the circuit path, and the choice of that point determines whether a real world outage counts against them or not. The typical path runs from your customer edge router to the demarcation point, through the local loop, into the carrier’s point of presence, and across its core backbone, and the measurement point a carrier prefers is the boundary of its own network, where the telemetry is cheap and the coverage is complete.

If availability is verified at the carrier’s point of presence, the local loop between the demarc and the POP is invisible to the measurement, which matters because the local loop is the segment that fails most often: the fiber that gets dug up, the copper that corrodes, the aerial plant that a truck takes down. A carrier measuring at its own edge can post a clean availability number while your site sits through a multi hour local loop outage, because the event never registers in its measurement. Insist that availability be verified at your customer edge, or at minimum that both sides keep independent records and reconcile them monthly.

The same scrutiny applies to who declares an outage and when the clock starts. Some contracts count only outages the carrier detects or acknowledges through its network management system, which leaves a silent failure, a link that is down on your side but invisible to the carrier’s monitoring, uncounted even though your users experienced it. Ask how the carrier detects failures, whether proving a fault requires a truck roll, and whether raw performance reports are available on request. A carrier that streams raw reports is a carrier whose SLA you can audit; a carrier that offers only monthly summaries is a carrier you should audit harder.

Service Credits: What a Breach Actually Pays

The remedy for a missed SLA is a service credit applied to your bill, and reading the credit mechanics is where most of the surprise lives. The typical formula credits a percentage of the monthly recurring charge per outage hour, with steeper tiers for failures that blow past the MTTR commitment, but three details decide what the clause is worth in practice: the de minimis threshold, commonly excluding incidents shorter than 15 to 30 minutes from any credit, the cap, usually between 30 and 100 percent of one month’s recurring charge, and the claim window, often 30 to 90 days from the event.

The de minimis threshold deserves particular attention because it changes the meaning of availability. A circuit with a monthly budget of four minutes can pass through ten brief two minute flaps, exceed its budget fourfold, and still produce zero credit under a contract that excludes incidents under 15 minutes, because the threshold swallows every individual event. The headline percentage and the credit outcome are only connected when you read the incident boundary, and on circuits running applications sensitive to short outages, a lower threshold or per event credits are worth negotiating before you sign rather than after the first bad month.

Credits are also rarely automatic. Most carriers require a claim with evidence of the outage start and end times, which means your outage log, not the carrier’s, is what stands between you and the money. Industry estimates have long suggested that a majority of eligible SLA credits go unclaimed because teams lack the records or miss the filing window, and the difference between collecting and not collecting is almost always documentation discipline rather than carrier goodwill.

Build the Baseline Before You Need It

None of these clauses can be verified retroactively. The work happens when the circuit is ordered, when the tests are run, and when the commitments are written into the record your team will consult on the day of the incident. At turn up, run the activation test and keep the result. At contract signing, write down what matters: the availability tier, the MTTR schedule, the latency, jitter, and loss targets, the CIR, the measurement point, the threshold, the cap, and the claim window. Across the life of the circuit, maintain a continuous outage log with timestamps, so the next quarterly review starts from your evidence rather than the carrier’s summary.

Review every circuit against the carrier’s reports on a regular cadence, quarterly is a reasonable rhythm, and treat a gap between what you measured and what the carrier reported as the beginning of a claim rather than the end of a discussion. Circuits that meet their metrics are negotiation material too, because a path with years of clean records supports asking for a price reduction at renewal, while a path with undisputed breaches supports asking for credits before renewal.

The discipline is an inventory problem as much as a monitoring problem, because a claim needs the commitment sitting next to the incident, per circuit, on record. On a multi site WAN that means recording SLA commitments, circuit identifiers, and outage history beside each circuit and link in the topology, in the same place the rest of the circuit data lives, rather than in another spreadsheet nobody reconciles. With the commitments stored on the circuit record, a quarterly compliance review becomes a reading of the page instead of a reconstruction of the contract, and the credits you collect are simply the compensation for keeping that record straight.

Frequently Asked Questions

How many minutes of downtime is 99.99 percent uptime?
On a 30 day month, 99.99 percent availability allows about 4.4 minutes of downtime, and over a full year the same tier allows 52.6 minutes. The next step down, 99.9 percent, allows 43 minutes per month and 8.8 hours per year, which is why the difference between three nines and four nines matters more than it sounds.
What is the difference between MTTR and MTBF in a circuit SLA?
MTTR, mean time to restore, is how long the carrier takes to fix an outage once the repair clock starts, while MTBF, mean time between failures, is how often outages occur in the first place. The two describe different risks: MTTR tells you how long a single incident can hurt, MTBF tells you how often you are exposed to one.
What packet loss percentage is acceptable on a WAN link?
Business circuit SLAs typically commit to packet loss at or below 0.1 percent, with premium products offering tighter targets, and real time applications like voice start to suffer visibly with sustained loss above about 1 percent. The more important question is how the metric is measured, since a monthly average can hide loss bursts, so ask for per interval reporting such as five minute buckets.
Do telecom carriers actually meet their availability SLAs?
Carriers usually meet headline availability because the definition is narrower than customers assume: it covers loss of connectivity at a specific measurement point, excludes scheduled maintenance and customer side faults, and often ignores short incidents entirely. Verifying against your own monitoring and comparing it with the commitments you recorded when the circuit was ordered, which a circuit inventory like Obelinf stores as structured fields on each link, is the only way to know what the month actually contained.
How are SLA credits calculated for a circuit outage?
Most contracts credit a percentage of the monthly recurring charge for each outage hour, subject to a de minimis exclusion for short incidents, capped at a fraction of one month's bill, and claimable within a deadline of generally 30 to 90 days. The practical amount is modest by design, so the credit is compensation rather than punishment, and it nearly always requires you to file a claim with evidence of the outage times, which is where an Obelinf circuit record showing the SLA parameters and outage history pays for itself.
Is 99.99 percent uptime good for a business internet circuit?
Four nines is a strong commitment for a dedicated internet access or MPLS circuit, and many standard products officially promise no more than 99.9 percent. What matters as much is the rest of the contract: the measurement point, the MTTR schedule, and the incident exclusion threshold, because a four nines headline combined with a 30 minute exclusion can deliver noticeably less than it promises, so read the fine print before treating the percentage as the whole story.

Stop reaching for a spreadsheet

Obelinf keeps every subnet, device, circuit, and rack in one live source of truth, with audit logs and a topology view. Free for personal use.

Related Articles