Migration · cutover
Replacing an incumbent BNG

Everyone agrees the box is better.
Nobody tells you how to get there.

A phased route off a chassis BNG: what moves first, what cannot be made seamless and why, how the RADIUS you already run stays the contract throughout, and what you can put back at every single step.

Replacing a BNG is not a forwarding problem. Forwarding is the easy part. The hard part is that your subscribers are already connected to the box you are replacing.

80
gateways running this way in production
302,761
live subscriber sessions on them
0
flag days required
1
RADIUS — the one you already own

The objection every vendor skips

Read any BNG comparison and it argues the destination: throughput, cost per subscriber, latency under load, what a chassis costs to grow. All of that can be true and still leave you exactly where you started, because the question that actually stops the project is a different one.

How do I move a live subscriber base off the thing that is currently carrying it? That question has a real answer, and it is not "carefully, over a weekend".

First, the part that cannot be made seamless

We would rather you hear this from us than discover it at 2am.

You cannot hand a live PPPoE session from one vendor's BNG to another's. There is no protocol for it, in either direction, between any two vendors. The session state — the PPP negotiation, the address, the accounting record — belongs to the box that established it. Any vendor who tells you their migration is invisible to subscribers is describing something else, or is wrong.

So subscribers reconnect. That is not a BNGSOFT limitation; it is the shape of the problem. The only thing under your control is how many of them reconnect at the same moment.

Which turns out to be the whole game. A reconnect is a few seconds for one subscriber. The same reconnect, arriving simultaneously from an entire subscriber base, is a thundering herd against your RADIUS and your address pools — and that is what people actually mean when they say a migration "went badly".

Reconnect load: one flag day versus a phased move A single cutover forces the entire subscriber base to re-authenticate at once, producing one very tall spike of RADIUS load. Moving a VLAN or a POP at a time produces a series of small, individually survivable steps spread over weeks. One flag day everyone one night · one spike · nothing to roll back to A VLAN at a time weeks · each step survivable · each step reversible The reconnect is unavoidable. Its concentration is a choice. Both routes end with the same subscribers on the same new gateway. Only one of them has a bad night in it, and only one of them lets you stop after the first two thousand and decide whether to continue.

Four phases, and you can stop after any of them

The route below is deliberately unheroic. Each phase leaves the network in a state you could stay in indefinitely, which is what makes the next one safe to attempt.

  1. Stand it up beside the incumbent — carrying no subscribers

    The new gateway joins the network in Border Mode: forwarding transit with the protection stack active, terminating nobody. It sees real traffic at real rates, and your existing BNG is untouched. If you learn something unwelcome about your NICs, your peering or your rack, you learn it here, where the blast radius is zero.

  2. Move one VLAN — the one you would least mind explaining

    Point a single access VLAN at the new box. A few hundred to a few thousand subscribers re-authenticate against the same RADIUS, receive addresses from the same pools, and are subject to the same per-plan policy. Now the migration is no longer a theory: it is a comparison, running side by side, with both estates visible in one console.

  3. Move the bulk — at whatever rate the evidence supports

    VLAN by VLAN, POP by POP, one operator's subscribers at a time. Nothing about this phase is novel once phase 2 has run; it is the same operation repeated, and the pace is set by what you saw rather than by a maintenance window that was booked before anyone knew.

  4. Retire the chassis last, and only when it is carrying nothing

    The incumbent is drained rather than switched off — its subscriber count reaches zero because they left, not because it failed. That is a decommission you can schedule in daylight.

Every phase is reversible by pointing the VLAN back. That is not a feature we built; it is a consequence of never having both estates depend on each other. The new gateway does not need the old one, the old one does not know about the new one, and your RADIUS is talking to whichever of them currently holds the subscriber.

Your RADIUS is the contract — and it does not change

This is the part that makes phased migration possible at all. A subscriber's plan, speed, address assignment, policy and accounting are defined by attributes your AAA already returns. They are not defined by the chassis.

What people fear

A new BNG means a new AAA integration, a new attribute dictionary, a re-write of provisioning, and a period where the billing system and the network disagree about who is connected.

That fear is well earned. It is what happens when a box expects to own subscriber policy rather than enforce it.

What actually happens

The same RADIUS, the same attributes, the same CoA and Disconnect-Message flows, the same accounting records — now arriving from a different NAS. Provisioning does not learn a new vocabulary.

The control plane authorises; the data plane enforces. Moving the data plane does not move the contract.

Where an incumbent has been doing something genuinely proprietary — a vendor-specific attribute that encodes policy no standard expresses — that is the work item, and it is discovered in phase 2 with a few thousand subscribers, not in phase 4 with all of them.

Running two estates at once, without pretending you aren't

For the duration of the migration you are operating a hybrid network, and the honest problem with a hybrid network is that it is very easy to lose track of what is in it. A gateway that was drained last month and a gateway that failed this morning both look "down" to anything counting availability.

NOC2 Fleet Inventory separating real outages from drained gateways, gateways never in service, and gateways with no data centre recorded.
The console distinguishes the two. A gateway that stopped reporting while subscribers were connected to it is a real outage. A gateway whose traffic was moved off first is drained — which is exactly what a planned decommission looks like, and is not a fault. During a migration that distinction is the difference between a meaningful availability figure and a meaningless one. The panel also surfaces gateways that were enrolled and alerting but never carried a single session — the classic residue of a migration that was started and abandoned.

Turn the extra capabilities on afterwards, not during

There is a temptation, having installed something more capable, to switch everything on at once. Resist it. Move the subscribers first with the new gateway doing exactly what the old one did; enable the things the old one could not do once the traffic is boring.

NOC2 Gateway Capabilities page: select gateways, then choose which capabilities to enable on them.
Capabilities are enabled per gateway, not per fleet. Pick the gateways, pick the functions, and the change is applied to those alone — so a new capability can be proven on the two gateways you migrated first before it reaches the other seventy-eight. Each entry states what it does and what to watch, because a capability switched on without a reason to switch it on is just a new variable in an incident.

What the migration actually costs you

ConcernThe honest answer
Subscriber downtimeOne reconnect per subscriber, once, at the moment their VLAN moves. Seconds, not minutes — and scheduled by you, VLAN by VLAN.
RADIUS loadA burst proportional to the VLAN you move, not to the subscriber base. This is the entire reason for phasing.
Provisioning changesNone for standard attributes. Vendor-specific attributes that encode proprietary policy are the real work, and they surface in phase 2.
Address poolsUnchanged. The new gateway draws from the same pools; subscribers keep receiving addresses from the ranges you already announce.
RollbackPoint the VLAN back. The incumbent has not been modified and has not lost the ability to serve those subscribers.
The chassis itselfStays racked and powered until it is carrying nothing. Its retirement is a separate, unhurried decision.

Why this is worth doing at all

Cost

Growth stops being a purchase order

Adding subscribers to a chassis eventually means a line card, a chassis, or a forklift. Adding them here means a commodity server — and the sizing is measured rather than quoted.

Capability

Functions the incumbent never had

CGNAT, low-latency queueing, per-subscriber protection and abuse containment run in the same data path, on the same box, rather than as three more appliances to buy and cable.

Operations

You can see inside it

The running state of every gateway — not the config you believe is loaded — is visible in one console, with the daemon's own account of which features are actually active.

The bottom line

The migration you are being sold elsewhere is a single night with a rollback plan nobody has tested. The one described here is a sequence of small, individually boring steps, each of which can be stopped or reversed, and the first of which carries no subscribers at all.

Eighty gateways run this way in production today, carrying just over three hundred thousand live subscriber sessions between them. None of them arrived there on a flag day.

Start with phase 1. Put the box in beside what you have, carrying transit and terminating nobody, and let it run for a fortnight. It costs you a rack unit and it answers more questions about your network than any evaluation document will.

About the screenshots. Every screen shown is a real production console. Gateway hostnames, operator names and site names are replaced with placeholders, and every address is replaced with one from the RFC 5737 documentation range, before the image is taken; no operator, site, region or subscriber identifier appears in any image.