BNGsoft/ OrionOS · bngxdpd · dtvbras/ Carrier BNG Platform

You are not buying throughput. You are buying what your subscribers feel.

Every BNG vendor will quote you a price per gigabit. It is the easiest number to compare and the least useful one to decide on — because the box is almost never what limits your network, and the costs that actually hurt never appear on the quote.

How to read this document. Every figure attributed to BNGsoft is measured on live production and lab systems, and tagged measured. Anything shown for illustration is tagged illustrative and carries no measurement claim. Competitor columns describe publicly documented architecture — confirm them with the vendor before making a decision. We would rather you verify than take our word.
5.4%
CPU utilisation on our busiest production BNG. The forwarding engine itself: 4.0%.
0.83s
Control-plane gap during an in-service software upgrade, down from 13.3 s.
1604
Subscribers held through a live datapath replacement — no detach, no restart, no redial.
67.1%
Of downstream bytes are QUIC — invisible to every TCP-based quality metric.
The price trap

The number on the quote is the smallest number in the decision

A BNG licence is a visible, annual, negotiable line item. It is also, for most operators, a minority of what the platform actually costs over five years. The rest is spent in the network operations centre, in the field, and in churn — and it is spent as a direct consequence of what the platform can and cannot do.

Consider what a single unresolved subscriber complaint costs you. The subscriber reports that streaming is bad in the evening. Your BNG reports the line is up, the session is authenticated, and the byte counters are climbing. Everything looks healthy. So you send a technician. The technician finds nothing wrong, because there is nothing wrong with the line — the loss is happening somewhere past your edge. You have paid for a truck roll, and the subscriber is still unhappy.

Multiply that by the number of complaints you cannot resolve from the office. That is the real bill, and no line on any quote shows it.

Where the money goes over a platform lifetime Structure only — no cost figures claimed
WHAT YOU COMPARE ON THE QUOTE Licence & hardware visible · annual · negotiable WHAT YOU ACTUALLY PAY Truck rolls for faults that were never on the line Churn from experience you could not measure or defend Night maintenance windows, and the staff who work them Support escalation, SLA credits, reputation
illustrative  The proportions here are not a cost model and are not measured — the diagram makes a structural point only: the item you compare vendors on is the one you have the most visibility into and the least leverage over.
Architecture

Three ways to build a BNG, and what each one costs you later

The difference between these platforms is not a feature list. It is where the packets are processed — and that single decision determines what you can see, what you can upgrade, and what you can integrate with for the rest of the platform's life.

Datapath placement across the three approaches Architecture
BNGsoft — XDP in kernel DPDK vBNG (e.g. NetElastic) Router OS (e.g. MikroTik) NIC (ice / bond) XDP datapath QoS · CGNAT · AQM · telemetry at the driver, before skb alloc Linux kernel stack bonds · VLAN · FRR · netlink dtvbras control plane PPPoE · IPoE · RADIUS CONSEQUENCE Line rate and the whole kernel ecosystem. Standard tooling still works. NIC bound to userspace PMD DPDK forwarding plane dedicated poll-mode cores spinning at 100% by design Kernel stack — bypassed parallel config & tooling world Control plane (often CUPS) separate node to size & operate CONSEQUENCE Fast, but cores and NICs are dedicated, and your existing Linux tooling does not apply. NIC OS forwarding path general-purpose router features cost CPU directly Queues, NAT, firewall shared CPU budget PPPoE / IPoE server built for routers, not BNGs CONSEQUENCE Excellent value at small scale. Every feature you enable takes headroom.
XDP runs our forwarding logic inside the kernel at the driver, before a socket buffer is even allocated — so we get the speed without leaving the operating system behind. Your bonds, your VLANs, your FRR routing, your monitoring agents and your engineers' existing Linux knowledge all keep working.
Why this matters commercially

A DPDK platform is genuinely fast. The trade is that the network interfaces are removed from the operating system and handed to a userspace process, and cores are dedicated to spinning on them. You gain throughput and you inherit a second, parallel operational world: separate configuration, separate counters, separate troubleshooting, separate staff knowledge.

Why cheap scales badly

A general-purpose router OS is superb value for a small deployment, and we say that plainly. But on that architecture every capability you switch on — queueing, NAT, filtering, accounting — spends the same CPU budget that forwards packets. Growth and features compete with each other, and the ceiling arrives without warning.

End-user experience

Most BNGs count bytes. Ours measures whether the subscriber is actually having a good time.

This is the single largest functional gap between our platform and everything else on your shortlist, and it is the one that shows up on your support desk every single evening.

Byte counters tell you a session is passing traffic. They cannot tell you that the video is buffering, that the game is unplayable, or — most valuable of all — whose fault it is. Answering that last question is what stops a truck from being dispatched.

The QUIC blind spot Measured across the production fleet
SHARE OF DOWNSTREAM BYTES 67.1% QUIC invisible to every TCP-based quality metric 32.9% TCP — visible 3 in 4 subscribers moving real download traffic cannot be judged by TCP measurement at all. Quiet boxes are the blindest — an overnight sweep is the least trustworthy of all.
measured  If a platform's quality reporting is built on TCP retransmission and TCP round-trip time, then for two thirds of the traffic your subscribers actually care about, it is reporting on nothing. We measure QUIC handshake round-trip time directly on the wire — using only the cleartext long-header fields that RFC 9312 permits an on-path observer to read. No decryption, no interception, no proxy.
Validation: injected latency vs recovered measurement Against known ground truth
INJECTED RECOVERED BY THE DATAPATH ERROR 10 ms 10.13 ms +1.3% 50 ms 50.33 ms +0.7% 150 ms 150.16 ms +0.1%
measured  We do not ask you to trust a quality metric because it produces a plausible-looking number. Known delays were injected as real frames and recovered by the live datapath to within 1.3%. Negative controls — non-handshake and already-established traffic — moved the metric by nothing, which is how you prove a probe is measuring the thing it claims to measure rather than reporting a comfortable default.
Loss localisation: is it the line, or is it past your edge? One production fleet snapshot
442 SUBSCRIBERS SHOWING PACKET LOSS 434 — loss is happening past your network no line fault exists to find · a technician would be sent for nothing 8 — genuine line-side evidence these are the ones actually worth a truck 98.2% of loss complaints on this snapshot were not the access line — and the platform can say so, with evidence, before anyone is dispatched.
measured  This is the clearest example of why functional capability beats unit price. A platform that cannot separate a line fault from downstream congestion will send technicians to both. The difference is not a rounding error in your operations budget.
Continuity

Upgrades that do not cost you a single subscriber session

Ask any vendor what happens to live subscribers when you deploy new software. The answer tells you more about the platform than any datasheet.

On most platforms the honest answer is: the sessions drop and the customer premises equipment redials. That is why upgrades happen at 3 a.m., why they need a change window and staff to sit through it, and why operators quietly defer them — running known-vulnerable software for months because the upgrade itself is the risk.

We rebuilt our upgrade path specifically to remove that trade-off. The forwarding program is replaced in place: the attachment is preserved, the state tables are reused, and subscribers are never told anything happened.

Subscriber sessions through a software upgrade BNGsoft measured · comparison illustrative
upgrade starts +5 min 100% 0% every session dropped CPE redial storm follows 1604 subscribers held — no detach, no restart BNGsoft in-place datapath update conventional restart cycle
measured The BNGsoft line: 1604 live subscribers carried through a full datapath replacement, with the control-plane gap reduced from 13.3 s to 0.83 s. At larger scale, 2089 of 2199 sessions survived.  illustrative The comparison curve is a generic full-restart-and-redial shape, not a measurement of any named product.
Survives the unplanned, too

The same state-preservation machinery works when the software does not exit politely. After a crash, sessions are restored with identical interface names, usernames, rate limits and accounting session identifiers — so RADIUS accounting stays continuous and billing does not develop a hole. Both PPPoE and IPoE.

Maintenance without a window

Nodes run as a cluster. To work on one, you drain it: new subscribers are steered to its peers, existing ones migrate, and the node empties on its own schedule. The work happens in daylight, with the network up, because no single node has to be the one that stays alive.

Headroom on the busiest box in production Measured
CPU UTILISATION — BUSIEST PRODUCTION NODE 5.4% used the forwarding engine itself accounts for 4.0% 94.6% spare The platform is memory-bound, not compute-bound. Buying a faster CPU to fix a BNG problem is, on this architecture, solving the wrong problem.
measured  This is why comparing platforms on price-per-gigabit is the wrong axis. On our architecture the box is not the constraint and has not been for a long time — so the question worth asking is not how cheaply you can push packets, but what the platform can tell you and do for you while it pushes them.
Abuse handling

Punish the attack, not the customer

When a subscriber line is the source or target of an attack, the conventional response is blunt: rate-limit or disconnect the subscriber. The abuse stops, and so does the service of a paying customer who, in the common case of a compromised device, has done nothing wrong and cannot understand why their internet died.

Our detection identifies the specific offending flow and drops only that flow. The subscriber's other traffic is untouched and the session is never penalised. Detection arms itself automatically — the operator does not have to predict thresholds in advance.

Surgical mitigation versus blunt mitigation Thresholds measured in lab
CONVENTIONAL Attack detected on a line → the whole subscriber is limited or cut support call · angry customer · possible churn BNGSOFT Attack flow identified precisely → only that flow is dropped subscriber stays online · no penalty · no call AUTOMATIC ARMING — NO THRESHOLD GUESSWORK ~1k SYN/s rate-limit engages ~10k SYN/s flow blocked outright subscriber session unaffected throughout
measured  Thresholds validated in lab against generated attack traffic, including the negative control that proves the detector is scoring real input rather than reporting an empty result as a clean one.
Comparison

What actually differs

Compare on the rows that change your operating costs, not the row that changes your purchase order.

Capability BNGsoft DPDK vBNG Router OS
Datapath location In-kernel XDP, at the driver Userspace, kernel bypassed OS forwarding path, on CPU
Keeps standard Linux tooling Yes — bonds, FRR, netlink, your monitoring NICs leave the OS; parallel tooling Vendor OS and its own tooling
Dedicated cores required No — 5.4% on the busiest node Yes, poll-mode cores spin by design No, but features consume forwarding CPU
Upgrade without dropping sessions Yes — 1604 held, 0.83 s gap verify with vendor verify with vendor
Session restore after a crash Yes — PPPoE and IPoE, accounting preserved verify with vendor verify with vendor
QUIC-aware quality measurement Yes — on-wire handshake RTT, RFC 9312 compliant verify with vendor verify with vendor
Line fault vs downstream loss Yes — loss localisation per subscriber verify with vendor verify with vendor
Per-subscriber experience scoring Yes — and it reports “unknown” rather than inventing a score verify with vendor verify with vendor
Surgical abuse mitigation Yes — drops the flow, not the customer verify with vendor Typically per-subscriber limiting
Carrier CGNAT with port blocks Yes — deterministic blocks, NAT64, DS-Lite Generally yes Basic NAT; limited at carrier scale
Modern AQM / L4S Yes verify with vendor Classic queueing disciplines
Clustering with drain for maintenance Yes — daylight maintenance, no window verify with vendor verify with vendor

Rows marked “verify with vendor” are deliberately not filled in. We will not characterise a competitor's capability we have not tested ourselves — and any brochure that confidently marks every rival row with a cross is telling you something about the brochure, not about the rivals. Take this table to them and ask. The BNGsoft column we will demonstrate on request, on your traffic.

How to evaluate

Eight questions that separate platforms better than any price list

Ask these of us and of everyone else you are considering. They are all answerable in a live demonstration, and the answers are difficult to fake.

01
Deploy new software right now, with my sessions on the box. What happens?
The most revealing question you can ask. If the answer involves a maintenance window, you will be paying for that window several times a year, forever.
02
This subscriber says video is bad. Is it their line, or is it past your edge?
If the platform cannot answer, every such complaint is a potential truck roll. On one fleet snapshot, 98.2% of subscribers showing loss had no line fault at all.
03
Two thirds of my downstream traffic is QUIC. What can you tell me about it?
A quality metric built on TCP is blind to the majority of modern traffic. Ask specifically what is measured, and how it was validated against known ground truth.
04
Show me a subscriber the platform admits it cannot score.
A metric with no way to say “I don't know” will hand you a confident number for subscribers it never measured. We learned this the hard way and fixed it: ours reports unknown.
05
Kill the forwarding process. What happens to my subscribers and my billing?
Session restore should preserve interface names, rate limits and accounting identifiers — otherwise a crash quietly becomes a billing discrepancy.
06
A customer's device is compromised and flooding. What does the platform do to that customer?
Blunt mitigation disconnects a paying subscriber who has done nothing wrong. Ask whether only the offending flow is dropped.
07
What is CPU utilisation at my real peak, and what is the actual bottleneck?
If a vendor cannot tell you what limits their platform, they have not looked. Ours is memory-bound, not compute-bound, and runs at 5.4% on the busiest node.
08
Do my existing bonds, routing daemon, and monitoring keep working?
Architectures that take the network interfaces out of the operating system also take your existing tooling and your team's knowledge with them.
What comes with it

The question is never "how fast". It is "how would you know?"

Every argument in this document depends on being able to answer questions about your own network — which subscribers are suffering, what an upgrade would cost you, whether a restart drops sessions. NOC2 is the system that answers them, and it is part of the platform rather than a separate purchase or a per-node licence.

Before you take an upgrade, it tells you what it would do

Not "version 3.8.35 is available" — that tells an operator nothing. Per gateway: whether the datapath can be swapped live, whether it needs the image and a reboot, and how many of your subscribers would drop if it did. A gateway that cannot be asked is listed as unknown, never as up to date.

It knows whether a restart keeps your subscribers

Four things have to be true at once for a BNG restart to preserve sessions. Miss any one and it drops everything while the box looks configured for the opposite. NOC2 checks all four per gateway and prints the fix in the order it has to be done.

It reads the traffic your other metrics cannot

QUIC is 67.1% measured of downstream bytes and it is invisible to every TCP-based quality metric — which means roughly three in four subscribers cannot be judged at all by the usual tooling. NOC2 reads the one part of QUIC that is not encrypted and reports it, clearly labelled as what it is.

It tells you when your alerts are reaching nobody

A monitoring system that cannot say whether its own warnings arrive is only writing things down. NOC2 checks each delivery route separately — mail, chat, on-call — and says plainly when an alert would leave the screen and go nowhere. An alert addressed to nobody looks exactly like one that was delivered.

It restricts who can reach your firmware

Release artifacts and the full system image are served to your gateways over plain URLs, because that is what an on-box updater can fetch. NOC2 generates the address allowlist from your own inventory, so a gateway added tomorrow is admitted without anyone remembering to edit a web server.

Per subscriber, not just per box

Line faults localised to the customer's own leg or your transit, subscribers at risk before they call, and what actually changed on a gateway in the last week — because "the link is fine" and "this subscriber is fine" are different questions.

The discipline behind it

It will tell you when it does not know

This is the part that is hard to demonstrate on a datasheet and obvious within a week of running it. Most monitoring renders a missing measurement as a zero. A gateway that has stopped reporting shows 0 Mbps. A counter that could not be read shows 0. A check that never ran shows green. Each of those is a lie told in the reassuring direction, and every one of them is indistinguishable from good news.

The same NOC2 panel in two states: a measured handshake RTT of 10.1 ms over 30 samples, and the same panel reading 'not measured' when nothing has been sampled.
The same panel, twice. Above: a real measurement, shown with its sample count — because 10.1 ms over 30 samples and over 30,000 are different claims. Below: the same panel when nothing has been sampled. It says not measured. It does not say 0 ms, and the counters beside it stay visible because those are still real.
Absent is not zero

Every value that can be missing carries a companion saying whether it was read. A session count that could not be taken is never rendered as an empty gateway; a cluster member that has not been heard from is never rendered as a healthy one.

"Cannot be asked" is its own answer

A gateway NOC2 could not reach is listed as unknown — not as passing, not as failing. Rounding it into "fine" is how estates end up with boxes nobody upgraded because no list ever showed them.

Measurements keep their conditions

A number is shown with what it was measured on and how many samples stand behind it. Figures in this document are tagged measured or illustrative for the same reason: a true number can acquire a false setting the moment it is quoted without its conditions.

Nothing dangerous happens without being named

Restarting the BNG daemon on a loaded gateway is refused unless the operator is shown the subscriber count and the actual consequence first. An upgrade that would drop sessions cannot be requested without a maintenance window attached.

None of this is a feature list you will find on a competitor's comparison chart, because it is not the kind of thing that fits in a row. It is the difference between a dashboard that is always green and a system that tells you which of its greens it actually checked.

In short

Choose the platform that makes your network answerable

We are not the cheapest way to move a gigabit, and we are not trying to be. On our architecture the box stopped being the constraint some time ago — the busiest node in production runs at 5.4%. Competing on price-per-gigabit would mean competing on the one dimension that has already been solved.

What is not solved, at most operators, is the evening support call nobody can answer, the technician sent to a line that was never broken, the upgrade deferred because it means dropping every session, and the customer disconnected for an attack their compromised device launched without their knowledge. Those are functional problems. They are what we build against, and they are where the money actually goes.

Bring us your traffic and your hardest subscriber complaint. We will show you the measurement, and we will show you how we proved the measurement is real.

On evidence. Figures tagged measured come from live BNGsoft production and lab systems and are reproducible on request; each is a specific observation rather than a modelled or projected figure. Figures tagged illustrative convey structure only and assert nothing measured. Individual numbers reflect specific fleet snapshots and configurations; your results will depend on your network.

On competitors. NetElastic and MikroTik are the trademarks of their respective owners and are referenced here for comparison only. Statements about their architecture describe publicly documented design at time of writing, and we have deliberately left capability rows unfilled rather than assert limitations we have not tested. Verify independently before you decide — including everything we have claimed about ourselves.