Sunday, August 8, 2021

IOS-XR Layer2 Interconnect

 1.        Data Center trends

 

Data Center Interconnects (DCI) products are targeted at the Edge or Border leaf of Data Center environments, joining Data Centers to each other in a Point-to-Point or Point-to-Multipoint fashion, or at times extending the connectivity to Internet Gateways or peering points. Cisco has two converged DCI solutions; one is with integrated DWDM and another with advanced L3 routing and L2 switching technologies. A recent Dell’Oro report, forecasts the aggregate sales of equipment for DCI will grow by 85 percent over the next five years. This is driving strong demand for Ethernet Data Center Switch, and Routing technologies.

The emerging need for simplified DCI offering spans four core markets.

 

·       Mega Scale DC

·       Cloud DC

·       Telco Cloud

·       Large Enterprises

The emergence of cloud computing has seen a rush of traffic being centralized in regional and global data center as the Data Center emergence to being the core of many service deliveries, more recently ‘far edge’ compute in 5G has reemphasized the trend, with DC’s now being at the core of 5G build outs, as Web companies and SPs embark on using automation and modern DC tools to turn up 5G sites at unprecedent rates and look at micro data centers at the edge to enhance the user experience.

 

DCI’s newest architectures is drive by massive DCs that need connecting by either leased lines from SPs or by deploying their own or leasing dark fiber.

 

Inside the DC they often deploy a mix of their home-grown applications over and defined technologies, mostly L2 type services to reach compute hosts at the peripherals, although we have seen recent trends of L3 being expended all the way to compute with Segment Routing (SR).

 

Outside the DC fiber is less abundant and inter DC solutions are fairly standardized with SP class products providing the richest functionality at the most optimal scale and price point. A motivation in the last 2 years for further DCI upgrades has been the migrating to MacSec for Inter DCI links.

 

A most recent trend is of 100GE and 400GE Data center build outs, driving DCI upgrades, we’re seeing customers migrate to higher speed links at different inflection points, with 100GE being the sweet spot current, creating catalyst for Terabit platforms that support advanced L2/L3 VPN services and Route and Bridge functions, case in point ASR9000 and NCS5500.

2.        ASR 9000 L2 DCI GW feature overview

 

Ethernet VPN (EVPN) and Virtual Extensible LAN (VXLAN) have become very popular technologies for Data Center (DC) fabric solution. EVPN is used as a control plane for the VXLAN-based fabric and provides MAC addresses advertisements via MP-BGP. It eliminates the use of the flood and learn approach of original VXLAN standard (RFC7348.)  As a result, DC fabric allows to reduce unwanted flooding traffic, increase load sharing, provide faster convergence and detection of link/device failures, and simplify DC automation.

ASR9k as a feature reach platform can be used in DC fabric as a DC edge router. With one leg in DC and other in WAN, ASR9k is the gateway for traffic leaving/entering DC.  At DC fabric facing, ASR9k operates as border leaf. At MPLS WAN facing, ASR9k operates as WAN edge PE. Such type of router is commonly referred to as DC Interconnect (DCI) Gateway router or EVPN-VXLAN L2/L3 gateway.

 

There are two main use-cases that are widely known in the industry:  L3 DCI gateway as a solution for L3VPN service on VXLAN fabric and L2-based gateway which provides L2 stitching between VXLAN fabric and MPLS -based Core.

The first use-case was available in XR release 5.3.2. It was the 1st phase of ASR9K based DCI GW solution

In phase 2, beginning from 6.1.1 release, EVPN Route Type 2 (MAC route) was integrated with L2 MAC learning/forwarding on the data plane. This functionality is called EVPN-VXLAN L2 gateway. Multicast routing was used as an underlay option for distribution of BUM (Broadcast, Unknown unicast, Multicast) traffic within VXLAN Fabric. Starting from release 6.3.1 ASR9K-based DCI solution supports Ingress Replication underlay capability on the VXLAN fabric side.

EVPN-VXLAN L2 gateway functionality will be described in the following sections.

The reference topology for this solution is shown at the picture below.

 

 

A close up of text on a white background

Description automatically generated

 

Nexus 9000 plays role of ToR/Leaf switches for this DCI solution. N9K ToRs should provide Per-VLAN first-hop L3 GW for Hosts/VMs behind the ToRs.

ASR9K DCI GW acts as an L2 EVPN-VXLAN GW within the fabric and participates in fabric-side EVPN control plane to learn local fabric MAC routes advertised from ToRs, and distribute external MAC routes learnt from remote DCI GWs, towards the ToRs. VXLAN data plane is used within the fabric. Fabric side BGP-EVPN sessions between DCI GWs and ToRs can be eBGP or iBGP.

On the core side, DCI GW will do BGP-EVPN peering with remote DCI GWs to exchange MAC routes together with host IP bindings (needed for ARP suppression on the ToRs). MPLS data plane is used on the core side. Core side BGP-EVPN session can be eBGP or iBGP as well.

On the WAN or external core side, ASR9K DCI GW acts as an L2 EVPN-MPLS GW that participates in EVPN control plane with remote DCI GWs to learn external MAC routes from them and distribute fabric MAC routes, learnt locally towards them. MPLS data plane is used with remote PODs, outside the fabric. In essence, the L2 DCI GW stitches the fabric side and WAN side EVPN control planes and in data plane, it bridges traffic between VXLAN tunnel bridge-port and MPLS tunnel bridge-port.

 

2.1.      Distribution of BUM traffic

 

In DC switching network, L2 BUM traffic flooding between leaf nodes (including border leaf on DCI GW) is necessary. There are 2 operational modes for BUM traffic forwarding. The first mode is called egress replication. In this mode, the VXLAN underlay is capable of both L3 unicast and multicast routing. L2 BUM traffic is forwarded using underlay multicast tree. Packet replication is done by L3 IP multicast - an egress replication scheme.

The second mode of operation is called VXLAN Ingress Replication (VXLAN IR). This mode is used when VXLAN underlay transport network is not capable of L3 multicasting. In this mode, the VXLAN imposition node maintains a per VNI list of remote VTEP nodes which service the same tenant VNI. The imposition node replicates BUM traffic for each remote VTEP node. Each copy of VXLAN packet is sent to destination VTEP by underlay L3 unicast transport.

 

3.        Multihoming deployment models

 

Depending on the capability of ToRs, ASR 9000 DCI supports two multi-homing deployment models

Anycast VXLAN L2 GW model

ESI based multi-homing VXLAN GW model

These models are based on different mechanisms inside DC fabric for multi-homing and load-balancing between ToRs and DCI GWs. On the MPLS WAN side both models use the same implementation approach.

 

3.1.      Anycast VXLAN L2 Gateway

 

DC gateway redundancy and load sharing is a critical requirement for modern high-scalable data centers. Today, in every new DC deployment, multi-homing DCI gateway is a must-have requirement. Anycast VXLAN gateway is a simple approach of multi-homing. It requires multi-homing gateway nodes to use a common VTEP IP. Gateway nodes in the same DC advertise the common VTEP IP in all EVPN routes from type 2 to 5. N9k ToR nodes in the DC see one DCI GW VTEP located on multiple physical gateway nodes. Each N9k forwards traffic to the closest gateway node via IGP routing.  Closest gateway is identified by shortest distance metric or ECMP algorithm. 

Among multi-homing DCI gateway nodes, an EVPN Ethernet Segment is created on VXLAN facing interface NVE. One of the nodes is elected as DF for a tenant VNI. The DF node is responsible for flooding BUM traffic from Core to DC.

All DCI GW/PE nodes discover each other via EVPN routes advertised at WAN. EVPN L2 tunnels are fully meshed between DCI PE nodes attach to MPLS Core. See at the blue lines on the picture below.

 

A close up of a logo

Description automatically generated

 

This picture describes a topology of anycast VXLAN gateway between DC and Core. In this topology, both ASR9k DCI GW nodes share a source VTEP IP address. N9k runs vPC mode in pairs.

Redundant DCI GWs use anycast VTEP loopback to advertise towards the VXLAN fabric. This loopback IP is used by ToR to forward traffic to DCI GW.

 

In the pictures below the BUM traffic distribution is shown. Flooded BUM traffic is dropped by the non-DF DCI node in both directions to cut the loop. BUM packet received from VXLAN fabric side will be dropped on the VNI port on ingress. DF will flood on MPLS side, which will also go to the peer non-DF DCI. However, this non-DF DCI will drop it on the VNI port in egress direction.

 

A close up of a logo

Description automatically generated

 

 

In the direction from DC-2/DC-3 to DC-1, both ASR9k DCI GW nodes receive the same BUM traffic from MPLS WAN. The DF for the tenant VNI forwards traffic to DC-1. Non-DF drops BUM traffic from WAN.

The N9k leaf nodes work in VPC pair and VPC is responsible for prevention of duplication traffic toward VM/Host

 

A close up of a logo

Description automatically generated

 

3.2.      All-Active Multi-Homing VXLAN L2 Gateway

 

Although anycast VXLAN provides a simple multi-homing solution for gateway, traffic going out of DC may not be properly load balanced on DCI GW nodes.  This is due to load balance relying on IGP shortest path metric. A N9k often sends traffic to one DCI gateway only. To overcome the limitation, Ethernet Segment based all-active multi-homing VXLAN L2 gateway is introduced. 

 

Figure below shows the topology of all-active multi-homing VXLAN L2 gateway. In this scenario, all leaf nodes and DCI nodes have a unique VTEP IP. Each N9k leaf node creates EVPN Ethernet Segment (ES1) for a dual-homed Host/Server. ASR9k border leaf nodes create an Ethernet Segment (ES3) for VXLAN facing NVE interface. Traffic from DCI GWs is load balanced between ToRs nodes. The same happens in opposite direction.

Each Leaf sends BUM traffic to both ASR9k nodes. To prevent traffic duplication, only one of the ASR9k nodes can accept VXLAN traffic from N9k leaf. This is done by DF rule. DF election is done at per tenant VNI level. Half of the VNIs elect top DCI GW as DF. The other half elect bottom DCI GW. DF accepts traffic both from DC and WAN. Non-DF drops traffic from DC and WAN. Load balance across VNI is thus achieved on the two DCI L2 gateway nodes.  

 

In all-active multi-homing topology, the data plane must perform ingress DF at VXLAN fabric facing. For the ASM-based underlay option, it’s not a problem, non-DF recognizes multicast encapsulated packets and drop them. But if Ingress Replication is used, DCI GW should have additional identifications to recognize unknown unicast traffic. It hence requires a data plane to implement Section 8.3.3 of RFC 8365 – BUM flag in the VXLAN header. The flag is used to identify L2 flood traffic received from VXLAN, thus a non-DF node can perform ingress drop operation to prevent duplicated traffic sent to a destination.

A picture containing object

Description automatically generated

 

Traffic flow on all-active multi-homing VXLAN L2 gateway is illustrated in figure below.  DC1 outbound BUM traffic arrives on a leaf first. Leaf replicates the traffic to two ASR9k DCI nodes.  DF DCI nodes flood traffic to WAN. Non-DF node drops traffic from DC fabric. Traffic flooded to WAN goes to DC-2 and DC-3. One copy comes back to DC-1 via bottom DCI node. The bottom DCI node compares the split horizon label in the received MPLS packet. Drops the packet with split horizon rule.

A close up of text on a white background

Description automatically generated

 

In the reverse direction (picture below), DC-1 inbound traffic from DC-2/DC-3 arrives on both top and bottom DCI nodes. The bottom one drops traffic with DF rule. The top node forwards 2 copies to remote leaf nodes. The N9k leaf nodes apply DF rule before forwarding traffic to Host.

 

 

 

Sources:

Jiri Chaloupka’s presentation from Cisco Live “BRKSPG-3965 EVPN Deep Dive”

Introduction to Seamless BFD (S-BFD)

 

Seamless BFD Introduction

 

Bidirectional Forwarding Detection (BFD) was introduced as part of OAM functionality for path continuity check that helps with rapid failure detection that plays a key role in enhancing fast convergence and traffic redirection within few 10s of milliseconds to abide with end user applications expectation. BFD is a light weight, less control plane overhead protocol that creates a session with neighbor after a 3-way handshake and exchange control packets at regular interval (few msec) and notify any failure to the client protocols like IGP, PIM, RSVP within few milliseconds.

 

Segment Routing is key enabler for Software Defined Networking and plays an essential role to dynamically instantiate tunnel between end points by encoding the state/path entries only on the tunnel head end. Such Traffic Engineering tunnels instantiated with Segment Routing are required to validate the path before using it for data traffic. Applying BFD for such requirement uncovers several limitations due to the base characteristics of BFD:

  • BFD creates state entries on both headend and tailend. Any node acting as tail end for 1000s of headend will result in creating 1000 state entries for uni-directional path validation. This results in scaling problems.
  • BFD requires exchanging the discriminators that introduces a delay before validating the path.
  • BFD can become resource intensive with respect to memory utilization and CPU processing due to 1-1 session between end points.

 

The above challenges directly imposed a need for a simple mechanism to monitor the path and virtual resources for any failure detection. In this document, we introduce Seamless BFD (S-BFD) and explain how S-BFD addresses the challenges listed above.

 

Seamless BFD Overview

 

RFC 7880 proposes the S-BFD architecture.  S-BFD was designed by extending the BFD and addressing the challenges faced with BFD. S-BFD introduces the below:

  1. Concept of pre-assigning domain wide unique discriminator and advertise the same using IGP protocol extension.
  2. Concept of Reflector session that reflects the control packet if the Your Discriminator matches the local reflector session discriminator value.

 

Figure 1. SBFD ComponentsFigure 1. SBFD Components

There are two components of S-BFD as below:

  • S-BFD Discriminators
  • S-BFD Reflector session

 

S-BFD Discriminator

 

A pair of Discriminator negotiated between BFD neighbors identifies a BFD session. In traditional BFD, the Discriminator is negotiated as part of 3-way handshake. With Seamless BFD, each network entity within same administrative domain will be pre-assigned with a minimum of one domain wide unique Discriminator and advertised using IGP protocol extensions.

 

The uniqueness of S-BFD Discriminator being the key aspect, it is the Operators responsibility to maintain the domain wide uniqueness as like the System or Router ID of IGP protocol. While the simple option is to arbitrarily allocate from a pool of Discriminator manually by Operators, there are other various options that can be considered. For example, when the underlying IGP protocol is OSPF, the uniqueness of 32-bit Router Identifier can be leveraged and use the same as S-BFD Discriminator. Alternately, the centralized intelligence like SDN can also be leveraged to dynamically assign from a pool. Cisco implementation supports leveraging the loopback IPv4 address as Discriminator.

 

 In Figure 1, it can be noted that each entity that needs to be monitored will be assigned with a unique S-BFD Discriminator. As shown in Figure 1, Host H1 assigns Discriminator D1 for Entity E1 and Discriminator D2 for Entity E2. When any node within the domain is required to monitor Entity E1 within Host H1, will send a S-BFD Control packet to H1 and set the control packet details in S-BFD as Your Discriminator (YD)=D1 and My Discriminator (MD)=Local-Random-value.

 

This enhanced characteristic of S-BFD completely eliminates the need for any handshake or Discriminator negotiation between end points to monitor any path or local resources. This drastically reduces the time taken by the Initiator to create a session and, consequently, augments the swiftness for path and resource liveliness validation.

 

S-BFD Reflector session

 

Each network entity participating in S-BFD architecture will create one or more reflector sessions. The intention of reflector session is to respond back to a received S-BFD control packet with “Your Discriminator” (YD) as any one of the locally assigned S-BFD Discriminator.

 

The routers are allowed to have multiple reflector session with associated S-BFD Discriminator. This is required when S-BFD Discriminator is used for services other than basic connectivity. For example, a router can assign a S-BFD Discriminator to any local service and create a reflector session for the service S-BFD Discriminator. The router will respond back to control packet till the associated service is up. In case of the service failure, the associated reflector session will be disabled and no response will be sent back causing the session to fail on Initiator to take any necessary action.

 

This aspect of S-BFD eliminates the need for a per session state entry in each BFD neighbors.

 

S-BFD Functional Operation

 

Figure 2. SBFD Functional OperationFigure 2. SBFD Functional Operation

 

Initiator behavior

     Any node, upon instructed to check the liveliness of a remote entity by local client or manually, will trigger a S-BFD session by querying the local database to fetch the respective Discriminator advertised by the remote node for the entity. This Discriminator value will be used in Your Discriminator field of S-BFD control packet.

 

 

Figure 3. SBFD Control PacketFigure 3. SBFD Control Packet 

 

In Figure 2, Initiator is instructed to check the liveliness of Entity E1 on Responder. Accordingly, it will trigger a BFD session to Responder and will use the below details:

  • “My Discriminator” as any random value that does not overlaps with local SBFD Discriminator.
  • “Your Discriminator” as “02020202” - the value assigned and advertised by Responder for Entity E1.
  • “State” to a value describing the local state.
  • Set “Demand” bit in control packet.

 

Figure 3 shows the control packet format where the above information will be populated. Upon receiving a reply back from the remote node (S-BFD Responder), the Initiator should declare that the session is UP. If there is no reply from remote node till the local timer expired, Initiator should intimate the failure to respective local client associated with this session.

 

Responder behavior

 

Figure 4. SBFD Responder BehaviorFigure 4. SBFD Responder Behavior

Any node, upon receiving S-BFD control packet with “Your Discriminator” value matching one of the locally assigned S-BFD Discriminator, should forward it to S-BFD Reflector session after following the traditional BFD packet validation. In Figure 2, Responder on receiving S-BFD with Your Discriminator of “02020202” will forward to S-BFD Reflector session. The Reflector session will generate a reply back with below details:

  • “My Discriminator” as “02020202”
  • “Your Discriminator” as value copied from received BFD control packet.
  • “State” to value describing the local state.
  • Clear the “Demand” bit.

 

Experimental Testing and analytical result

 

Test Environment

 

We performed the testing in a test bed with two Cisco routers back-to-back connected. For performance accuracy, we didn’t include any transit devices We tried the testing on 2 different Cisco IOS-XR software versions. We used MPLS-TE as BFD Client. This study is to compare the resource utilization between traditional BFD and S-BFD.

 

Figure 5. Test EnvironmentFigure 5. Test Environment

  • Number of Sessions: ~1000
  • Metrics studied: Memory consumption, CPU Utilization for BFD process, Time taken for session establishment.
  • Hardware Type: Cisco ASR9000 Routers
  • Software: Cisco IOS-XR 5.3.4 and Cisco IOS-XR 6.1.2.

 

Analytical Result

 

Below is the observed analytical result while comparing the performance between traditional BFD and S-BFD using MPLS-TE as the client.

 

Traditional BFD (Cisco IOS-XR version 5.3.4)

No. of Sessions

Memory Consumed (in KB)

Time taken for session establishment (in msec)

CPU Utili

200

2138

4012

0.18%

400

4437

9120

0.85%

600

6679

12238

1.02%

960

10672

17331

2.03%

 

 

 

  

 

 

 

 

 

 

 

 

 

 

Table 1. Statistics for Traditional BFD (IOS-XR 5.3.4 Version)

 

Traditional BFD (Cisco IOS-XR version 6.1.2)

No. of Sessions

Memory Consumed (in KB)

Time taken for session establishment (in msec)

CPU Utili

200

2617

12231

0.20%

400

5718

13141

0.80%

600

8721

16998

1.00%

960

14914

23124

2.00%

 

Table 2. Statistics for Traditional BFD (IOS-XR 6.1.2 Version)

 

Seamless BFD (Cisco IOS-XR version 5.3.4)

No. of Sessions

Memory Consumed (in KB)

Time taken for session establishment (in msec)

CPU

Util

200

122

200

0.12%

400

245

2510

0.68%

600

369

3011

0.94%

960

590

4000

1.92%

 

Table 3. Statistics for Seamless BFD (IOS-XR 5.3.4 Version)

 

Seamless BFD (Cisco IOS-XR version 6.1.2)

No. of Sessions

Memory Consumed (in KB)

Time taken for session establishment (in msec)

CPU

Util

200

140

1009

0.50%

400

260

2110

0.70%

600

380

2560

1.00%

960

610

3000

2.00%

 

Table 4. Statistics for Seamless BFD (IOS-XR 6.1.2 Version)

 

Graphical Comparision is as below:

 

Mmeory Util Graph

 

Time Graph

Conclusion:

 

As summarized in the above tables and Graphs, below are the observed metrics while testing BFD and S-BFD:

 

  • The Memory consumption for traditional BFD is directly proportional to the number of sessions while it does not increase drastically for Seamless BFD.
  • As the number of session increases, the time taken for session establishment is more for traditional BFD due to the Discriminator negotiation while it is minimal for Seamless BFD as there is no Discriminator negotiation.
  • The CPU utilization is nearly the same for both Traditional BFD and Seamless BFD with a very negligible difference.

 

Disclaimer: The statistics are collected from testing environment and is not absolute value but relative values.

 

* There is no difference in rate at which packets are exchanged as it is based on the configured interval