CONTACT MLE
Please fill in the form and your requirements below, and our team will contact you soon.


    AMD (Xilinx)Altera (Intel)LatticeMicrochip (MicroSemi)Other


    *By submitting this form you are consenting to being contacted by the MLE via email and receiving marketing information.

    X
    CONTACT MLE

    NPAP Datasheet

    TCP/UDP/IPv4 Network Protocol Accelerator Platform (NPAP) Datasheet

    For the stand-alone TCP/UDP/IPv4 Stack Full-Accelerator Subsystem allowing communication at full line rate and low latency

    NPAP Summary

    MLE’s Network Protocol Accelerator Platform (NPAP) is a proven TCP/UDP/IPv4 network protocol Full-Accelerator to enables high-bandwidth, low-latency communication solutions for FPGA- and ASIC-based systems for 1G / 2.5G / 5G / 10G / 25G / 40G / 50G / 100G / 400G Ethernet links.

    1. Features

    With a focus on reliability and low, deterministic latency, HHI designed this for embeddable FPGA and ASIC systems with the following features:

    • Interface to 1 / 2.5 / 5 / 10 / 25 / 40 / 50 / 100 / 200 / 400 Gigabit Ethernet 
    • Full-duplex with 128 bit wide bidirectional datapath 
    • Full line rate of 70 Gbps, or more, per instance in FPGA
    • Full line rate of over 100 Gbps per instance in ASIC
    • Low one-way latency NPAP-to-NPAP (600 nanoseconds for 100 Bytes)
    • Network diagnostics functions (optional)
    • TCP session priority management (optional)
    • Transport Layer Security (TLS) (optional)
    • Time-Sensitive Networking (TSN) (optional)
    • Network Impairment Generators (optional)

    Designed for maximum flexibility, NPAP implements in programmable logic the relevant network communication protocols:

    • IPv4 The core of the most standards-based networking protocols
    • TCP Reliable connectivity for direct secured connectivity
    • UDP Widespread protocol to enable simple direct or multicast communication
    • RRRRP Reliable, Rapid Request-Response Protocol based on Stanford HOMA
    • ICMPv4 Diagnostic protocol to validate connections
    • IGMPv4 Enables joining of multicast groups (optional)

    Due to the modularity NPAP can easily be enhanced by application specific protocols.

    The 128 bit wide datapath in combination with a pipelined architecture allows to scale throughput to line-rates of 50 GbE, and beyond, when using modern FPGA fabric, and up to 100 Gbps for ASIC implementations. NPAP is available in versions which recently have been merged:

    • Version 2 (currently 2.9.0) for new ASIC and AMD/Xilinx Versal, Ultrascale+, Ultrascale, Altera/Intel Agilex-7, Agilex-5E and Stratix-10, Microchip PolarFire and Lattice Avant G/X development, includes many recent resource optimizations plus timing optimizations focused on pipelines in FPGA, including Altera/Intel HyperFlex and Intel HyperFlex2
    • Version 1 (currently 1.10.1) back-ports bug fixes from Version 2 for long-term customer support. Not recommended for new design starts!

    1. NPAP Features

    1. NPAP Features With a focus on reliability and low, deterministic latency, HHI designed this for embeddable FPGA and ASIC systems with the following features: Interface to 1 / 2.5 / 5 / 10 / 25 / 40 / 50 / 100 / 200 / 400 Gigabit Ethernet (depending on the FPGA device and speedgrade.) Full-duplex with 128 bit wide bidirectional datapath  Full line rate of 70 Gbps, or more, per instance in FPGA Full line rate of over 100 Gbps per instance in ASIC Low one-way latency NPAP-to-NPAP (600 nanoseconds for 100 Bytes) Network diagnostics functions (optional) TCP session priority management (optional) Transport Layer Security (TLS) (optional) Time-Sensitive Networking (TSN) (optional) Network Impairment Generators (optional) Designed for maximum flexibility, NPAP implements in programmable logic the relevant network communication protocols: IPv4 The core of the most standards-based networking protocols TCP Reliable connectivity for direct secured connectivity UDP Widespread protocol to enable simple direct or multicast communication RRRRP Reliable, Rapid Request-Response Protocol based on Stanford HOMA ICMPv4 Diagnostic protocol to validate connections IGMPv4 Enables joining of multicast groups (optional) Due to the modularity NPAP can easily be enhanced by application specific protocols. The 128 bit wide datapath in combination with a pipelined architecture allows to scale throughput to line-rates of 50 GbE, and

    2. NPAP Applications

    2. NPAP Applications NPAP enhances your networked application with fast, scalable and reliable data connectivity. The powerful architecture of the underlying TCP/UDP/IPv4 Stack allows it to transfer data at line-rate with low processing latency without using any CPUs in the data path. The ubiquitous TCP/IPv4 and/or UDP/IPv4 communication protocol suite uses industry standard network infrastructure to address a wide-range of applications: High-Speed connectivity for distributed systems and Systems-of-Systems Scale-out datacenter connectivity Reliable, long-range chip-to-chip connectivity with backpressure FPGA-based SmartNICs High-Bandwidth Security with FPGA-based Smart Data Diodes  In-Network Compute Acceleration (INCA) Hardware-only implementation of TCP/IPv4 in FPGA PCIe Long Range Extension Networked storage, such as iSCSI or NVMe/TCP Test & Measurement connectivity Automotive backbone connectivity based on open standards High-speed, low-latency camera interfaces Video-over-IP for 3G / 6G / 12G transports Bring full TCP/UDP/IPv4 connectivity to FPGAs High-speed sensor data acquisition:stream data out of FPGAs into Network-Attached Storage (NAS)  High-speed robotics control and machine-to-machine:Stream data from servers via FPGA into actuators Hyper-converged computational storage acceleration for “over-Fabric” NVMe/TCP Deterministic low-latency, high-bandwidth, secure alternative to lwIP or Linux on embedded CPU 🌐 www.missinglinkelectronics.com MLE (Missing Link Electronics) is offering technologies and solutions for Domain-Specific Architectures, which focus on heterogeneous computing using FPGAs. MLE is headquartered in Silicon Valley with offices in Neu-Ulm and Berlin,

    3. NPAP Functionality Description

    3. NPAP Functionality Description NPAP is a complete subsystem of a high-performance programmable-logic based, standalone network stack featuring transparent handling of complete TCP/IPv4 and UDP/IPv4 protocol tasks, e.g. packet encoding, packet decoding, acknowledge generation, link supervision, timeout detection, retransmissions and fault recovery.  NPAP supports complete automatic connection control including tear up and tear down. Compute and manage retransmission timers as in RFC 6298. Transparent checksum generation and checksum checking, integrated flow control. RFC 9293 compatibility (TCP/IPv4 stack for Windows and Linux).  Depending on the project’s needs, deliverables can be: HDL source code or netlist Integrated FPGA system implementation Testbenches and scripts for real-life testing Comprehensive documentation and interfacing guide Development & design-in support NPAP has been optimized to ensure the best bandwidth-delay product performance for your application. To guarantee delivery of full performance and reliability Team MLE will support you in all engineering aspects: System-level architecture design where aspects such as mapping ingress / egress data streams to TCP sessions, or handling TCP’s congestion control, or optimizing the bandwidth-delay-product are handled in order to meet system-level bandwidth and latency requirements.  Chip-design, i.e. managing chip resources, interfacing with the Multi-Gigabit Transceivers (MGT), handling on-chip streaming (such as AXI beats) while integrating NPAP into your FPGA or ASIC device on your target hardware. Network

    3.1. NPAP Technical Features

    3.1. NPAP Technical Features FeatureSpecificationSupported on-chip Interfaces128 bit wide AXI4-StreamCompatibility with 3rd party Ethernet PHY interfacesStandard IEEE Ethernet PHYs with RMII, GMII, XGMII, etc via PCS/PMA via ASIC/FPGA Ethernet SubsystemCompatibility with 3rd party Ethernet Media Access ControllersFraunhofer HHI 10G/25G Low-Latency MACAMD/Xilinx 10G/25G Ethernet Subsystem (PG210)AMD/Xilinx 100G Ethernet Subsystem (PG165, PG203, PG314)Altera 10G / 25G Ethernet FPGA IPMicrochip PolarFire FPGA 10G Ethernet (UG0727)Lattice 10G / 25G Ethernet IP (FPGA-IPUG-02245)Supported protocols (Hardware based)Ethernet, ARP, IPv4, ICMPv4 (response only), IGMPv4, UDP & TCP, DHCP (client only)Number of simultaneous connectionsOne per TCP Core instantiation – see “Architecture Choices” below, a TCP Core in NPAP relates to a TCP socket in LinuxMessage SizesSupport for Ethernet Jumbo Frames of arbitrary lengthInterface to applicationDatapath via AXI4-Stream 128-bitand separate custom TCP command interfaceSupported FPGAs  Complete stack uses generic VHDL code (IEEE-1076 2002 or 2008, depending on NPAP version)AMD/Xilinx Virtex 4 to Virtex UltraScale+AMD/Xilinx Kintex to Kintex UltraScale+AMD/Xilinx Artix UltraScale+AMD/Xilinx Zynq-7000AMD/Xilinx Zynq UltraScale+ MPSoCAMD/Xilinx Zynq UltraScale+ RFSoCAMD/Xilinx Versal ACAP SeriesAltera Cyclone IV seriesAltera Cyclone 10 GX seriesAltera Stratix VAltera Stratix 10 GX seriesAltera Agilex 5 D, E SeriesAltera Agilex 7 F, I, M SeriesLattice Avant-G, Avant-XMicrochip Polarfire and PolarFire SoCPerformance70 Gbps line rate, or more, for single TCP/IPv4 session (depending on clock rate, see below)Typ. 600 ns transport delay (depending on

    3.2. NPAP Implementation Details

    3.2. NPAP Implementation Details

    NPAP implements a full accelerator, hence all network protocol processing is running as digital logic. Because NPAP does not rely on “soft” CPUs nor on external CPUs, NPAP shows very low and deterministic latency. Tradeoff cost vs performance over chip resources.

    • Dataflow is full duplex 128 bits wide using AXI4 Stream. This enables high data throughput without “FPGA bloat” nor timing issues.
    • Control-flow uses AXI4 Lite register interfaces, along with a hardware abstraction layer (HAL), Linux device drivers and Python scripts for NPAP administration.
    • NPAP is intensively tested for performance and interoperability against many other TCP/UDP/IPv4 network stacks.
    • NPAP brings its own 10G / 25G Low-Latency Ethernet MAC,
      but can interface with many Ethernet subsystems from the FPGA vendors.
    • NPAP implements a complete TCP/UDP/IPv4 stack including functions like ARP, ICMPv4, IGMPv4, DHCP.
    • NPAP is delivered with “Support IP blocks” including reference designs, design examples for setting MAC addresses and/or IPv4 addresses and/or TCP port numbers either from Programmable Logic / ASIC or via software running on a (Linux) host, either ARM or x86 based.
    • For each TCP connection that remains open at the same time, there shall be a dedicated instance of a TCP Core. An AXI4-Lite interface may be used to prioritize TCP sessions during runtime.
    • One and only one single instance of a UDP Core must be instantiated when UDP support is required. If there is no need to process UDP, then the UDP Core can be removed completely.
    • NPAP is highly parameterizable to optimize for lowest FPGA resource needs while still delivering full functionality and performance: Number of UDP Cores (zero or one), number of TCP Cores (zero or many), for each TCP Core: Rx Buffer Size and, separately, Tx Buffer Size
    • Other than Rx and Tx buffers and some buffers for clock domain crossings, NPAP hardly uses any buffers at all, which results in very low and deterministic latency.
    • Your “Layer 7” Application can directly be connected to NPAP via AXI4 Stream which gives you the option of keeping all traffic inside the Programmable Logic / ASIC, and/or to interface with “software” running on (Linux) host, either ARM based or x86 based, via DMA.

    A hierarchical design philosophy is used to integrate MLE NPAP together with the auxiliary blocks. Related blocks are packaged together in a “support wrapper”. These are nested and joined within a top-level wrapper. Please refer to the Section 8. NPAP Developer Documentation.

    At the highest level, a single NPAP based subsystem with NPAP Support IPs but with external MAC looks like this:

    NPAP Subsystem Top-Level

    Throughout this datasheet and our documentation we use the following color coding legend for certain blocks of functionality:

    NPAP Color Coding Legend

    Besides the number of TCP Cores, NPAP is highly parameterizable: Some parameters can be set at Compile-Time, others at Runtime, and some both ways.

    3.2.1. Compile-Time Parameters

    NPAP makes use of VHDL Generics to parameterize certain functionality as well as the NPAP design structure (such as the number of TCP cores) at compile-time. 

    Here some examples:

    3.3. NPAP Dataflow and Block Diagram

    3.3. NPAP Dataflow and Block Diagram

    The following shows the dataflow view of an exemplary design integrating NPAP with one UDP core and multiple TCP cores (3 for user-level plus 2 for Netperf), each with an example user application, plus Netperf (for bandwidth and latency benchmarking), plus network impairment (for Bit Error Rate Testing), plus diagnostics counters:

    MLE TCP IP NPAP Block Diagram Dataflow (fully loaded)

    The example user applications serve as an example on how to send and/or receive data from programmable logic via TCP/UDP/IPv4. For TCP this logic is inside one (or more) TCP Wrappers which contain HDL code for handling the control and data flow:

    3.4. NPAP Control-Flow View and Hardware Abstraction Layer

    3.4. NPAP Control-Flow View and Hardware Abstraction Layer

    For administration and control at run-time, NPAP implements so-called Runtime Parameterization and administration via an AXI-Lite register space. Access to this interface is exported through the so-called NPAP Hardware Abstraction Layer (HAL). Along with a Python library, NPAP HAL provides a high-level API for Linux software, utilizing swappable backends to communicate with the hardware across diverse environments, from SOC processing systems to remote workstations.

    MLE TCP IP NPAP Block Diagram Control Flow

    Here a list of connectivity choices for NPAP HAL:

    1. Via USB-UART or USB-JTAG or USB-IIC
    2. Via the FPGA-integrated Processing System which can be ARM, RISC-V, MicroBlaze, NIOS, etc
    3. Via PCIe connection with the host CPU 
    4. Via out-of-band Ethernet and UDP (on the roadmap)
    5. Via in-band Ethernet and UDP (on the roadmap)

    Besides NPAP HAL, MLE further provides a python based command-line tool, called npap-admin, that abstracts complex runtime parameterization into simple configuration file editing. Customers have been using npap-admin during evaluation and development, for Continuous Integration or in-the-field when NPAP-based products have been deployed.

    Good design examples for Runtime parameterization and administration of NPAP are the so-called NPAP Evaluation Reference Designs (ERD) running on many off-the-shelf FPGA platforms. Please read 3.5. NPAP Evaluation Choices for more information.

    3.4.1. NPAP Admin via USB-UART or USB-JTAG

    This mode of administration is very useful when using off-the-shelf FPGA development kits which feature a USB-to-UART connection which is then connected to a Linux workstation running Python.

    MLE TCP IP NPAP Admin via USB

    In this setup, the NPAP HAL utilizes the Extensible FPGA Control Platform (XFCP) to bridge the USB connection to the AXI4-Lite register space on the FPGA.

    Here an outline of the connectivity stack:

    3.5. NPAP Evaluation Choices

    3.5. NPAP Evaluation Choices MLE offers multiple ways to evaluate and benchmark NPAP: Free-of-charge we provide so-called NPAP Evaluation Reference Designs (ERD) implemented on off-the-shelf hardware. These are great for evaluating the functionality and the performance of MLE NPAP, when running in off-the-shelf FPGA hardware. These are also great as a reference design when integrating MLE NPAP into your system. A “Developers License” is a highly discounted extended evaluation license which gives you full source code access to integrate and run NPAP and NPAP ERDs within your target hardware. For evaluating the functionality and the performance of NPAP MLE provides Evaluation Reference Designs (ERD) for several FPGA Development Kits: ERD #FPGA Board / Development KitFeatures0AMD/Xilinx ZCU102 with Zynq Ultrascale+ MPSoC ZU9EG2x 1G1AMD/Xilinx ZCU102 with Zynq Ultrascale+ MPSoC ZU9EG1x 10G2MLE’s Network Protocol Accelerator Card “Ketch” with Altera Stratix 10 GX2x 10G3AMD/Xilinx ZCU111 with Zynq Ultrascale+ RFSoC ZU28EG1x 25G5AMD/Xilinx ZCU111 with Zynq Ultrascale+ RFSoC ZU28EG (early access)1x 100G7Microchip PolarFire MPF300-EVAL-KIT1x 10G8Trenz Electronics TE0950 with AMD Versal AI Edge VE23021x 25G11AMD/Xilinx Alveo U55C with Virtex UltraScale+4x 25G12Lattice Avant-G Versa Board1x 10G13Arrow AXE5-Eagle with Altera Agilex 5E1x 10G14AMD/Xilinx ZCU111 with Zynq Ultrascale+ RFSoC ZU28EG1x 10G15Tria AUB15p with AMD/Xilinx Artix Ultrascale+1x 10G101AMD/Xilinx ZCU102 with Zynq Ultrascale+ MPSoC ZU9EG1x 10G In MLEs Developer Zone you will find an

    3.6. IP Core Deliverables

    3.6. IP Core Deliverables Deliverables include IEEE 1685 IP-XACT packages which include NPAP Kernel plus NPAP Support IP blocks plus non-IP-XACT FPGA reference design projects. The following design blocks are part of a typical delivery package: Low-Latency Ethernet MAC for 10G/25G NPAP Kernel with IPv4, TCP, UDP NPAP Support IP Data Generator Checker(great for system-level testing and for performance tuning) TCP Command Application TCP Demo Application UDP Demo Application Netperf Control Application Hardware Abstraction Layer (HAL) python package Network Statistics and Diagnostics (optional) Network Impairment and Bit Error Insertion (optional) Evaluation Reference Design (as FPGA Design Project Archive) 🌐 www.missinglinkelectronics.com MLE (Missing Link Electronics) is offering technologies and solutions for Domain-Specific Architectures, which focus on heterogeneous computing using FPGAs. MLE is headquartered in Silicon Valley with offices in Neu-Ulm and Berlin, Germany.