ISSUE 2 - JANUARY 2026
Meet
H0G The opensource system for handling your HDL code on Git
+
The move of post-quantum cryptography to the testbench
The case for pre-layout signal integrity
Building resilient FPGAs and SoCs for every radiation environment
Bridging FPGAs with the Vitis System Device Tree flow
US EAST Join
and more
April 29th-30th, 2026 DCU Convention Center, 50 Foster Street, Worcester, MA 01608
Learn more at
fpgahorizons.com/ us-east-26
FOREWORD
Publisher: Adam Taylor Editor: Matt Hilbert Designer/Cover art: Susie Hinchliffe Marketer: Louise Paul
CONTRIBUTORS: Adam Taylor, Adiuvo Engineering Dan Binnun, E3 Designers Espen Tallaksen, EmLogic Dr Francesco Gonnella, Birmingham University Jeff Johnson, Opsero Inc. Matt Hilbert, Editor, FPGA Horizons Journal Matthew Holder, XJTAG Dr Pierre Maillard, AMD Published by Adiuvo Events. ©️ Adiuvo Events. All rights reserved. No part of this publication may be reproduced in whole or in part in any medium without the express permission of the publisher. For editorial enquiries email contribute@fpgahorizons.com For advertising enquiries email advertise@fpgahorizons.com
Welcome to Issue 2 of FPGA Horizons Journal One of the great strengths of FPGA technology is the wide range of applications and disciplines it brings together. From applications such as embedded vision, to robotics, networking and communication to disciplines such as board-level decisions, verification and software integration, FPGA development is never just about logic in isolation. That range was clearly reflected at FPGA Horizons London 25, which proved to be a standout event for the community. With 290 attendees, 28 exhibitors, and 21 technical talks, the one-day conference delivered a concentrated programme of high quality, engineer-to-engineer discussion. The level of engagement throughout the sessions, and in the conversations that followed, reinforced the demand for technicallyfocused events that prioritize real-world challenges over marketing noise. This journal is a natural extension of that same philosophy. In this issue, we look across the FPGA design lifecycle. We explore why PCB testing and signal integrity need to be addressed earlier in the design process, how requirements tracking and HDL project management can be made more robust and reproducible, and how modern tools and methodologies support these goals. We also examine how today’s FPGAs and SoCs are being deployed in increasingly demanding roles. From radiation-resilient designs spanning terrestrial to space environments, to multi-FPGA systems that bridge hardware and software using flows such as the Vitis System Device Tree, the articles highlight how architecture, verification, and software must evolve together. Alongside this, we address longer term pressures including postquantum cryptography and architectural techniques like logic folding that extract more performance from modern devices. The momentum from London now carries forward to our upcoming FPGA Horizons US event, where we will continue these conversations with a growing international audience. As always, FPGA Horizons exists to inform, challenge, and connect the FPGA community. We hope this issue provides practical insights and sets the direction as we look toward the next horizon. Adam Taylor Publisher
In
Issue 2
6
Industry roundup
8
Matt Hilbert (Editor) explains The move
of post-quantum cryptography from the classroom to the testbench
Taylor (Adiuvo Engineering & 12 Adam Training / Publisher) shows How
logic folding boosts FPGA speeds and reduces footprints
18 Matthew Holder (XJTAG) clarifies Why PCB testing needs to shift left to the design stage
22 COVER STORY Dr Francesco Gonnella (Birmingham University) talks about the advantages of introducing Hog, The open-
source system for handling your HDL code on Git
4
CONTENTS
Pierre Maillard (AMD) illustrates the 28 Dr. importance of Building resilient FPGAs and
SoCs for every radiation environment
Tallaksen (EmLogic) walks through 32 Espen Simplifying requirements tracking
with the UVVM Requirements Traceability Matrix
Binnun (E3 Designers) makesThe case for 38 Dan pre-layout signal integrity
Johnson (Opsero) demonstrates a novel 42 Jeff approach for Bridging FPGAs with the
Vitis System Device Tree flow
Disclaimer The content published in the FPGA Horizons Journal is contributed by independent authors and researchers. While we strive to ensure accuracy and maintain a standard of quality, the views and opinions expressed in individual articles are those of the respective contributors and do not necessarily reflect the views of the FPGA Horizons Journal editorial team or its affiliates. We are committed to using neutral and inclusive language wherever possible. However, given the diversity of voices and topics, variations in tone and expression may occur. The FPGA Horizons Journal does not accept responsibility for any errors, omissions, or differing viewpoints presented in the submitted content. Readers are encouraged to critically engage with the material and consult additional sources where appropriate.
5
Industry roundup AMD starts shipping its Versal™ RF Series, redefining high-performance RF systems in a single chip
SEGGER Flashers get FPGA programming capabilities with Flasher BitStreamer
The AMD Versal RF Series brings together high-performance RF-sampling ADCs and DACs, hard IP DSP compute, and AI Engines for DSP in a single adaptive SoC, for nextgeneration radar, EMSO, and test platforms.
SEGGER has announced the introduction of the Flasher BitStreamer, a new embedded software solution that expands the programming capabilities of its Flasher family of insystem programmers (ISPs).
cocotb 2.0 has landed
This is the next major milestone for the Python-based verification framework, and puts the spotlight on developer experience by making testbenches easier to understand and extend.
Vitis functional simulation using Python
To accelerate verification of HLS and AI Engines, Vitis 2025.1 has introduced the ability to perform functional simulation of the HLS or AIE application using MATLAB or Python.
Altera expands its strategic partnership with AXISCADES for mission-critical defense applications development
The collaboration aims to accelerate development time by combining Altera’s advanced programmable logic technologies with AXISCADES’ deep expertise in aerospace, defense, and artificial intelligence system design.
Unofficial FPGA Sega Neptune gets pushed into 2026
The launch of the GF1 Neptune, an FPGA recreation of the long-awaited Sega Genesis / Mega Drive and 32X combination in a single system that never actually hit store shelves, has been delayed until this year.
Shrike-Lite combines Raspberry MCU RP2040 with Renesas FPGA for US$4 For just $4, the open-source FPGA development board, Shrike-Lite, combines Raspberry Pi’s microcontroller RP2040 with a 1120-LUT FPGA from Renesas for powerful, cost-effective prototyping.
QuickLogic selected by Chipus for 12 nm high-performance data center ASIC QuickLogic’s eFPGA Hard IP has been chosen for a data center production ASIC that will be fabricated on an industry-proven 12 nm process technology. The embedded FPGA was shown to meet strict performance and connectivity requirements while optimizing the fabric to minimize silicon area.
Lattice brings post-quantum cryptography to low-power FPGAs
The MachXO5-NX TDQ FPGA platform is the first to be fully compliant with the Commercial National Security Algorithm (CNSA) 2.0 postquantum cryptography (PQC) standard, and includes advanced cryptography capabilities and hardware root of trust (RoT).
6
LONDON 25
Watch all 21 talks online
Visit
fpgahorizons.com/london-25/ london-25-talks
Registered to attend London 25? Access to the talks is free. Didn’t register? There’s a one-off fee of £40.
Post-quantum cryptography
GEAR UP:
Post-quantum cryptography has moved from the classroom to the testbench Matt Hilbert, Editor, FPGA Horizons Journal
Ever since 1980 when American physicist Paul Benioff published the first quantum mechanical model of a Turing machine and proved that computation could be described within the laws of quantum mechanics, people have been reading and talking about quantum computing. No surprise there. It promises to unlock a level of computational power far beyond what today’s computers can achieve by harnessing the unique quantum properties of superposition and entanglement. Instead of processing information in binary bits, quantum computers use qubits that can represent multiple states at once, enabling them to explore vast solution spaces in parallel. The result is the potential to solve problems that are currently intractable, from simulating complex molecules for new medicines, to optimizing global logistics, to breaking the cryptographic codes that underpin digital security everywhere. At its core, the promise of quantum computing is to transform fields where exponential complexity overwhelms classical systems, offering breakthroughs that could reshape science, industry, and security.
8
That’s the promise and we’re now moving close to reality with major companies investing billions in their efforts. Google, for example, is focusing on error correction as the ultimate milestone and aims to resolve the fragility of qubits and make quantum computing credible at scale. Quantinuum is addressing the fragmentation problem by evolving trapped‑ion machines and algorithms together for end‑to‑end solutions in finance, pharma, and cybersecurity. And rather than racing to build standalone quantum hardware, Alibaba is emphasizing cloud-accessible quantum services, integrating experimental quantum processors into its massive cloud ecosystem. And that’s the problem. By around 2035, quantum computers with the power to crack current encryption methods are expected to arrive, and the standard cryptographic primitives we rely on today will no longer be sufficient. Public key cryptography approaches such as RSA, ECC, ECDSA and EdDSA which depend on the hardness of integer factorization or discrete logarithm problems will be broken, forcing a shift to post-quantum algorithms. Symmetric crypto and bitstream protection will remain but will need stronger parameters and integration with new quantum-safe standards. The most urgent changes will be in secure boot, firmware signing, and key exchange protocols,
It’s time to rethink every encryption standard you thought you knew At first glance, it might appear all of this talk is a theoretical exercise that can be left on the shelf until some time in the 2030s. For FPGA engineers, however, it’s not. FPGAs are designed for use over decades and some FPGA families are now guaranteed into the 2040s. Er, after quantum computing lands. The engineering efforts to introduce postquantum cryptography (PQC) need to start now because they involve more than implementing stronger key exchanges and digital signature schemes in firmware, secure boot processes and other cryptographic routines. It’s also about ensuring the FPGA hardware can handle the new cryptographic primitives efficiently. That could mean adapting the hardware architecture to support larger key sizes, or incorporating dedicated modules designed specifically for quantum‑safe algorithms. It’s a tandem effort, updating the software for new encryption standards and making sure the hardware can run those standards effectively.
How Shor’s algorithm can help us
Fortunately, the NIST initiative provided a practical entry point into PQC in August 2024 when it published the first three Federal Information Processing Standards (FIPS):
Fortunately, we’ve been a step ahead of this challenge since 1994 when American mathematician Peter Shor published an algorithm which showed how a quantum computer could break RSA and ECC cryptography in polynomial time, making current FPGA secure boot and update signatures vulnerable. Shor’s paper sent shockwaves through the cryptographic world and, by 2001, experimental demonstrations of his algorithm factoring small numbers proved the threat was more than academic.
FIPS 203: ML-KEM – Module-Lattice-Based Key-Encapsulation Mechanism Standard The primary standard for key encapsulation (secure key exchange) based on the CRYSTALS-Kyber algorithm, with comparatively small encryption keys compared to other PQC alternatives but still significantly larger than RSA/ECC FIPS 204: ML-DSA – Module-Lattice-Based Digital Signature Standard The primary standard for protecting digital signatures using the CRYSTALS-Dilithium algorithm
Through the 2000s and 2010s, as quantum hardware steadily advanced, governments and industry realized that RSA and ECC could not be relied upon for long-term protection. In 2016 the US National Institute of Standards and Technology (NIST) responded with a global competition for proposals to standardize new public key cryptographic algorithms resistant to quantum computer attacks. It culminated in 2022 with the selection of lattice-based algorithms such as Kyber for key exchange and Dilithium for digital signatures.
FIPS 205: SLH-DSA – Stateless Hash-Based Digital Signature Standard A backup method for protecting digital signatures which uses the hash-based Sphincs+ algorithm, should ML-DSA prove vulnerable
9
These standards come with recommended parameter sets, interoperability guidance and transition advice, providing a path from current systems to quantum-resistant ones. ML‑KEM is directly relevant to FPGA workflows because its polynomial arithmetic and number‑theoretic transforms can be mapped efficiently onto Digital Signal Processing (DSP) slices, Look‑Up Tables (LUTs) and Block RAM (BRAM), enabling hardware‑accelerated quantum‑safe key exchange. ML‑DSA extends this protection to authentication and secure boot, giving FPGA designers a standards‑compliant way to verify bitstreams and firmware. SLH‑DSA offers a conservative, hash‑centric alternative that trades speed for long‑term trust, making it valuable for root‑of‑trust anchors and certification processes.
How this will change the way FPGA engineers approach encryption The arrival of PQC will reshape FPGA engineering practices in several concrete ways. First, design priorities will need to accommodate the wide variation in key, signature and ciphertext sizes that different PQC schemes impose. Where RSA or ECC implementations often worked within predictable size ranges, PQC can require much larger buffers and different memory layouts, so you’ll need to plan memory and I/O early in board and device selection. Secondly, a new set of computation kernels becomes central: parameterizable, pipelined engines for polynomial multiplication, highly-optimized NTT implementations, modular reduction circuits and Keccak/SHAKE acceleration will move from research prototypes into standard IP libraries. These kernels demand careful trade-offs between area and throughput. An FPGA implementation for an edge device, for example, will favor compact, low power data paths while a gateway or server accelerator will prioritize parallel lanes and low latency. Equally important is the imperative for cryptoagility. Because PQC is still evolving and additional parameter sets or algorithms may become recommended over time, FPGA crypto subsystems must be modular, with discrete algorithm blocks behind a clear, upgradeable interface so firmware or secure boot routines can switch algorithms without an entire hardware redesign.
That modularity pairs with lifecycle concerns, so field upgrade mechanisms, secure firmware delivery and robust fallback modes will become standard planning items for devices expected to operate through the 2030s. Sidechannel and fault injection countermeasures also gain prominence. PQC primitives have new leakage profiles and you need to design constant time, masked or otherwise hardened implementations rather than rely on purely software mitigations. Finally, verification and compliance work will expand. Engineers should adopt PQC-ready testbenches which integrate NIST reference vectors and measure PQC-specific metrics such as key-generation cost, encapsulation and decapsulation latency, signer and verifier throughput, and memory footprint. These measurements will inform procurement tradeoffs and system-level decisions, and will be necessary to demonstrate interoperability with software stacks and conformance to evolving standards.
Start hybrid, migrate later We’ve been talking about the need to introduce PQC now, right now, but it’s the longterm destination, not the one we want to reach immediately. Instead, full PQC adoption should be gradual. For FPGA engineers, the need now is hybrid cryptography with protocols that combine classical RSA/ECC with PQC algorithms in the same handshake or signature. This dual‑mode approach ensures backward compatibility with existing infrastructures while introducing quantum‑safe primitives. It also buys time for toolchains, IP libraries, and hardware modules to mature. FPGA crypto subsystems should therefore be designed to support both classical and PQC blocks behind modular interfaces, so that firmware and secure boot routines can switch algorithms without requiring a wholesale redesign. Over the next decade, hybrid deployments will dominate secure boot, firmware signing, and key exchange, before systems transition fully to PQC standards.
01100101110001010101111010.012 0.2349 0.6725 0.2 0.9 0.112 0.145 0.133 0.2 0.6725 0.211 0.987 0.323 0.124 10
9 4
Post-quantum cryptography
Practical steps to stay ahead
CONCLUSION
FPGA engineers should begin by logging where public-key cryptography is used in products and assessing which secrets or communications are long-lived. For designs that require PQC acceleration, prioritize building modular, synthesizable IP for the common computational kernels identified across the standards. Think Number‑Theoretic Transform (NTT) engines and polynomial multipliers, optimized hash and Pseudo‑Random Number Generator (PRNG) units, and configurable memory architectures for large coefficient arrays. Boards and FPGA selections should reserve BRAM or Ultra RAM (URAM) capacity and plan Direct Memory Access (DMA) bandwidth for moving larger keys and ciphertexts efficiently. Security engineering must be integrated from the start, selecting and implementing side-channel countermeasures that match the chosen PQC primitive, and developing test harnesses that exercise algorithmic edge cases and validate against official test vectors. Finally, plan for field upgradeability so algorithm swaps and parameter updates can be delivered securely over the product lifecycle rather than requiring hardware recalls.
NIST’s post-quantum standards mark a major, practical step toward preserving cryptographic security when quantum computation moves from Shor’s algorithm to everyone’s in-box. For FPGA engineers this transition is both an opportunity and a challenge. While FPGAs offer excellent acceleration for many PQC kernels, successful adoption requires rethinking memory budgets, adding new arithmetic IP, strengthening sidechannel protections and designing for crypto-agility and lifecycle updates. Teams that invest now in modular, parameterized implementations and in measurement-driven design will be best positioned to meet the NIST timelines and deliver post-quantum secure FPGAs.
SECURITY IP CORES FOR FPGA-BASED SYSTEMS
Designed in-house. Ready for tomorrow. • Quantum-secure and classical protection • Rapid time-to-market and easy integration • No hidden CPUs or software
Trusted in aerospace, defence, telecom & industry. 11
www.xiphera.com
Logic folding
How logic folding boosts FPGA speeds and reduces footprints
Adam Taylor, Publisher & Founder/Principal Consultant, Adiuvo Engineering & Training
At the design stage, FPGA engineers face several core challenges, from achieving timing and power closure to minimizing device size and optimizing resource utilization. One option is to leverage the higher-performance logic provided by the AMD Spartan™ UltraScale+™ FPGA to implement designs in a smaller logic footprint while still achieving the desired performance. This uses a technique called logic folding, which collapses the required logic resources due to the device’s highspeed fabric. Logic folding is effectively temporal multiplexing enabled by higher operating frequencies. The high-performance fabric means we can reduce bus widths while still achieving the processing throughput required. For example, an existing design might leverage a 32-bit bus operating at 200 MHz, whereas we can leverage a 16-bit bus at 400 MHz. To get the best performance from any AMD FPGA, the UltraFast™ design methodology is a great starting point.
12
It ensures the developed Register Transfer Level (RTL) can be correctly mapped to the target FPGA resources. When performing logic folding, especially in an existing design, it is a good idea to ensure the RTL aligns with the UltraFast methodology requirements prior to making any detailed code updates.
Key elements of the UltraFast design methodology include: Minimize reset use – When possible, minimize the use of resets because, for an SRAM-based FPGA, a global set / reset (GSR) is performed at the end of the configuration. This means variable initialization within the RTL will be implemented when the bit file is generated. The use of resets prevents the synthesis tool from being able to leverage the Shift Register Lookup Table (SRL) as storage elements. Reset style – If resets are necessary, it is preferable to use a synchronous reset, not an asynchronous reset. Register I/O – Ensure when interfacing with hard macros, such as block RAMs and Digital Signal Processing (DSP) blocks, to implement the necessary pre- and post-registers. For example, use two pre-adder registers for DSP48E2 alignment, and registered inputs and outputs for block RAM.
Spartan UltraScale+ architecture Based on a 16 nm process node, the UltraScale architecture blends logic density, deterministic timing, high-speed connectivity, and memory bandwidth in a scalable platform. The Spartan UltraScale+ family extends these capabilities to cost-optimized applications with high I/O counts, transceiver support, and hard IP for an LPDDR5-class memory interface, while maintaining low power consumption and offering long product lifecycles. The architecture is built around high-performance, programmable logic coupled with embedded memory, DSP resources, and high-speed I/O. At its core, it uses configurable logic blocks (CLBs) that contain 6-input Look-Up Tables (LUTs), fast carry chains, and flip-flops, enabling efficient implementation of both combinational and sequential logic. Spartan UltraScale+ devices incorporate several tiers of memory, including 36 Kb block RAM with builtin ECC and FIFO modes, 288 Kb UltraRAM blocks in larger device sizes offering deeper, more powerefficient storage, and distributed RAM in the LUTs for fast, localized memory. DSP48E2 slices deliver compute capability with 27×18 multipliers, pre/post adders, accumulators, and wide XOR logic, allowing the architecture to meet demanding signal processing workloads. High-speed serial connectivity is delivered through GTH transceivers, which operate at up to 16.3 Gb/s and appear in devices larger than the Spartan UltraScale+ SU35P devices.
These transceivers include sophisticated equalization, pre-emphasis, and adaptive receiver features designed to support long PCB traces, backplanes, and modern high-speed protocols. Devices with transceivers also provide hard IP for PCIe® Gen4 blocks. Clocking resources are organized into clock management tiles, each containing a MixedMode Clock Manager (MMCM) and two PhaseLocked Loops (PLLs). Global and regional clock routing networks ensure low skew and flexible distribution. The architecture’s segmented clocking scheme reduces power and improves determinism, while the ability to derive clocks from PLLs, MMCMs, or transceiver outputs provides system-level flexibility. Integration between the clocking network and memory interfaces ensures stable timing for Double Data Rate (DDR) memory and high-performance physical layers (PHYs). Memory interfaces are another major strength of the UltraScale+ portfolio. While traditional DDR4 interfaces remain widely supported, Spartan UltraScale+ devices introduce hard IP controllers for LPDDR4X and LPDDR5, delivering data rates up to 4266 Mb/s via the XP5IO banks. These hard IP controllers reduce both logic utilization and power, and simplify PCB design compared to soft memory controllers. Devices larger than the SU35P in the Spartan UltraScale+ range support PCIe Gen4 through integrated hard blocks, which also enable a reduction in power dissipation compared to soft implementations.
Getting started with logic folding
Pipelining – Provide pipeline registers at the beginning or end of the datapath module. This allows the synthesis tool to perform retiming, inserting registers in the datapath as needed to achieve the highest performance.
When it comes to implementing designs on a Spartan UltraScale+ FPGA using logic folding, we can leverage the capabilities of the programmable logic to implement a solution that is not only faster but also operates at a higher clock frequency.
Placement wrapper – Create wrappers around a module to register its I/O. This provides the implementation tool with the ability to locate the registers as needed in the routing path to improve timing closure.
A good way to demonstrate this is to migrate a module from a Spartan 7 XC7S100 in a -2 speed grade to a Spartan UltraScale+ XCSU35P FPGA, also in a -2 speed grade with a voltage supply of Vnom. The module combines several features common to FPGA applications: a CIC filter implementing a rolling average, high-speed AXI-Stream interfaces, and Cyclic Redundancy Check (CRC) protection.
Write flexible code – Leverage generics and parameters to enable the module to be easily configurable for different bus widths, etc. Leverage capabilities – Take advantage of the device features, such as the Single Instruction, Multiple Data (SIMD) pattern detection capabilities of the DSP48E2.
13
Conversion rules To ensure a conversion is as representative as possible, the following rules are used:
Figure 1: The architecture of the Spartan 7 FPGA
The architecture of the original Spartan 7 FPGA implementation is outlined above. The design is based around a 32-bit streaming architecture. When implemented within the Spartan 7 FPGA, the maximum speed achieved was 215 MHz. Resource requirements after implementation were 742 LUTs and 487 FFs. This design utilization and clock rate, therefore, define the baseline performance of the system, and the floorplan of the device when implemented is shown in figure 2.
Determining maximum frequency
Both devices will be targeted at the same speed grade Synthesis and Implementation settings will be identical in both projects Performance shall be the same for both modules, e.g., dynamic range Latency through the module may be modified to enable a higher clock rate
Spartan UltraScale+ FPGA migration The first step is creating a project targeting the Spartan UltraScale+ FPGA with the unchanged original module. This allows validation of architectural compatibility and identifies any macros or primitives requiring updates. This also lets us determine the new design’s maximum frequency. In the Spartan UltraScale+ device, the original module requires 715 LUTs and 487 FFs and operates at 308 MHz, nearly a 100 MHz improvement. This comes from the faster CLB architecture, improved routing, and more efficient DSP/block RAM interfacing. However, this is not the full benefit. By re-architecting the logic to exploit higher fabric speed, we can dramatically improve both Fmax and resource utilization.
As we begin the conversion to the Spartan UltraScale+ FPGA, it is key to understand the maximum clock rate we can achieve with our design in the target device. The best way to determine this is to implement the design while iteratively increasing the target clock frequency until a small negative slack is reported (WNS<0). We can, therefore, calculate the maximum clock frequency using the equation: FMAX (MHz) = max(1000/(T - WNS))
Re-architected folded implementation
Figure 2: The floorplan of the Spartan 7 FPGA
Figure 3 shows we get a speed increase of nearly 100 MHz for the same module just by changing to a later generation of the device. However, this is not the maximum benefit we can get from changing a generation. By re-architecting the module slightly, we can achieve an implementation that operates at a higher clock rate and significantly reduces the logic footprint. The most efficient way to achieve a speed update is to reduce the size of the AXI streaming interconnect and the CRC associated with it. Within the original design, the interface used a 32-bit AXI stream and a 32-bit CRC in the Tuser field to protect the input and output data.
14
Logic folding
Figure 3: The architectural compatibility of the Spartan UltraScale+ FPGA
Figure 4: Re-architecting the design reduces resource requirements
Within the Spartan 7 device, the 32-bit implementation is required to achieve a throughput of 6,880 Mb/s when operating at the 215 MHz clock.
However, when comparing the power dissipation in the original design against the folded design, we do see a slight reduction in power.
Re-architecting the design to use a 16-bit interface and an 8-bit CRC in the Tuser field enables the clock frequency to be increased to 450 MHz. This is more than double the original clock frequency within the Spartan 7 device. This re-architecting means we can leave the 32 bits of the CIC filter unchanged, which ensures the migrated design maintains the dynamic range of the filter. Of course, this approach requires two 16-bit transfers over the AXI Stream to receive the 32 bits required to be passed through the CIC filter. The 32 bits are retained at the filter to keep the dynamic range. With the clock rate of the Spartan UltraScale+ device being over twice that achieved in the Spartan 7 FPGA, the algorithm is still capable of achieving the required throughput, reaching 7200 Mb/s, a comfortable margin over requirements. This re-architecting of the design reduces the resource requirements significantly, to 531 LUTs and 497 FFs (see figure 4). This equates to a small increase in registers but leads to a significant decrease in the number of LUTs required with a reduction of approximately 30% against the base case. The floor plan of the device in figure 5 shows the reduction in area, clearly from the baseline.
The original design required 0.034 watts to power the design, while the updated folded design requires 0.03 watts of power. While small, this demonstrates the folding technique does not adversely affect the demands of existing power architectures. Of course, if folding the design enables us to deploy the design in a small device, there will be a significant reduction in static power required.
Complexities of rearchitecting This reduction in area over the baseline design required a modification to the original RTL. The main modification of this code was to change the AXI Data and User widths to accommodate a narrower bus. Because this code follows the UltraFast design methodology guidelines, the RTL was parameterized. Therefore, the initial change in bus width was simply the change to a parameter or generic. The largest change in this re-architecting was the creation of an 8-bit CRC in place of the original 32-bit CRC. This redevelopment required the design and test of a new module and test benching. Because CRC-8 can be sourced from internal IP libraries, open-source IP, or generated using AI tools, the implementation effort was minimal. The 8-bit CRC block was completed within a few hours. Figure 5: The floorplan of the Spartan UltraScale+ FPGA
Power differences Along with the reduction in resource requirements, the power dissipation of the device will also be impacted by the design changes. There are several factors which come into play when considering the power dissipation. While the advanced fabric enables a more power-efficient solution, we are clocking the folded design at a much faster frequency. When folding a design, we therefore should not expect to see such a dramatic change in the power dissipation as we did with the logic resources.
15
Logic folding Of course, each design is unique and requires careful consideration as to the level of rearchitecting required. However, this example has shown a significant reduction can be achieved by leveraging higher performance, narrower data buses, something which can be achieved easily with parameterization. More complex applications may require more in depth re-architecting, though as this article has demonstrated, the rewards for the time invested can be significant.
Logic folding summary and generic rules Logic folding is a technique where a designer reduces the width, parallelism, or resource duplication in a digital design and compensates by increasing the operating frequency or reusing hardware across multiple cycles. Modern FPGAs with higher-performance logic fabric make this trade-off attractive because time-multiplexed designs can achieve equivalent throughput while using fewer logic resources. Folding enables smaller devices, lower power consumption, and simpler routing, while still meeting system-level performance requirements.
Summary
Define the performance requirements Identify required throughput, latency, precision, and any system-level constraints. Determine which metrics must remain unchanged when folding. Determine the available clock frequency headroom Estimate or measure the maximum clock rate achievable in the target device. Compare this to the original design’s clock rate to understand how much folding is possible. Reduce parallelism or data width Narrow datapaths reduce vector widths, or reuse operators instead of duplicating them. Ensure the new multi-cycle structure still meets throughput requirements given the increased clock rate. Time-multiplex functional units Replace replicated arithmetic or logic blocks with shared units used across several cycles. Add scheduling or control logic to coordinate multicycle operations.
Logic folding is a technique that leverages the higher performance of newer FPGA architectures, such as in the AMD Spartan UltraScale+ family, to reduce logic footprint while still meeting (or exceeding) throughput requirements. Because Spartan UltraScale+ devices can operate at significantly higher clock frequencies than earlier families like Spartan 7, designers can reduce bus widths and time-multiplex operations without sacrificing data rate or algorithmic performance. This enables smaller, faster, more resourceefficient implementations. The process typically involves understanding the original design’s throughput, determining the maximum achievable clock frequency in the new device, and re-architecting datapaths (such as AXI-Stream widths) to match performance requirements, while minimizing LUT and routing resources. When combined with the UltraFast design methodology, logic folding results in cleaner timing, reduced area, and improved robustness.
Add or increase pipelining Insert pipeline registers to support higher operating frequencies. Ensure pipeline depth does not violate system latency constraints.
Consider memory and buffering needs Multi-cycle operations often require additional temporary storage or buffering. Ensure that memory access patterns still meet timing and bandwidth needs.
Preserve algorithmic accuracy Maintain internal precision, state widths, and accumulator sizes unless mathematically safe to change. Folding should alter when operations occur, not what they compute.
Revalidate timing, throughput, and functional behavior Run functional simulations and timing analysis. Confirm that the folded design meets timing, achieves required throughput, and behaves identically to the original algorithm.
Update control and handshaking Modify Finite State Machines (FSMs), enables, and valid/ready mechanisms to accommodate multi-cycle behavior. Ensure downstream modules tolerate any added per-sample latency.
Document the folding strategy clearly Capture the updated cycle counts, interface behavior, data formatting, and multi-cycle timing. Provide clear guidance for integration and future maintenance.
16
Cost effective verification solutions built for modern design challenges RTL linting Reset & CDC analysis High-performance RTL simulation Hardware-assisted verification
TR:
Discover the new cutting-edge hardware and software verification solutions for ASIC, FPGA and IP development teams. Visit bluepearlsolutions.com
17
Why PCB testing needs to shift left
Why PCB testing needs to
shift left
to the design stage Matthew Holder, Hardware Engineer, XJTAG Rather than being designed in from the start, PCB testing is often something bolted on at the end of a design as an afterthought. Yet too many PCB faults are only discovered at the release stage or in the field. By then the cost of correcting those faults is much higher, involving rework, scrapped boards, missed deadlines, and damaged customer trust.
Conversely, designing with testability in mind from the beginning reduces these risks by turning otherwise hidden problems into predictable checks that are simple to automate. Often called shift-left testing, it moves test considerations back along the development pipeline into the schematic, PCB layout, and design reviews rather than waiting until prototypes or production. This early test planning lets you ensure nets are accessible, choose appropriate test connectors, reserve test points, and partition subsystems so faults can be isolated quickly. It also makes production tests faster and more deterministic, shortens debug time during bring-up, and gives support engineers clear diagnostics when failures occur in the field. A shift-left mindset doesn’t prescribe a single method, however. Instead, it means embedding test requirements into the design so you can later apply the right mix of structural, functional, and automated tests. By integrating testing into the design process, failures can be traced back to their root causes rather than leaving you chasing symptoms later. Put simply, the earlier you plan how to test a board, the fewer surprises you’ll face down the line, reducing testing costs, shortening the time to market, improving product quality, and enhancing fault detection at every stage.
18
Practical ways to track test requirements On the production side, test steps should be embedded in the manufacturing bill of materials (MBOM) or shop floor work orders, rather than in engineering only files. This ensures that test routing and operations are documented where process details belong, making test execution part of the standard workflow.
Ensuring that test requirements remain visible and aligned with design starts with a disciplined approach to documentation and configuration control. One effective method is to maintain dedicated test plans and procedures in version controlled repositories. By keeping project files, fixture lists, and pass/fail criteria with a clear document ID and revision history, teams can avoid ambiguity and ensure that everyone is working from the same baseline.
Integration with test management or Manufacturing Execution System (MES) tools adds another layer of traceability. Assigning test IDs, controlling which version each board should run, and capturing results centrally allows analytics to be performed across the product lifecycle, strengthening confidence in the data.
Version control also plays a critical role in managing test repositories. Storing scripts, firmware, and project data in source control keeps test assets synchronized with board revisions, reducing the risk of mismatches between hardware and software.
Finally, linking test issues to Engineering Change Orders (ECOs) ensures that lessons learned are not lost. By tying failures to corrective actions, design rules and future test plans are updated in response to real-world experience.
Equally important is linking test artefacts directly within product lifecycle management (PLM) systems. When test plans and results are attached to the product record, design revisions and test requirements remain connected and auditable, preventing the disconnect that often arises between engineering and manufacturing.
These methods keep design and procurement artefacts clean while ensuring test requirements remain discoverable, traceable, and under configuration control – which is the whole point of shifting test planning left: including it from the start.
For design reviews and Design for Testability (DFT) sign-off, many teams rely on a testability matrix. This structured mapping of modules, nets, and components to expected coverage and methods, like associating a power rail with ICT or a bus with boundary-scan, makes test coverage explicit and easy to communicate.
Traditional Testing - Starts at the Release Stage Development Stage REQUIREMENTS
Release Stage
DESIGN
PROTOTYPE & PRODUCTION
DEVELOPMENT
Shift Left Testing - Starts at the Development Stage Development Stage REQUIREMENTS
DESIGN
Release Stage DEVELOPMENT
PROTOTYPE & PRODUCTION
Figure 1: Moving to a Shift-Left testing approach
19
TIME SAVED
Common pitfalls: Case studies from the field The value of shift-left test design becomes obvious when you look at common failure modes. The following real-world examples show how seemingly minor oversights can block testing and drive up costs across manufacturing and field service.
Missing boundary-scan access for power sequencing
Symptom:
During manufacturing test, the board fails to power up consistently. As a result, the JTAG boundary scan chain cannot initialize, and downstream structural tests are blocked.
Detection:
During bring-up, structural JTAG tests fail to run. Investigation shows the device never enters boundary-scan mode because its rails are not enabled in the correct order. Manual jumper wiring of the rails allows partial testing, confirming the sequencing dependency.
Root cause:
Power sequencing for key rails is controlled by a device that itself requires boundary-scan mode for configuration. However, no provision is made in the design to allow external control or override of the power supplies. As a result, the scan chain remains inaccessible at startup.
Design lesson:
Always design an external override or bypass path for power sequencing so that test equipment can force rails on independently. Boundary-scan access is only useful if the scan chain can be reached — power domains and enable signals must be testable without relying on the devices under test.
Compliance pins not considered for boundary-scan entry Root cause: The JTAG chain does not initialize, and devices are reported as absent or stuck in BYPASS.
Compliance pins on a key IC are neither connected out to test pads nor tied to the correct logic levels, preventing the device from entering boundary-scan mode. With one device stuck, the whole chain becomes inaccessible.
Detection:
Design lesson:
Symptom:
During prototype testing, boundary-scan does not detect several expected devices. Pintracing in the schematic reveals that some compliance pins are pulled to incorrect logic levels and lack any board-level access to override them during test.
Designs must treat compliance pin handling as a first-class DFT requirement. Their expected states should be documented, correct pull-up/pull-down resistors must be in place, and accessible pads or jumpers should be provided where overrides may be required. Missing these simple connections can block the entire JTAG chain.
20
Why PCB testing needs to shift left
Inaccessible TAP location on server board
Conclusion Symptom:
Root cause:
Detection:
Design lesson:
Field diagnostics require disassembly of production servers, adding time to troubleshooting and increasing service cost.
Service technicians report extended turnaround times for boundary-scan diagnostics. The mean time to repair (MTTR) is significantly higher than forecast because even simple checks require a major system teardown.
The Test Access Port (TAP) is placed mid-board for convenience during layout but is not routed to an accessible connector on the front or rear panel. This means every in-field boundaryscan or firmware test requires removing covers, boards, and cabling.
Designs must consider service and fieldtest access during layout. TAP headers should be placed at the chassis edge, or JTAG signals brought to accessible connectors, even if not populated in production. The minor cost of adding a connector footprint is far outweighed by the recurring operational cost of inaccessible diagnostics.
Shifting test planning left isn’t just about avoiding rework, it’s about building confidence into every stage of the product lifecycle. The cost of a few extra pads, connectors, or overrides is tiny compared to the recurring cost of inaccessible or incomplete testing later. By embedding testability at the design stage, you can reduce costs, accelerate schedules, and ensure more reliable products reach your customers.
Link Issues to ECOs Track failures and tie fixes to engineering change orders.
Integrate with Tools Use test management and MES tools for control and analytics.
Version-Control Repositories Store test data and scripts in source control.
Create Testability Matrix Map modules, nets, components to test coverage and methods.
Include Test Steps Document test routing in the MBOM or work orders
Link Test Artefacts Attach test plans and results to the product record for traceability.
Maintain Test Plans Keep test procedures and criteria in version-controlled documents. Figure 2: Achieving Comprehensive PCB Testing
21
Why you need Hog
Meet
H0G
Introducing the open-source system for handling your HDL code on Git Dr Francesco Gonnella, Birmingham University
Managing HDL code for FPGAs on Git needs particular attention, mainly because the tools to develop it are proprietary and they tend not to be Git friendly. As a result, when you copy HDL projects to Git and try to recreate them, questions often arise like “I haven’t touched anything, so why doesn’t it work?”.
22
Time to think about Hog What you want instead is certainty. You want to be absolutely sure that when people clone your repository from Git, they get exactly what they expect, without hidden changes. And the reverse is just as important: knowing what’s actually running on the device. A lab build may be working today, but tomorrow it fails and no one is sure which version it was. You need a way to guarantee traceability, while keeping overhead as low as possible and letting engineers work in the way they’re used to.
All of the problems I’ve talked about are why we created Hog at CERN. We started by using a Tcll script to recreate projects locally but then we started looking for a better approach. One where we could automatically go from repository 4 files 4 build 4 chip in a reproducible way, and back again from the binary file on the FPGA to the exact source that produced it. We also wanted registers in the device to embed identifiers that link a binary back to its repository version, so we always know what’s running and where to start if changes are needed.
Why not just commit everything to Git?
Welcome to Hog which does all of this by embedding the Git commit Simple Hashing Algorithm (SHA) into the binary file and checking that nothing has been touched.
You could, but even Xilinx recommends against committing project files directly. Instead, you need to create a Tool Command Language (Tcl) script that recreates the project locally and many engineers already have custom scripts to do this. But what’s in the project files and why can’t we just put them on Git?
Every binary file produced is also traceable to its source commit. In Continuous Integration (CI) pipelines this happens automatically, but even local builds are checked. If files were changed without being committed, Hog will detect it and prevent the binary from being claimed as clean. In this way, reproducibility is guaranteed. If you commit, everything is fine. If you don’t, Hog will flag it. And importantly, it requires no extra installations. Engineers often ask, “Why do I need Python or other tools if I only use Vivado?” With Hog, you don’t. All you need is Git and your FPGA toolchain. It runs inside the Tcl shell of Vivado and other IDEs like Altera, Libero, and Diamond. Another principle is ease of collaboration: when you invite someone to join a project, they shouldn’t spend days setting up. They should be able to clone the repository, run the script, and start working immediately.
The problem is hard-coded paths which you’ll see in files from vendors like AMD, Altera, Lattice and Microchip. Sometimes they’re self-repairing and Vivado, for example, will fix them and in doing so change them. We don’t want that. We want nothing to be touched and the repository to remain pristine by committing to Git only those files that are needed, not everything from the project folder, which makes it easier to track and debug. The same principle applies to IP cores. Xilinx IP, for example, generates many files that shouldn’t be committed and sometimes can’t be because they require licenses to regenerate. We don’t want this. We want a way to automatically ignore self-generated files and ensure they can be regenerated consistently.
What you need to use Hog At its core, Hog is just a set of Tcl scripts less than a megabyte in size and a methodology, a set of prescriptions to follow. You only need what you would need anyway: your source files to build the project (VHDL, Verilog, IPs, and constraints), plus a list of the files that are needed for a project. Since a repository may contain multiple projects or extra files, you also need to list the ones that you want to use. These lists are simple text files. They can also include properties such as “VHDL 2008” or “not used in implementation,” along with projectwide settings like the target FPGA or maximum fan-out. Hog defines where this information should be stored, and ensures it is committed to the repository.
The second problem is certifying that files are truly untouched compared to the repository. Even with good intentions, people say, “I only changed this one thing so it shouldn’t matter.” But it does. That’s why the process must be automatic.
23
The main feature of Hog is that it runs inside the Tcl shell of the FPGA tools you already use, so there are no extra requirements. You simply add Hog as a Git submodule, no installation is needed, and collaborators don’t need to install anything either. If Hog is updated, you choose when to update your submodule, so your project isn’t broken by changes. This guarantees reproducibility by controlling the files, and traceability by marking each binary with the SHA file that you can follow back to the repository. Equally important, Hog provides ready-made YAML scripts for CI with GitLab or GitHub Actions. With the FPGA tools available on the CI machine, setting up CI is just a matter of writing a few lines of YAML. In a typical installation, Hog runs on a standard PC with only Vivado installed. You simply clone the repository, enter it, and use the ./Hog/Do VIEW command to see all the project files.
Figure 1: Cloning and viewing project files using Hog on a standard PC
24
If you want to create the project, you then use the ./Hog/Do CREATE command and Hog rebuilds it locally from the list files and you can then open it in Vivado and work in the GUI as usual. It also works on Linux and Windows, since it runs inside the IDE itself. Git works locally, of course, but you can also use GitHub or GitLab. Hog mainly supports AMD tools, but also works with Altera, Lattice, and Microchip. In some cases you need additional installs because those tools don’t embed Tcl libraries, but the principle is the same. Finally, for the repository structure Hog requires a particular set of folders. The HDL code itself can be placed wherever you like, and each project sits under a top folder, with subfolders containing the list files and configuration. The idea is to make the workflow correct and reproducible.
Why you need Hog
}
}
IDE to use
Project Variables
Project GENERICS
Figure 2: A sample hog.conf script showing the basic variables required to generate an HDL project
How to get started with Hog
Clicking the update button will refresh the hog. conf, after which you must commit it. If you don’t, Hog will produce a ‘dirty’ bitstream and flag it with a warning. The rule is simple: commit before you build. Even for test runs, committing takes only a second, and it guarantees traceability because Git generates a SHA that links the binary back to the exact source.
Hog is free to use and open-source, and there are two ways to get started: writing the text files by hand or automatically producing them using the Hog buttons in the Vivado GUI. If you create the txt files by hand, you need to list the files used in your project in “.src” text files contained in ./Top/<my_project>/list/. To specify your project’s property you need to create a TOML-style configuration file called ./Top/<my_ project>/hog.conf. This is expected to contain a few basic variables with the information needed to build your project. You can define properties such as the target FPGA and, at its simplest, you only need to specify the device, although you can add other settings to match your project (see Figure 2).I
If you forget to commit, the binary will contain uncommitted changes that cannot be traced. Hog does save a diff file against the repository, so recovery is technically possible, but the principle is clear: always commit before synthesis. The same applies when adding new files or making changes to the project in the GUI: commit them manually or using the Hog buttons.
If you’re using Vivado (and, increasingly, other supported tools) and you don’t want to write the list files and hog.conf yourself, Hog can generate them automatically within the IDE. I still recommend keeping these files updated by hand, but the automation makes setup easier.
In the CI workflow that is automatically provided, the builds only start when a merge request is opened and marked as non-draft. This avoids triggering CI on every push. Once the merge request is approved by the librarian, another pipeline begins and copies the files, optionally renames them with the version, and creates a tag that matches the version embedded in the bitstream.
When Hog creates a project, it integrates automatically and runs Tcl scripts hooked into all the build stages, for example, pre-synthesis, preimplementation, post-bitstream, etc. These scripts interact with your repository, checking whether it is clean or if anything has changed. If changes are detected, Hog issues a critical warning. It also performs other tasks, such as copying files into a versioned location. The process is automatic and guarantees nothing is touched without being tracked.
The numerical version is handled automatically in CI pipelines using the standard Major.Minor.Patch (M.m.p) format. As all the workflow needs to be reproducible, special branch names can be used to increase the M or m version number. What makes numerical versioning of FPGA gateware more difficult than software, is that the version number is required before synthesis. This is because the version number is embedded into the build itself. Hog solves this problem, calculating the future version in advance by increasing the proper value of M, m, or p.
For example, if you go to your project and select Segmented Configuration in Vivado and it isn’t listed in your hog.conf, Hog will compare the project against the configuration and warn you of the mismatch.
25
Why you need Hog
Where to go next Hog is already used by several academic and industrial projects including ATLAS, CMS Phase-I and Phase-II upgrades, ESRF, GAPS, FOOT, NASDAQ and NOKIA. It’s available at gitlab.com/hog-cern/hog and the documentation can be found at cern.ch/hog. If you’d like to try it:
Hog is completely free and developed mainly as a pet project by Davide Cieri and Francesco Gonnella. If you use it, please cite our paper in your articles and proceedings. Here is the citation in BibTeX format.
> git clone --recursive https://gitlab.com/ hog-cern/hog-examples.git > cd hog-examples > ./Hog/Do CREATE vivado/fifo > vivado ./Projects/vivado/fifo/fifo.xpr
Learn more by watching the original FPGA Horizons Conference talk online This article is an edited narrative of a talk at the FPGA Horizons Conference in London in October 2025. If you’d like to see Dr Francesco Gonnella from Birmingham University presenting the session and discover a lot more about Hog – A system to handle your HDL code on Git, you can access a recording of the talk online. It’s available at fpgahorizons.com/london-25/ london-25-talks/ along with the 20 other fascinating and informative talks given at the conference.
26
APPLICATIONS
DIGITALLY AGILE RADAR WITH THE ANDROMEDA XRU50 RFSOC Modern radar systems require higher bandwidth, faster response, and precise synchronization. Learn how direct RF sampling and multi-device synchronization (MDS) simplify phasedarray architectures while improving accuracy, reducing latency, and minimizing system complexity.
https://www.enclustra.com/en/products/system-on-chip-modules/andromeda-xru50/
Read the Full Radar Article HERE A qr code on a white background AI-generated content may be incorrect.
Connect with us! https://www.enclustra.com/en/home/
Visit Our Website www.enclustra.com
https://www.linkedin.com/company/enclustra
https://x.com/enclustra
https://www.youtube.com/@Enclustra
Building resilient FPGAs and SoCs for every radiation environment
From space to server:
Building resilient FPGAs and SoCs for every radiation environment Dr. Pierre Maillard, Radiation Effects and RAS Solution Team Lead, AMD Radiation-induced single-event effects (SEEs), potentially leading to loss of functionality, high current states, etc., pose major reliability challenges for all electronic systems, from ground-based to space platforms. Modern FPGAs and SoCs are increasingly susceptible due to higher integration densities, smaller feature sizes, and growing on-chip resources. As device complexity and transistor counts rise, so does the likelihood of radiation-induced soft errors, even at ground level, from neutron particles [1]. To ensure reliable operation, robust SEE mitigation at both hardware and system levels is essential to reduce failures and system downtime. However, these protections must be validated through accelerated radiation testing (e.g., heavy-ion or proton beams) to confirm real-world effectiveness. Such tests provide critical insights into system resilience, guiding the design of nextgeneration radiation-tolerant FPGAs, SoCs, and electronics [1 - 3].
28
In this article and as an example, we share results from the AMD Versal™ portfolio, including a radiation-tolerant AI-ML platform that has 1 functional interrupt every 60,000 years per AI Engine tile at 40 kft relative to NYC sea level and reduces datapath errors by 3.7X versus the baseline platform [3]. We also highlight a processing system that shows improved single-event resilience by architecture and has 1 event every 6 years in GEO orbit [2].
Why radiation still matters for modern FPGAs and SoCs From satellites orbiting Earth to high-altitude avionics and autonomous cars at sea level, radiation-induced errors are unavoidable. As process geometries shrink, the susceptibility of SRAM-based FPGAs and adaptive SoCs to single-event effects becomes more pronounced. These events can flip bits, freeze state machines, or disrupt deep learning inference pipelines. Yet, with proper design, detection, and mitigation, radiation effects can be managed rather than feared. For all environments, terrestrial and space designs must consider radiation. Cosmic particles can strike any device at any altitude, and as memory density increases, the statistical likelihood of such events grows.
Real-world examples of radiation-induced events Over the past two decades, several high-profile incidents have underscored the impact of radiation on electronics. In October 2025, an Airbus A320 flying from Cancún to Newark, NJ suddenly lost altitude due to a radiation-induced computer malfunction. This event led to the grounding of more than 6,000 Airbus aircraft, one of the largest aviation recalls to date, thus causing major travel disruptions [4]. Similarly, in 2021, the Crew Dragon Resilience capsule triggered false emergency alarms aboard the International Space Station, likely caused by a single-event upset (SEU) in an avionics power unit [5]. Even terrestrial systems are not immune: a 2003 voting machine error in Belgium added 4,096 phantom votes, later attributed to a suspected cosmic-ray bit flip [6]. These examples illustrate that radiation-induced upsets can affect systems across environments, from orbit to the ground, and reinforce the need for robust mitigation and validation strategies.
The SEE Taxonomy: Speaking a common language
SET (Single-Event Transient): A temporary voltage glitch induced in combinational logic, which may propagate through sequential elements and manifest as an SEU if latched. SEFI (Single-Event Functional Interrupt): A more severe, system-level event that disrupts normal functionality and requires a reset or user intervention to recover. SEL (Single-Event Latchup): A highcurrent condition caused by the activation of parasitic structures within the device, potentially leading to permanent damage if not mitigated. Radiation-effects engineers, reliability experts, and functional safety architects often use different terminologies to describe similar fault mechanisms. To align these perspectives, AMD has developed a unified classification framework that integrates traditional radiation-effects terminology with Reliability, Availability, and Serviceability (RAS) concepts and Functional Safety (FuSa) metrics. By correlating particleinduced effects with failure modes, mean time to failure (MTTF), and probabilistic safety metrics, this framework supports the development of robust, strategically resilient FPGA and SoC products that address the diverse reliability and safety requirements across markets such as automotive, aerospace, industrial, and data center applications.
Three-layer mitigation: From silicon to system Our SEE Mitigation Strategy is structured across three complementary layers: Silicon level: At the silicon level, AMD employs a combination of custom SEU-hardened memory cells, strategic bit interleaving, memory errorcorrecting codes (ECC), and optimized layout rules to minimize susceptibility to ion strikes. These techniques, together with selective use of triple modular redundancy (TMR), reduce the likelihood of severe events such as SEFIs and SELs.
Single-event effects are disturbances in the normal operation of a circuit or system caused by the interaction of a charged particle with a sensitive node within the device. Depending on their impact and recoverability, SEEs can manifest as soft errors, transient, non-destructive faults, or as hard errors, which may lead to permanent damage. The most common SEE categories include: SEU (Single-Event Upset): A soft error resulting from a bit flip in a memory cell, register, or logic element. SEUs are typically correctable through error-detection and correction (EDAC) mechanisms.
29
Device level: Integrated software-based safety mechanisms enhance fault detection and correction capabilities. For example, configuration memory (CRAM) scrubbing and built-in self test (BIST) are used to detect, correct, and isolate soft errors in real time, ensuring system integrity and reducing the risk of error accumulation.
System level: At the system level, system designers or end users can implement additional mitigation. Techniques such as external TMR, voting logic, and configuration scrubbing provide higher-level resilience against cumulative or latent upsets, complementing the underlying silicon and device protections. I want to emphasize that effective mitigation begins with understanding what must be protected. CRAM typically dominates overall vulnerability, while compiled memories, registers, and control logic require targeted protection strategies. Continuous scrubbing remains essential to prevent the accumulation of silent errors that could otherwise compromise system reliability.
Data-driven validation: Why accelerated beam testing matters While fault-injection simulations are valuable, they cannot replicate the complex behavior of a charge particle going through silicon, especially multi-cell upset patterns. With uncertainty comes conservativeness and potentially over-engineering for SEE mitigation. This in turn impacts power, performance, and/or area. Therefore, beam testing using sources such as neutrons, protons, and heavy ions remains critical and the gold standard in industry practice for validating and characterizing the single-event response of FPGA and SoC devices. The AMD test methodology evaluates both the silicon architecture and the effectiveness of mitigation strategies under accelerated radiation exposure, following established industry standards such as JESD89, MIL-STD-883, and ISO 26262. Experimental results often reveal gaps between predictive models and real-world behavior, particularly within complex blocks such as highspeed SerDes, processor systems, and AI Engines. These empirical insights are crucial for refining mitigation strategies and optimizing the power, performance, and area (PPA) trade-offs in nextgeneration architectures.
Examples of beam test results: AMD Versal platforms under terrestrial and space radiation Processing System (PS) Since no standardized benchmark currently exists to characterize a processor’s susceptibility to soft errors (e.g., SEU, SET, SEFI), we introduced a new top-down processor validation methodology in the 16 nm UltraScale+™ architecture. This approach leverages the AMD System Validation Tool (SVT), a selfhosted, OS-like validation framework that generates focused random test vectors to thoroughly exercise the full processing system (PS) and associated IPs. SVT provides extensive coverage and was used to demonstrate that the PS fault coverage exceeds 99% [2], including the programmable logic (FPGA) portion of the Versal adaptive SoC. Accelerated radiation testing using heavy-ion and proton beams demonstrated the robustness of this methodology. Results showed that more than 99.99% of cache and RAM SEU events are correctable under both low Earth orbit (LEO) and geostationary Earth orbit (GEO) radiation environments. Furthermore, for the complete PS, the GEO SEFI rate was measured at approximately 1 event every 6 years, while in LEO at 500 km and 51.6° inclination, the PS SEFI rate was observed at roughly 1 event per year as shown in “Heavy-Ion and Proton Evaluation of AMD 7nm Versal™ Multicore Scalar Processing System (PS)”.[2] These results confirm that the SVT-based top-down methodology provides an effective and repeatable framework for quantifying and improving soft-error resilience from both reliability and functional safety perspectives.
AI Engine and ML datapaths
End users are strongly encouraged to validate their own designs (“test as you fly”), since applicationspecific testing provides more representative results than generic vendor-provided estimates, which may not accurately capture unique architectural configurations, workloads, or system environments.
30
Three primary types of radiation-induced error signatures can be observed following a particle strike in a portion of the FPGA or SoC design corresponding to the deep learning model deployment: 1) SEFI, 2) image misclassification errors, and 3) degradation in classification probability/accuracy as shown in “Protons Evaluation of 7nm Versal™ AI Engine (AIE) Based Radiation Tolerant Platform for Deep Learning Applications”. [3] SEFIs are usually manifested as uncorrectable silicon events like system hangs or resets, while the other two types of errors manifest as corruption in the content of the datapath.
Building resilient FPGAs and SoCs for every radiation environment
To minimize uncorrectable SEFI events at the silicon and datapath level, AMD developed a radiation-tolerant platform for deep learning applications for terrestrial and space systems using a combination of a silicon-level solution such as the CRAM scrubber and a fault-aware training (FAT) methodology. The results showed: 1 SEFI per AI Engine every 60,000 years at 40 kft relative to NYC sea level and a 3.7X decrease in datapath error while eliminating errors with >5% probability degradation. These findings underscore that software-level techniques (like FAT) complement hardware mitigations effectively, achieving radiation tolerance without major performance sacrifices. Details implemented and results from the 64 MeV protons test are available in “Protons Evaluation of 7nm Versal™ AI Engine (AIE) Based Radiation Tolerant Platform for Deep Learning Applications”.[3]
References 1. Weber, Cavin; and Maillard, Pierre. “The Bring-Up: AMD in Space (Part 2).” YouTube, 2024. (www.youtube.com/watch?v=QN1W-DaRRQo)
2. Maillard, Pierre; Chen, Yanran Paula; Arver, Jue; Merugu, Venkatesh; Shui, Ava; and Dhavlle, Abhijitt. “Heavy-Ion and Proton Evaluation of AMD 7nm Versal™ Multicore Scalar Processing System (PS).” IEEE, 2023. (ieeexplore.ieee.org/document/10265824)
3. Maillard, Pierre; Dhavlle, Abhijitt; Chen, Yanran Paula; Fraser, Nicholas; Chen, Yushan; and Vacirca, Nicholas. “Protons Evaluation of 7nm Versal™ AI Engine (AIE) Based Radiation Tolerant Platform for Deep Learning Applications.” IEEE, 2024.
Looking forward: Evolving architectures and methodologies As modern architectures integrate increasingly complex functions, such as advanced processing systems, AI Engines, and heterogeneous computing fabrics, radiation-effects mitigation strategies must evolve in parallel. Innovation in this domain must progress at the same pace as advancements in power, performance, and efficiency. To accelerate time to market and ensure mission readiness, beam testing should begin early in the R&D cycle, enabling rapid validation of design assumptions and mitigation techniques. Given the limited availability of traditional radiation test facilities, emerging tools such as laser-based fault injection systems (e.g., NRL and third-party platforms) are becoming valuable complements to conventional proton, neutron, and heavy-ion testing. These techniques provide faster feedback during development while maintaining correlation with physical radiation mechanisms.
(ieeexplore.ieee.org/document/10759203)
4. Baraniuk, Chris. “Bit flips: How cosmic rays grounded a fleet of aircraft.” BBC, 2025. (https://www.bbc.com/future/article/20251201-howcosmic-rays-grounded-thousands-of-aircraft)
5. McGlaun, Shane. “SpaceX Crew Dragon Capsule Attached to the ISS Sounds False Alarms.” SlashGear, 2021. (https://www.slashgear.com/spacex-crew-dragoncapsule-attached-to-the-iss-sounds-falsealarms-27665801/)
6. Nappa, Antonio; Hobbs, Christopher; Lanzi, Andrea. “Deja-Vu: A Glimpse on Radioactive SoftError Consequences on Classical and Quantum Computations.” arXiv:2105.05103, 2021.
Consequently, validation and characterization of next-generation FPGAs and SoCs must continually adapt to increasing architectural complexity. Adopting precise error classification schemes, standardized testing methodologies, and consistent reliability metrics is essential both to developing strategic SEE mitigation solutions and to assessing the suitability of FPGA and SoC devices for specific radiation environments.
(arxiv.org/abs/2105.05103)
31
The UVVM Requirements Traceability Matrix
Simplifying the requirements tracking journey with the
UVVM Requirements Traceability Matrix Espen Tallaksen, CEO, EmLogic
For FPGA engineers, requirements tracking often feels like a balancing act between precision and practicality.
So, what is requirements tracking and how can it be made simpler? At its most basic, it’s the process of ensuring that all requirements in a specification have been properly verified. A prerequisite is, of course, to have clearly specified requirements. For a very simple motor controller, some of the requirements could be as shown in figure 1. There would normally be a lot more but reducing it to five makes it easier to explain. This example illustrates four normal operation requirements and one for what should happen in the event of an illegal module setup:
Specifications may be vague, inconsistent, or scattered across multiple tools or documents, making it difficult to ensure complete coverage. Verification teams struggle with inefficient one-to-one mapping between requirements and test cases. For mission-critical and safetyrelated projects, regulations add another layer of complexity, mandating strict traceability and verification in predefined test cases. As a result, the process quickly becomes time-consuming, error-prone, and hard to maintain across large as well as small projects.
- The acceleration shall be *** - The top speed shall be maximum *** - The deceleration shall be *** - The final position shall be *** - An illegal setup shall result in *** Figure 1: Requirements
Requirement specification Once you have the requirements, the tracking part starts to get more complicated the further you go into the process. For example, the requirements must be specified somewhere, like Excel, Word, Jira, etc, and should be properly labelled in order to easily recognize and reference the various requirements.
32
A common but very inefficient approach
There is no standard way of doing this labelling, and companies do it in many different ways, from long prosaic names like Req_acceleration_max, via various types of abbreviations like Req_Acc_3, to pure numeric labels like R_14. There is no right or wrong here, but some notations and approaches make the development process a bit easier, especially for the requirements definition phase.
Quite a few projects apply a one-to-one approach to Requirement vs Test, where they run a dedicated test case per requirement. This way, they can get the total overview and generate a report from the test case label (or name or summary) from all test cases executed. So 100 test cases are needed for 100 requirements. This is often extremely inefficient. In the case of our four normal operation motor requirements, we would need to run four test cases as shown in figure 4. But this is a major waste of time when we could have run the tests in series as shown in figure 5.
Examples of this could be to use names like R_UART_ TX_2_STOP_BITS or R_UART_TX_5, rather than just R_251, but there could be good reasons for most approaches. The important point is to have a welldefined system that suits your needs. Let us assume that the motor controller requirements are defined as in figure 2:
TC1
Requirement Label
Description
Motor_R1
The acceleration shall be ***
TC3
Motor_R2
The top speed shall be given by ***
TC4
Motor_R3
The deceleration shall be ***
Motor_R4
The final position shall be ***
Motor_R5
An illegal setup shall result in ***
R1
TC2
TC_A
R4
R1
R2
R3 R4
Figure 5: Efficient test case execution
The Requirements Traceability Matrix
This is a very simplified example. In most projects, there could be 50, 100, 200 or more requirements. So what if, in the motor controller example, there were 100 requirements out of which 40 could be run in sequence, and what if the majority of these requirements could only be properly verified close to the end of the deceleration? The inefficient approach might take potentially 30 times the simulation time compared to the efficient approach. For most projects using the inefficient approach, it seems the reason is just tool restrictions.
There are two ways of seeing the tracking of requirements vs the test case in which they are verified. Forward traceability shows which test cases verify each requirement. Backward traceability shows which requirements are verified by each test case. The combination of these yields the Requirements Traceability Matrix (RTM). This could be shown as a real matrix, as indicated in figure 3, but very often it is just shown as two separate tables showing forward and backward traceability respectively. Test B
R3
Figure 4: Inefficient test case execution
Figure 2: Requirements list
Test A
R2
Predefined requirement vs Test case relation
Test C
Req 1
Very often, it’s not important in which test case a requirement is verified, as long as it is verified, but for mission-critical and safety-related projects, regulations come into play. A good example is DO-254 from the Radio Technical Commission for Aeronautics (RTCA), which defines the design assurance process for airborne electronic hardware such as FPGAs and ASICs.
Req 2 Req 3 Req 4 Figure 3: Requirements Traceability Matrix
33
Under such regulations, each requirement must be verified in the test case specified in the verification plan. Verification of a requirement in a non-specified test case will be ignored and not count towards compliance. Specifying up-front in which test case a requirement should be verified, prior to implementing the testbenches and test cases, is an old-fashioned and rather inefficient approach to requirements tracking. This probably made sense when requirements tracking was a manual job, but with modern automated tools for requirements tracking, this is no longer the case. With such tools, you can easily generate an RTM along with other useful reports. However, there are still many applications and projects, including all of those which are missioncritical or safety-related, where it is mandatory to specify in which test case a requirement must be executed. This must be supported by the development process and preferably also by the tools. For such an approach, the requirement list given at the start must be extended with the required test case, as shown in figure 6: Requirement
Description
To be tested in
Motor_R1
The acceleration shall be ***
TC_A
Motor_R2
The top speed shall be given by *** TC_A
Motor_R3
The deceleration shall be ***
TC_A
Motor_R4
The final position shall be ***
TC_A
Motor_R5
An illegal setup shall result in ***
TC_B
It would be far more efficient to add corresponding sub-requirements in the requirements tracking tool rather than the Requirement Specification. However, to achieve this, we need support for sub-requirement handling in the tool itself.
Other important aspects So far, we’ve discussed standard simple scenarios, but there are other aspects that often apply: 1. A requirement must be tested in one specific test case, but also tested in other test cases, with or without errors 2. A requirement must be tested in multiple specified test cases 3. A requirement may be tested in only one out of multiple specified test cases 4. Reuse of verified modules and testbenches from a previous project may fail compliance if requirement labels or groupings differ in the new project. For example, even if the same UART design and testbench are reused, differing requirement labels or groupings mean the verification results no longer map correctly to the new specification, and compliance is lost. All of this could be handled manually, but that would be time-consuming, error-prone and boring.
Step in UVVM
Figure 6: Requirement list with test cases
Compound requirements Very often, the tracking requirements from sales, higher-level system requirements or other sources are vague or a combination of multiple, more or less related requirements. A requirement may sometimes just refer to a complete protocol like UART or AXI-stream, or a table of various modes. This is quite common but it can be difficult to ensure we’ve covered all relevant functionality in our test cases. So how can this be handled? From general feedback, just remembering to check it all seems to be the most common approach, but it would be better to modify the Requirement Specification to extend or split the initial requirements. This would mean modifying the potentially already approved Requirement Specification document, however, which could be bad for both manpower usage, schedule and deliveries.
34
This is where the Universal VHDL Verification Methodology (UVVM) can help. The free and open‑source framework designed to make FPGA and ASIC verification more structured, efficient, and reusable is already used by 27 percent of FPGA engineers and provides a standardized architecture with ready‑made verification components, utilities, and libraries that simplify the creation of testbenches. First launched in 2015, UVVM has evolved over time in collaboration with major partners like the European Space Agency (ESA). One of our first projects with ESA was to enhance the functionality of UVVM yet further with essential improvements that introduced specification coverage, or requirements tracking, to the project. Importantly, the functionality is fully compliant with the needs of mission-critical ESA projects. Very soon after the release, we got feedback from users that they had successfully used the UVVM Specification Coverage capability for various safety applications including DO-254. It meets all of the requirements and provides a ready-made roadmap to handle and unify every step of the requirements tracking journey in one place.
The UVVM Requirements Traceability Matrix
Getting started with UVVM Specification Coverage
The Requirements Traceability Matrix
It’s easy to introduce requirements tracking to your workflow with UVVM, and starts with four steps:
Following the workflow for the motor controller, let’s assume that in test case TC_B, where we’re checking for illegal module setup (Motor_R5), we also check requirement Motor_R2. This would yield the partial coverage files shown in figures 7 and 8:
1. Make a CSV file with labelled requirements – with or without test cases specified. Often this can be generated automatically from Word, Excel or Jira. 2. Implement your testbench and test case(s) as before with the following commands written inside the test cases:
Partial coverage
‘partial cov tc a.csv’ TESTCASE_NAME: tc_a Motor_R1, tc_a,PASS Motor_R2, tc_a,PASS Motor_R3, tc_a,PASS Motor_R4, tc_a,PASS SUMMARY, tc_a, PASS
a. initialize_req_cov() at the start of your test case b. tick_off_req_cov() for each requirement you have verified c. finalize_req_cov() at the end of your test case
Figure 7: Partial coverage TC_A
3. Run (simulate) your test case(s). This will generate a partial coverage file per test case 4. Run a simple Python script to accumulate the coverage from all partial coverage files (Not strictly needed if there is only one test case)
Partial coverage
‘partial cov tc b.csv’ TESTCASE_NAME: tc_b Motor_R2, tc_b,PASS Motor_R5, tc_b,PASS SUMMARY, tc_b, PASS
This will generate all of the report files needed.
Figure 8: Partial coverage TC_B
Then we could run the Python script (run_spec_cov.py) to accumulate the result over all the partial coverage files (here, only two) to generate the RTM files. The following two files would normally be the most interesting and yield a complete RTM via the forward and backward traceability files, as shown in figures 9 and 10. *.req_compliance_minimal.csv Motor_R1, tc_a, COMPLIANT Motor_R2, tc_a, COMPLIANT Motor_R3, tc_a, COMPLIANT Motor_R4, tc_a, COMPLIANT Motor_R5, tc_b, COMPLIANT Figure 9: Forward traceability
*.testcase_list.csv tc_a, PASS, Motor_R1 & Motor_R2 & Motor_R3 & Motor_R4 tc_b, PASS, Motor_R2 & Motor_R5
Figure 10: Backward traceability
35
The UVVM Requirements Traceability Matrix
Taking it one step further
Summary
UVVM can also generate additional files like:
It is unfortunately quite common during FPGA development to tick off somewhere, at some time, that a particular requirement has been tested, often just as a mental exercise. It is always better to use a written, repeatable and automated approach to guarantee requirements tracking has been conducted in a demonstrable and auditable manner. The UVVM Specification Coverage capability significantly simplifies such an approach.
*.req_compliance_extended.csv which is similar to the minimal above, but includes all test cases where the relevant requirement has been verified. *.req_non_compliance.csv which lists all requirements (if any) that are non-compliant or not tested, and the reasons for any NON_COMPLIANCE. *.warnings.csv which shows issues that the user should be aware of. The example above showed a situation where all test cases pass. If a test case does not pass, the partial coverage file will show it as FAIL rather than PASS, and the coverage summary files will show NON_COMPLIANT rather than COMPLIANT, with the second parameter showing a reference to the noncompliance file rather than the test case. The result of these reports is a very good overview of the RTM, the status and what potentially needs to be fixed.
This article is the second in a series from Espen Tallaksen about UVVM. You can read the first article, ‘How UVVM can result in faster and better FPGA verification’ in Issue 1 of the FPGA Horizons Journal online. In the next article, Espen will look at how UVVM enhanced randomization gives users a strong randomization capability and makes the functionality more understandable and readable.
THE BROADEST RFSOM PORTFOLIO • proven designs, shipping in volume since 2020 • ideal for communications systems, SIGINT, radar, T&M, and instrumentation of physics experiments
knowres.com
knowledgeresourcesgmbh
36
Debug FPGA with ease! Use the FPGA vendor flow Visualise gigabytes from inside FPGA
out of box experience was straightforward. “ The I could insert the core simply and then connect to and debug the application with ease. Contact us !
www.exostivlabs.com
Avoiding costly respins
Avoiding costly respins: The case for pre-layout signal integrity
Dan Binnun, Founder, E3 Designers While sometimes overlooked, signal integrity has an important role to play in the design of modern FPGA and System-on-Chip (SoC) printed circuit board assemblies (PCBAs). Neglecting the principles from the outset can lead to significant design issues, a problem which has become more prominent as data rates rise and the complexity of embedded systems follows in kind.
A far better approach is a proactive one which allows for the back end of the process to act more as a confirmation step and remain far less disruptive to the workflow. The necessary trade-off, and often a pain-point for management, is the front-loading of signal integrity effort into the early stages of design. It shouldn’t be because the critical work begins long before the first trace is routed on the PCB. This pre-layout phase is where we make foundational decisions and do critical design work that will ultimately determine the success of our designs.
Too often, engineering teams tackle these challenges in a reactive mode, only looking for problems after the core design work is done. A common workflow is to follow various collections of application notes, rules of thumb, and self-styled best practices that the team may have accumulated over their careers. After finishing the entire design, including schematic capture, PCB layout and the internal review steps required, the team will reach out to an internal or external engineer to perform post-layout signal integrity simulations.
The cornerstones of pre-layout signal integrity
This is where alarm bells should be ringing loudly because the expectation (or rather, the hope) seems to be that the signal integrity exercise will result in a stamp of approval. It’s a matter of checking the box and moving onto fabrication and procurement. There are times, however, when serious issues are discovered that cannot be addressed with simple or straightforward changes. Sometimes there is a fundamental problem with something as basic as a connector choice, pin assignment, PCB material, or stackup design. A problem that can result in significant and costly rework, redesign and lengthy project delays.
Stackup design and (dielectric) material selection come next, considering the number of layers needed to properly escape from high-density devices. It also means factoring in the frequencies and data rates of interest by choosing suitable dielectric material families.
38
There are four key elements in the pre-layout phase, beginning with component selection. This goes beyond selecting the main FPGA or SoC and includes a holistic view of the entire signal path. From connectors and cables to clock generators and their associated termination components, every aspect of the channel should be considered.
Simulating a full end-to-end channel with a prelayout mockup channel analysis is another valuable exercise in understanding the viability of a plan. If you have too much loss in a channel you may need something like a re-timer or re-driver to have a viable channel.
This is becoming more common in interfaces such as PCIe Gen 4 and PCIe Gen 5, as frequencies increase and insertion loss rises proportionally.
Ultimately, every component constituting the highspeed channel must be selected with simulation in mind. Identifying all elements of the signal path, such as the main FPGA, SoC or ASIC connectors, and any off-board cable assemblies or backplane interconnect, and securing their corresponding models early is the foundation of a proactive signal integrity strategy. This approach enables meaningful analysis long before the layout is complete.
Finally, we need to consider pre-layout structure optimization. The areas where signals transition from a chip package or connector’s land pattern into the main PCB, the breakout region, are increasingly critical. Before full board layout, we can optimize the specific via structures and trace geometries in these dense regions, minimizing degradation in signal integrity caused by these necessary but disruptive transitions.
Stackup design Once key components are identified, the design of the physical PCB structure, the stackup, is the next critical pre-layout activity. Every electrical characteristic of a trace, from its impedance to the attenuation it will introduce, is defined by the stackup. Getting this wrong can make a design impossible to salvage, while getting it right lays the groundwork for success.
Component selection The primary focus during component selection is on the central FPGA, SoC, or Application-Specific Integrated Circuit (ASIC). Verifying its support for the required protocols and data rates is a fundamental first step. However, a common and dangerous pitfall, especially under tight deadlines, is to assume that a chip supports all features of a given protocol. This is where a deeper analysis of layout-simplifying features is critical. Two prime examples are polarity inversion (P/N swap) and lane reversal (mapping lane 0…N to lane N…0). When available, these features can dramatically simplify complex routing out of dense packages, potentially saving board layers and reducing manufacturing costs. However, their availability is not guaranteed. A feature may be allowed by a protocol like PCIe, but it must also be explicitly implemented in the silicon of your chosen components. Verifying both levels of support during the component selection phase is essential; an incorrect assumption here can lead to a nonfunctional interface discovered only after the board has been built. An even more complex topic is signal equalization, the technique modern transceivers use to compensate for signal degradation. The big takeaway for pre-layout design is the critical importance of simulation models like those using the Input/Output Buffer Information Specification – Algorithmic Modeling Interface (IBIS-AMI). These are the industry standard and the only reliable way to simulate how a chip’s unique equalization scheme will perform in your specific channel. When a vendor does not provide these models, we face a significant challenge. The solution is not to skip the analysis, but to select a suitable proxy model to stand in for the missing component. While this introduces a degree of uncertainty, a simulation based on a reasonable assumption is infinitely better than no simulation at all.
Figure 1: A PCB stackup design in Altium Designer
For example, the insulating material between copper layers is far more than just a physical separator, it’s an active participant in signal propagation. One primary differentiator between material grades is the loss tangent tan(δ). Standard materials, commonly referred to as ‘FR4’, are inexpensive but absorb a significant amount of high-frequency signal energy. Use of low-loss materials (e.g., Megtron 6, Tachyon 100G) becomes mandatory when dealing with modern high-speed interfaces. Choosing the material is a fundamental cost vs. performance trade-off that must be modeled and decided upon before any layout begins.
39
Beyond pure electrical properties like loss, the physical construction of the material is also critical. The woven glass fibers that reinforce the dielectric can create localized inconsistencies. This phenomenon, known as the weave effect, can create a velocity difference between the two traces within a differential pair, introducing timing skew that can corrupt the signal. To combat this, low-loss materials often feature mechanically spread glass weaves to create a more uniform medium. This materiallevel solution can be further supported by layout techniques such as zigzag routing, where traces are intentionally routed at a slight angle to average out the inconsistencies of the weave. The other key property to be aware of is the dielectric constant (Dk). This property determines the speed at which signals travel and is a key variable in calculating trace impedance. A consistent dielectric constant across the frequency domain is crucial for consistent performance at high frequencies. The number and order of layers of copper built around this dielectric material dictates routing capacity, signal isolation, and the quality of the signal return path. High-pin-count FPGAs and SoCs require numerous layers simply to escape all the signals from the Ball Grid Array (BGA) footprint. However, each additional layer adds cost and manufacturing complexity. High-speed signals require an adjacent, uninterrupted reference plane to provide a clean, low-inductance return path. A stackup that sandwiches signal layers between solid reference planes is the gold standard for controlling impedance and minimizing noise and crosstalk. Vias are necessary to transition signals between layers, but they are also a major source of signal degradation if not planned for. A standard throughhole via can create a long, unused barrel, or stub, which presents a discontinuity at high frequencies, potentially causing signal reflections. The stackup design must account for this, often by specifying that critical high-speed vias be back-drilled (where the stub is mechanically drilled out after fabrication). This adds cost but is often the lowest cost approach to make a high-speed link viable. For extremely dense designs, High-Density Interconnect (HDI) technology using techniques such as via-in-pad, laser-drilled microvias, blind vias, and sequential lamination steps can be necessary. These vias and additional process steps can result in better performance, but they fundamentally change the stackup structure and fabrication process, typically trading cost for increased flexibility and improved performance.
40
Pre-layout mockup channel analysis With our key components selected, models gathered and our stackup defined, we can now simulate a full end-to-end high-speed channel. This mockup analysis is, in effect, a dress rehearsal for our design, allowing us to test our foundational choices and ensure they work in concert to deliver a robust and reliable link. First, we must properly define the channel. For modern protocols like PCI Express, it is the entire signal path from the transmitter’s silicon die to the receiver’s silicon die. This includes the chip packages, the motherboard (or host) traces and vias, one or more connectors, the add-in card traces and vias, and potentially a cable assembly (like a riser cable). Simulating this complete system is the only way to accurately predict real-world performance. This brings us to a crucial question: how do we simulate a channel that hasn’t been physically laid out or may consist of unknown components? The answer lies in using a combination of vendor-provided models and standardized proxy models. An excellent example of this is the VITA 68.3 standard, which provides a library of S-parameter models for common VPX backplane profiles. A designer can take the model for their proposed PCB trace (based on the stackup design), combine it with the IBIS-AMI model for their chosen FPGA, and connect it to a VITA 68.3 backplane model. This allows us to approximate the system with a reasonable level of accuracy well before it is designed. The ultimate goal of this exercise is not simply to achieve a pass/fail result, but to quantify our design’s robustness by measuring its margin. One approach for this is utilizing Channel Operating Margin (COM). COM is a sophisticated figure of merit, expressed in decibels (dB), that calculates the signal-to-noise ratio of the link after the effects of the receiver’s equalization are included. It provides a single number that tells us not just if the link works, but how well it works. The ideal conditions of a simulation – perfect trace geometries, nominal material properties, and quiet power supplies – will never be realized in physical hardware. Manufacturing tolerances and other realworld effects will degrade performance. A healthy margin in pre-layout simulation (for instance, achieving a COM of 6dB where the specification only requires 3dB) provides the necessary buffer to ensure the final, fabricated board meets its performance targets in a real-world environment. Tools such as Siemens HyperLynx support what-if simulations of COM, eye diagrams, eye masks, and bit-error rate (BER).
Avoiding costly respins
Figure 2: Part of a full-channel mockup in Siemens HyperLynx (recreation)
Pre-layout structure optimization While mockup channel analysis provides a systemlevel view of our design, structure optimization is where we zoom in with a microscope. The highspeed signals we implement will not live in ideal traces. They must navigate a series of complex physical transitions from a chip onto the board, through vias to other layers, and through connectors off board. Each of these transitions is an impedance discontinuity that can ultimately cause a link to fail. The goal of pre-layout structure optimization is to design and validate these critical transitions in isolation before they are used in the PCB layout. Using 3D electromagnetic (EM) field solver tools like Ansys Electronics Desktop or Siemens HyperLynx, we can model these small, complex regions with a high degree of accuracy. The objective isn’t to create a perfect transition; it is to manage the discontinuity, minimize its impact, and characterize its performance.
Figure 3: A 3D EM solved solution for the optimized geometry for the 3D CAD model in Ansys Electronics Desktop
Once a structure is optimized, its model can be saved and handed to the PCB layout team. This creates a library of pre-validated, high-performance building blocks. The layout designer can then confidently place these structures, knowing their performance has already been simulated and approved, drastically reducing design risk.
Over to you Signal integrity issues extend beyond the PCB layout alone. They affect the workflows of FPGA, SoC, and ASIC engineers and can disrupt entire projects. By shifting signal integrity analysis and optimization into the pre‑layout phase, engineering teams can make informed choices about components, stackup, and critical structures before costly mistakes are baked into the design. This proactive approach transforms simulation from a reactive diagnostic into a predictive design tool, ensuring signal integrity margins are built in rather than hoped for, ultimately reducing risk and minimizing the likelihood of expensive respins.
This article is the second in a series from Dan Binnun about power and signal integrity. You can read the first article in Issue 1 of the FPGA Horizons Journal online.
41
Bridging FPGAs with the Vitis System Device Tree flow
Bridging FPGAs with the Vitis System Device Tree flow A universal MIPI camera FMC PoC Jeff Johnson, Owner and Lead Designer, Opsero
Vision has always been one of the most compelling and natural application areas for FPGAs. Their ability to perform massively parallel computations and interface directly with high-speed image sensors makes them ideal for realtime image processing. Yet despite these strengths, it is still difficult to find modular solutions for connecting FPGAs to image sensors when it comes to rapid prototyping and experimentation.
42
In an ideal prototyping or Proof-of-Concept (PoC) scenario, a developer would simply pick an FPGA board and an image sensor and connect them through a standard interface. While many FPGA platforms can interface with USB cameras which are often sufficient for low-bandwidth applications, most vision systems demand a higherperformance link. For these, the Mobile Industry Processor Interface Camera Serial Interface 2 (MIPI CSI-2) has become the most widely used interface between image sensors and FPGAs. MIPI CSI-2 delivers high-bandwidth pixel data over just a few differential pairs, making it ideal for compact, high-speed camera links. Each lane can run at multiple gigabits per second (Gbps), and with up to four lanes per sensor, CSI-2 can easily support multi-megapixel, high-frame-rate video streams. Its low pin count, excellent signal integrity, and broad sensor compatibility have made it the standard for everything from smartphone cameras to industrial vision modules. I’ve long felt that there should be an FPGA Mezzanine Card (FMC) based solution for connecting MIPI CSI-2 image sensors to FPGA development boards.
The solution
A few options do exist including the Opsero RPi Camera FMC but these are essentially custom designs, each tied to one or two specific carrier boards rather than being broadly compatible. This runs counter to the goal of the VITA 57.1 FMC standard which is to modularize I/O and make hardware interchangeable across platforms.
A different approach is to break that limitation by adding a small FPGA directly on the FMC card. This way, the MIPI CSI-2 pin assignments are fixed and can be chosen to suit the mezzanine card’s FPGA. On the other side, the interface with the carrier board’s FPGA can be a generalpurpose, easily reproducible link such as Aurora, ensuring compatibility with a wide range of carrier boards. The image in Figure 1 illustrates the concept and uses the Artix UltraScale+ for the mezzanine FPGA and two Raspberry Pi cameras for the MIPI CSI-2 image sensors.
So why isn’t there a universal MIPI CSI-2 FMC on the market? The answer lies in pin restrictions. The AMD MIPI CSI-2 RX and equivalent implementations from vendors such as Altera and Lattice, requires specific FPGA pins that support the electrical and timing characteristics of the MIPI D-PHY physical layer standard interface. The FMC standard, while it defines the locations and electrical specifications for I/Os, clocks, gigabit transceivers and power supplies, does not constrain pins at the level of the byte-group or specialized functions which are important for MIPI implementations (e.g., Dedicated Byte Clocks (DBCs) and Quad Byte Clocks (QBCs) in AMD UltraScale/ UltraScale+ devices). As a result, a MIPI CSI-2 FMC can be designed to match one or two FPGA carriers, but it’s not possible to assign the pins to make it broadly compatible.
This architecture not only solves the compatibility problem but comes with an added benefit: the mezzanine FPGA provides more resources that can be used for image processing, freeing up resources on the carrier board. For example, a basic video pipe for a single MIPI CSI-2 image sensor can require 20,000 Lookup Tables (LUTs), 20,000 registers and 40 Block RAMs, and these numbers can grow exponentially when higher resolutions, frame rates or more complex processing need to be supported. By offloading the video pipe to the mezzanine card, either partially or fully, the main device can dedicate more resources to higher level functions such as AI inference.
Figure 1: A small FPGA on an FMC card with two Raspberry Pi cameras for the MIPI CSI-2 image sensors
43
Mezzanine FPGA
MIPI Camera
MIPI CS12 RX
Demosaic
Gamma LUT
Video Processing Subsystem
AXI-Lite
Frame Buffer Write IP
AXI Chip2Chip
FMC GT
AXI-MM
Figure 2: A typical video pipe implemented in a mezzanine FPGA
The diagram in Figure 2 illustrates a typical video pipe that could be implemented in the mezzanine FPGA. It starts with a MIPI camera on the left and ends with the FMC connector on the right. The video pipe contains the basic elements that convert the MIPI camera’s RAW10 output to an AXI-Stream of RGB pixels, a preferred format for processing images in an FPGA.
The AXI Chip2Chip IP bridges this gap again by passing an AXI4 memory-mapped interface across the same Aurora link, allowing the mezzanine’s video pipeline to write directly to the carrier’s memory subsystem.
Here, the AXI Chip2Chip IP core is used to configure and control the processing blocks. This enables an AXI4-Lite interface to be tunneled between the two FPGAs over a serial Aurora link, effectively extending the control bus across the FMC connection.
To demonstrate the concept works as expected, we built a demonstration platform (Figure 3) with an AMD ZCU106 Multi-Processor System-on-Chip (MPSoC) serving as the main board and a Tria AUBoard standing in for the mezzanine card. The AUBoard hosts the RPi Camera FMC, allowing two standard Raspberry Pi cameras to be connected. The Aurora serial link is implemented over an SFP+ DAC cable connected between the two boards.
At the end of the video pipeline, a Frame Buffer Write IP typically stores pixel data into Double Data Rate (DDR) memory. In our design, the DDR memory resides on the carrier board not the mezzanine.
The PoC
Figure 3: The demonstration platform using an AMD ZCU106 MPSoC as the main board and a Tria AUBoard as the mezzanine card
44
Bridging FPGAs with the Vitis System Device Tree flow
Unifying the software platform
With enhanced hardware description, it also allows additional hardware elements such as I²C devices or AXI Chip2Chip links to be represented even if they are not defined in the XSA file. As a result, the build system can become fully board-aware. Finally, the low-level interrupt controller setup is abstracted away, simplifying interrupt and sleep handling. Instead, interrupt IDs and related parameters are automatically made available through SDTpopulated config structures.
Even with most of the hardware architecture defined for a truly universal MIPI CSI-2 FMC and a prototype to prove the concept, one important factor remains: how to make the main FPGA aware of the hardware that resides in the mezzanine FPGA. In a typical single-FPGA system, this isn’t an issue. Software tools such as the Vitis IDE already know the details of the hardware design because the information is passed on via a hardware handoff file (XSA). In multi-FPGA designs, however, the situation is more complicated. Each FPGA has its own hardware design and its own XSA file. The software environment usually references only the XSA from the main FPGA, leaving the tools in the dark regarding any IP or peripherals instantiated in the secondary FPGA. To develop a coherent software platform, we need a way to describe the remote hardware like its IP blocks, base addresses, interrupts, and required drivers so that the tools can treat it as part of the overall system.
Key differences for bare metal developers From a user perspective, several practical differences stand out when working with the new SDT flow compared to the legacy flow: The DEVICE_ID macro is no longer generated or used. Devices are now identified and initialized by base address (BASEADDR). Interrupt setup is greatly simplified through the helper function XSetupInterruptSystem(), provided in xinterrupt_wrap.h.
This is where AMD’s System Device Tree (SDT) flow comes in. Introduced in the 2023.2 release of the Vitis Unified IDE, SDT provides a way to describe externally connected hardware using a device tree. Exactly what is needed to overcome the final hurdle in a universal MIPI camera FMC design.
Driver binding uses YAML files instead of MDD files. Drivers are matched to hardware using the compatible strings defined in the SDT, resolved by the Lopper framework. Config structures (typically generated in *_g.c files) are populated from SDT node properties like reg, interrupts, and vendor-specific attributes instead of from xparameters.h.
The basics of the SDT flow In Vitis Classic, the build system extracts the hardware metadata directly from the XSA file to generate the Board Support Package (BSP) and supporting files such as xparameters.h, the linker script, and the driver config files. In the Vitis Unified IDE, this process changes significantly. The tools now begin by creating an SDT, a structured, open-source hardware description that serves as the foundation for both Linux and bare-metal domains. Hardware metadata is no longer parsed directly from the XSA. Instead, it is extracted from the SDT using a tool called Lopper, which reads device-tree nodes and their properties to populate driver and library configurations.
Adding the video pipeline with a User DTS overlay The SDT flow enables us to add the hardware details of a video pipeline running on the AUBoard by supplying Vitis with a User Device Tree Source (DTS) overlay that contains the nodes for all the elements in our video pipe. To help determine what this overlay should look like, a Vivado project is created for the main board and the video pipe elements are added to the design. After generating an XSA file from this project and importing it into Vitis, the tools automatically create a System Device Tree that includes nodes for each of the IP blocks in the video pipe. Once the SDT is generated, the devicetree nodes of the video pipeline can be copied and pasted into the User DTS overlay. Finally, we need to include each of these devices in our memory map.
The advantages of the SDT flow The SDT-based flow provides several important advantages. Firstly, it’s built on open standards, using Device Tree, CMake, and Lopper to create a more transparent and maintainable build system.
45
Bridging FPGAs with the Vitis System Device Tree flow A simplified example of such an overlay is shown in Figure 4. Note that the first part involves adding the nodes to amba_pl, while the second part involves adding devices to the address-map. Since there is no way to cleanly append to the address map, we first delete the address-map property and redescribe the entire address map with the additional devices included. At this stage, we can create a Vivado project for the AUBoard, playing the role of the mezzanine card in our PoC. The block design will include the video pipeline IP and the AXI Chip2Chip IP. Note that it does not require a processor since it will be controlled remotely by the main board through the AXI Chip2Chip IP. It is important to ensure that the base addresses assigned to the IP match the addresses specified in the User DTS. These addresses should be chosen based on the main board’s memory map and the available address space in that design. Once the design is built, we can program the AUBoard either via JTAG using Vivado Hardware Manager or by storing the bitstream in flash memory for standalone boot. To complete the system, we create a new Vivado design for the ZCU106 that includes the processor (PS) and the AXI Chip2Chip IP core. From this design, we generate an XSA file and use it in Vitis to create a new platform component.
In the Advanced Options, Vitis allows us to specify our User DTS overlay. When the platform is built using both the XSA and the User DTS, Vitis automatically pulls in the drivers and generates the configuration data for all the IP blocks defined in our video pipeline. This means that when we begin developing the baremetal application for the ZCU106, the build system already recognizes the video pipeline running on the external FPGA (the Artix UltraScale+ of the AUBoard). When the AUBoard is connected to the ZCU106 via the SFP link, the complete system becomes operational. We can now develop and debug it as a single system from Vitis through a JTAG connection to the ZCU106 board.
Conclusion As FPGA designers, we are all familiar with singleFPGA systems but some problems are now best solved through multi-FPGA architectures, although they bring additional challenges. It’s important to recognize when this approach is appropriate and to understand the tools that can help us to work with these systems. As we’ve seen, the Vitis SDT flow provides a clean and maintainable way to deal with multi-FPGAs, offering a path for simple integration of an FMC with a carrier board. If you would like to reproduce the PoC, the source code is available at fpgadeveloper.com.
/ { }; &amba_pl {
...
...
...
...
... };
mipi_csi2_rx_subsyst_0: mipi_csi2_rx_subsystem@84a00000 { compatible = "xlnx,mipi-csi2-rx-subsystem-6.0"; status = "okay";
Figure 4: A simplified example of a User DTS overlay
}; demosaic_0: v_demosaic@84a10000 { compatible = "xlnx,v-demosaic-1.1"; status = "okay"; }; v_gamma_lut: v_gamma_lut@84ae0000 { compatible = "xlnx,v-gamma-lut-1.1"; status = "okay"; }; v_proc: v_proc_ss@84a40000 { compatible = "xlnx,v-proc-ss-2.3" , "xlnx,vpss-scaler-2.2" , "xlnx,v-vpss-scaler-2.2" , "xlnx,vpss-scaler"; status = "okay"; }; v_frmbuf_wr: v_frmbuf_wr@84ad0000 { compatible = "xlnx,v-frmbuf-wr-2.5" , "xlnx,axi-frmbuf-wr-v2.2"; status = "okay"; };
&cpus_a53 { /delete-property/ address-map; address-map = <0x0 0xf0000000 &amba 0x0 0xf0000000 0x0 0x10000000>, <0x0 0xf9000000 &amba_apu 0x0 0xf9000000 0x0 0x80000>, <0x0 0x0 &zynqmp_reset 0x0 0x0 0x0 0x0>, <0x0 0x0 &psu_ddr_0_memory 0x0 0x0 0x0 0x7FF00000>, ... <0x0 0x84a00000 &mipi_csi2_rx_subsyst_0 0x0 0x84a00000 0x0 0x1000>, <0x0 0x84a10000 &demosaic_0 0x0 0x84a10000 0x0 0x10000>, <0x0 0x84ae0000 &v_gamma_lut 0x0 0x84ae0000 0x0 0x10000>, <0x0 0x84a40000 &v_proc 0x0 0x84a40000 0x0 0x40000>, <0x0 0x84ad0000 &v_frmbuf_wr 0x0 0x84ad0000 0x0 0x10000>, ... };
46
Cost-Optimized PolarFire® Core FPGAs and SoCs Performance With a 30% Lower Price Tag
As Bill of Material (BOM) costs are rising and other FPGA vendors announce price increases, Microchip is offering a new cost-optimized solution with PolarFire Core FPGAs and SoCs. The new device families provide the same industry-leading low-power consumption, proven security and dependability, and reduce customer costs by up to 30 percent by optimizing features and removing integrated transceivers. PolarFire Core FPGAs and SoCs provide savings without sacrificing functionality, processing capability or quality. Designed for automotive, industrial automation, medical, communication, defense and aerospace markets, PolarFire Core devices are designed to be pin-to-pin compatible with the full line of PolarFire FPGAs to accommodate various design SKUs, enhancing value for applications that prioritize cost efficiency.
Key Features •
Architecture and process optimizations for 25K–500K LE devices
•
Best-in-class defense-grade security for intelligent, connected systems
•
Deterministic, coherent RISC-V CPU cluster for Linux® and real-time applications
•
1.6 Gbps I/Os supporting DDR4/DDR3/LPDDR3, LVDS-hardened I/O gearing logic with CDR (supports SGMII/GbE links on GPIOs)
•
Small form factor 11 × 11 mm package option
microchip.com/polarfire
Discover how PolarFire Core FPGAs and SoC FPGAs can help power your next innovation.
The Microchip name and logo and the Microchip logo are registered trademarks of Microchip Technology Incorporated in the U.S.A. and other countries. All other trademarks are the property of their registered owners. © 2025 Microchip Technology Inc. All rights reserved. MEC2625A-UK-07-25
ENGINEERING AND TRAINING, LTD.
“Houston, we have a problem.” Or maybe you’re on the verge of a breakthrough. Let’s give your project the engineering edge it needs to succeed...
Launch with a Galaxia® Space Tile
Bring the SpaceWire Codec onboard
developed for missions or advanced prototyping / testing solutions
developed for use across a range of FPGAs targeting space
Learn the FPGA development skills
Discover mission-critical design
– from architecting FPGAs to debugging – to become an effective FPGA designer
and explore what the environmental challenges mean to the logic designer
Continue your journey at adiuvoengineering.com