Author Topic: Best MCU for the lowest input capture interrupt latency  (Read 46512 times)

0 Members and 3 Guests are viewing this topic.

Offline SiliconWizard

  • Super Contributor
  • ***
  • Posts: 17799
  • Country: fr
Re: Best MCU for the lowest input capture interrupt latency
« Reply #150 on: April 15, 2022, 04:44:24 pm »
 
The following users thanked this post: hans

Offline jemangedeslolosTopic starter

  • Frequent Contributor
  • **
  • Posts: 386
  • Country: fr
Re: Best MCU for the lowest input capture interrupt latency
« Reply #151 on: April 22, 2022, 05:22:18 pm »
Hello everyone  :)

I received my PCB and I gave a try to a dsPIC33CK ( exact ref:dsPIC33CK128MP503 ).
There are more parameters than with dsPIC33EP so I don't know if I was doing something wrong with dsPIC33EP but I was able to make it work with this little dsPIC33CK  8)



The latency is not as low as other settings but at leat It works  >:D

 

Offline NorthGuy

  • Super Contributor
  • ***
  • Posts: 3530
  • Country: ca
Re: Best MCU for the lowest input capture interrupt latency
« Reply #152 on: April 22, 2022, 06:20:58 pm »
The latency is not as low as other settings but at leat It works  >:D

Congratulations!

I'm sure, on EP, there was some sort of small thing you somehow overlooked.

Roughly 60 ns (if judging by the 1/2 voltage level). I would expect it to be a little bit better. Is the CPU running at 100 MHz (200 MHz Fosc)? Have you used the dedicated INT pin (pin #19) for the signal?

Anyway, this is way faster than you would get with the interrupt.
 

Offline SpacedCowboy

  • Frequent Contributor
  • **
  • Posts: 439
  • Country: gb
  • Aging physicist
Re: Best MCU for the lowest input capture interrupt latency
« Reply #153 on: September 23, 2022, 06:36:57 pm »
Resurrecting an old thread, but I happened to be searching for IO expanders and the results I found reminded me of this. The SX150x chips are IO expanders that embed a PLD, and you can use the I2C interface to determine what happens when an IO changes state. The max response time looks to be 25ns if you're running your VDD/VIO at 5v, going down to 125ns if you're running at 1.2v VDD/IO. 25ns is, I suspect, a reasonably good figure for anything that's not a dedicated FPGA/CPLD; probably way too late to be useful but thought it was worth mentioning...
 

Offline PCB.Wiz

  • Super Contributor
  • ***
  • Posts: 3358
  • Country: au
Re: Best MCU for the lowest input capture interrupt latency
« Reply #154 on: September 24, 2022, 02:48:47 am »
Resurrecting an old thread, but I happened to be searching for IO expanders and the results I found reminded me of this. The SX150x chips are IO expanders that embed a PLD, and you can use the I2C interface to determine what happens when an IO changes state...

Interesting parts, the Logic looks to be very simple, non registered, with maximum of 3 line to 1 line any pattern defined by a single byte bit pattern. ( ie a 3 input universal gate level )

If you were using external parts for this, the SiLego GreenPAK parts have more advanced logic capabilities
https://www.renesas.com/us/en/products/programmable-mixed-signal-asic-ip-products/greenpak-programmable-mixed-signal-products#parametric_options


They started with OTP parts, but have expanded recently to include MTP parts config over i2c
 

Offline niconiconi

  • Frequent Contributor
  • **
  • Posts: 388
  • Country: cn
Re: Best MCU for the lowest input capture interrupt latency
« Reply #155 on: May 06, 2026, 07:54:51 pm »
Necrobumping an old thread for the record.

Timing

On the STM32H743 (SYSCLK = 480 MHz from PLL1, Revision V), the absolute best pure GPIO input-to-output IRQ latency I was able to achieve was 70 ns, WFE busy-polling latency was 60 ns. The test code was written in C, HAL was used, but without calling high-level HAL functions in the handler, SYSTICK is also turned off. Both the vector table and the IRQ handler running in ITCM, and the handler does nothing but to toggle the GPIO's ODR register, clear the IRQ pending bit, and return after a slight NOP delay loop (because the cleared flag takes time to propagate to NVIC). All code complied by GCC with -O3 and Link Time Optimization (LTO) enabled to maximize inlining.

Source code available upon request (because it no longer exists, and takes time to replicate).

If the vector table was not properly relocated by reprogramming VTOR, the latency is increased to 100 ns. This can be solved by either enabling the cache (not recommending), or by relocating the vector table. When both the vector table and the handler are relocated in ITCM, ICache and DCache have no effect on the latency. In fact enabling the cache becomes a useful debugging tool - if latency decreases, Flash reads are performed by mistake, which must be corrected.

Instead of using an IRQ, busy-polling the EXTI event in the main loop via _WFE() instead of an IRQ reduced it further to 60 ns.

In comparison, a naive HAL example would take around 300 to 500 ns because of the slow Flash and the dynamic decision-making code in the HAL.

I also found a GPIO pulse can’t be shorter than 25 ns, which means software-controlled GPIO toggling is limited to a 20 MHz square wave. Perhaps there's a way, but so far I’m not able to overcome this limitation.

Phase-Shifted Clock Generation

If the external pulse is periodic, such as a clock, I found it’s possible to hide the IRQ latency for all but the first pulse by triggering an internal timer and wait for the timer IRQ instead. The timer is programmed to fire slightly before next edge of the real pulse arrives. For example, if the pulse has a period of 1000 ns and the IRQ latency is 100 ns, you can use the GPIO pulse to trigger a timer to count for 900 ns (100 ns earlier), effectively predicting the next pulse and eliminating the IRQ latency for all subsequent IRQs.

The timer is triggered from GPIO via TIMx_CHx, or TIMx_ETR on every rising edge, so it's always phase-locked with the source with no long-term drift.

Unsuccessful Peripheral Exploitation

There's a general agreement online that STM32H7 has a notoriously slow GPIO controller, because it's on the D3 AHB4 bus, far away from the CPU's D1 domain, requests travel from the AXI port to the AXI matrix, to the AXI-to-D3 AHB bridge, to the D3 AHB matrix, to the GPIO controller. The GPIO pins themselves are much faster (> 100 MHz) in alternate function modes, but they are not designed for bitbanging.

In theory, we can speed it up by two methods.

BDMA

The first idea is to try driving it via the BDMA controller in domain D3, so the traffic stays locally in the same domain and matrix. Unfortunately, the 25 ns barrier remains, even if the transactions are initiated by BDMA. This makes me suspect that the long delay from the CPU is not the main source of latency, instead, the culprit is perhaps the synchronization or wait states within the GPIO controller itself or the AHB4. Alternatively, BDMA and SRAM4 themselves are too slow, so the gain of data locality is nearly canceled out.

Code: [Select]
static void dma_config(void)
{
__attribute__((section (".sram4_data")))
static const uint32_t gpio_data[8] = {
0xFFFFFFFF, 0x00000000, 0xFFFFFFFF, 0x00000000,
0xFFFFFFFF, 0x00000000, 0xFFFFFFFF, 0x00000000,
};

LL_AHB4_GRP1_EnableClock(LL_AHB4_GRP1_PERIPH_BDMA);

static LL_BDMA_InitTypeDef bdma_ctx = {
.PeriphOrM2MSrcAddress = (uint32_t) &gpio_data,
.MemoryOrM2MDstAddress = (uint32_t) &GPIOA->ODR,
.Direction = LL_BDMA_DIRECTION_MEMORY_TO_MEMORY,
.Mode = LL_BDMA_MODE_NORMAL,
.PeriphOrM2MSrcIncMode = LL_BDMA_PERIPH_INCREMENT,
.MemoryOrM2MDstIncMode = LL_BDMA_MEMORY_NOINCREMENT,
.PeriphOrM2MSrcDataSize = LL_BDMA_PDATAALIGN_WORD,
.MemoryOrM2MDstDataSize = LL_BDMA_MDATAALIGN_WORD,
.NbData = sizeof(gpio_data) / sizeof(gpio_data[0]),
.PeriphRequest = LL_DMAMUX2_REQ_MEM2MEM,
.Priority = LL_BDMA_PRIORITY_VERYHIGH,
.DoubleBufferMode = LL_BDMA_DOUBLEBUFFER_MODE_DISABLE,
/* don't care */
.TargetMemInDoubleBufferMode = LL_BDMA_CURRENTTARGETMEM0,
};

LL_BDMA_Init(BDMA, LL_BDMA_CHANNEL_0, &bdma_ctx);
}

static void mainloop(void)
{
        dma_config();
while (true) {
LL_BDMA_EnableChannel(BDMA, LL_BDMA_CHANNEL_0);
LL_mDelay(1000);
LL_BDMA_DisableChannel(BDMA, LL_BDMA_CHANNEL_0);
LL_BDMA_SetDataLength(BDMA, LL_BDMA_CHANNEL_0, 8);
}
}

I also tried to use BDMA's "data unpacking" feature, with the hope that the BDMA can read once from SRAM and write 4 times.

Code: [Select]
.PeriphOrM2MSrcDataSize = LL_BDMA_PDATAALIGN_WORD,
.MemoryOrM2MDstDataSize = LL_BDMA_MDATAALIGN_BYTE,

But it also doesn't overcome the 25 ns barrier, suggesting either the BDMA doesn't support optimized data unpacking (generating 4 fetches or 3 dummy cycles instead), or AHB4 bus write / GPIO controller itself is bottlenecked.


Peripheral Abuse

Another way is to abuse peripherals on the D1 AXI bus instead for bitbanging, such as the SDMMC, LTDC, FMC, QuadSPI. I decided to try the QSPI because it was easy to use, but without success. 

I found the QSPI controller has a "dual Flash" mode, each half is 4-bit, which is potentially a useful 8-bit bus transmitter. But latency also did not improve by abusing QSPI controller at 133 MHz to generate a write request. I got a < 10 ns pulse train, which broke the 25 ns pulse width barrier, but in fact the latency slightly degraded, even if the QSPI controller is physically close to the CPU. I blame the half-cycle chip enable delay. To write Flash in "indirect mode" (e.g. register-based, not memory-mapped), the first half-clock cycle is used for enabling /CE before the first bit is transmitted, I don't see anyway to skip this step. The /CE is held low only if the SPI Flash is in memory-mapped mode. If I read the datasheet correctly, this is used for Flash prefetching, meaning that the controller will keep generating commands on the 8-bit bus, writing garbage on the data bus. I don't see a way to have no prefetching and no /CE delay, perhaps there's a hack (potential solution that I didn't try: if we're using MMIO writes only, the controller is allowed to prefetch garbage, but we can deselect the alternate function to tristate the pins).

But I ran out of patience. Perhaps it can be improved through alternative solutions, such as the QSPI MMIO mode, or by exploiting SDMMC/LTDC/FMC.
« Last Edit: May 10, 2026, 07:43:26 pm by niconiconi »
 

Online Siwastaja

  • Super Contributor
  • ***
  • Posts: 11231
  • Country: fi
Re: Best MCU for the lowest input capture interrupt latency
« Reply #156 on: May 07, 2026, 03:03:33 pm »
Necrobumping an old thread for the record.

On the STM32H743 (SYSCLK = 480 MHz from PLL1, Revision V), the absolute best pure GPIO input-to-output IRQ latency I was able to achieve was 70 ns.

Sanity check:
AHB4 bus maximum speed 200MHz (which I assume you obviously would have used for the test), standard synchronization logic delay of 2 clock cycles:
10ns has passed when the peripheral register sees the input

Re-synchronize to 480MHz bus -> again 2 clock cycles of delay (at destination bus speed):
10ns + 4.2ns has passed when the NVIC sees it.

Now the actual interrupt latency of 12 cycles at 480MHz:
10ns + 4.2ns + 25ns

Now execute the instructions - load the peripheral register address, load the value to be written, write to peripheral. 2+2+2 cycles:
10ns + 4.2ns + 25ns + 12.5ns

Now synchronize to the 200MHz AHB4 clock domain: 2 cycles of the target:
10ns + 4.2ns + 25ns + 12.5ns + 10ns ...

and we are at 61.7ns, pretty close to the 70ns.

So I forgot something, or got some number off - but as you can see, with quick analysis, pretty close. So I asked AI what I'm missing, and it says:
Input pad schmitt trigger 1-2ns, another 2-4ns for output slew rate / capacitance
EXTI (edge detection) logic takes 1-2 cycles
AXI-AHB bridge handshake is slower than my estimated 2 target cycles

All believable. No surprises.

Now compare to some XMOS etc. and you get rid of interrupt latency and the internal bridge (due to simpler clock domain design), but input pad delay, initial synchronization, edge detection, and output pad delay still all apply, so the delay is maybe one third, and jitter probably too (still have the original synchronization jitter). Half an order of magnitude improvement.
« Last Edit: May 07, 2026, 03:11:03 pm by Siwastaja »
 
The following users thanked this post: niconiconi

Offline niconiconi

  • Frequent Contributor
  • **
  • Posts: 388
  • Country: cn
Re: Best MCU for the lowest input capture interrupt latency
« Reply #157 on: May 07, 2026, 03:36:27 pm »
Necrobumping an old thread for the record.

On the STM32H743 (SYSCLK = 480 MHz from PLL1, Revision V), the absolute best pure GPIO input-to-output IRQ latency I was able to achieve was 70 ns.

Sanity check:
AHB4 bus maximum speed 200MHz (which I assume you obviously would have used for the test), standard synchronization logic delay of 2 clock cycles:

240 MHz to be exact. Pre-V silicon was limited to SYSCLK = 400 MHz, AHBCLK's PLL is DIV2, so 200 MHz. Revision V (since 2019) ICs are rated for SYSCLK = 480 MHz (if Tj < 105 °C can be guaranteed), so AHBCLK is 240 MHz.

« Last Edit: May 07, 2026, 03:41:29 pm by niconiconi »
 

Offline SpacedCowboy

  • Frequent Contributor
  • **
  • Posts: 439
  • Country: gb
  • Aging physicist
Re: Best MCU for the lowest input capture interrupt latency
« Reply #158 on: May 07, 2026, 05:59:15 pm »

Now compare to some XMOS etc. and you get rid of interrupt latency and the internal bridge (due to simpler clock domain design), but input pad delay, initial synchronization, edge detection, and output pad delay still all apply, so the delay is maybe one third, and jitter probably too (still have the original synchronization jitter). Half an order of magnitude improvement.

I'm not sure that's actually true - the improvement I mean. Some time ago I wanted to pair an FTDI FT600Q to an XMOS device and I was chatting to one of their engineers on xcore.com. My reading of the datasheet was that I was limited to 60MHz as a synchronised input clock rate, and he was gently pointing out that they run 133MHz, in software, to do their QSPI implementation, which is somewhat tested since that's how they read the chip's program on boot.

That's reading new bits every ~7.5ns, which is really down to how the ports on an XMOS chip are different to those on (say) an STM32. The ports can synchronise to an input clock to help them latch their data, and they have SERDES built in, so (I'm assuming) in this case the 4 bits are clocked into a 32-bit word and *that* being sent down the wire to the CPU is far less than the max 60MHz rate. You don't have to use the SERDES of course, but then you're going to hit that 60MHz cap faster.

Since the XMOS chips are inherently parallel, you don't incur any interrupt cost to respond to the signal, you tell the CPU core to wait on i/o. it's sitting there waiting for the data to arrive - like a busy-wait but without the busy-part - and it's prepared to go as soon as the data does arrive. If what you want to do is toggle a LED or drive an output to measure on a scope, the next instruction can do that.

Bear in mind that the CPU (although it can be running at up to 800 MHz) is round-robinning between a minimum of 5 instances, so your real MHz is more like a max of 160MHz. Given that it's ready to go, though, that is still a very rapid response.

The code:
Code: [Select]
#include <xs1.h>
#include <xcore/port.h>

port_t oneBit  = XS1_PORT_1A;
port_t counter = XS1_PORT_4A;

int main(void) {
    int x;
    int i = 0;
    port_enable(oneBit);
    port_enable(counter);
    x = port_in(oneBit);
    while (1) {
        port_set_trigger_in_not_equal(oneBit, x);
        x = port_in(oneBit);
        port_out(counter, ++i);
    }
}

sets a hardware trigger on the CPU to wait on a transition on the 'oneBit' port, and immediately output a counter value to the 'counter' port.

XMOS parts are actually very good at i/o, that's one of their strengths. They suffer in terms of clock speed, memory (though the newer 'ai' cores can have DDR3) and a curiously laid out set of ports which seem almost intentionally engineered to be as annoying as possible (oh, you want a 32-bit port, how about 27 of those bits but not the other 5?)

Still, I did get it to read that FT600Q, pretty simply in the end, at a 66MHz synchronous clock. That's the fastest I've synced input on an XMOS device, but having seen how it worked, I quite easily believe that it could go twice as fast, as long as you can use the SERDES in the port hardware itself.
 
The following users thanked this post: Siwastaja

Offline niconiconi

  • Frequent Contributor
  • **
  • Posts: 388
  • Country: cn
Re: Best MCU for the lowest input capture interrupt latency
« Reply #159 on: May 08, 2026, 12:32:57 am »
Keep optimizing and benchmarking the IRQ latency of the STM32H743 @ 480 MHz. Ultimately, I was able to achieve a WFE polling latency of 38 ns, and an IRQ latency of 45 ns. I think these numbers are close to the theoretical limits.

2812809-0
2812941-1

The slow AHB bus and the GPIO controller prevented me from measuring the true IRQ latency of the core itself accurately. But I found a way out: EVENTOUT. Inside the ARM IP core, there's an internal signal called TXEV, originally meant for synchronizing multicore systems (by connecting it to the RXEV input of another core). On the STM32, TXEV can be mapped to any GPIO pin, allowing single-cycle pulse generation directly from the core via the SEV instruction without any controller overhead.

Full project files: https://codeberg.org/niconiconi/stm32h7-exti-latency
« Last Edit: May 08, 2026, 03:53:08 am by niconiconi »
 
The following users thanked this post: PCB.Wiz

Offline niconiconi

  • Frequent Contributor
  • **
  • Posts: 388
  • Country: cn
Re: Best MCU for the lowest input capture interrupt latency
« Reply #160 on: May 11, 2026, 03:54:57 am »
Keep optimizing and benchmarking the GPIO latency of the STM32H743 @ 480 MHz. :box: I was able to find a way to break the GPIO latency barrier using DMA1 with FIFO and burst mode, now generating a pulse train with a latency of 45 ns and a frequency of 127 MHz.

2814919-0

Triggering a DMA from GPIO is complicated on the STM32H7. So there are some notes on implementing it.

BDMA Triggering

Remark 1: The SYSCFG_EXTICRx registers provide 16 configuration 4-bit bitfields, each bitfield select a GPIO bank from A to K, the input pin must have the name number as the bitfield number. For example, bitfield 0 select the input GPIO bank of pin 0. This provides 16 multiplexed signal outputs EXTI[0:15].

Remark 2: Syscfg_extiX_mux is the raw output of EXTI[0:15]. Since only Syscfg_exti0_mux and Syscfg_exti2_mux are available as input, you can only use the signal EXTI[0] and EXTI[2], which correspond to GPIO[A:K] pin 0 and GPIO[A:K] pin 2. You cannot use other pins to trigger the DMA because of the lack of connections, but they don't need EXTI or NVIC.

Even better, because it's in the D3 domain, it keeps working autonomously even if the CPU core is halted in the debugger, or if the D1 domain is powered off completely. Unfortunately, only Pin 0 and Pin 2 of a GPIO bank can be used for that. For other pins you need to do it via a real peripheral.

2814911-1

Code:

Code: [Select]
static void bdma_config(void)
{
/*
* Select GPIO bank B as the source of the EXTI input line 0.
* GPIOB Pin 0 would generate a signal on EXTI Line 0.
*/
LL_SYSCFG_SetEXTISource(LL_SYSCFG_EXTI_PORTB, LL_SYSCFG_EXTI_LINE0);

LL_AHB4_GRP1_EnableClock(LL_AHB4_GRP1_PERIPH_BDMA);

/* Use EXTI line 0's rising edge to generate DMA Request 0 */
LL_DMAMUX_SetRequestSignalID(
DMAMUX2, LL_DMAMUX_REQ_GEN_0, LL_DMAMUX2_REQ_GEN_EXTI0
);
LL_DMAMUX_SetRequestGenPolarity(
DMAMUX2, LL_DMAMUX_REQ_GEN_0, LL_DMAMUX_REQ_GEN_POL_RISING
);

/* Trigger 2 DMA transfers per request */
LL_DMAMUX_SetGenRequestNb(DMAMUX2, LL_DMAMUX_REQ_GEN_0, 2);
LL_DMAMUX_EnableRequestGen(DMAMUX2, LL_DMAMUX_REQ_GEN_0);

__attribute__((section (".sram4_data")))
static const uint32_t gpio_data[] = { 0xFFFFFFFF, 0x00000000 };

static LL_BDMA_InitTypeDef bdma_ctx = {
.PeriphOrM2MSrcAddress = (uint32_t) &GPIOD->ODR,
.MemoryOrM2MDstAddress = (uint32_t) &gpio_data,
.Direction = LL_BDMA_DIRECTION_MEMORY_TO_PERIPH,
/* auto-reload if NbData goes to 0, otherwise DMA stops */
.Mode = LL_BDMA_MODE_CIRCULAR,
.PeriphOrM2MSrcIncMode = LL_BDMA_PERIPH_NOINCREMENT,
.MemoryOrM2MDstIncMode = LL_BDMA_MEMORY_INCREMENT,
.PeriphOrM2MSrcDataSize = LL_BDMA_PDATAALIGN_WORD,
.MemoryOrM2MDstDataSize = LL_BDMA_MDATAALIGN_WORD,
.NbData = sizeof(gpio_data) / sizeof(gpio_data[0]),
.PeriphRequest = LL_DMAMUX2_REQ_GENERATOR0,
.Priority = LL_BDMA_PRIORITY_VERYHIGH,
.DoubleBufferMode = LL_BDMA_DOUBLEBUFFER_MODE_DISABLE,
/* don't care */
.TargetMemInDoubleBufferMode = LL_BDMA_CURRENTTARGETMEM0,
};

LL_BDMA_Init(BDMA, LL_BDMA_CHANNEL_0, &bdma_ctx);
LL_BDMA_EnableChannel(BDMA, LL_BDMA_CHANNEL_0);
}

static void mainloop(void)
{
/*
* Set GPIO pin A9 as OUTPUT, this is connected to an LED on my devboard
* for indicating the firmware status. Can be removed.
*/
LL_GPIO_SetPinMode(GPIOA, LL_GPIO_PIN_9, LL_GPIO_MODE_OUTPUT);
LL_GPIO_SetOutputPin(GPIOA, LL_GPIO_PIN_9);

while (true) {
__NOP();
}
}

DMA1 Triggering

Remark 3: exti_exti0_it is the output of the EXTI interrupt controller, the same signal also goes to the CPU's NVIC. The same signal is connected to the DMAMUX controllers, which can use the interrupt's edge to initiate DMAs. But to use these signals to trigger DMA, EXTI's interrupt output to NVIC must also be enabled. This creates a problem: if you don't call LL_EXTI_ClearFlag, the EXTI's IRQ output line remains high, which means the DMA stops after the first rising edge.

So, after each DMA triggering, you either have to also enable an IRQ for the sole purpose of clearing this flag, or you have to clear the EXTI interrupt pending bit by polling, or by another DMA. This is why it's not recommended by ST.

2814915-2

Why should you use DMA1/DMA2 instead of BDMA? The D2 domain can access any memory, the D3 domain can only access SRAM4. Furthermore, the D2 domain's SRAM1 and DMA1 are faster than D3 domain's SRAM4 and BDMA (even though D2 still needs to access D3's GPIO controller). I'm not sure about the exact timing, but from the examples above, BDMA generates a GPIO pulse of 33 ns, but DMA1 generates a GPIO pulse of only 20 ns.

Code:

Code: [Select]
static void dma_config(void)
{
/*
* Select GPIO bank B as the source of the EXTI input line 0.
* GPIOB Pin 0 would generate a signal on EXTI Line 0.
*/
LL_SYSCFG_SetEXTISource(LL_SYSCFG_EXTI_PORTB, LL_SYSCFG_EXTI_LINE0);

/* Use EXTI Line 0 input's rising edge to trigger an IRQ. */
static LL_EXTI_InitTypeDef exti_ctx = {
.Line_0_31 = LL_EXTI_LINE_0,
.Line_32_63 = LL_EXTI_LINE_NONE,
.Line_64_95 = LL_EXTI_LINE_NONE,
.LineCommand = ENABLE,
.Mode = LL_EXTI_MODE_IT,
.Trigger = LL_EXTI_TRIGGER_RISING,
};
LL_EXTI_Init(&exti_ctx);

LL_AHB1_GRP1_EnableClock(LL_AHB1_GRP1_PERIPH_DMA1);

/* Use EXTI line 0's rising edge to generate DMA Request 0 */
LL_DMAMUX_SetRequestSignalID(
DMAMUX1, LL_DMAMUX_REQ_GEN_0, LL_DMAMUX1_REQ_GEN_EXTI0
);
LL_DMAMUX_SetRequestGenPolarity(
DMAMUX1, LL_DMAMUX_REQ_GEN_0, LL_DMAMUX_REQ_GEN_POL_RISING
);

/* Trigger 4 DMA transfers per request */
LL_DMAMUX_SetGenRequestNb(DMAMUX1, LL_DMAMUX_REQ_GEN_0, 4);
LL_DMAMUX_EnableRequestGen(DMAMUX1, LL_DMAMUX_REQ_GEN_0);

__attribute__((section (".sram1_data")))
static const uint32_t gpio_data[] = {
0xFFFFFFFF, 0x00000000, 0xFFFFFFFF, 0x00000000
};

static LL_DMA_InitTypeDef dma_ctx = {
.PeriphOrM2MSrcAddress = (uint32_t) &GPIOD->ODR,
.MemoryOrM2MDstAddress = (uint32_t) &gpio_data,
.Direction = LL_DMA_DIRECTION_MEMORY_TO_PERIPH,
.Mode = LL_DMA_MODE_CIRCULAR,
.PeriphOrM2MSrcIncMode = LL_DMA_PERIPH_NOINCREMENT,
.MemoryOrM2MDstIncMode = LL_DMA_MEMORY_INCREMENT,
.PeriphOrM2MSrcDataSize = LL_DMA_PDATAALIGN_WORD,
.MemoryOrM2MDstDataSize = LL_DMA_MDATAALIGN_WORD,
.NbData = sizeof(gpio_data) / sizeof(gpio_data[0]),
.PeriphRequest = LL_DMAMUX1_REQ_GENERATOR0,
.Priority = LL_DMA_PRIORITY_VERYHIGH,
.FIFOMode = LL_DMA_FIFOMODE_ENABLE,
.FIFOThreshold = LL_DMA_FIFOTHRESHOLD_FULL,
.MemBurst = LL_DMA_MBURST_INC4,
.PeriphBurst = LL_DMA_PBURST_SINGLE,
.DoubleBufferMode = LL_DMA_DOUBLEBUFFER_MODE_DISABLE,
/* don't care */
.TargetMemInDoubleBufferMode = LL_DMA_CURRENTTARGETMEM0,
};

LL_DMA_Init(DMA1, LL_DMA_STREAM_0, &dma_ctx);
LL_DMA_EnableStream(DMA1, LL_DMA_STREAM_0);
}

/*
 * Use "noinline" to prevent this function from being inlined to main(),
 * which is in Flash.
 */
__attribute__((section (".itcm_text"), noinline))
static void mainloop(void)
{
/*
* Set GPIO pin A9 as OUTPUT, this is connected to an LED on my devboard
* for indicating the firmware status. Can be removed.
*/
LL_GPIO_SetPinMode(GPIOA, LL_GPIO_PIN_9, LL_GPIO_MODE_OUTPUT);
LL_GPIO_SetOutputPin(GPIOA, LL_GPIO_PIN_9);

/*
* Because DMAMUX is edge-triggered by EXTI's IRQ output, not input, we
* must clear the EXTI flag LL_EXTI_ClearFlag_0_31(LL_EXTI_LINE_0); after
* every IRQ, otherwise the output flatlines and DMA stops. This can be
* done in a real ISR by enabling the IRQ.
*
     * NVIC_SetPriority(EXTI0_IRQn, 0);
     * NVIC_EnableIRQ(EXTI0_IRQn);
*
* Because we don't want the actual IRQ, NVIC is not enabled on this
* signal. But we still need to simulate the IRQ flag clearing
* by polling (or potentially another DMA transaction). This convoluted
* procedure is the reason that GPIO DMA triggering via EXTI is not
* recommended. Use another peripheral such as a timer if possible.
*/
while (true) {
/*
* Use WFI, not polling, because polling itself consumes bus
* bandwidth, which makes DMA transfer themselves if we're still
* polling during the DMA. If you must use polling, it's better
* to pull the NVIC, not the EXTI.
*
*     NVIC_GetPendingIRQ(EXTI0_IRQn)
*
*     // after LL_EXTI_ClearFlag_0_31 + 20 _NOP()
*     NVIC_ClearPendingIRQ(EXTI0_IRQn)
*/
__WFI();

/* Clear IRQ flags immediately, otherwise DMA stops. */
LL_EXTI_ClearFlag_0_31(LL_EXTI_LINE_0);

/*
* The STM32H7 has a notorious "spurious IRQ" limitation because
* LL_*_ClearFlag() must travel across multiple buses and buffers
* with a long delay, causing the IRQ to retrigger. It can be worked
* around by clearing the IRQ flag early and relying on the natural
* ISR code delay. Adjust this delay if the CPU clock frequency or
* emulator performance changes. Always verify with an oscilloscope.
*
* http://efton.sk/STM32/gotcha/g7.html
*
*/
for (uint8_t i = 0; i < 20; i++) {
__NOP();
}
}
}

DMA1 Triggering With FIFO

Furthermore, DMA1/DMA2 has a FIFO, which means the DMA doesn't need to access memory for every transfer, it can preload data and write that the destination in a burst. In this process, the DMA uses the "burst transfer" type in the AHB protocol, reducing the latency between consecutive words.

Here's an example of bitbanging GPIO via DMA1 using 8 16-bit burst, generating 8 transitions with a period of 7.8 ns, which is a frequency greater than 127 MHz (the real pulse width is below my oscilloscope's rise time, so I can't measure it). I think this is likely the fastest way to bitbang the GPIO on the STM32H7.

2814919-3  2814923-4

The number of pulses in the pulse train is limited to 4, 8, or 16. Note that time is needed to reload the DMA after each burst, the reloading time depends on the burst settings. In this example, it's 50 ns.

2814927-5

This shows the raw GPIO controller can be fairly fast, it's just challenging to feed it with data due to the lack of a fast master on its local domain.

Code:

Code: [Select]
static void dma_config(void)
{
/*
* Select GPIO bank B as the source of the EXTI input line 0.
* GPIOB Pin 0 would generate a signal on EXTI Line 0.
*/
LL_SYSCFG_SetEXTISource(LL_SYSCFG_EXTI_PORTB, LL_SYSCFG_EXTI_LINE0);

/* Use EXTI Line 0 input's rising edge to trigger an IRQ. */
static LL_EXTI_InitTypeDef exti_ctx = {
.Line_0_31 = LL_EXTI_LINE_0,
.Line_32_63 = LL_EXTI_LINE_NONE,
.Line_64_95 = LL_EXTI_LINE_NONE,
.LineCommand = ENABLE,
.Mode = LL_EXTI_MODE_IT,
.Trigger = LL_EXTI_TRIGGER_RISING,
};
LL_EXTI_Init(&exti_ctx);

LL_AHB1_GRP1_EnableClock(LL_AHB1_GRP1_PERIPH_DMA1);

/* Use EXTI line 0's rising edge to generate DMA Request 0 */
LL_DMAMUX_SetRequestSignalID(
DMAMUX1, LL_DMAMUX_REQ_GEN_0, LL_DMAMUX1_REQ_GEN_EXTI0
);
LL_DMAMUX_SetRequestGenPolarity(
DMAMUX1, LL_DMAMUX_REQ_GEN_0, LL_DMAMUX_REQ_GEN_POL_RISING
);
LL_DMAMUX_SetGenRequestNb(DMAMUX1, LL_DMAMUX_REQ_GEN_0, 2);
LL_DMAMUX_EnableRequestGen(DMAMUX1, LL_DMAMUX_REQ_GEN_0);

__attribute__((section (".sram1_data")))
static const uint16_t gpio_data[] = {
0xFFFF, 0x0000, 0xFFFF, 0x0000,
0xFFFF, 0x0000, 0xFFFF, 0x0000,
0xFFFF, 0x0000, 0xFFFF, 0x0000,
0xFFFF, 0x0000, 0xFFFF, 0x0000,
};

static LL_DMA_InitTypeDef dma_ctx = {
.PeriphOrM2MSrcAddress = (uint32_t) &GPIOD->ODR,
.MemoryOrM2MDstAddress = (uint32_t) &gpio_data,
.Direction = LL_DMA_DIRECTION_MEMORY_TO_PERIPH,
.Mode = LL_DMA_MODE_CIRCULAR,
.PeriphOrM2MSrcIncMode = LL_DMA_PERIPH_NOINCREMENT,
.MemoryOrM2MDstIncMode = LL_DMA_MEMORY_INCREMENT,
.PeriphOrM2MSrcDataSize = LL_DMA_PDATAALIGN_HALFWORD,
.MemoryOrM2MDstDataSize = LL_DMA_MDATAALIGN_HALFWORD,
.NbData = sizeof(gpio_data) / sizeof(gpio_data[0]),
.PeriphRequest = LL_DMAMUX1_REQ_GENERATOR0,
.Priority = LL_DMA_PRIORITY_VERYHIGH,
.FIFOMode = LL_DMA_FIFOMODE_ENABLE,
.FIFOThreshold = LL_DMA_FIFOTHRESHOLD_FULL,
.MemBurst = LL_DMA_MBURST_INC8,
.PeriphBurst = LL_DMA_PBURST_INC8,
.DoubleBufferMode = LL_DMA_DOUBLEBUFFER_MODE_DISABLE,
/* don't care */
.TargetMemInDoubleBufferMode = LL_DMA_CURRENTTARGETMEM0,
};

LL_DMA_Init(DMA1, LL_DMA_STREAM_0, &dma_ctx);
LL_DMA_EnableStream(DMA1, LL_DMA_STREAM_0);
}

Just make sure gpio_data is located in SRAM 1 (0x30000000) for DMA1, or SRAM 4 (0x38000000) for BDMA, may require custom startup and linker code. How to implement this is out of scope of this thread. I'm using a bare-metal environment, your mileage may vary.

What I do at startup:

Code: [Select]
void flash_to_mem(
char *flash_start,
volatile char *mem_start,
volatile char *mem_end
)
{
size_t len = mem_end - mem_start;

#ifdef SEMIHOSTING
printf(
"relocate %p-%p to %p, %d bytes\n",
flash_start,
flash_start + len,
mem_start,
len
);
#endif

/*
* Language Lawyering.
*
* Don't use memcpy(). This is widely used in embedded libraries, but it's
* a theoretical Undefined Behavior as C compilers can remove memcpy()
* when they don't have any visible effects to C code under the "as-if"
* rule. Mark destination variables as "volatile char *", and copy manual-
* ly. Keyword "volatile" ensures it's always executed, and "char *" is
* the only data type in C that is safe to cast into without breaking
* aliasing or alignment rules (even "uint8_t *" does not enjoy this
* exception).
*/
for (size_t i = 0; i < len; i++) {
mem_start[i] = flash_start[i];
}
}

void relocate_to_itcm(void)
{
extern char _si_isr_vector;
extern volatile char __isr_vector_start, __isr_vector_end;
flash_to_mem(&_si_isr_vector, &__isr_vector_start, &__isr_vector_end);

extern char _si_itcm_text;
extern volatile char __itcm_text_start, __itcm_text_end;
flash_to_mem(&_si_itcm_text, &__itcm_text_start, &__itcm_text_end);

SCB->VTOR = D1_ITCMRAM_BASE;

/*
* Test memory access latency by declaring arrays with:
*
*     __attribute__((section (".axisram_data")))
*     __attribute__((section (".sram1_data")))
*     __attribute__((section (".sram2_data")))
*     __attribute__((section (".sram3_data")))
*     __attribute__((section (".sram4_data")))
*
* Note that different SRAMs have different access latency due
* to their bus locations. Some peripherals can only read from
* a specific SRAM bank (e.g. BDMA can only read from SRAM4).
*
* Not used in this example, can be removed.
*/
extern char _si_axisram_data;
extern volatile char __axisram_data_start, __axisram_data_end;
LL_AHB3_GRP1_EnableClock(LL_AHB3_GRP1_PERIPH_AXISRAM);
flash_to_mem(
&_si_axisram_data, &__axisram_data_start, &__axisram_data_end
);

extern char _si_sram1_data;
extern volatile char __sram1_data_start, __sram1_data_end;
LL_AHB2_GRP1_EnableClock(LL_AHB2_GRP1_PERIPH_D2SRAM1);
flash_to_mem(&_si_sram1_data, &__sram1_data_start, &__sram1_data_end);

extern char _si_sram2_data;
extern volatile char __sram2_data_start, __sram2_data_end;
LL_AHB2_GRP1_EnableClock(LL_AHB2_GRP1_PERIPH_D2SRAM2);
flash_to_mem(&_si_sram2_data, &__sram2_data_start, &__sram2_data_end);

extern char _si_sram3_data;
extern volatile char __sram3_data_start, __sram3_data_end;
LL_AHB2_GRP1_EnableClock(LL_AHB2_GRP1_PERIPH_D2SRAM3);
flash_to_mem(&_si_sram3_data, &__sram3_data_start, &__sram3_data_end);

extern char _si_sram4_data;
extern volatile char __sram4_data_start, __sram4_data_end;
LL_AHB4_GRP1_EnableClock(LL_AHB4_GRP1_PERIPH_SRAM4);
flash_to_mem(&_si_sram4_data, &__sram4_data_start, &__sram4_data_end);
}

With the linker file:

Code: [Select]
/* SPDX-License-Identifier: Apache-2.0 AND (0BSD OR CC0-1.0) */
/*
******************************************************************************
**

**  File        : LinkerScript.ld
**
**
**  Abstract    : Linker script for STM32H7 series
**                2048Kbytes FLASH, 64Kbytes ITCMRAM, 128Kbytes DTCMRAM
**
**                Set heap size, stack size and stack location according
**                to application requirements.
**
**                Set memory bank area and size if external memory is used.
**
**  Target      : STMicroelectronics STM32
**
**  Distribution: The file is distributed as is without any warranty
**                of any kind.
**
*****************************************************************************
** @attention
**
** Copyright (C) 2026 niconiconi.
**
** Modified from stm32h745xx_flash_CM7.ld to stm32h743xx_flash_CM7.ld,
** for STM32H7 support, including project-specific ITCM/DTCM customizations.
**
** This file is free software: you may copy, redistribute and/or modify it
** under the terms of the BSD Zero Clause License, or (at your option)
** Creative Commons Zero v1.0 Universal license. See "SPDX-License-Identifier"
** for more details.
**
** This file is distributed in the hope that it will be useful, but WITHOUT ANY
** WARRANTY; without even the implied warranty of MERCHANTABILITY or FITNESS
** FOR A PARTICULAR PURPOSE.
**
** This file incorporates work covered by the following copyright and
** permission notice:
**
**   Copyright (c) 2019 STMicroelectronics.
**   All rights reserved.
**
**   This software is licensed under terms that can be found in the LICENSE
**   file in the root directory of this software component. If no LICENSE
**   file comes with this software, it is provided AS-IS.
**
** LICENSE file in the root directory of the origin:
**
**   Component: CMSIS Device
**   Copyright: ARM Limited, STMicroelectronics
**   License: Apache License 2.0
**
******************************************************************************
*/

/* Entry Point */
ENTRY(Reset_Handler)

/* Highest address of the user mode stack */
_estack = 0x20020000;    /* end of RAM */
/* Generate a link error if heap and stack don't fit into RAM */
_Min_Heap_Size = 0x200;      /* required amount of heap  */
_Min_Stack_Size = 0x400; /* required amount of stack */

/* Specify the memory areas */
MEMORY
{
FLASH (rx) : ORIGIN = 0x08000000, LENGTH = 2048K
ITCMRAM (xrw) : ORIGIN = 0x00000000, LENGTH = 64K
DTCMRAM (xrw) : ORIGIN = 0x20000000, LENGTH = 128K
AXISRAM (rw)    : ORIGIN = 0x24000000, LENGTH = 512K
SRAM1 (rw)      : ORIGIN = 0x30000000, LENGTH = 128K
SRAM2 (rw)      : ORIGIN = 0x30020000, LENGTH = 128K
SRAM3 (rw)      : ORIGIN = 0x30040000, LENGTH = 32K
SRAM4 (rw)      : ORIGIN = 0x38000000, LENGTH = 64K
}

/* Define output sections */
SECTIONS
{
  /* The startup code goes into ITCM (relocated from FLASH)  */
  .isr_vector :
  {
    . = ALIGN(4);
    __isr_vector_start = .;
    KEEP(*(.isr_vector)) /* Startup code */
    . = ALIGN(4);
    __isr_vector_end = .;
  } >ITCMRAM AT> FLASH

  .itcm_text :
  {
    . = ALIGN(4);
    __itcm_text_start = .;
    *(.itcm_text)
    *(.itcm_text*)
    . = ALIGN(4);
    __itcm_text_end = .;
  } > ITCMRAM AT> FLASH

  /* used by the startup to initialize data */
  _si_isr_vector = LOADADDR(.isr_vector);
  _si_itcm_text = LOADADDR(".itcm_text");

  /* The program code and other data goes into FLASH */
  .text :
  {
    . = ALIGN(4);
    *(.text)           /* .text sections (code) */
    *(.text*)          /* .text* sections (code) */
    *(.glue_7)         /* glue arm to thumb code */
    *(.glue_7t)        /* glue thumb to arm code */
    *(.eh_frame)

    KEEP (*(.init))
    KEEP (*(.fini))

    . = ALIGN(4);
    _etext = .;        /* define a global symbols at end of code */
  } >FLASH

  /* Constant data goes into FLASH */
  .rodata :
  {
    . = ALIGN(4);
    *(.rodata)         /* .rodata sections (constants, strings, etc.) */
    *(.rodata*)        /* .rodata* sections (constants, strings, etc.) */
    . = ALIGN(4);
  } >FLASH

  .ARM.extab (READONLY) : /* The READONLY keyword is only supported in GCC11 and later, remove it if using GCC10 or earlier. */
  {
    . = ALIGN(4);
    *(.ARM.extab* .gnu.linkonce.armextab.*)
    . = ALIGN(4);
  } >FLASH
  .ARM (READONLY) : /* The READONLY keyword is only supported in GCC11 and later, remove it if using GCC10 or earlier. */
  {
    . = ALIGN(4);
    __exidx_start = .;
    *(.ARM.exidx*)
    __exidx_end = .;
    . = ALIGN(4);
  } >FLASH

  .preinit_array (READONLY) : /* The READONLY keyword is only supported in GCC11 and later, remove it if using GCC10 or earlier. */
  {
    . = ALIGN(4);
    PROVIDE_HIDDEN (__preinit_array_start = .);
    KEEP (*(.preinit_array*))
    PROVIDE_HIDDEN (__preinit_array_end = .);
    . = ALIGN(4);
  } >FLASH
  .init_array (READONLY) : /* The READONLY keyword is only supported in GCC11 and later, remove it if using GCC10 or earlier. */
  {
    . = ALIGN(4);
    PROVIDE_HIDDEN (__init_array_start = .);
    KEEP (*(SORT(.init_array.*)))
    KEEP (*(.init_array*))
    PROVIDE_HIDDEN (__init_array_end = .);
    . = ALIGN(4);
  } >FLASH
  .fini_array (READONLY) : /* The READONLY keyword is only supported in GCC11 and later, remove it if using GCC10 or earlier. */
  {
    . = ALIGN(4);
    PROVIDE_HIDDEN (__fini_array_start = .);
    KEEP (*(SORT(.fini_array.*)))
    KEEP (*(.fini_array*))
    PROVIDE_HIDDEN (__fini_array_end = .);
    . = ALIGN(4);
  } >FLASH

  /* used by the startup to initialize data */
  _sidata = LOADADDR(.data);
  _si_sram1_data = LOADADDR(.sram1_data);
  _si_sram2_data = LOADADDR(.sram2_data);
  _si_sram3_data = LOADADDR(.sram3_data);
  _si_sram4_data = LOADADDR(.sram4_data);
  _si_axisram_data = LOADADDR(.axisram_data);

  /* Initialized data sections goes into RAM, load LMA copy after code */
  .data :
  {
    . = ALIGN(4);
    _sdata = .;        /* create a global symbol at data start */
    *(.data)           /* .data sections */
    *(.data*)          /* .data* sections */

    . = ALIGN(4);
    _edata = .;        /* define a global symbol at data end */
  } >DTCMRAM AT> FLASH

  .axisram_data :
  {
    . = ALIGN(4);
    __axisram_data_start = .;
    *(.axisram_data)
    *(.axisram_data*)
    . = ALIGN(4);
    __axisram_data_end = .;
  } >AXISRAM AT> FLASH

  .sram1_data :
  {
    . = ALIGN(4);
    __sram1_data_start = .;
    *(.sram1_data)
    *(.sram1_data*)
    . = ALIGN(4);
    __sram1_data_end = .;
  } >SRAM1 AT> FLASH

  .sram2_data :
  {
    . = ALIGN(4);
    __sram2_data_start = .;
    *(.sram2_data)
    *(.sram2_data*)
    . = ALIGN(4);
    __sram2_data_end = .;
  } >SRAM2 AT> FLASH

  .sram3_data :
  {
    . = ALIGN(4);
    __sram3_data_start = .;
    *(.sram3_data)
    *(.sram3_data*)
    . = ALIGN(4);
    __sram3_data_end = .;
  } >SRAM3 AT> FLASH

  .sram4_data :
  {
    . = ALIGN(4);
    __sram4_data_start = .;
    *(.sram4_data)
    *(.sram4_data*)
    . = ALIGN(4);
    __sram4_data_end = .;
  } >SRAM4 AT> FLASH

  /* Uninitialized data section */
  . = ALIGN(4);
  .bss :
  {
    /* This is used by the startup in order to initialize the .bss section */
    _sbss = .;         /* define a global symbol at bss start */
    __bss_start__ = _sbss;
    *(.bss)
    *(.bss*)
    *(COMMON)

    . = ALIGN(4);
    _ebss = .;         /* define a global symbol at bss end */
    __bss_end__ = _ebss;
  } >DTCMRAM

  /* User_heap_stack section, used to check that there is enough RAM left */
  ._user_heap_stack :
  {
    . = ALIGN(8);
    PROVIDE ( end = . );
    PROVIDE ( _end = . );
    . = . + _Min_Heap_Size;
    . = . + _Min_Stack_Size;
    . = ALIGN(8);
  } >DTCMRAM



  /* Remove information from the standard libraries */
  /DISCARD/ :
  {
    libc.a ( * )
    libm.a ( * )
    libgcc.a ( * )
  }

  .ARM.attributes 0 : { *(.ARM.attributes) }
}
« Last Edit: May 16, 2026, 02:50:26 pm by niconiconi »
 
The following users thanked this post: SpacedCowboy


Share me

Digg  Facebook  SlashDot  Delicious  Technorati  Twitter  Google  Yahoo
Smf

 

-->