Necrobumping an old thread for the record.
Timing
On the STM32H743 (SYSCLK = 480 MHz from PLL1, Revision V), the absolute best pure GPIO input-to-output IRQ latency I was able to achieve was 70 ns, WFE busy-polling latency was 60 ns. The test code was written in C, HAL was used, but without calling high-level HAL functions in the handler, SYSTICK is also turned off. Both the vector table and the IRQ handler running in ITCM, and the handler does nothing but to toggle the GPIO's ODR register, clear the IRQ pending bit, and return after a slight NOP delay loop (because the cleared flag takes time to propagate to NVIC). All code complied by GCC with -O3 and Link Time Optimization (LTO) enabled to maximize inlining.
Source code available upon request (because it no longer exists, and takes time to replicate).
If the vector table was not properly relocated by reprogramming VTOR, the latency is increased to 100 ns. This can be solved by either enabling the cache (not recommending), or by relocating the vector table. When both the vector table and the handler are relocated in ITCM, ICache and DCache have no effect on the latency. In fact enabling the cache becomes a useful debugging tool - if latency decreases, Flash reads are performed by mistake, which must be corrected.
Instead of using an IRQ, busy-polling the EXTI event in the main loop via _WFE() instead of an IRQ reduced it further to 60 ns.
In comparison, a naive HAL example would take around 300 to 500 ns because of the slow Flash and the dynamic decision-making code in the HAL.
I also found a GPIO pulse can’t be shorter than 25 ns, which means software-controlled GPIO toggling is limited to a 20 MHz square wave. Perhaps there's a way, but so far I’m not able to overcome this limitation.
Phase-Shifted Clock Generation
If the external pulse is periodic, such as a clock, I found it’s possible to hide the IRQ latency for all but the first pulse by triggering an internal timer and wait for the timer IRQ instead. The timer is programmed to fire slightly before next edge of the real pulse arrives. For example, if the pulse has a period of 1000 ns and the IRQ latency is 100 ns, you can use the GPIO pulse to trigger a timer to count for 900 ns (100 ns earlier), effectively predicting the next pulse and eliminating the IRQ latency for all subsequent IRQs.
The timer is triggered from GPIO via TIMx_CHx, or TIMx_ETR on every rising edge, so it's always phase-locked with the source with no long-term drift.
Unsuccessful Peripheral Exploitation
There's a general agreement online that STM32H7 has a notoriously slow GPIO controller, because it's on the D3 AHB4 bus, far away from the CPU's D1 domain, requests travel from the AXI port to the AXI matrix, to the AXI-to-D3 AHB bridge, to the D3 AHB matrix, to the GPIO controller. The GPIO pins themselves are much faster (> 100 MHz) in alternate function modes, but they are not designed for bitbanging.
In theory, we can speed it up by two methods.
BDMA
The first idea is to try driving it via the BDMA controller in domain D3, so the traffic stays locally in the same domain and matrix. Unfortunately, the 25 ns barrier remains, even if the transactions are initiated by BDMA. This makes me suspect that the long delay from the CPU is not the main source of latency, instead, the culprit is perhaps the synchronization or wait states within the GPIO controller itself or the AHB4. Alternatively, BDMA and SRAM4 themselves are too slow, so the gain of data locality is nearly canceled out.
static void dma_config(void)
{
__attribute__((section (".sram4_data")))
static const uint32_t gpio_data[8] = {
0xFFFFFFFF, 0x00000000, 0xFFFFFFFF, 0x00000000,
0xFFFFFFFF, 0x00000000, 0xFFFFFFFF, 0x00000000,
};
LL_AHB4_GRP1_EnableClock(LL_AHB4_GRP1_PERIPH_BDMA);
static LL_BDMA_InitTypeDef bdma_ctx = {
.PeriphOrM2MSrcAddress = (uint32_t) &gpio_data,
.MemoryOrM2MDstAddress = (uint32_t) &GPIOA->ODR,
.Direction = LL_BDMA_DIRECTION_MEMORY_TO_MEMORY,
.Mode = LL_BDMA_MODE_NORMAL,
.PeriphOrM2MSrcIncMode = LL_BDMA_PERIPH_INCREMENT,
.MemoryOrM2MDstIncMode = LL_BDMA_MEMORY_NOINCREMENT,
.PeriphOrM2MSrcDataSize = LL_BDMA_PDATAALIGN_WORD,
.MemoryOrM2MDstDataSize = LL_BDMA_MDATAALIGN_WORD,
.NbData = sizeof(gpio_data) / sizeof(gpio_data[0]),
.PeriphRequest = LL_DMAMUX2_REQ_MEM2MEM,
.Priority = LL_BDMA_PRIORITY_VERYHIGH,
.DoubleBufferMode = LL_BDMA_DOUBLEBUFFER_MODE_DISABLE,
/* don't care */
.TargetMemInDoubleBufferMode = LL_BDMA_CURRENTTARGETMEM0,
};
LL_BDMA_Init(BDMA, LL_BDMA_CHANNEL_0, &bdma_ctx);
}
static void mainloop(void)
{
dma_config();
while (true) {
LL_BDMA_EnableChannel(BDMA, LL_BDMA_CHANNEL_0);
LL_mDelay(1000);
LL_BDMA_DisableChannel(BDMA, LL_BDMA_CHANNEL_0);
LL_BDMA_SetDataLength(BDMA, LL_BDMA_CHANNEL_0, 8);
}
}
I also tried to use BDMA's "data unpacking" feature, with the hope that the BDMA can read once from SRAM and write 4 times.
.PeriphOrM2MSrcDataSize = LL_BDMA_PDATAALIGN_WORD,
.MemoryOrM2MDstDataSize = LL_BDMA_MDATAALIGN_BYTE,
But it also doesn't overcome the 25 ns barrier, suggesting either the BDMA doesn't support optimized data unpacking (generating 4 fetches or 3 dummy cycles instead), or AHB4 bus write / GPIO controller itself is bottlenecked.
Peripheral Abuse
Another way is to abuse peripherals on the D1 AXI bus instead for bitbanging, such as the SDMMC, LTDC, FMC, QuadSPI. I decided to try the QSPI because it was easy to use, but without success.
I found the QSPI controller has a "dual Flash" mode, each half is 4-bit, which is potentially a useful 8-bit bus transmitter. But latency also did not improve by abusing QSPI controller at 133 MHz to generate a write request. I got a < 10 ns pulse train, which broke the 25 ns pulse width barrier, but in fact the latency slightly degraded, even if the QSPI controller is physically close to the CPU. I blame the half-cycle chip enable delay. To write Flash in "indirect mode" (e.g. register-based, not memory-mapped), the first half-clock cycle is used for enabling /CE before the first bit is transmitted, I don't see anyway to skip this step. The /CE is held low only if the SPI Flash is in memory-mapped mode. If I read the datasheet correctly, this is used for Flash prefetching, meaning that the controller will keep generating commands on the 8-bit bus, writing garbage on the data bus. I don't see a way to have no prefetching and no /CE delay, perhaps there's a hack (potential solution that I didn't try: if we're using MMIO writes only, the controller is allowed to prefetch garbage, but we can deselect the alternate function to tristate the pins).
But I ran out of patience. Perhaps it can be improved through alternative solutions, such as the QSPI MMIO mode, or by exploiting SDMMC/LTDC/FMC.