Author Topic: STM32: finally data sampling delay option for SPI master has arrived!  (Read 3716 times)

0 Members and 3 Guests are viewing this topic.

Offline KarelTopic starter

  • Super Contributor
  • ***
  • Posts: 2544
  • Country: 00
Quote
SPI/I2S configuration register 1 (SPI_CFG1)

Bit 24DRDS: data sampling delay on master input (MISO)

  This delay is used to compensate the delay on the MISO line (for example due to the signal optical isolation).
 
  0: Sampling is performed on the configured SCK edge (no delay is applied)
  1: Sampling is postponed by a fixed delay equal to a half SCK cycle
     (the sampling is performed just at the end of the bit validity interval at MISO line).

https://www.st.com/resource/en/reference_manual/rm0481-stm32h52333xx-stm32h56263xx-and-stm32h573xx-armbased-32bit-mcus-stmicroelectronics.pdf

Why did it take so long for them to implement this  :-//
 

Offline mark03

  • Frequent Contributor
  • **
  • Posts: 782
  • Country: us
Re: STM32: finally data sampling delay option for SPI master has arrived!
« Reply #1 on: September 06, 2026, 09:41:17 pm »
Assuming it's the "new" SPI (or a minor revision of the one in the H7 series), it hardly matters.  I defy anyone to produce a working driver based on the nonsense that passes for a theory of operation in the Reference Manual.  OK, that is a slight exaggeration, but only just.
 
The following users thanked this post: pardo-bsso

Offline KarelTopic starter

  • Super Contributor
  • ***
  • Posts: 2544
  • Country: 00
Re: STM32: finally data sampling delay option for SPI master has arrived!
« Reply #2 on: September 06, 2026, 10:02:07 pm »
I did. No HAL (not even peeking at it), purely based on the ref manual.
I admit I struggled at times and I only used master mode, no slave mode.
It also works with DMA, TX only in order to drive an oled screen
(Newhaven Display NHD-3.12-25664UCB2 with SSD1322 controller).
I also admit that I was cursing the manual sometimes because of lack of clarity and lack of examples.

Speaking of DMA, the "new" linked-list for their GPDMA caused me some headache as well but in the end
I managed.
After all, I'm very content with it (the STM32H5xx series that is).
 
The following users thanked this post: pardo-bsso

Offline mark03

  • Frequent Contributor
  • **
  • Posts: 782
  • Country: us
Re: STM32: finally data sampling delay option for SPI master has arrived!
« Reply #3 on: September 06, 2026, 10:07:57 pm »
I did as well, but it took me much longer than it should have, and I only got there by ruthlessly ignoring most of the supposed advantages of the new peripheral.  For example, I am only transferring one byte at a time with the DMA.  I imagine that I could make a fancier driver if I threw enough time at it, but at that point it would be well into the realm of reverse engineering, and I already feel bad enough for being an enabler of their (ST's) poor behavior.

It would probably be worth starting a community documentation project, but a part of me feels the same way about that: at the end of the day it just gives ST a free pass.  I would want to find some way of stealing from them as they have stolen from us, and as I can't think of one, here we are  |O
 

Offline NorthGuy

  • Super Contributor
  • ***
  • Posts: 3527
  • Country: ca
Re: STM32: finally data sampling delay option for SPI master has arrived!
« Reply #4 on: September 06, 2026, 11:57:54 pm »
Why did it take so long for them to implement this  :-//

Such a simple and useful feature. They did have that in their QUAD SPI flash modules from the beginning, but for some reason held it away from SPI. Now you can do much faster SPI. It is in C5 too.

 

Offline Jeroen3

  • Super Contributor
  • ***
  • Posts: 4577
  • Country: nl
  • Embedded Engineer
    • jeroen3.nl
Re: STM32: finally data sampling delay option for SPI master has arrived!
« Reply #5 on: September 07, 2026, 05:55:37 am »
This does not give you fine control of the delay...
If you enable this when you're not running fast enough to have phase shift due to isolators, you're sampling at the transition?
 

Online hans

  • Super Contributor
  • ***
  • Posts: 1968
  • Country: 00
Re: STM32: finally data sampling delay option for SPI master has arrived!
« Reply #6 on: September 07, 2026, 07:44:42 am »
Assuming it's the "new" SPI (or a minor revision of the one in the H7 series), it hardly matters.  I defy anyone to produce a working driver based on the nonsense that passes for a theory of operation in the Reference Manual.  OK, that is a slight exaggeration, but only just.
I did. And even on STM32U5 with LPTIM1 + LPGPIO + LPDMA and MSIK! Or LPBAM! Or SRD! Or whatever they called it this week! :P

From this autonomous perspective with chip-select automata in hardware, I get why they designed it this way.
Its been a few years since I wrote that, but I put a linked descriptor chain of 2x the length compared to the transfer size of the buffer I wanted to sample. Each transfer could only do 1 SPI frame at a time (which was 16-bits), as after a frame the SPI deasserts CS and needed to be rearmed.
So the DMA was a chain consisting of alternating RXDR reads and then a IFC write to re-arm the DMA again. To grab 64x16-bit external ADC samples, I had to have a DMA descriptor chain of 128 blocks, each taking 4 uint32's, or 2kiB of RAM..
 :clap: :clap:

Code: [Select]
static void _lpdmaChainGo(uint8_t ch, uint_fast8_t pos) {
    assert(ch < 4);
    assert(pos < MAX_CHAIN_LEN);

    DMA_Channel_TypeDef* io = reinterpret_cast<DMA_Channel_TypeDef*>(LPDMA1_BASE_NS + 0x50 + 0x80*ch);

    io->CTR2 = lp_chain[pos].CTR2;
    io->CSAR = lp_chain[pos].SAR;
    io->CDAR = lp_chain[pos].DAR;
    io->CLLR = lp_chain[pos].LLR;
    io->CBR1 = 2;

    assert(io->CDAR != 0);
    assert(io->CSAR != 0);

    // GO
    BitsSet(io->CCR, DMA_CCR_EN);
}

/*** LPDMA as chain ***/
struct LPDMA_MultiTransferChain_Single {
    uint32_t CTR2;
    uint32_t SAR;
    uint32_t DAR;
    uint32_t LLR;
};
#define MAX_CHAIN_LEN 128
ATTR_DATA_SRAM4 volatile LPDMA_MultiTransferChain_Single lp_chain[2*MAX_CHAIN_LEN] __attribute((aligned(16)));
ATTR_DATA_SRAM4 volatile uint16_t lp_spi_ifc;
ATTR_DATA_SRAM4 volatile uint16_t lp_tim3_arr;

static void _lpdmaChainSpiRx(uint32_t ch, uint16_t* dst, size_t length) {
    assert(length <= MAX_CHAIN_LEN);
    assert((length & 1) == 0);

    // The chain consists of Nx2 requests:
    // Each request is served in bursts.
    // The first request, waits for SPI3_RX request to operate. The SPI3 is set to trigger on LPTIM3_Ch1
    // After the SPI3 is triggrered, it needs to be rearmed as the EOT flag is set. This is normally a CPU activity..
    // but, this is what the 2nd DMA request is for. The 2nd DMA request writes the EOT-clear word to the SPI3 peripheral

    // The requests are linked in a chain of Nx2 long. The last request is configured as NULL and needs to be re-armed
    // by the software IRQ. The half/full parts are handled by software. Therefore, the chain is replicated twice
    // so all DAR values to the dst buffer are preloaded.
    // So:
    //                +-----+-----+-----+-----+---------+-----+-----+-----+-----+-----+-----+---------+-----+-----+
    // Chain idx      |   0 |   1 |   2 |   3 | ... ... |  62 |  63 |  64 |  65 |  66 |  67 | ... ... | 126 | 127 |
    // Peripheral     | SPI | SPI | SPI | SPI |   ...   | SPI | SPI | SPI | SPI | SPI | SPI |   ...   | SPI | SPI |
    // Register       | RDR | IFC | RDR | IFC |   ...   | RDR | IFC | RDR | IFC | RDR | IFC |   ...   | RDR | IFC |
    // SAR            | Reg | Mem | Reg | Mem |   ...   | Reg | Mem | Reg | Mem | Reg | Mem |   ...   | Reg | Mem |
    // DAR            | Mem | Reg | Mem | Reg |   ...   | Mem | Reg | Mem | Reg | Mem | Reg |   ...   | Mem | Reg |
    // ULL (^=update) | ^ 1 | ^ 2 | ^ 3 | ^ 4 |   ...   | ^63 | IRQ | ^65 | ^66 | ^67 | ^68 |   ...   | ^127| IRQ |
    //                +-----+-----+-----+-----+---------+-----+-----+-----+-----+-----+-----+---------+-----+-----+
    // At idx=63 and idx=127, a software IRQ is generated to re-arm the peripheral and work on data.
    // Note that IDX=0 and IDX=64 are written to the buffer for convenience, but are then copied to hardware..
    assert(ch < 4);

    __HAL_RCC_LPDMA1_CLK_ENABLE();
    _lpdmaClr(ch);

    auto tfer_size = DMA_CTR1_DDW_LOG2_0 | DMA_CTR1_SDW_LOG2_0; // 0b00=8-bit, 0b01=16-bit, 0b10=32-bit, 0b11=illegal
    auto tfer_bytes = BitsAnySet(tfer_size, DMA_CTR1_DDW_LOG2_0) ? 2 :
                      (BitsAnySet(tfer_size, DMA_CTR1_DDW_LOG2_1) ? 4 : 1);
    assert(tfer_bytes != 4);

    DMA_Channel_TypeDef* io = reinterpret_cast<DMA_Channel_TypeDef*>(LPDMA1_BASE_NS + 0x50 + 0x80*ch);

    BitsSet(io->CCR, DMA_CCR_TCIE/* | DMA_CCR_DTEIE | DMA_CCR_ULEIE | DMA_CCR_USEIE | DMA_CCR_SUSPIE | DMA_CCR_TOIE*/); // full transfer complete interrupts
    BitsSet(io->CCR, DMA_CCR_PRIO_0 | DMA_CCR_PRIO_1); // high priority
    BitsWrite(io->CTR1, DMA_CTR1_DDW_LOG2 | DMA_CTR1_SDW_LOG2, tfer_size);
    BitsWrite(io->CLBAR, DMA_CLBAR_LBA, reinterpret_cast<intptr_t>(lp_chain));
    BitsWrite(io->CBR1, DMA_CBR1_BNDT, tfer_bytes);

    uint32_t ctr2_rdr = 0;
    uint32_t ctr2_ifc = 0;
    uint32_t llr_flags = 0;

    // Req 2 = SPI_RX:
    BitsWrite(ctr2_rdr, DMA_CTR2_REQSEL, 2);
    BitsWrite(ctr2_rdr, DMA_CTR2_TCEM, DMA_CTR2_TCEM); // TCME 0b11 => transfer complete IRQ at end of channel (all LLIs processed)

    // SWreq:
    BitsSet(ctr2_ifc, DMA_CTR2_SWREQ);
    BitsWrite(ctr2_ifc, DMA_CTR2_TCEM, DMA_CTR2_TCEM); // TCME 0b11 => transfer complete IRQ at end of channel (all LLIs processed)

    // Transfer link-register, source/destination and CTR2 register
    BitsSet(llr_flags, DMA_CLLR_ULL | DMA_CLLR_USA | DMA_CLLR_UDA | DMA_CLLR_UT2);

    lp_spi_ifc = SPI_IFCR_EOTC | SPI_IFCR_TXTFC;

    auto lli_head = lp_chain;
    for (auto i = 0u; i < length; i++) {
        auto lli_rdr = lli_head++;
        auto lli_ifc = lli_head++;

        lli_rdr->CTR2 = ctr2_rdr;
        lli_ifc->CTR2 = ctr2_ifc;

        lli_rdr->SAR = reinterpret_cast<intptr_t>(&(SPI3->RXDR));
        lli_rdr->DAR = reinterpret_cast<intptr_t>(&(dst[i]));

        lli_ifc->SAR = reinterpret_cast<intptr_t>(&(lp_spi_ifc));
        lli_ifc->DAR = reinterpret_cast<intptr_t>(&(SPI3->IFCR));

        lli_rdr->LLR = llr_flags | (0xFFFC & reinterpret_cast<intptr_t>(lli_ifc));
        if (((i + 1) % (length / 2)) == 0) { // last word!
            lli_ifc->LLR = 0;
        } else {
            lli_ifc->LLR = llr_flags | (0xFFFC & reinterpret_cast<intptr_t>(lli_head));
        }
    }

    // TODO: flush dcache/SRAM4 blk to ensure it is written to SRAM4
}

void initSPI() {

    // Initialize TIM3 as sampling clock
    uint32_t period = 32768 / symbolrate - 1;

    // LPTIM3: dma source (LPDMA1 req: 14/16)
    __HAL_RCC_LPTIM3_CLK_ENABLE();

    BitsSet(LPTIM3->CR, LPTIM_CR_ENABLE);
    LPTIM3->ARR = period;
    LPTIM3->CCR1 = 0;
    BitsSet(LPTIM3->DIER, LPTIM_DIER_UEDE);
    BitsClear(LPTIM3->CCMR1, LPTIM_CCMR1_CC1SEL);
    BitsClear(LPTIM3->CCMR1, LPTIM_CCMR1_CC1E);

        __HAL_RCC_SPI3_CLK_ENABLE();
        SPI3->CR1 = 0;
        SPI3->IFCR = 0xBF8; // clear all IRQs;

        SPI3->CR1 = SPI_CR1_MASRX;
        SPI3->CFG1 = SPI_CFG1_DSIZE_3 | // 0b01xxx = 16-bits transfers for SPI3 (limited feature set)
                     SPI_CFG1_RXDMAEN |
                     SPI_CFG1_BPASS;
        SPI3->CFG2 = SPI_CFG2_AFCNTR |
                     SPI_CFG2_COMM_1 | // 0b10= Rx only
                     SPI_CFG2_MASTER |
                     SPI_CFG2_SSM |
                     SPI_CFG2_SSOE |
                     SPI_CFG2_AFCNTR |
                    (0x1 << SPI_CFG2_MSSI_Pos);
        SPI3->AUTOCR = 7 << SPI_AUTOCR_TRIGSEL_Pos; // TRG7 = Lptim3_ch1
        BitsSet(SPI3->AUTOCR, SPI_AUTOCR_TRIGEN);
        SPI3->CR2 = 1; // burst-size
        BitsSet(SPI3->CR1, SPI_CR1_SPE);


        _lpdmaChainSpiRx(LPDMACH_SPI3, adcBuffer.getHwCircPtr(), adcBuffer.getHwCircSize());
        _lpdmaChainGo(LPDMACH_SPI3, 0);

        NVIC_EnableIRQ(LPDMA1_Channel1_IRQn);
}

But hey, stepping asides my cynical tone.. it is possible with this newer SPI peripheral! It wasn't with the old version, because as soon as SPI went active it would keep NSS asserted.. or it needs to be under software control which bypasses the whole point of using autonomous operation.
So this code can run on the U5 whilst the CPU subsystem is off. It samples an external low-power differential 12-bit ADC here at 1sps-16.384ksps. I did some measurements on the power consumption of the chip.. at a few dozen samples per second its right down at a few uA of sleep current. Iirc at 8ksps the frequent MSIK restarts did push it up to 30-40uA though.

The H5 has the same linked-list DMA peripheral, but not the SRD feature I think. We waited on a linked-list descriptor function for quite a while too... the H7 chips had it, but not on all DMA's.

Pretty much all STM32s are copy-paste on peripherals side, depending what is the latest IP version in stock. I presume they develop all their DMA, SPI, etc. IP in house, and I get it why they don't want to push too many new features at once to new silicon (avoids a PIC32MZ disaster). But honestly, chips like C5 vs H5 look very much like an economical decision of "how much MCU do you want"..
 

Offline KarelTopic starter

  • Super Contributor
  • ***
  • Posts: 2544
  • Country: 00
Re: STM32: finally data sampling delay option for SPI master has arrived!
« Reply #7 on: September 07, 2026, 07:50:03 am »
This does not give you fine control of the delay...
If you enable this when you're not running fast enough to have phase shift due to isolators, you're sampling at the transition?

When you use optical or digital isolators at a relatively high speed e.g. 20 MHz, you simply enable it. No need to have fine control.
If you don't use any isolators or you use a low clockspeed, you simply disable it. By default it's disabled.
The PIC32MX series from Microchip had this feature for ages, I used it for an EKG device that uses the ADS1298 and digital
isolators. It avoids the need the use a second spi interface as slave with an extra isolator channel in order to compensate
for the delay.

It's an extremely useful feature.

 

Offline Jeroen3

  • Super Contributor
  • ***
  • Posts: 4577
  • Country: nl
  • Embedded Engineer
    • jeroen3.nl
Re: STM32: finally data sampling delay option for SPI master has arrived!
« Reply #8 on: September 07, 2026, 08:01:05 am »
Yeah, I recognize that. But this feels a bit like a absolute minimal workaround for this problem. But if it works  :-//
 

Online langwadt

  • Super Contributor
  • ***
  • Posts: 5784
  • Country: dk
Re: STM32: finally data sampling delay option for SPI master has arrived!
« Reply #9 on: September 07, 2026, 10:07:20 am »
Yeah, I recognize that. But this feels a bit like a absolute minimal workaround for this problem. But if it works  :-//

using the same edge for setting and sampling maximizes setup time, hold time needed is often close to zero, and as long as time doesn't go backwards that is a given
 

Offline peter-h

  • Super Contributor
  • ***
  • Posts: 6020
  • Country: gb
  • Doing electronics since the 1960s...
Re: STM32: finally data sampling delay option for SPI master has arrived!
« Reply #10 on: September 07, 2026, 10:37:57 am »
It sounds like you get an option of a delay in sampling of half the clock period so at say 10MHz you can select a delay of 50ns.

The Q is when is this useful.

I've been using optoisolation for many years. The main issue is making it work fast enough without drawing too much power. And then the delay becomes CTR dependent. I've build some fast stuff using slow optos, with fancy signal handling around it. But it needs more components... The whole issue is solved by using those GHz internal oscillator couplers e.g. SI8641 which are very fast.

What would be much more useful would be a sampling point programmed in CPU clock ticks :)
Z80 Z180 Z280 Z8 S8 8031 8051 H8/300 H8/500 80x86 90S1200 32F417
 

Online hans

  • Super Contributor
  • ***
  • Posts: 1968
  • Country: 00
Re: STM32: finally data sampling delay option for SPI master has arrived!
« Reply #11 on: September 07, 2026, 11:37:32 am »
The MOSI/MISO will change data output on 1 clock edge, and sample them on the other edge. SCK has to travel to devices, setup/hold time of device, MISO responds, travel back to STM32, setup/hold time of STM32. And because the data changes relative to edge 1 and is sampled at edge 2, the timing has to be met within 50% of the clock period time. With delayed sampling, this becomes 75%, because it samples half way between the 2nd edge and the 1st edge of the next clock cycle..

The OCTOSPI and SDMCC peripheral actually have a DLYB peripheral to shift the output and input clocks apart, so it has a wider margin. But iirc the OCTOSPI also had a mid-sampling bit within it.
 

Offline NorthGuy

  • Super Contributor
  • ***
  • Posts: 3527
  • Country: ca
Re: STM32: finally data sampling delay option for SPI master has arrived!
« Reply #12 on: September 07, 2026, 12:51:40 pm »
This does not give you fine control of the delay...
If you enable this when you're not running fast enough to have phase shift due to isolators, you're sampling at the transition?

This is not a fine control, but it does give you improvements. Roughly, it'll work if:

- hold requirement is less than round-trip delay
- the sum of setup requirement and round-trip delay is less than clock-period

For example, for H5F chips hold time is 1 ns and setup time is 3.5 ns.

To satisfy the hold time you need a round-trip delay more than 1 ns. This is practically guaranteed and doesn't depend on the frequency.

To satisfy setup time the round trip delay must be less than clock period minus 3.5 ns. If your round-trip delay is 10 ns, the shortest clock period you can master is 10 + 3.5 = 13.5 ns - roughly 75 MHz.

With regular way of sampling, you sample half-period earlier. The signal now have only half of the clock period to arrive. This half must be at most 13.5 ns, hence the clock period must be at least 27 ns - roughly 38 MHz.

That is how the new feature lets you go from 40 MHz to 75 MHz.

Similarly, if you're using slow isolators with 200 ns round-trip delay, you can do 4.9 MHz with delayed sampling compared to 2.4 MHz with regular sampling. So, this benefits slowpokes just as well.
 

Offline Jeroen3

  • Super Contributor
  • ***
  • Posts: 4577
  • Country: nl
  • Embedded Engineer
    • jeroen3.nl
Re: STM32: finally data sampling delay option for SPI master has arrived!
« Reply #13 on: September 07, 2026, 01:04:32 pm »
Yes I can see it working. But it just feels like the wrong fix for the problem at hand. But sometimes that is engineering  :P

A programmable MOSI gate clock skew would be much more preferable rather than a clock invert. But skew sees side effects in the implementation, how many ticks you can skew depends on the ratio spi_pclk to sck, probably requiring significant changes in the dividers elsewhere in the parts. Raising the complexity.

I suspect this invert was a really small and cheap fix though, most if the logic is already there.
Catering to many, so it's an easy win.
 

Offline NorthGuy

  • Super Contributor
  • ***
  • Posts: 3527
  • Country: ca
Re: STM32: finally data sampling delay option for SPI master has arrived!
« Reply #14 on: September 07, 2026, 01:28:47 pm »
A programmable MOSI gate clock skew would be much more preferable rather than a clock invert. But skew sees side effects in the implementation, how many ticks you can skew depends on the ratio spi_pclk to sck, probably requiring significant changes in the dividers elsewhere in the parts. Raising the complexity.

This feature doesn't affect MOSI, only MISO.
 

Online SiliconWizard

  • Super Contributor
  • ***
  • Posts: 17795
  • Country: fr
Re: STM32: finally data sampling delay option for SPI master has arrived!
« Reply #15 on: September 07, 2026, 03:58:18 pm »
Quote
SPI/I2S configuration register 1 (SPI_CFG1)

Bit 24DRDS: data sampling delay on master input (MISO)

  This delay is used to compensate the delay on the MISO line (for example due to the signal optical isolation).
 
  0: Sampling is performed on the configured SCK edge (no delay is applied)
  1: Sampling is postponed by a fixed delay equal to a half SCK cycle
     (the sampling is performed just at the end of the bit validity interval at MISO line).

https://www.st.com/resource/en/reference_manual/rm0481-stm32h52333xx-stm32h56263xx-and-stm32h573xx-armbased-32bit-mcus-stmicroelectronics.pdf

Why did it take so long for them to implement this  :-//

That's something you can have on QSPI peripherals, and SDIO, but for SPI, it's still pretty rare. Just a matter of not that many use cases. Also, most peripherals in STM32 MCUs are third-party IPs (such as Synopsys) and upgrading to a more complex one has a cost they need to justify. Don't assume they implement everything in-house, that's not the case.

But yes, that could be useful. I not so long ago implemented a serial isolated link using SPI up to 30 MHz and that feature would have proven useful at the higher bus frequencies, with the digital isolators introducing a small delay, but enough to cause potential issues.

 

Online langwadt

  • Super Contributor
  • ***
  • Posts: 5784
  • Country: dk
Re: STM32: finally data sampling delay option for SPI master has arrived!
« Reply #16 on: September 07, 2026, 10:40:06 pm »
A programmable MOSI gate clock skew would be much more preferable rather than a clock invert. But skew sees side effects in the implementation, how many ticks you can skew depends on the ratio spi_pclk to sck, probably requiring significant changes in the dividers elsewhere in the parts. Raising the complexity.

This feature doesn't affect MOSI, only MISO.

and there using the same edge make sense so you get almost a full cycle for setup, the STM has close to zero hold and since data can't possibly change before the clock that will always be more than zero

 

Offline Jeroen3

  • Super Contributor
  • ***
  • Posts: 4577
  • Country: nl
  • Embedded Engineer
    • jeroen3.nl
Re: STM32: finally data sampling delay option for SPI master has arrived!
« Reply #17 on: September 08, 2026, 06:46:33 am »
A programmable MOSI gate clock skew would be much more preferable rather than a clock invert. But skew sees side effects in the implementation, how many ticks you can skew depends on the ratio spi_pclk to sck, probably requiring significant changes in the dividers elsewhere in the parts. Raising the complexity.

This feature doesn't affect MOSI, only MISO.
I'm sorry, classic mistake.
 

Online asmi

  • Super Contributor
  • ***
  • Posts: 3334
  • Country: ca
Re: STM32: finally data sampling delay option for SPI master has arrived!
« Reply #18 on: September 08, 2026, 02:02:29 pm »
SPI was never designed for high speed operation, hence the lack of source-synchronous receive clock for master. It does specify that data capture should be on the "inactive" clock edge, which gives 1/2 SCLK period for data to arrive and stabilize. With 6ps/mm average propagation delay in PCB you need a VERY long traces for trace delay to become an issue, or to run an interface at insane frequency. Practically speaking, not many SPI slave devices can go faster than about 50 MHz, so in practice this usually is not a problem. The one thing which changes this equation a lot is if you need to have a voltage translator between a master and a slave. These things can easily have a few ns of propagation delay, eating your timing budget in no time. I actually had this case when I was working on my Artix board with SODIMM - due to pinout limitations, I had to use an IO bank containing configuration-related balls for DDR3 SODIMM interface, which means it had to run at 1.5 V, while my config QSPI flash was 1.8 V, so I had to use a voltage translator. It reduced the max frequency I can run flash by half from ~100 MHz to about 50 MHz.

Offline KarelTopic starter

  • Super Contributor
  • ***
  • Posts: 2544
  • Country: 00
Re: STM32: finally data sampling delay option for SPI master has arrived!
« Reply #19 on: September 08, 2026, 03:14:48 pm »
The one thing which changes this equation a lot is if you need to have a voltage translator between a master and a slave.

Or if you need to have an optical/digital isolator...
 

Online asmi

  • Super Contributor
  • ***
  • Posts: 3334
  • Country: ca
Re: STM32: finally data sampling delay option for SPI master has arrived!
« Reply #20 on: September 08, 2026, 03:29:05 pm »
Or if you need to have an optical/digital isolator...
It's worth noting that some specialized voltage translators (for example NVT4858HK for SD/SDIO bus) provide a feedback clock, which is input clock which went through the voltage translation in both directions. But I've never seen such feature for regular SPI.

Offline peter-h

  • Super Contributor
  • ***
  • Posts: 6020
  • Country: gb
  • Doing electronics since the 1960s...
Re: STM32: finally data sampling delay option for SPI master has arrived!
« Reply #21 on: September 08, 2026, 03:35:21 pm »
Quote
SPI was never designed for high speed operation

Not gigabit speeds, no :)

But I use it very nicely at 42MHz for a ST7789 LCD controller.
Z80 Z180 Z280 Z8 S8 8031 8051 H8/300 H8/500 80x86 90S1200 32F417
 

Offline NorthGuy

  • Super Contributor
  • ***
  • Posts: 3527
  • Country: ca
Re: STM32: finally data sampling delay option for SPI master has arrived!
« Reply #22 on: September 08, 2026, 04:01:19 pm »
I can run flash by half from ~100 MHz to about 50 MHz.

Xilinx can do late sampling when loading the configuration:

Code: [Select]
set_property BITSTREAM.CONFIG.SPI_FALL_EDGE YES [current_design]
 

Offline Jeroen3

  • Super Contributor
  • ***
  • Posts: 4577
  • Country: nl
  • Embedded Engineer
    • jeroen3.nl
Re: STM32: finally data sampling delay option for SPI master has arrived!
« Reply #23 on: September 09, 2026, 05:35:51 am »
Do we have an alternative for high speed on board interconnects that is unaffected by this delay effect?
 

Online AndyC_772

  • Super Contributor
  • ***
  • Posts: 4562
  • Country: gb
  • Professional design engineer
    • Cawte Engineering | Reliable Electronics
Re: STM32: finally data sampling delay option for SPI master has arrived!
« Reply #24 on: September 09, 2026, 08:38:27 am »
In my experience the limitations are usually:

- device has a maximum specified operating frequency (SCLK max)
- STM32 SPI peripheral has a coarse, limited range of operating frequencies to choose from (HCLK / n)

Pick the highest frequency available that's below the maximum for the device, and that's the best you're going to get. It probably won't be nearly as fast as you'd like, and you need to drive CS# (NSS) in software anyway.

Fortunately I don't often have large blocks of data to transfer, so it's not that big a deal.


Share me

Digg  Facebook  SlashDot  Delicious  Technorati  Twitter  Google  Yahoo
Smf

 

-->