Author Topic: GCC compiler optimisation  (Read 100977 times)

0 Members and 19 Guests are viewing this topic.

Offline brucehoult

  • Super Contributor
  • ***
  • Posts: 6464
  • Country: nz
Re: GCC compiler optimisation
« Reply #300 on: February 11, 2023, 07:12:12 am »
I don't think so at all. This is a perfect example. Normal stupid people don't think about DMA vs. CPU bus arbitration at all, and everything just works out because the engineers at ARM / ST did a sane design, for those stupid people to use.

The same situation applies, after all, to multiple CPUs sharing a memory system.

A DMA unit is just a special purpose CPU with a very limited instruction set.
 

Offline Nominal Animal

  • Super Contributor
  • ***
  • Posts: 8349
  • Country: fi
    • My home page and email address
Re: GCC compiler optimisation
« Reply #301 on: February 11, 2023, 07:35:17 am »
Words are hard, too.  Well, actually fuzzy and vague and not well defined or "solid" at all.

You also haven't met stupid until you meet a student who excitedly and happily tells you they managed to chop off all their fingers, so they wouldn't have to go to shop class, and could now concentrate better on their dream: becoming a professional concert pianist.
 

Online peter-hTopic starter

  • Super Contributor
  • ***
  • Posts: 6020
  • Country: gb
  • Doing electronics since the 1960s...
Re: GCC compiler optimisation
« Reply #302 on: February 11, 2023, 07:46:11 am »
Quote
because you came up with your own search term which does not make sense ("interruptible")

Hmm, no.

Quote
I suggest you learn to instrument your code

Hmm, have been doing that for 40+ years.
Z80 Z180 Z280 Z8 S8 8031 8051 H8/300 H8/500 80x86 90S1200 32F417
 

Offline wek

  • Frequent Contributor
  • **
  • Posts: 591
  • Country: sk
Re: GCC compiler optimisation
« Reply #303 on: February 11, 2023, 08:23:24 am »
The STM32 dual-port DMA used in STM32F4xx is described in AN4031. Chapter 2 deals with both arbitration at the bus matrix, and arbitration within the DMA itself.

Both are round-robin, so DMA won't hog the system.

There's also timing information there. M2M is nothing special, it works exactly the same as P2M, except the triggers are provided internally so that when one transfer finishes the next is started (in fact you can achieve the same effect by triggering transfers from sources which are not cleared by the transfer, e.g. using a SPI-RXNE-triggered transfer which does read from given SPI's data register (how do I know that...)). In ideal situation, one transfer would IMO take 3 or 4 cycles (I'm not ARM/ST insider, not exactly sure about the arbitration delays) if source and destination lay in different memories. The ideal situation should be easy to benchmark. How far real world is from the ideal situation, is of course dependent on dozens of pesky details.

JW
 
The following users thanked this post: peter-h

Online peter-hTopic starter

  • Super Contributor
  • ***
  • Posts: 6020
  • Country: gb
  • Doing electronics since the 1960s...
Re: GCC compiler optimisation
« Reply #304 on: February 11, 2023, 09:11:22 am »
Quote
one transfer would IMO take 3 or 4 cycles

That's interesting, since it suggests that software is no slower, especially if you do 4 bytes in one go.

DMA, byte at a time, at 168MHz, I would expect to be similar to an optimised loop.

Getting back to optimisation, I found that even -Og (my default mode for the whole project, with -O0 used for some functions) removes the four consecutive moves found in some memcpy sources. Stepping through assembler, I had one case where 2 of the 4 were removed (and the loop counter halved). Quite weird. I think to get a real unrolled loop you would need to use -O0 and perhaps "register" on some values.
Z80 Z180 Z280 Z8 S8 8031 8051 H8/300 H8/500 80x86 90S1200 32F417
 

Online ataradov

  • Super Contributor
  • ***
  • Posts: 12473
  • Country: us
    • Personal site
Re: GCC compiler optimisation
« Reply #305 on: February 11, 2023, 09:23:45 am »
How are you doing copy in software 3 or even 4 cycles? Load and store instructions on CM4 take 2 cycles. Then you need to decrement the counter and branch, which is 2-4 cycles if taken, since it does pipeline flush and there is no branch predictor in CM4. And this is for zero wait state code memory or flash accelerator working out of its mind.

Note that consecutive load/store instructions of the same type may pipeline, so if you are going to do optimized software version, it is better to do 4 or 8 consecutive loads and then 4 or 4 stores. This amortizes the loop counter decrement and branch penalty.

DMA will always be faster. Also if you transfer between different SRAMx slaves, then it would pipeline and provided no other masters want to access the same SRAM, it would be 1 transfer/cycle sustained. This is not a likely scenario though, especially placement of buffers in different SRAMs.

I don't personally want to participate in optimization stuff anymore. I can't take that anymore.
« Last Edit: February 11, 2023, 09:34:08 am by ataradov »
Alex
 

Offline Nominal Animal

  • Super Contributor
  • ***
  • Posts: 8349
  • Country: fi
    • My home page and email address
Re: GCC compiler optimisation
« Reply #306 on: February 11, 2023, 10:10:53 am »
How are you doing copy in software 3 or even 4 cycles? Load and store instructions on CM4 take 2 cycles.
Consider
Code: [Select]
loop:
    ldr r0, [r1], #4 ; 1 cycle (pipelined with following str)
    str r0, [r2], #4 ; 2 cycles
    cmp r1, r3       ; 1 cycle
    blt loop         ; 2 cycles
Each loop iteration takes 6 cycles, right?  As described in Cortex M4 Technical Reference Manual, 3.3.2 Load/Store Timings.  (I might be wrong, though, because it does not explicitly describe the address post-increment case, [register],#increment, and only refers to the [register,offset] forms, with paired ldr-str taking three cycles.)

If the loop is unwound, say
Code: [Select]
loop:
    ldr r0, [r1], #4 ; 1 cycle (pipelined with following str)
    str r0, [r2], #4 ; 2 cycles
    ldr r0, [r1], #4 ; 1 cycle (pipelined with following str)
    str r0, [r2], #4 ; 2 cycles
    ldr r0, [r1], #4 ; 1 cycle (pipelined with following str)
    str r0, [r2], #4 ; 2 cycles
    ldr r0, [r1], #4 ; 1 cycle (pipelined with following str)
    str r0, [r2], #4 ; 2 cycles
    cmp r1, r3       ; 1 cycle
    blt loop         ; 2 cycles
then we do four copies for every 15 cycles in the optimum case, averaging to 3.75 cycles per word copied, or about 0.9375 cycles per byte copied.

I don't personally want to participate in optimization stuff anymore. I can't take that anymore.
Micro-optimizations like these are almost certainly wasted time, although optimizing memcpy() for ones own hardware might be worth the effort. 

What really bugs me is that the compiler always calls the generic version, even when it knows the pointer alignment and size beforehand.  It could use say memcpy_alignedN_sizeM() for a few different N and M, and just have them as weak symbols that map to memcpy(). Optimizing those might make a difference.

Cramming everything into calling a single function, ignoring all the knowledge about the pointers and size, and then trying to optimize that single function is a bit lunatic, if you think about it carefully.
 

Offline Kalvin

  • Super Contributor
  • ***
  • Posts: 2175
  • Country: fi
  • Embedded SW/HW.
Re: GCC compiler optimisation
« Reply #307 on: February 11, 2023, 10:28:10 am »
Using DMA for a generic memcpy will complicate things in systems with a preemptive scheduler (as the DMA needs to be shared & synchronized between all tasks in the system), and will make code less portable. I would just keep things simple, and do some manual loop unrolling, resulting simple, portable implementation. For example, manual unrolling by 8 will reduce the loop checking and jump from N to N/8, keeping the pipeline penalty very low.

Code: [Select]
#define MEMCPY_BLOCK_SIZE (8)

void *memcpy_my(void *restrict dest, const void *restrict src, int len)
{
    char *dp = (char *restrict)dest;
    const char *sp = (const char *restrict)src;

    while(len >= MEMCPY_BLOCK_SIZE)
    {
        *dp++ = *sp++; // 1
        *dp++ = *sp++; // 2
        *dp++ = *sp++; // 3
        *dp++ = *sp++; // 4
        *dp++ = *sp++; // 5
        *dp++ = *sp++; // 6
        *dp++ = *sp++; // 7
        *dp++ = *sp++; // 8
        COMPILETIME_ASSERT(MEMCPY_BLOCK_SIZE == 8);
        len -= MEMCPY_BLOCK_SIZE;
    }

    while( len > 0)
    {
        len--;
        *dp++ = *sp++;
    }

    return dest;
}

macro COMPILETIME_ASSERT is just there as a place-holder for the actual compile-time assert, and will warn if the MEMCPY_BLOCK_SIZE and the actual implementation will get out of sync.
 

Online peter-hTopic starter

  • Super Contributor
  • ***
  • Posts: 6020
  • Country: gb
  • Doing electronics since the 1960s...
Re: GCC compiler optimisation
« Reply #308 on: February 11, 2023, 11:44:58 am »
Code: [Select]

        *dp++ = *sp++; // 1
        *dp++ = *sp++; // 2
        *dp++ = *sp++; // 3
        *dp++ = *sp++; // 4
        *dp++ = *sp++; // 5
        *dp++ = *sp++; // 6
        *dp++ = *sp++; // 7
        *dp++ = *sp++; // 8

The problem I found is that for most optimisation levels most of these eight (or four, etc) get removed (and the loop counter is adjusted accordingly). So to be sure you are getting what you want, and the code is proof against future compiler versions, you would have to do it in assembler.

And it is no use saying that if your code breaks with say -O3 then it is broken. Lots of people here have been posting that. This is not true at all if talking to hardware and having to e.g. meet CS timings. The reality of embedded is that these things do matter, and real detail often matters.

Very good point about a DMA version of memcpy needing a mutex. So one would do a DMA version only for a particular thread context. In this case I was trying to optimise the low_level_* functions which interface ETH PHY to LWIP. Some versions of this use zero-copy but that is way too convoluted for me to understand (and lots of people evidently have trouble with it).

« Last Edit: February 11, 2023, 11:54:57 am by peter-h »
Z80 Z180 Z280 Z8 S8 8031 8051 H8/300 H8/500 80x86 90S1200 32F417
 

Offline Kalvin

  • Super Contributor
  • ***
  • Posts: 2175
  • Country: fi
  • Embedded SW/HW.
Re: GCC compiler optimisation
« Reply #309 on: February 11, 2023, 12:06:19 pm »
The problem I found is that for most optimisation levels most of these eight (or four, etc) get removed (and the loop counter is adjusted accordingly). So you would have to do it in assembler.

Yes, I noticed that also while playing with that particular piece of code in godbolt.org. It is possible to use a suitable #pragma in the source code, or isolate the function into its own compilation unit so that the compiler can then use the best optimization for that particular function. Both solutions will probably require target- and compiler-specific pragmas / compiler options. But it is doable with little effort.

For GCC, the following seems to work with different compiler optimization levels so that the unrolled loop will not get optimized away:
Code: [Select]
#pragma GCC push_options
#pragma GCC optimize "-Os"
#pragma GCC optimize "-fno-tree-loop-distribute-patterns"
void *memcpy_my(void *restrict dest, const void *restrict src, int len)
{
   // body here
}
#pragma GCC pop_options
 

Offline Siwastaja

  • Super Contributor
  • ***
  • Posts: 11224
  • Country: fi
Re: GCC compiler optimisation
« Reply #310 on: February 11, 2023, 02:31:57 pm »
Very good point about a DMA version of memcpy needing a mutex.

To circumvent the difficulty of parallel programming, you can always just set up DMA and poll for its completion, doing nothing useful on CPU in the meantime. Wasting the parallel nature of the DMA seems like counterintuitive waste, but if and when DMA is faster than CPU copy, you still save time; even if wasting the possibility of saving even more time.

There can be useful combinations, like start copying aligned words on DMA first, and while that is running, let the CPU copy the few odd bytes.
 

Online ataradov

  • Super Contributor
  • ***
  • Posts: 12473
  • Country: us
    • Personal site
Re: GCC compiler optimisation
« Reply #311 on: February 11, 2023, 05:56:03 pm »
then we do four copies for every 15 cycles in the optimum case, averaging to 3.75 cycles per word copied, or about 0.9375 cycles per byte copied.
2 cycles per branch is optimistic. But otherwise I agree. At the same time nothing stops DMA from doing 4 bytes/cycle in the same setup assuming addresses are aligned. And if addresses are not aligned, then timings for ldr/str are wrong, since they would be converted into multiple aligned accesses internally, so will take longer.

What really bugs me is that the compiler always calls the generic version, even when it knows the pointer alignment and size beforehand.  It could use say memcpy_alignedN_sizeM() for a few different N and M, and just have them as weak symbols that map to memcpy(). Optimizing those might make a difference.
But if you don't mess with compiler setting to much and let it use builtin memcpy(), then it will do almost the same. It does not exactly use different versions, but the version you get quickly checks the alignment and proceeds to do fast transfers if possible.
Alex
 
The following users thanked this post: boB

Offline wek

  • Frequent Contributor
  • **
  • Posts: 591
  • Country: sk
Re: GCC compiler optimisation
« Reply #312 on: February 11, 2023, 07:16:07 pm »
Some experimental results - 0x1000 word transfers on Disco F4, reset clock settings.

Code: [Select]
                           FIFO   P(src) M(dst)
EXPERIMENT  transfer       thrsh. burst  burst  => t
         0  SRAM1->SRAM2   full     1      1       0x4c2a
         1  SRAM2->SRAM2   full     1      1       0x4c2a
         2  SRAM1->SRAM2   1/4      1      1       0x401e
         3  SRAM2->SRAM2   1/4      1      1       0x401e
         4  SRAM1->SRAM2   full     4      4       0x3819
         5  SRAM1->SRAM2   full     4      1       0x581d
         6  SRAM1->SRAM2   full     1      4       0x4c1d

(I took some old blinky experiment, that's what the timer stuff is there, just ignore it)

JW
 
The following users thanked this post: peter-h

Online ataradov

  • Super Contributor
  • ***
  • Posts: 12473
  • Country: us
    • Personal site
Re: GCC compiler optimisation
« Reply #313 on: February 11, 2023, 08:01:01 pm »
So, about 4 cycles per transfer.

With 4K words test block you are violating the requirement to not cross 1K boundary for burst transfers. I don't think this would affect timings, but the data is likely to wrap around to 1K boundary, so the code is only useful for time testing. Although with completely aligned blocks bursts would not cross the boundary, so in this case the test is valid even for data.

And I think you might get better performance doing single transfers with 1/2 FIFO threshold. In theory in this scenario the only masters requesting the bus access to SRAM are DMA interfaces, so they would be granted on each cycle. So if you use FIFO as an actual FIFO instead of accumulator for a burst transfer, you should get the state where reading master always has space to place the next value and the writing master always has data in the FIFO. This would not work if DMA does not try to pipeline its requests without bursts though. Which is entirely possible if it is oriented at peripheral transfers.

And just for a teat, I would try to add a block of 8-16 nops in the register poll loop. Just to see if the core generating bus requests somehow affects the result. It should not, but who knows.

« Last Edit: February 11, 2023, 08:11:56 pm by ataradov »
Alex
 

Offline wek

  • Frequent Contributor
  • **
  • Posts: 591
  • Country: sk
Re: GCC compiler optimisation
« Reply #314 on: February 11, 2023, 08:08:27 pm »
> With 4K words test block you are violating the requirement to not cross 1K boundary for burst transfers.

These are 4-beat word bursts, so it's enough to align them at 32-byte boundaries to be sure not to cross the 1k boundary.

> And I think you might get better performance doing single transfers with 1/2 FIFO threshold.

a moment please...

JW

 
The following users thanked this post: peter-h

Offline wek

  • Frequent Contributor
  • **
  • Posts: 591
  • Country: sk
Re: GCC compiler optimisation
« Reply #315 on: February 11, 2023, 08:26:32 pm »
Code: [Select]
                           FIFO   P(src) M(dst)                disturb
EXPERIMENT  transfer       thrsh. burst  burst  => t           1=SRAM2wr 2=SRAM1wr 3=SRAM1rd
         0  SRAM1->SRAM2   full     1      1       0x4c2a      0x502b    0x502b    0x502b
         1  SRAM2->SRAM2   full     1      1       0x4c2a      0x502b    0x4c25    0x4c25
         2  SRAM1->SRAM2   1/4      1      1       0x401e      0x4023
         3  SRAM2->SRAM2   1/4      1      1       0x401e      0x4023
         4  SRAM1->SRAM2   full     4      4       0x3819      0x4016    0x401a    0x401a
         5  SRAM1->SRAM2   full     4      1       0x581d      0x601a    0x601a
         6  SRAM1->SRAM2   full     1      4       0x4c1d      0x4c1c
         7  SRAM1->SRAM2   1/2      1      1       0x4024
----
SW loop   => t2 = 0x6005

  for (uint32_t i = 0; i < 0x1000; i++) {
    *(volatile uint32_t *)0x2001C000 = *(volatile uint32_t *)0x20001000;
  }
 80002ac: f44f 5280 mov.w r2, #4096 ; 0x1000
 80002b0: 6829      ldr r1, [r5, #0]
 80002b2: 6021      str r1, [r4, #0]
 80002b4: 3a01      subs r2, #1
 80002b6: d1fb      bne.n 80002b0 <main+0xdc>

Instead of nops in the DMA-done-wait loop, I added read/write to the SRAMs. In cases, where it made no difference, tried to add/remove a couple (literally) of NOPs but that didn't make any difference either.

One of the things which is sort of a mystery to me is the non-integer number of cycles per transfer.

The software loop is 6-cycle.

JW
 
The following users thanked this post: peter-h

Online ataradov

  • Super Contributor
  • ***
  • Posts: 12473
  • Country: us
    • Personal site
Re: GCC compiler optimisation
« Reply #316 on: February 11, 2023, 08:36:17 pm »
The only explanation I have for this is that DMA does not pipeline its bus access. It waits for the completion of the transfer and then starts a new one. In this case maximum length bust would give the best performance, since at least data phase of the transfer would be pipelined. This is not unexpected from a general purpose DMA, I guess. Most of the time it works with peripherals where pipelined access makes no sense.

As far as non-integer number of cycles, I would exclude the register setup from measurement. Just measure the time from channel enable to the loop end. Different constants may generate slightly different versions of the code, which changes alignment.  And may be try different transfer size to see how it scales per transfer overhead vs fixed overhead.

And the code one is unfair, since it is just a word transfer with no increments. I don't expect write-back versions of the instructions to be slower, but who knows. At the very least, ldr/str instructions would be 32-bit instructions, which might matter for the code size and it fitting into the flash accelerator buffers, which would be an issue at real clock speeds.
« Last Edit: February 11, 2023, 08:44:50 pm by ataradov »
Alex
 
The following users thanked this post: wek

Online peter-hTopic starter

  • Super Contributor
  • ***
  • Posts: 6020
  • Country: gb
  • Doing electronics since the 1960s...
Re: GCC compiler optimisation
« Reply #317 on: February 11, 2023, 09:19:07 pm »
So DMA is roughly 3/2 i.e. 1.5x faster. Very useful to know!

The application area where this is really important is going to be very narrow, especially given that it is a single-thread proposition.

In my product I have moved all SPI transfers to DMA. Even the very slow ones, for uniformity. Runs very nicely. But all the SPI peripherals are mutex protected for thread safety, and the SPI is auto re-initialised (inside the mutex region, obviously) when peripherals are switched. Took a while to get this to work right. One could do the same with mem-mem DMA. But then my SPI can run up to 21MHz clock and - some past threads - software can't keep up with that, largely because peripherals run very slowly so their register access is slow.
« Last Edit: February 11, 2023, 10:07:29 pm by peter-h »
Z80 Z180 Z280 Z8 S8 8031 8051 H8/300 H8/500 80x86 90S1200 32F417
 

Offline wek

  • Frequent Contributor
  • **
  • Posts: 591
  • Country: sk
Re: GCC compiler optimisation
« Reply #318 on: February 11, 2023, 10:32:50 pm »
> The only explanation I have for this [...]

I'm afraid we don't have enough information to make definitive conclusions - and, more importantly, predictions. What would be impact of changing transfer width at either side? What's the impact of slow memories (FLASH,  external)? What's the impact of other streams of the same DMA running concurrently? All this could be benchmarked but the data would be overwhelming and even harder to interpret.

Of course, it would be far better if an insider would explain the exact working of this DMA clearly and concisely; but ST is not willing to engage in anything beyond clicking in CubeMX. I digress.

> As far as non-integer number of cycles, I would exclude the register setup from measurement

That's quite obviously the cca 0x20 portion of the timing, the true timing would be most probably 0xXY00. But still, Y is nonzero, and we are talking 0x1000 transfers (e.g. in first case, 0x4c00 cycles per 0x1000 transfers, so its 4 and 3/4 of cycle per transfer, or, 3 transfers out of 4 take 5 cycle and the fourth 4 cycle).

> And the code one is unfair

True.

I couldn't coax gcc into doing what I wanted, it kept incrementing only one of the pointers and adding a constant to it to get the other, this resulted in 8 cycles per transfer. I concocted this (please don't laugh too loudly, I am not expert at gcc inline asm):
Code: [Select]
    uint32_t a1 = 0, a2 = 0, a3 = 0, a4 = 0;

    __asm__ volatile (
       "ldr  %[src], =0x20001000"  "\n\t"
       "ldr  %[dst], =0x2001C000"  "\n\t"
       "mov  %[cnt], 0x1000"       "\n"
"1:\t" "ldr  %[tmp], [%[src]], #4" "\n\t"
       "str  %[tmp], [%[dst]], #4" "\n\t"
       "subs %[cnt], #1"           "\n\t"
       "bne  1b"
       : [src] "=r" (a1)
       , [dst] "=r" (a2)
       , [cnt] "=r" (a3)
       , [tmp] "=r" (a4)
    );
and this yields 7 cycles per transfer. More precisely 0x7006, yes some cycles for setup etc.

This runs at the default 16MHz clock so no waitstates. Without going through the hassle of firing up the crystal oscillator and PLL (which is irrelevant for timing), I just set the latencies and enable the jumpcache a.k.a. "ART accelerator"
Code: [Select]

#if (1) // just mimic high gear by setting latency
  FLASH->ACR =  0
    | FLASH_ACR_ICEN          // meantime, configure flash for maximum performance - enable both caches
    | FLASH_ACR_DCEN
    | FLASH_ACR_PRFTEN        // and enable prefetch
    | FLASH_ACR_LATENCY_5WS   // for VCC>2.7V and FSYS>150MHz, 5 waitstate is appropriate
  ;
  RCC->CFGR = (RCC->CFGR & (~(0
    | RCC_CFGR_PPRE1
    | RCC_CFGR_PPRE2
  ))) | (0
    | RCC_CFGR_PPRE1_DIV4   // APB1 prescaler set to 4 -> APB1 clock = 45MHz (required to be <= 45MHz)
    | RCC_CFGR_PPRE2_DIV2   // APB2 prescaler set to 2 -> APB2 clock = 90MHz (required to be <= 90MHz)
  );
#endif
and... the result is 0x7016, so the literal fetches and the first jump suffered the latency, the remaining 0xFFF jumps were served from the jumpcache; plus maybe some of the setup code exhausted the prefetch or something similar.

JW
« Last Edit: February 11, 2023, 11:21:03 pm by wek »
 
The following users thanked this post: peter-h

Offline abyrvalg

  • Frequent Contributor
  • **
  • Posts: 898
  • Country: es
Re: GCC compiler optimisation
« Reply #319 on: February 12, 2023, 11:07:11 am »
Assembly implementations could utilize LDM/STM instructions to move more data per loop iteration. Some common library memcpy() (ARMCC’s? Not sure which one) chooses between 3 loops based on data size/alignment - 16/4/1.
 

Online peter-hTopic starter

  • Super Contributor
  • ***
  • Posts: 6020
  • Country: gb
  • Doing electronics since the 1960s...
Re: GCC compiler optimisation
« Reply #320 on: February 12, 2023, 01:33:22 pm »
An asm version of the memcpy I posted above (which moves 32 bits at a time until the last few (or 0) residual bytes) would be very useful. I have done too little arm32 asm to work it out myself though.

On the wider topic I am surprised that Cube doesn't come with this stuff. After all, you specify the CPU in a config file, and that is used to throw in different libs and all kinds of stuff.

Predictably, this was discussed before
https://www.eevblog.com/forum/programming/should-memcmp-be-faster-than-a-loop/

This is the asm from my above C function, with -Og

Code: [Select]
         
memcpy_fast:
0804ff4c:   cmp     r2, #3
0804ff4e:   bhi.n   0x804ff6c <memcpy_fast+32>
0804ff50:   mov     r3, r0
0804ff52:   b.n     0x804ff8e <memcpy_fast+66>
0804ff54:   ldrb.w  r2, [r1], #1
0804ff58:   strb.w  r2, [r3], #1
0804ff5c:   mov     r2, r12
0804ff5e:   add.w   r12, r2, #4294967295
0804ff62:   cmp     r2, #0
0804ff64:   bne.n   0x804ff54 <memcpy_fast+8>
0804ff66:   ldr.w   r4, [sp], #4
0804ff6a:   bx      lr
0804ff6c:   mov     r3, r0
0804ff6e:   cmp     r2, #3
0804ff70:   bls.n   0x804ff8e <memcpy_fast+66>
0804ff72:   push    {r4}
0804ff74:   ldr.w   r4, [r1], #4
0804ff78:   str.w   r4, [r3], #4
0804ff7c:   subs    r2, #4
0804ff7e:   cmp     r2, #3
0804ff80:   bhi.n   0x804ff74 <memcpy_fast+40>
0804ff82:   b.n     0x804ff5e <memcpy_fast+18>
0804ff84:   ldrb.w  r2, [r1], #1
0804ff88:   strb.w  r2, [r3], #1
0804ff8c:   mov     r2, r12
0804ff8e:   add.w   r12, r2, #4294967295
0804ff92:   cmp     r2, #0
0804ff94:   bne.n   0x804ff84 <memcpy_fast+56>
0804ff96:   bx      lr

With -O0 you get this

Code: [Select]
          memcpy_fast:
0804ff4c:   cmp     r2, #3
0804ff4e:   bhi.n   0x804ff6c <memcpy_fast+32>
395        char *dst = dst0;
0804ff50:   mov     r3, r0
0804ff52:   b.n     0x804ff8e <memcpy_fast+66>
420        *dst++ = *src++;
0804ff54:   ldrb.w  r2, [r1], #1
0804ff58:   strb.w  r2, [r3], #1
419        while (len0--)
0804ff5c:   mov     r2, r12
0804ff5e:   add.w   r12, r2, #4294967295
0804ff62:   cmp     r2, #0
0804ff64:   bne.n   0x804ff54 <memcpy_fast+8>
424       }
0804ff66:   ldr.w   r4, [sp], #4
0804ff6a:   bx      lr
404        aligned_dst = (uint32_t*)dst;
0804ff6c:   mov     r3, r0
407        while (len0 >= 4)
0804ff6e:   cmp     r2, #3
0804ff70:   bls.n   0x804ff8e <memcpy_fast+66>
393       {
0804ff72:   push    {r4}
409        *aligned_dst++ = *aligned_src++;
0804ff74:   ldr.w   r4, [r1], #4
0804ff78:   str.w   r4, [r3], #4
410        len0 -= 4;
0804ff7c:   subs    r2, #4
407        while (len0 >= 4)
0804ff7e:   cmp     r2, #3
0804ff80:   bhi.n   0x804ff74 <memcpy_fast+40>
0804ff82:   b.n     0x804ff5e <memcpy_fast+18>
420        *dst++ = *src++;
0804ff84:   ldrb.w  r2, [r1], #1
0804ff88:   strb.w  r2, [r3], #1
419        while (len0--)
0804ff8c:   mov     r2, r12
0804ff8e:   add.w   r12, r2, #4294967295
0804ff92:   cmp     r2, #0
0804ff94:   bne.n   0x804ff84 <memcpy_fast+56>
0804ff96:   bx      lr

I've done megabytes of asm but struggle with the weird ways of the GCC compiler. Obviously stuff like this

Code: [Select]
aligned_dst = (uint32_t*)dst;
aligned_src = (uint32_t*)src;

is just for the benefit of C and if writing in asm you hardly consider it; a register value can be used as an address directly.
« Last Edit: February 12, 2023, 02:03:54 pm by peter-h »
Z80 Z180 Z280 Z8 S8 8031 8051 H8/300 H8/500 80x86 90S1200 32F417
 

Offline abyrvalg

  • Frequent Contributor
  • **
  • Posts: 898
  • Country: es
Re: GCC compiler optimisation
« Reply #321 on: February 12, 2023, 03:04:32 pm »
The library memcpy mentioned earlier:
Code: [Select]
__rt_memcpy                            ; unaligned version
                CMP             R2, #3
                BLS             loc_59F00
                ANDS            R12, R0, #3
                BEQ             loc_59ED2
                LDRB            R3, [R1],#1
                CMP             R12, #2
                ADD             R2, R12
                IT LS
                LDRBLS          R12, [R1],#1
                STRB            R3, [R0],#1
                IT CC
                LDRBCC          R3, [R1],#1
                SUB             R2, R2, #4
                IT LS
                STRBLS          R12, [R0],#1
                IT CC
                STRBCC          R3, [R0],#1

loc_59ED2                               ; CODE XREF: __rt_memcpy+A↑j
                ANDS            R3, R1, #3
                BEQ             __aeabi_memcpy8
                SUBS            R2, #8

loc_59EDC                               ; CODE XREF: __rt_memcpy+54↓j
                BCC             loc_59EF0
                LDR             R3, [R1],#4
                SUBS            R2, #8
                LDR             R12, [R1],#4
                STM             R0!, {R3,R12}
                B               loc_59EDC
; ---------------------------------------------------------------------------

loc_59EF0                               ; CODE XREF: __rt_memcpy:loc_59EDC↑j
                ADDS            R2, R2, #4
                ITT PL
                LDRPL           R3, [R1],#4
                STRPL           R3, [R0],#4
                NOP 

loc_59F00                               ; CODE XREF: __rt_memcpy+2↑j
                LSLS            R2, R2, #0x1F
                ITT CS
                LDRBCS          R3, [R1],#1
                LDRBCS          R12, [R1],#1
                IT MI
                LDRBMI          R2, [R1],#1
                ITT CS
                STRBCS          R3, [R0],#1
                STRBCS          R12, [R0],#1
                IT MI
                STRBMI          R2, [R0],#1
                BX              LR
; End of function __rt_memcpy

__aeabi_memcpy8          ; 8-byte aligned version       
                PUSH            {R4,LR}
                SUBS            R2, #0x20
                BCC             loc_59FC6

loc_59FB0                               ; CODE XREF: __aeabi_memcpy8+1A↓j
                ; 32 bytes per iteration loop
                LDM             R1!, {R3,R4,R12,LR} 
                SUBS            R2, #0x20
                STM             R0!, {R3,R4,R12,LR}
                LDM             R1!, {R3,R4,R12,LR}
                STM             R0!, {R3,R4,R12,LR}
                BCS             loc_59FB0

loc_59FC6                               ; CODE XREF: __aeabi_memcpy8+4↑j
                MOVS            R12, R2,LSL#28
                ; 16-byte remainder
                ITT CS
                LDMCS           R1!, {R3,R4,R12,LR}
                STMCS           R0!, {R3,R4,R12,LR}
                ; 8-byte remainder
                ITT MI
                LDMMI           R1!, {R3,R4}
                STMMI           R0!, {R3,R4}
                POP             {R4,LR}
                MOVS            R12, R2,LSL#30
                ; 4-byte remainder
                ITT CS
                LDRCS           R3, [R1],#4
                STRCS           R3, [R0],#4
                IT EQ
                BXEQ            LR
                LSLS            R2, R2, #0x1F
                ; 2-byte remainder
                IT CS
                LDRHCS          R3, [R1],#2
                ; 1-byte remainder
                IT MI
                LDRBMI          R2, [R1],#1
                IT CS
                STRHCS          R3, [R0],#2
                IT MI
                STRBMI          R2, [R0],#1
                BX              LR
; End of function __aeabi_memcpy8

The compiler emits direct __aeabi_memcpy8 calls when it knows the data alignment and __rt_memcpy in all other cases.
« Last Edit: February 12, 2023, 03:06:30 pm by abyrvalg »
 

Offline cv007

  • Super Contributor
  • ***
  • Posts: 1061
Re: GCC compiler optimisation
« Reply #322 on: February 12, 2023, 03:40:47 pm »
Simple question - why do you have to move big chunks of memory-memory rather than use something like a description block to prevent the need to do so?
 

Online peter-hTopic starter

  • Super Contributor
  • ***
  • Posts: 6020
  • Country: gb
  • Doing electronics since the 1960s...
Re: GCC compiler optimisation
« Reply #323 on: February 12, 2023, 04:57:53 pm »
Because I am moving between two "blocks" which I don't understand and which almost nobody else understands.

There is in this case a zero-copy approach but despite having been banging about for ~10 years is complicated and I have never seen code known to be working and robust (for the 32F4).
Z80 Z180 Z280 Z8 S8 8031 8051 H8/300 H8/500 80x86 90S1200 32F417
 

Offline Siwastaja

  • Super Contributor
  • ***
  • Posts: 11224
  • Country: fi
Re: GCC compiler optimisation
« Reply #324 on: February 12, 2023, 06:52:08 pm »
Yeah, zero-copy solutions are nice when you can properly engineer the whole thing. In practice, one needs to just memcpy things, but OTOH, thankfully, people also tend to overestimate the cost of doing so.

Zero-copy has not only the advantage of saving time; it also saves memory. All kind of stupid buffers can take a lot of RAM in embedded system, when every layer wants to have their own buffers.
 


Share me

Digg  Facebook  SlashDot  Delicious  Technorati  Twitter  Google  Yahoo
Smf

 

-->